Mechanical arm control model training method and system

By constructing a difference matrix to identify restricted sections and implement reverse correction, combined with energy accumulation identification and inertia attenuation compensation, the energy overload problem of the robotic arm under complex working conditions is solved, the trajectory continuity and system stability are improved, and efficient energy management is achieved.

CN120755889AInactive Publication Date: 2025-10-10SHENZHEN DEYI MEDICAL TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511211624.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-10-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing robotic arm control strategies lack a flexible dynamic adjustment mechanism when facing complex changes in working states, resulting in weakened trajectory segments and abnormal accumulation of control energy consumption, causing local short-term energy overload, and causing problems such as drive current saturation, torque jump, and path oscillation.

Method used

By constructing a difference matrix to identify the restricted section and implementing reverse correction, a reverse recursive correction function of the control input sequence is constructed by combining energy accumulation identification with inertia attenuation compensation. The control input sequence is corrected, and the control model parameters are optimized using the error function.

Benefits of technology

It realizes the progressive buffer reconstruction of the robot arm trajectory, improves the execution fluency, trajectory smoothness and posture stability, reduces the risk of abnormal energy accumulation, enhances the robustness and safety of the system, and has good energy prediction accuracy and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120755889A_ABST
    Figure CN120755889A_ABST
Patent Text Reader

Abstract

The invention discloses a mechanical arm control model training method and system, and relates to the technical field of industrial automation. Progressive buffer reconstruction of limited actions in a mechanical arm track is realized by constructing a parameter difference matrix before and after limiting, recognizing continuous limited data segments and combining a directional dynamic correction mechanism; the correction is fused with the limiting amplitude, the position of the maximum deviation point and a distance weight factor, and a direction sign is used for guiding a correction trend, so that the action parameters which are suddenly changed originally are gradually recovered to a reasonable interval for fitting an original track. And the continuity and controllability of the limited track section in the dynamic execution process are enhanced, the problems of shaking, pause, nonlinear oscillation and the like in the motion of the mechanical arm are effectively avoided, and the execution smoothness, the track smoothness and the attitude stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of industrial automation, in particular to a mechanical arm control model training method and system. BACKGROUND

[0002] With the wide application of mechanical arms in precise assembly, collaborative manufacturing and complex trajectory execution, the control system of the mechanical arm has higher requirements for the comprehensive adjustment ability of execution stability, trajectory following accuracy and energy response. Under this background, how to effectively correct the weakened trajectory segment while following the system limit constraint, dynamically control the energy release rate, and continuously optimize the control model based on the execution feedback has become a key technology direction to realize efficient and safe operation.

[0003] Most of the existing mechanical arm control strategies rely on fixed model structure and static control parameters, and lack flexible dynamic adjustment mechanism when facing complex working state changes (such as limit clipping and short-time energy peak). For example, in the actual trajectory execution process, a joint of the mechanical arm may be compressed due to the limit strategy trigger action, resulting in a significant weakening of the original control instruction, and this control change cannot be fed back to the control model in a timely manner, thereby causing structural deviation between the model output and the actual system response. Taking the typical current limiting strategy as an example, when the joint current output by the control model exceeds the driver limit value, the system can clip the instruction hard, but the control model does not know the existence of the weakening event, so it continues to output high-power instructions, causing abnormal accumulation of control energy, forming a local short-time energy overload phenomenon, and then triggering problems such as drive current saturation, torque jump and even path oscillation. SUMMARY

[0004] In view of the deficiencies of the prior art, the present application provides a mechanical arm control model training method and system, which solves the problems in the above background art.

[0005] To achieve the above purpose, the present application is realized by the following technical scheme: a mechanical arm control model training method, comprising,

[0006] Constructing a difference matrix and detecting a limited section, based on the limited section, performing reverse correction on the parameters after the limit to realize trajectory buffer reconstruction, the difference matrix is used to measure the difference between the mechanical arm before and after the limit;

[0007] After reverse correction, execute the action sequence of the mechanical arm, and judge whether there are early warning signs of energy accumulation exceeding the standard by synchronously recording the control input sequence of the upper controller issued to the driver and the execution data feedback by the driver; the action sequence refers to the updated limited section;

[0008] If there are, identify the risk level in the current action cycle and locate the energy abnormal trigger time;

[0009] Based on the energy anomaly triggering moment, the previous control time index is traversed in reverse to construct a reverse recursive correction function of the control input sequence to compensate for the sharp change of the continuous control input with inertia attenuation, and the control input sequence is corrected;

[0010] According to the corrected control input sequence, the relative reduction degree of the local transient energy index before and after correction is compared to obtain an energy consumption reduction feedback reinforcement sample set;

[0011] The energy consumption reduction feedback reinforcement sample set is taken as the input of the control model, and the control model is trained by back propagation of the model parameters using the error function to iteratively update the model parameters.

[0012] Preferably, the difference matrix is constructed and the restricted section is detected, including:

[0013] The behavior data in the corresponding action period of the mechanical arm is obtained in advance, including the pre-limit data set and the post-limit data set;

[0014] Based on the pre-limit data set, a two-dimensional array is constructed, and the columns of the two-dimensional array are the parameter dimensions and the behavior sampling time;

[0015] According to the limit strategy of each parameter, the difference amount before and after the limit at each element position in the two-dimensional array is determined to construct a difference matrix;

[0016] Each column in the difference matrix is traversed to identify a continuous non-zero data section, which is recorded as a restricted section;

[0017] Based on the restricted section, the corresponding post-limit parameter is corrected in reverse to quantify the correction condition of the gradual buffering of the corresponding parameter under the premise of maintaining the limit strategy of the corresponding parameter, and the trajectory buffering reconstruction is realized.

[0018] Preferably, based on the restricted section, the corresponding post-limit parameter is corrected in reverse to quantify the correction condition of the gradual buffering of the corresponding parameter under the premise of maintaining the limit strategy of the corresponding parameter, including:

[0019] For each restricted section, the maximum value is extracted, and the row number of the maximum value in the difference matrix and the positive or negative sign of the difference between the maximum value and other parameters in the corresponding restricted section are marked;

[0020] The row numbers corresponding to the parameters in the restricted section are recorded to obtain a row number set;

[0021] The minimum row number recorded in the row number set is taken as the starting row;

[0022] S101: For the starting row in each restricted section, the row number difference is calculated according to the distance between the starting row corresponding row number and the maximum value row number in the column;

[0023] S102: performing reverse correction on the corresponding limited parameter by taking the inverse form of the linear superposition of the limit front-back difference amount and the standardized row number difference value, and combining the positive and negative signs, to obtain a dynamic correction amount;

[0024] S103: based on the dynamic correction amount, correcting the corresponding limited parameter in the limited data set to obtain a corrected parameter;

[0025] Taking the next row in the corresponding limited section, repeating S101 to S103 until the last row in the row number set, to update the corresponding limited section;

[0026] The updated limited section is taken as the action sequence of the completed staged limiting correction in the current action period.

[0027] Preferably, the early warning signs of excessive energy accumulation in the current execution action include:

[0028] S201: In the process of each mechanical arm executing the action sequence, the control input sequence issued by the upper controller to the driver and the execution data fed back by the driver are recorded synchronously, and the execution data includes the actual driving current and joint torque;

[0029] S202: In the process of action execution, to dynamically quantify the energy response characteristics of the driving system within a short time window, the local transient energy index is constructed by integrating the product of the actual driving current and joint torque in the execution data, which is used to reflect the total amount of energy release in the current local stage of the mechanical arm;

[0030] If the local transient energy index exceeds 50% of the preset safety threshold, it is determined that the current execution action has early warning signs of excessive energy accumulation, the control time index at this time is recorded, and the input adjustment mechanism is activated, wherein the control time at this time is recorded as the energy anomaly triggering time.

[0031] Preferably, after activating the input adjustment mechanism, the risk level of the local transient energy index in the current action period is identified, including:

[0032] Based on the ratio relationship between the current local transient energy index and the safety threshold, three energy risk levels are divided, which are low risk interval, medium risk interval and high risk interval;

[0033] According to the corresponding energy risk level, the corresponding inertia dispersion factor value interval is configured, and the inertia dispersion factor corresponding to the current local transient energy index is obtained from the inertia dispersion factor value interval.

[0034] Preferably, based on the energy anomaly triggering moment that has been located, the reverse recursive correction function of the control input sequence is constructed by traversing the previous control time index, so as to compensate the sharp change of the continuous control input by inertia attenuation, including:

[0035] Suppose that the control time index corresponding to the energy anomaly triggering moment extracted from the original control input sequence is the local sequence {u k,k+Δt}, wherein u k,k+Δt is the system control quantity at the corresponding time t in the corresponding sliding window, k is the energy anomaly triggering moment, and k+Δt is the end time in the corresponding sliding window;

[0036] The energy anomaly triggering moment is recursively propagated to the nearest stable energy boundary point, and the compensation interval [k s , k] is constructed, wherein k s is the nearest stable energy boundary point, and the stable energy boundary point refers to the latest time point that meets the low-risk interval;

[0037] In the compensation interval [k s , k], the increments between the original local sequence at the corresponding time and the original local sequence at the previous time are fused by taking the inertia dispersion factor as the weight in a recursive manner, so as to generate the corrected local sequence;

[0038] The corrected local sequence is spliced with the original control input sequence segment that is not disturbed, so as to obtain the corrected control input sequence;

[0039] According to the corrected control input sequence, S201 and S202 are re-executed to compare the relative reduction degree of the local transient energy indicators before and after correction, and the energy consumption reduction feedback reinforcement sample set is obtained.

[0040] Preferably, according to the corrected control input sequence, S201 and S202 are re-executed to compare the relative reduction degree of the local transient energy indicators before and after correction, and the energy consumption reduction feedback reinforcement sample set is obtained, including:

[0041] S201 and S202 are re-executed to obtain the corrected local transient energy indicators;

[0042] The relative reduction degree of the local transient energy indicators before and after correction is compared to obtain a relative reduction ratio;

[0043] If the relative reduction ratio exceeds a preset proportion threshold, it is determined that the inertia compensation is effective, and the corresponding corrected control input sequence and the corrected local transient energy indicators are taken as a set of execution mapping samples, otherwise it is determined that the inertia compensation is ineffective, and the inertia dispersion factor is re-adjusted until the inertia compensation is effective;

[0044] Corresponding execution mapping samples in a plurality of action periods are counted to obtain an energy consumption reduction feedback reinforcement sample set.

[0045] Preferably, the energy consumption reduction feedback reinforcement sample set is taken as an input of the control model, and a target energy index is taken as a training output.

[0046] An error between a predicted energy index of the control model and the target energy index is taken as a loss function.

[0047] The control model is trained by using the error function to perform back propagation of model parameters, wherein a gradient descent method is used in the training process to iteratively update the model parameters until the control model can generate a control input sequence with stable energy-saving characteristics under different action periods.

[0048] A mechanical arm control model training system comprises,

[0049] A continuous correction subsystem constructs a difference matrix and detects a limited section, and based on the limited section, performs reverse correction on the corresponding limited parameters to realize trajectory buffer reconstruction, and the difference matrix is used to measure the difference amount of the mechanical arm before and after limiting.

[0050] An accumulation subsystem, after reverse correction, executes an action sequence on the mechanical arm, and judges whether there is an early warning sign of energy accumulation exceeding the standard by synchronously recording the control input sequence issued by the upper controller to the driver and the execution data feedback by the driver; the action sequence refers to the updated limited section.

[0051] A risk identification subsystem identifies the risk level in the current action period and locates the energy abnormality triggering time if there is any.

[0052] A correction subsystem, based on the energy abnormality triggering time, reversely traverses the previous control time index, constructs an inverse recursive correction function of the control input sequence, and performs inertia attenuation compensation on the sharp change of the continuous control input to correct the control input sequence.

[0053] A sample collection subsystem compares the relative reduction degree of the local transient energy index before and after correction according to the corrected control input sequence, and obtains an energy consumption reduction feedback reinforcement sample set.

[0054] A training subsystem takes the energy consumption reduction feedback reinforcement sample set as an input of the control model, and trains the control model by using the error function to perform back propagation of model parameters to iteratively update the model parameters.

[0055] The present application provides a mechanical arm control model training method and system, which has the following beneficial effects:

[0056] (1) The application realizes gradual buffering reconstruction of the limited action in the trajectory of the mechanical arm by constructing a limit front and back parameter difference matrix, identifying a continuous limited data segment, and combining a directional dynamic correction mechanism. The correction amount fuses the limit amplitude, the position of the maximum deviation point, and the distance weight factor, and guides the correction trend with the positive and negative signs of the direction, so that the originally mutated action parameters are gradually restored to the reasonable interval of the fitted original trajectory. This method not only retains the safety constraints of the original limit strategy of the system, but also enhances the continuity and controllability of the limited trajectory segment in the dynamic execution process, effectively avoiding problems such as shaking, pausing, and nonlinear oscillation in the movement of the mechanical arm, and improving the execution fluency, trajectory smoothness, and posture stability.

[0057] (2) The risk level identification and inertia correction mechanism based on the local transient energy index can trigger high-sensitivity response and adjustment action at the initial stage of energy accumulation abnormality in the control process of the mechanical arm. By introducing a multi-level energy risk division and a proportionally driven inertia dispersion factor regulation method, the system can flexibly generate control input adjustment strength of corresponding amplitude according to different energy consumption levels. The proposed reverse recursive correction function has low-pass filter type suppression characteristics, can smooth and attenuate the inertia of the mutated control instruction, effectively suppresses the risk of abnormal energy accumulation in a short time window, thereby significantly reduces the probability of joint saturation, current impact and control abnormality, and improves the robustness and safety of the system under complex working conditions.

[0058] (3) By dynamically comparing the control input and energy index before and after correction, an enhanced sample set is constructed with energy reduction effectiveness as the screening condition, only retaining the action sequence that actually triggers energy consumption reduction and does not lead to system instability as the model training sample, further enhancing the relevance and effectiveness of the training data. Cooperating with the model back propagation optimization method based on error function, the control model can gradually learn the energy-saving response strategy under different working conditions and different input change trends, and form an efficient control generation capability for energy consumption constraints through iterative training, so that the trained control model has good energy prediction accuracy, stability and generalization ability, and can be widely used in energy optimal path planning and efficient control output in various mechanical arm execution tasks. BRIEF DESCRIPTION OF DRAWINGS

[0059] Figure 1 It is a mechanical arm control model training method flowchart of the application;

[0060] Figure 2 It is a mechanical arm control model training system block diagram of the application. DETAILED DESCRIPTION

[0061] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0062] Embodiment 1

[0063] Please refer to Figure 1 The present application provides a mechanical arm control model training method, comprising,

[0064] A difference matrix is constructed and a restricted section is detected, and based on the restricted section, reverse correction is performed on the corresponding parameters after limiting to achieve trajectory buffer reconstruction. The difference matrix is used to quantify the difference between the mechanical arm before and after limiting.

[0065] By extracting the control parameters (such as joint target current, expected angle) before limiting and the parameters (actual allowed value when executed) after limiting, the difference between the two corresponding positions in the time dimension is calculated to form a two-dimensional difference matrix. Each column in the difference matrix represents a control parameter dimension (such as driving current, torque), and each row represents a sampling time. The difference matrix is used to measure the control difference between the original control intention and the actual execution permission.

[0066] Wherein, limiting refers to setting physical or software limits on the movement range, speed, acceleration, torque, etc. of the mechanical arm.

[0067] The restricted section refers to a time section with continuous significant difference, indicating that the control behavior in this time section is subject to limiting or force limiting clipping. For example, when the mechanical arm executes a high-speed grabbing task, due to the existence of joint current limiting strategy, the target current in a certain period of time is 4.5A, and the system will clip it to 3.0A due to protection limit. The continuous current difference in this period of time constitutes a restricted section. Through this step, the restricted control points in the control path can be accurately identified, providing a positioning basis for subsequent trajectory buffer correction, which helps to improve the trajectory continuity and motion controllability of the mechanical arm.

[0068] After reverse correction, the mechanical arm executes a motion sequence, and by synchronously recording the control input sequence issued by the upper controller to the driver and the execution data feedback by the driver, it is determined whether the current execution motion has early warning signs of energy accumulation exceeding the standard. The motion sequence refers to the updated restricted section.

[0069] After executing the reverse correction strategy based on the spatial gradual reduction characteristics on the above restricted section, it is input into the system as a new control instruction, executes the corrected motion, and synchronously collects the control input sequence issued by the upper controller and the execution data feedback by the driver, and records it as a control-execution data pair.

[0070] Wherein, the control input sequence refers to the current, position, speed, etc. expected instructions issued by the control system; the execution data refers to the actual current, joint torque and pose data fed back by the driving system. For example: for the section where the current is limited, adjust the target current to 3.2A, 3.4A, etc. gradually after re-execution, and the actual driving current and joint torque are collected in real time by the sensor, such as 3.1A. By constructing the control-execution data pair, the comparison record of system expectation-actual behavior is realized, which provides a data basis for subsequent judgment of whether there is a local energy overload risk.

[0071] If so, identify the risk level in the current action period, and locate the energy abnormal trigger time;

[0072] Wherein, the driving current and joint torque data recorded are multiplied and integrated to obtain the local transient energy index, which is compared with the safety threshold set by the system. If it exceeds a certain proportion (such as 50%), it is considered that there is a trend of energy accumulation, and the trigger time is recorded;

[0073] The local transient energy index reflects the short-term energy release intensity;

[0074] The energy abnormal trigger time refers to the time index point when the energy exceeds the limit for the first time.

[0075] For example: set the safety threshold to 300J, and the energy index in the sliding window is 180J (i.e. 60%) during the execution process, which indicates that there is a sign of energy accumulation, and the system records the control index at this moment t=320ms.

[0076] Through this section, early identification of rapid energy accumulation can be achieved to prevent actuator overload or trajectory disturbance caused by control signal jump in the short term, and improve system safety.

[0077] Based on the energy abnormal trigger time, the control input sequence is modified by constructing a reverse recursive correction function of the control input sequence by traversing the control time index of the previous section in reverse, to compensate for the sharp change of the continuous control input with inertia attenuation, and to modify the control input sequence.

[0078] Wherein, the historical time sequence of the control input is traced back from the energy abnormal trigger time as the starting point, and a recursive function with inertia dispersion factor is used to adjust the control input, smooth the jump slope, and reduce the sharp rising trend of energy.

[0079] The reverse recursive correction function refers to a mathematical model based on inertia dispersion factor to suppress the historical control change amplitude;

[0080] The inertia dispersion factor is used to reflect the proportional factor of the control input adjustment strength, and the value is determined according to the risk level partition.

[0081] For example, if an energy anomaly trigger is identified at t = 320 ms, the control input sequence from t = 270 ms to 320 ms is recursively corrected, so that the current change is changed from the original [2.5A→4.5A] to [2.5A→3.3A→3.9A→4.2A] and the like, which is used to prevent high amplitude mutation control instructions from being continuously applied to the driver, reduce short-time energy consumption peak and suppress system mechanical shock, and enhance the adaptability of the control system to complex working conditions.

[0082] According to the corrected control input sequence, the relative reduction degree of the local transient energy index before and after correction is compared, and an energy consumption reduction feedback reinforcement sample set is obtained;

[0083] The energy consumption reduction feedback reinforcement sample set is used as the input of the control model, and the error function is used for back propagation training of the model parameters of the control model to iteratively update the model parameters.

[0084] Specifically, the control input sequence after inertia compensation and effective reduction of the local transient energy index is used as the input and output pair to form an execution mapping sample, and the sample set is used to train the control model to optimize its prediction ability.

[0085] The energy consumption reduction feedback reinforcement sample set is a set of control-response mapping pairs with significantly reduced energy consumption after compensation and stable execution;

[0086] The error function is used to reflect the loss measure of the difference between the predicted energy value of the control model and the target energy value.

[0087] For example, the local energy consumption of 270J generated by the original input in a certain action period is reduced to 230J after compensation, so that the action period sample is classified as a reinforcement sample, and is used to optimize model parameters such as weight W and bias b, so as to improve the energy prediction accuracy and output robustness of the control model, so that the system can automatically generate energy-saving and high-stability control input sequences under multiple working conditions, realize system self-learning and adaptive ability evolution.

[0088] By constructing a difference matrix and detecting a limited section, the amplitude change of the control parameter before and after the limit can be identified, and gradual reverse correction can be implemented under the premise of maintaining the constraint of the limit strategy, so as to buffer and reconstruct the limited trajectory, solving the problem that the traditional model cannot perceive the deformation of the actual trajectory caused by the limit.

[0089] In the execution process of the action sequence, the upper control instruction and the driver feedback data are collected, and the energy accumulation anomaly is identified based on the local transient energy index, so as to realize a dynamic energy early warning mechanism of the control process, effectively improve the self-perception ability of the system to sudden energy consumption anomaly, and enhance the steady-state operation safety of the control system.

[0090] For the energy anomaly triggering moment, the inertia compensation function is constructed by backward recursion to flexibly compress and correct the sharply changing paragraphs in the control input sequence, avoiding the drive saturation and execution jump caused by the short-time energy burst, and improving the continuity of the control instruction and the stability of the system response.

[0091] By comparing and analyzing the local energy indicators before and after the correction, the control behavior samples with significant energy reduction effect are extracted to construct the energy consumption reduction feedback reinforcement sample set, ensuring that the training data highly fit the actual working conditions and enhancing the generalization ability of the model under complex energy response conditions.

[0092] After the training sample set is optimized by the error function, the model is iteratively updated using the parameter backpropagation mechanism, which can significantly improve the learning ability of the control model to energy feedback changes, so that the control output meets the execution accuracy while having higher energy saving and dynamic and steady-state control effect.

[0093] Embodiment 2

[0094] Please refer to Figure 1 , in particular: constructing a difference matrix and detecting a restricted section, including:

[0095] Obtaining behavior data within the corresponding motion period of the robot arm in advance, including a pre-limit data set and a post-limit data set;

[0096] The behavior data can be obtained by industrial bus listening, driver feedback acquisition, hardware sampling or controller export, etc., which constitutes the basis for subsequent inertia compensation, sample extraction and control model training;

[0097] Based on the pre-limit data set, a two-dimensional array is constructed, and the columns of the two-dimensional array are the parameter dimensions and the behavior sampling time;

[0098] Wherein, the parameter dimensions include but are not limited to the position, angular velocity, output current, end effector linear velocity and driver power density of each link of the robot arm;

[0099] According to the limit strategy of each parameter, the difference amount before and after the limit at each element position in the two-dimensional array is determined to construct a difference matrix; specifically: i,j =b i,j -a i,j , wherein c i,j is the difference amount before and after the limit at the i-th row and j-th column, b i,j is the pre-limit parameter at the i-th row and j-th column, a i,j is the post-limit parameter at the i-th row and j-th column, i is the column number of the two-dimensional array, and j is the row number of the two-dimensional array;

[0100] Limiting strategy refers to the range constraints, maximum allowable values, rate limits or absolute boundaries set for each parameter to ensure the physical safety, system stability and execution rationality of the robot arm during execution.

[0101] Traverse each column in the difference matrix, identify the continuous non-zero data segment, and record it as the restricted section (i.e. continuously exist limited motion feature points);

[0102] Based on the restricted section, the corresponding limited parameter is executed to correct the reverse, to quantify the correction condition of the gradual buffer of the corresponding parameter (restricted data) under the premise of maintaining the limiting strategy of the corresponding parameter (i.e. limiting boundary constraints), and to realize trajectory buffer reconstruction.

[0103] Each parameter here refers to the parameter type corresponding to each column in the two-dimensional array;

[0104] Motion cycle refers to the continuous control and motion process of the robot arm from the starting state to the target state in a complete trajectory execution unit, which is usually initiated by the upper task scheduling system or motion planner, and runs through a series of position, attitude, speed, acceleration and other change sequences, which is called a motion unit in the control cycle;

[0105] The limited data set before limiting refers to the original planning control data before the system executes limiting, limiting, limiting, etc. Strategy, usually generated by: path planning module, control model output module (such as trajectory interpolator, feedforward controller) ideal control parameters.

[0106] The limited data set after limiting refers to the actual delivery control instruction or state sampling value of the above control data before execution, which has been clipped by the limiting strategy of the system (including maximum angular velocity, maximum torque, current upper limit, joint limiting, etc.). It is the final allowed execution data, which is the "safe version" of the system due to safety, physical constraints and other factors.

[0107] The correction mechanism fuses the amplitude of the parameter difference before and after limiting with the distance of the relative extreme point in the time sequence dimension to generate a regulation amount with spatial decreasing characteristics, so as to realize the gradual buffer correction of the restricted data under the premise of guaranteeing the limiting boundary constraints.

[0108] Trajectory buffer reconstruction refers to the process of introducing correction or buffer factors, gradual correction mechanisms or nonlinear dynamic compensation strategies to locally adjust the limited trajectory data after the original planning trajectory is "cut down" or "mutated" due to limiting, saturation or hardware constraints, so that it gently returns to the original trajectory trend, gradually transitions and restores the structure, to realize the improvement of trajectory continuity, motion smoothness and control robustness.

[0109] Based on the limited section, the corresponding limited parameter is corrected in reverse to quantify the correction condition of the gradual buffer of the corresponding parameter (limited data) under the premise of maintaining the limiting strategy (i.e. limiting boundary constraint) of the corresponding parameter, including:

[0110] For each limited section, the maximum value in it is extracted, and the row number of the maximum value in the difference matrix and the positive and negative signs of the difference between the maximum value and other parameters in the corresponding limited section are marked;

[0111] The positive and negative signs of each column difference are to preserve the semantic information of the "correction direction" and prevent misadjustment in the reverse direction. For example, if the original instruction is to move in the positive direction (for example, a joint angle +5°), the limit becomes only +2°, and the difference is +3°, which means that it is less than the original instruction. At this time, the buffer correction should slowly compensate in the positive direction. If the positive and negative signs are not recorded, the subsequent correction may appear to be in the opposite direction. The positive sign indicates that the positive direction acceleration is cut off, indicating that the motion wants to accelerate execution. The negative sign indicates that the negative direction overshoot is cut off, indicating that there is a reverse overshoot trend;

[0112] The row numbers corresponding to the parameters in the limited section are recorded to obtain a row number set;

[0113] The minimum row number recorded in the row number set is taken as the starting row to be calculated;

[0114] S101: For the starting row in each limited section, the row number difference is calculated according to the distance between the starting row corresponding row number and the maximum value row number, specifically: d i = |r start -r max |, wherein d i is the row number difference corresponding to the i-th row, r start is the row number corresponding to the starting row, and r max is the row number of the maximum value;

[0115] S102: The reverse correction of the corresponding limited parameter is performed by linearly superimposing the difference before and after the limit with the standardized row number difference (i.e. distance term) to obtain a dynamic correction amount;

[0116] The dynamic correction amount is used to quantify the correction condition of the gradual buffer of the corresponding parameter (limited data) under the premise of maintaining the limiting strategy (i.e. limiting boundary constraint) of the corresponding parameter;

[0117] The dynamic correction amount is obtained in the following manner: Wherein, Z i,j is the dynamic correction amount at the i-th row and the j-th column;

[0118] c i,jThe limit difference before and after the i-th row and the j-th column (i.e., representing the locally weakened strength) is used to measure the weakening degree of the current row;

[0119] The sign(*) is a sign function (positive value is +1, negative value is -1), c extreme The maximum value of a certain limited section (i.e., representing the representative limited difference) is obtained.

[0120] The alpha is a decay coefficient constant, taking a value of 0.1, used for normalizing distance influence to maintain the scale consistency between the influence dimensions, ensuring that the correction behavior has gradient continuity and control adjustability in the amplitude direction and time sequence direction;

[0121] d i The row number difference corresponding to the i-th row is used to measure the distance from the corresponding position to the extreme point.

[0122] When the limit difference before and after is smaller and the distance is farther, the Z value is smaller (i.e., the correction is less), and when the point is closer to the maximum weakening point and the weakening degree is greater, the Z value is greater (i.e., the correction is more);

[0123] The sign(c extreme ) is used to record the direction consistency for subsequent adjustment.

[0124] Through distance adjustment, the row number difference drives the adjustment amount, and the farther the adjustment is greater.

[0125] The calculation of the Z value takes the representative maximum difference symbol in the limit difference before and after the matrix as the direction guide, combines the current difference amplitude and its distance relative to the maximum weakening point to construct a local decay coefficient, realizes a buffer correction mechanism with consistent direction, different strength, and position decay, which can restore the structural continuity and trend consistency of the trajectory without breaking the limit constraint, and improves the dynamic stability of the system and the authenticity of the learning sample.

[0126] The staged iteration means step-by-step local continuous adjustment from the starting point to the ending point.

[0127] S103: Based on the dynamic correction amount, the corresponding post-limit parameter in the post-limit data set is corrected to obtain a corrected parameter; specifically: a i,j ’ = a i,j -Z i,j , wherein a i,j ’ The corrected parameter of the i-th row and the j-th column is a i,j The post-limit parameter of the i-th row and the j-th column is a

[0128] The next row in the corresponding limited section is taken, and S101 to S103 are repeated until the last row in the row number set to update the corresponding limited section.

[0129] The updated restricted section is taken as the action sequence of the current action cycle, which has completed the phased limit correction, for subsequent driver energy evaluation and readjustment mechanism.

[0130] The dynamic correction amount is a trajectory correction strength factor with directionality and spatial gradual decrease.

[0131] The row number of the maximum value (the maximum value row number) is the row number of the maximum value extracted from the restricted section in the difference matrix;

[0132] Through this row-by-row, segment-by-segment, phased processing, the secondary disturbance introduced by the sudden correction due to the limit can be effectively reduced, and the continuous region is adjusted in an increasing distance, avoiding overcompensation or local divergence, and gradually pulling back to a reasonable interval, rather than a one-time hard pull back.

[0133] Traditional single limit will cause the data of the robot arm to appear a short time drastic callback, which may induce mechanical vibration, excessive acceleration peak value and other problems, the present application adopts: first overall limit, to form the limit data, then do phased fine adjustment based on the difference matrix, and according to the distance factor formed by the row number difference, the adjustment is gradually unfolded, so a more gentle, continuous and controllable fine correction process is obtained, which is very beneficial to the smoothness of the robot arm execution trajectory.

[0134] In the present embodiment, the purpose of constructing a two-dimensional array is to form a standardized data structure, which is convenient for subsequent difference calculation and trajectory correction. The pre-limit data set represents the trajectory control parameters that should be executed in the ideal state of the system, such as speed, position, joint current, etc.; the post-limit data set represents the safe version of these control parameters after being clipped by the current limiting and amplitude limiting strategy; the parameter dimensions include current, speed, angle, etc. to form the columns of the two-dimensional array; the sampling time is the timestamp in the control period, which constitutes the rows of the two-dimensional array.

[0135] Through the display of this content, a standard data structure can be provided for complex trajectory data modeling, making the subsequent difference matrix construction operable, and at the same time, it can be standard interfaced with the later control model training data.

[0136] The difference matrix measures the strength of the trajectory intervention by the limiting strategy, and by identifying the continuous non-zero data segment, it determines where the trajectory weakening or modulation occurs.

[0137] The restricted section identification is to find out the time period with continuous non-zero difference, indicating that this segment of trajectory is affected by amplitude limiting or current limiting.

[0138] By accurately identifying the space-time range where the actual trajectory of the system is interrupted, a basis for subsequent compensation is provided, avoiding indiscriminate processing throughout the cycle. At the same time, it can be used for graphical diagnosis to present the limit impact to the developer more intuitively. For example: the current in column 2 has non-zero differences from row 25 to row 30, indicating that the joint of the robot arm is limited by the system at the 25th to 30th sampling point. By combining the difference and distance (row number difference) to correct the limited parameter, the trajectory is smoothed without exceeding the limit, achieving trajectory continuity repair.

[0139] wherein the distance term is the time distance between the difference point and the maximum difference point, used to control the buffering degree.

[0140] The dynamic correction amount is the trajectory repair strength that fuses spatial decay and direction;

[0141] Specifically, by keeping the parameter limit strategy unchanged, the trajectory discontinuity is repaired, the impact and vibration risk during execution of the robot arm motion is reduced, and the influence of control input mutation caused by limit on execution stability is further prevented.

[0142] For example, if the difference in column 3 of row 20 is -2.5 and the maximum difference position is row 25, the row number difference is 5. After correction: which means that the original limit value should be adjusted by 0.33 units, and the corrected value is added to the original value after the limit to obtain the corrected trajectory parameter, and the entire section is updated to have a buffer transition capability.

[0143] The corrected parameter refers to the corrected control input close to the original planned trajectory;

[0144] Updating the limited section means that all corrected data are collected to form a new motion execution trajectory section, and finally a control command that takes into account safety constraints and trajectory continuity is formed, laying a data foundation for subsequent control-energy relationship evaluation, further improving the representativeness and coverage of input samples during model training.

[0145] In summary, trajectory continuity enhancement effectively solves the problem of trajectory discontinuity or mutation caused by continuous limit strategy, improves the coherence of robot arm motion process, solves the problem of system instability and reduced operation precision caused by frequent data over-limit problems during execution of complex operations. At the same time, it slows down the sharp control change, avoids short-term energy surge or impact, and the updated motion sequence is more representative than the data after the limit, which is conducive to building a more stable and universal control model.

[0146] The difference matrix not only reveals the numerical difference of the parameters before and after limiting, but also forms a constraint comparison with the limiting strategy to determine which parameters are suppressed due to physical boundaries (such as maximum joint torque and current limit). Based on this difference information, the system can respond to different amplitudes of trajectory "cutting" in stages, providing a criterion basis for subsequent correction mechanisms and enhancing the dynamic coordination ability of the control system under the hardware protection strategy.

[0147] The dynamic correction amount takes the "maximum difference point" as the reference source, constructs a dynamic correction coefficient with spatial decay through the "distance from the maximum limiting difference", and combines the direction symbol to build a trajectory compensation amount with physical reasonableness and decreasing trend, realizing point-by-point back-pushing type buffer repair of the original trajectory. This method is superior to the traditional translation type trajectory repair method, and can regress layer by layer without touching the limiting boundary, improving the controllability and flexibility of trajectory repair.

[0148] Due to the limited parameters, the trajectory may locally drop or break, causing the actuator output to jump instantaneously.

[0149] Due to the accumulation of fine tuning in a short time window, the driver current or torque saturates. Taking a mechanical arm on a high-precision 3C product automatic assembly line as an example, the mechanical arm is used to pick up and accurately insert flexible flat cables into small mainboard connection slots. In this process: the end effector positioning error is required to be ≤0.05mm, and the downward force is dynamically adjusted during insertion to avoid damaging the flat cable or connection slot. In order to prevent damage caused by excessive assembly speed or transient inertia accumulation, the system uses the phased limiting and row-by-row difference adjustment scheme in the present application to solve the impact problem caused by continuous over-limiting. However, in multiple batches of operation, the accumulation of multiple fine tuning in a short time window causes the peak value of the driver current to exceed the short-term allowance, triggering servo protection and causing temporary shutdown. Therefore, in order to solve the problem of the peak value of the driver current exceeding the short-term allowance caused by the accumulation of multiple fine tuning in a short time window, the present application sets Embodiment 3.

[0150] Embodiment 3

[0151] Please refer to Figure 1 Specifically: determine whether the current execution action has early warning signs of energy accumulation exceeding the standard, including:

[0152] S201: In the process of each mechanical arm execution action sequence, the control input sequence issued by the upper controller to the driver and the execution data feedback by the driver are recorded synchronously, and the execution data includes the actual driving current and joint torque, to form a complete control-execution data pair;

[0153] That is, the control input sequence is composed of the expected control instructions issued by the upper controller to the driver, including joint target current, expected position and speed information; the execution data (i.e. the execution feedback sequence) is collected by the driver, including the actual joint current, torque and end pose.

[0154] The actual driving current refers to the real-time current value flowing through the motor winding collected by the driver during the execution of the mechanical arm motion, which reflects the output strength of the electromagnetic force required by the driving system to realize the current control instruction (position / speed / torque). The greater the current, the stronger the current output torque of the motor, which may also mean that the system load is greater or the trajectory changes more sharply.

[0155] The joint torque refers to the actual output torque generated at each rotating or moving joint of the mechanical arm, which is fed back by the driver or torque sensor, representing the rotating force exerted by the driving motor or force control module on the joint at the current time.

[0156] The driver is the core execution unit between the control system and the actuator (such as the motor), which is used to convert the control instructions (such as voltage, current reference value) output by the upper controller into power signals that can actually drive the motor to act.

[0157] The actual driving current and the joint torque are one input variable of the motor and one mechanical output of the driving system to the load, and the product of the two can be used to approximately describe the instantaneous power (energy release rate) of the system.

[0158] S202: In the motion execution process, in order to dynamically quantify the energy response characteristics of the driving system in a short time window, the product of the actual driving current and the joint torque in the execution data is integrated to construct a local transient energy index, which is used to reflect the total energy release in the current local stage of the mechanical arm.

[0159] The local transient energy index is obtained by the following formula: Where E is the local transient energy index, t is the starting time of the current window, Δt is the length of the sliding window, I(t) is the actual driving current at time t, and T(t) is the joint torque at time t.

[0160] The local transient energy index reflects the total energy release in the current local stage of the mechanical arm by accumulating the product of the driving current and the corresponding output torque at each time in the time window, which is used for subsequent energy overload detection and dynamic model feedback correction.

[0161] The driving system is a power response and execution module between the control system and the mechanical arm body, mainly including servo drivers, motors and feedback sensors and other sub-components, and its function is to convert control commands into driving current and torque signals and act on the joints of the mechanical arm. In the process of constructing local transient energy indicators, the driving system is the main carrier of energy release and consumption, and its state parameters such as output current, voltage and joint torque constitute the basic variables for energy calculation.

[0162] The local transient energy indicator is compared with the safety threshold to determine whether the current execution action has early warning signs of excessive energy accumulation. If the local transient energy indicator exceeds 50% of the preset safety threshold, it is determined that the current execution action has early warning signs of excessive energy accumulation, the control time index at this time is recorded, and the input adjustment mechanism is activated, otherwise, it is not activated. The control time at this time is recorded as the energy abnormality triggering time;

[0163] The control time index represents the time point when the overload risk is first triggered, which is the start time stamp in the current sliding window.

[0164] The current peak value is the extreme value of the current signal in the current sliding window.

[0165] The local transient energy indicator is used to identify the problem of excessive energy accumulation in the process of continuous correction of micro-trajectory.

[0166] The control input sequence is the action command sequence sent by the upper controller to the driver during the execution of a specific action of the mechanical arm, specifically including target position instructions, speed instructions and expected current or torque reference values. The sequence can be obtained through real-time industrial bus interface collection, control strategy backstepping construction or hardware insertion layer method, as the core data source for subsequent model training, energy evaluation and execution behavior backtracking analysis.

[0167] By comparing the local transient energy indicator with the safety threshold, the purpose is to detect whether there is a short-term energy overload risk. Once it is established, the subsequent inertia correction mechanism is triggered.

[0168] The control time index refers to the corresponding sliding window.

[0169] In this embodiment, while the mechanical arm is executing an action, the control-execution pairing data between the upper and lower systems is recorded in full synchronization to form a data pair that can be used for subsequent energy calculation and model training, further realizing strong synchronization of action execution and data acquisition, ensuring complete and aligned data timing. Provide accurate basis for subsequent construction of energy indicators, diagnosis of abnormal behavior and training of control models.

[0170] The driving current and torque data collected in S201 are used to perform product integral calculation in each sliding window to obtain a local transient energy index in the window, which reflects the energy output degree of the control behavior to the physical system per unit time, is used to represent whether there is abnormal amplification of energy release in the local stage, can monitor and quantify the potential energy overload trend in the micro action in real time, provides an energy response characteristic curve in the continuous action state, facilitates risk warning, and avoids abnormal hardware energy consumption or unstable execution caused by excessive intensity of the control strategy in a short time window.

[0171] The sliding window is used to capture the time period of short-term energy fluctuation and realize dynamic tracking. For example, taking the sliding window length of 200 ms, in the process of "lifting the load → pausing → translating", if the energy index in the "lifting" paragraph suddenly rises to 80% of the threshold value, it indicates that the load is large or there is inertia impact, and the adjustment mechanism needs to be triggered in time to avoid continuous overload of the driver;

[0172] The obtained local transient energy index is compared with the safety threshold preset by the system. If it is determined that there are early signs of energy accumulation exceeding the standard, the time when the energy anomaly is triggered is recorded, and the current current peak in the sliding window is also recorded, and the input adjustment process is entered. Otherwise, the adjustment mechanism is not triggered, and the monitoring continues, which is helpful for identifying the energy overload signs caused by the cumulative control input in advance, realizing early warning and early adjustment, avoiding excessive system thermal load, aging or instability of the execution component caused by energy accumulation, and providing a trigger basis and abnormal positioning point for subsequent inertia compensation.

[0173] The safety threshold refers to the maximum local energy consumption upper limit allowed by the system in each control stage;

[0174] The current peak is used to locate the extreme point of the energy impact in a short period of time, which helps to quantify the abnormal degree;

[0175] For example, when a welding robot performs high-speed switching station action, the local energy index suddenly rises to 60% of the set safety value at a certain moment, and the current peak is as high as 90% of the design value. The system records this time as the trigger point, and performs inertia attenuation compensation in the subsequent correction step.

[0176] In summary, the logic chain in this embodiment is: collecting real control data, constructing a local energy index, judging whether to warn, recording the trigger point and preparing for correction. This series of logic ensures the closed-loop coupling and dynamic adjustment capability between the control behavior and the energy response, not only improves the stability of the system operation, but also provides high-quality energy response samples for subsequent model training.

[0177] On the basis of the trajectory has been staged limit correction, if still appear short time local energy rapid rise phenomenon, often suggest that the current correction strategy can not effectively alleviate the risk of execution overload, the present application by continuous monitoring of local energy index, can identify the implicit problem of surface limit but overload in advance, and its trigger time point and current peak as abnormal mark, further activate the subsequent input adjustment mechanism, weaken the energy accumulation trend from the source, avoid lagging behind remedy, enhance the feedforward adaptive ability of control chain.

[0178] Embodiment 4

[0179] Please refer to Figure 1 , specific: in the activation of input adjustment mechanism, the risk level of local transient energy index in the current action cycle is identified, including:

[0180] Based on the ratio relationship between the current local transient energy index and the safety threshold, three energy risk levels are divided, which are low risk interval, medium risk interval and high risk interval, for example, if the local transient energy index does not exceed 50% of the preset safety threshold, the current local transient energy index is in the low risk interval; if the local transient energy index exceeds 50% of the preset safety threshold, but does not exceed 80% of the safety threshold, the current local transient energy index is in the medium risk interval; if the local transient energy index exceeds 80% of the preset safety threshold, the current local transient energy index is in the high risk interval;

[0181] According to the corresponding energy risk level, the corresponding inertia dispersion factor value interval is configured, for example: if in the low risk interval, the corresponding inertia dispersion factor value interval is automatically configured as 0.1 to 0.2, indicating weak intervention; if in the medium risk interval, the corresponding inertia dispersion factor value interval is automatically configured as 0.21 to 0.6, indicating buffer; if in the high risk interval, the corresponding inertia dispersion factor value interval is automatically configured as 0.61 to 0.95, indicating strong dispersion;

[0182] By dynamically adjusting the adjustment amplitude, over-adjustment or under-adjustment is prevented;

[0183] The inertia dispersion factor corresponding to the current local transient energy index is obtained from the inertia dispersion factor value interval, specifically: Wherein, ε k is the inertia dispersion factor at the current energy abnormal trigger time k, ε max and ε min are the maximum value and the minimum value in the corresponding inertia dispersion factor value interval, E k is the local transient energy index at the current energy abnormal trigger time k, E th is the safety threshold;

[0184] The inertia dispersion factor refers to a proportional parameter used for smoothing the control rate of change and reducing the control shock speed when compensating the control input, and reflects the tolerance and diffusion degree of the system when the local energy upper limit is approached.

[0185] The proportion The inertia dispersion factor is linearly interpolated according to the current local transient energy index in the interval [0, 1], so that the system can realize the flexible dynamic adjustment capability of the control input sequence under different energy levels.

[0186] According to the interval of the local transient energy index, the inertia dispersion factor is selected;

[0187] Based on the located energy anomaly triggering moment, the control time index of the previous period is traversed in reverse to construct a reverse recursive correction function of the control input sequence to perform inertia attenuation compensation on the sharp change of the continuous control input, so as to prevent the energy from continuously accumulating in a short period to cause local saturation, including:

[0188] Suppose the control time index corresponding to the energy anomaly triggering moment extracted from the original control input sequence is the local sequence {u k,k+Δt}, wherein u k,k+Δt is the system control amount (which can be joint current, drive voltage or torque command) at the corresponding time t in the corresponding sliding window, k is the energy anomaly triggering moment, and k+Δt is the end time in the corresponding sliding window;

[0189] The energy anomaly triggering moment is recursively propagated to the nearest stable energy boundary point to construct a compensation interval [k s , k], wherein k s is the nearest stable energy boundary point, and the stable energy boundary point refers to the latest time point that meets the low-risk interval (i.e., the local transient energy index E is less than 0.5*security threshold E th );

[0190] In the compensation interval [k s , k], the increments between the original local sequence at the corresponding time and the original local sequence at the previous time are fused with the inertia dispersion factor as the weight in a recursive manner to generate the corrected local sequence (i.e., the control input after inertia compensation);

[0191] The corrected local sequence is spliced with the original control input sequence segment that is not disturbed to obtain the corrected control input sequence;

[0192] According to the corrected control input sequence, S201 and S202 are re-executed to compare the relative reduction degree of the local transient energy index before and after correction, and the energy consumption reduction feedback reinforcement sample set is obtained.

[0193] The modified control input sequence is used to form an inertia adjustment control input trajectory covering the whole motion cycle.

[0194] The inertia adjustment control input trajectory can be directly substituted for the original control input sequence to drive the system, and used in subsequent steps to reevaluate system energy stability and control precision response.

[0195] The system buffers and adjusts the increment between the current control input to be issued and the control input at the previous time, specifically, sets the inertia dispersion factor as ε, whose value range is 0<ε<1, and then updates the control input u q ‘ The following methods are used to obtain the inertia adjustment control input trajectory:

[0196] In the compensation interval [k s , k], the original local sequence is adjusted by the following recursive formula in reverse inertia attenuation: u q ‘ = u q - ε q *(u q - u q-1 ), where u q ‘ is the modified local sequence (i.e., the inertia-compensated control input) at time q, u q and u q-1 are the original local sequences at time q and time q-1, ε q is the inertia dispersion factor at time q, which dynamically changes with k, q is the time number in the compensation interval [k s , k], and (u q - u q-1 ) is the control change rate of the step at time q.

[0197] As can be seen, when ε q = 0, the original control input is kept unchanged, when ε q = 1, the control input at the previous time is completely replaced, achieving the maximum degree of inertia retention, and intermediate values correspond to different degrees of inertia weakening. This operation is essentially an input smoothing method for short-term energy mutation suppression, which can effectively reduce the problem of sudden increase of driving current and actuator saturation caused by too large short-time control input change.

[0198] u q ‘= u q - ε q *(u q - u q-1) is a reverse recursive correction function of the control input sequence, which has the effect of a low-pass filter, that is, it compresses the original control input changes with a lag, thereby slowing down the sharp jump of the control amplitude in the short time window without interfering with the overall control target of the system.

[0199] The original control input sequence refers to the control input sequence before inertia attenuation compensation is performed.

[0200] In this embodiment, when the system detects that the local transient energy index exceeds the set proportion of the safety threshold, the input adjustment mechanism is triggered. To ensure that the adjustment process has sensitive response capability and flexible adjustment capability, the system needs to evaluate the ratio between E and the safety threshold, and divide it into three energy risk levels according to the ratio. Through the division of the energy risk level, the severity of the current energy consumption accumulation can be effectively judged, so as to accurately match input adjustment strategies of different intensities.

[0201] In the high-risk state, a larger amplitude of inertia compensation can be activated in advance to avoid the system entering the saturation or hard stop interval. For example, when performing high-speed movement, if the transient energy index caused by the peak value of joint 2 driving current has reached 95%, it is classified as a high-risk area and needs to enter the inertia compensation step immediately.

[0202] According to the aforementioned risk level, the system constructs a proportion factor from the ratio of the local transient energy index to the safety threshold in the corresponding inertia dispersion factor interval, and calculates the inertia dispersion factor according to the proportion factor. The adjustment attenuation proportion in the control input adjustment function (i.e., the corresponding formula for obtaining the inertia dispersion factor) is used to control the input change amplitude. The introduction of the inertia dispersion factor makes the control input not too drastic when responding abnormally, which ensures system stability and does not lose response accuracy.

[0203] From the time when the energy anomaly is triggered, search forward until the nearest stable energy boundary point is found, mark this segment as the compensation interval, and perform inertia compensation on the original control input sequence in the compensation interval according to the recursive formula. The compensation formula essentially realizes the lag suppression of the input command change amplitude, that is, it constructs an input adjustment function similar to a low-pass filter.

[0204] The reverse recursive correction function is a function mechanism that adjusts the control command change trend in reverse order of time.

[0205] The hysteresis compression processing is achieved by controlling the input "change rate adjustment" to realize the smooth transition of the input curve, which avoids the sudden rise of energy caused by the sudden change of the control input, and is particularly suitable for the critical input jitter frequently occurring in the coordinated motion of multiple joints. For example, in the trajectory of joint 1 suddenly decreasing from 50 rad / s to 5 rad / s, the change can be smoothed into a decreasing sequence through inertia adjustment, so as to prevent the sudden change of the driving system energy.

[0206] In combination with the original and modified control instructions, the control data of the entire action cycle is updated, and through the modified control sequence, not only the core target of completing the action task is maintained, but also better dynamic performance indicators and lower transient energy consumption are obtained.

[0207] The reverse recursive correction function constructed in the application implements dynamic attenuation fusion processing on the change rate of the control input at adjacent time points by introducing an inertia dispersion factor, so that the sharp transition existing in the control sequence is suppressed. This operation is equivalent to a kind of time sequence smoothing operation with "low-pass filtering" characteristics on the original control input, and is particularly suitable for eliminating the local disturbance of the control signal introduced by the trajectory fine-tuning, limit clipping and other mechanisms, thereby enhancing the continuity of the driving response and the robustness of the system.

[0208] It should be noted that the application does not perform a one-size-fits-all correction processing on the entire control sequence, but only implements segmented compensation for the compensation interval before the energy abnormal triggering point, and takes the "most stable energy boundary point" as the compensation starting point, effectively avoiding invalid disturbance to the stable control segment. This compensation strategy of "limited range, dynamic attenuation and continuous splicing" further reduces the disturbance degree of system load adjustment and improves the local convergence ability and controllability of system response. Compared with the traditional rigid processing mode based on amplitude limiting or instruction truncation, the linear mapping design of the inertia dispersion factor in the application enables the system to dynamically adjust the control input buffer amplitude according to the local energy state, and establishes a flexible adjustment mechanism. This adjustment strategy not only retains the structural continuity of the original trajectory control trend, but also takes into account the execution stability and energy use safety, and is particularly suitable for the adaptive suppression demand of joint actuators to sudden load changes in high dynamic mechanical arm systems.

[0209] The control input sequence after inertia correction can effectively suppress local abnormal excitation without changing the macro structure of the action cycle, and generate training samples with better stability and representativeness.

[0210] Embodiment 5

[0211] Please refer to Figure 1 , specifically: according to the modified control input sequence, S201 and S202 are re-executed to compare the relative reduction degree of the local transient energy indicators before and after the modification, and an energy consumption reduction feedback reinforcement sample set is obtained, including:

[0212] Re-execute S201 and S202 to obtain a corrected local transient energy index;

[0213] Compare the relative reduction degree of the local transient energy index before and after correction to obtain the relative reduction ratio, specifically: Among them, ΔE is the relative reduction ratio, E ’ is the corrected local transient energy index;

[0214] If the relative reduction ratio exceeds a preset ratio threshold (i.e., the set minimum energy reduction ratio threshold (e.g., 10%)), the inertia compensation is determined to be effective, and the corresponding corrected control input sequence and the corrected local transient energy index are used as a set of execution mapping samples. Otherwise, it is determined to be invalid and the inertia dispersion factor is readjusted until it is effective.

[0215] The execution mapping samples corresponding to multiple action cycles are counted to obtain an energy consumption reduction feedback reinforcement sample set.

[0216] The specific method of re-adjusting the inertia dispersion factor is as follows: at the time point within the corresponding sliding window, the adjusted inertia dispersion factor is calculated. Among them, ε' is the inertia dispersion factor after adjustment, ε0 is the inertia dispersion factor before adjustment, β is the risk response sensitivity coefficient, E(t) is the local transient energy index at time t, is the rate of change of the local transient energy index;

[0217] The risk response sensitivity coefficient is used to control the response intensity of the local transient energy index change rate to the inertia dispersion factor correction. and is large), indicating that the system may be rapidly accumulating energy and there is a risk of overload or local saturation. In order to avoid instability caused by sudden changes in control input, it is necessary to enhance the system's "buffering" ability to inertia. At this time, the β amplification The contribution to ε' increases ε', that is, it increases the "inertia dispersion factor" and thus achieves a stronger "inertia attenuation effect"; in simple terms, β determines the system's sensitivity to the energy growth rate. It is a key dynamic response amplification factor in the inertia control system. Its value is observed using a large amount of historical execution data: under different action cycles, The relationship with system instability / oscillation, as well as the response of the system to energy consumption fluctuations after ε adjustment, is considered. The β range is sought so that ε' neither rises too quickly (resulting in delayed response and difficulty in system adaptation) nor changes too little (resulting in insufficient inertia adjustment and ineffective energy reduction). The typical range for β is between 0.05 and 0.5.

[0218] The execution mapping sample refers to a sample data pair that meets the following conditions: in a certain action period, the modified control input compared with the original control input causes a significant decrease (exceeding a proportional threshold) in the energy index within a unit time window, and its control behavior has no instability or execution failure phenomenon, that is, an effective energy reduction sample, which is used for subsequent control model reinforcement training to improve the model's ability to capture energy response changes and enhance its ability to generate low-energy and high-stability control instructions under complex working conditions.

[0219] The energy reduction feedback reinforcement sample set is used as the input of the control model, and the target energy index is used as the training output.

[0220] The error between the predicted energy index of the control model and the target energy index is used as the loss function.

[0221] The error function is used to perform model parameter backpropagation training on the control model, wherein the gradient descent method is used in the training process to iteratively update the model parameters until the control model can generate control input sequences with stable energy-saving characteristics under different action periods.

[0222] The control model is constructed based on a convolutional neural network, wherein the control model can be divided into five layers of structure modules.

[0223] The input layer is used to input a plurality of control input sequences, wherein each control input sequence includes control parameters such as control current, voltage, and desired torque.

[0224] The time convolution extraction layer is used to extract local time-dependent relationships in the control input sequence.

[0225] The local feature aggregation layer is used to fuse local features at multiple time scales to capture the impact of step changes on energy response.

[0226] The fully connected prediction layer is used to map the aggregated multi-dimensional features to the energy response prediction corresponding to the control behavior.

[0227] The energy residual error estimation output layer is used to use the energy residual error as the target function.

[0228] In this embodiment, based on the control input sequence after completing inertia compensation, the previously described local transient energy index calculation process is re-executed to obtain the compensated energy response. Through re-evaluation, it can be directly reflected whether the inertia adjustment operation actually reduces the energy consumption level within a unit time window. Through this process, the quantitative verification of the inertia compensation effect is realized, avoiding the errors brought by using the control curve change as the only judgment standard, which is helpful for constructing the feedback reinforcement sample set.

[0229] By calculating the relative reduction ratio of the original energy index and the modified energy index, the ratio is used as a criterion for whether to constitute an effective energy reduction behavior, a quantifiable threshold is provided as a reinforcement sample screening standard to ensure that only samples that truly achieve energy consumption reduction are included in model training. For example, if the original index is 80J and the modified index is 68J, then AE = 15%, which is greater than the preset 10%, and it is determined that the compensation strategy is effective, and a set of samples composed of the control input sequence before and after modification and the energy index is recorded; otherwise, the inertia dispersion factor needs to be adjusted, and the control input is modified again until it is effective, so as to screen out operation samples with significant energy reduction and no control instability, and to construct a high-quality training set to enhance the generalization ability of the model.

[0230] The inertia dispersion factor is recalculated to realize dynamic optimization of the inertia factor, so that it responds to different energy consumption change rates and avoids excessive compensation leading to trajectory disturbance. For example: if the current energy index rises at a fast rate, the original value should be quickly reduced to increase the inertia compensation strength.

[0231] Taking the constructed execution mapping sample as input and the target energy index as output, the loss function is defined as the difference between the predicted value and the target value, and the gradient descent method with adaptive step size is used for back propagation optimization of model parameters, which is beneficial to improve the energy response prediction accuracy of the control model through reinforcement learning mechanism, so as to automatically output control instruction sequence with low energy consumption and high stability. For example: the control model structure can adopt a three-layer convolutional neural network, in which the input layer is the control input vector and the output layer is the energy prediction value; the target is to realize that the prediction energy error is less than 5% under various working conditions.

[0232] In summary, through the verification of inertia compensation effect, the screening of sample set, the dynamic self-adjustment of inertia factor, and the training of control model, a complete energy consumption optimization path from control input modification to model capability improvement is realized.

[0233] Embodiment 6

[0234] Please refer to Figure 2 , specifically: a mechanical arm control model training system, comprising,

[0235] The continuous correction subsystem constructs a difference matrix and detects a restricted section, based on the restricted section, performs reverse correction on the corresponding limited parameters to realize trajectory buffer reconstruction, and the difference matrix is used to quantify the difference between the mechanical arm before and after limiting;

[0236] The accumulation subsystem executes a motion sequence on the mechanical arm after reverse correction, judges whether there are early warning signs of energy accumulation exceeding the standard in the current execution motion by synchronously recording the control input sequence issued by the upper controller to the driver and the execution data feedback by the driver; the motion sequence refers to the updated restricted section;

[0237] a risk identification subsystem, if present, identifies the risk level in the current action cycle and locates the energy anomaly triggering moment;

[0238] a correction subsystem, based on the energy anomaly triggering moment, traverses the previous control time index in reverse, constructs a reverse recursive correction function of the control input sequence, and performs inertia decay compensation on the sharp change of the continuous control input to correct the control input sequence;

[0239] a sample collection subsystem, according to the corrected control input sequence, compares the relative reduction degree of the local transient energy index before and after correction, and obtains an energy consumption reduction feedback reinforcement sample set;

[0240] a training subsystem, takes the energy consumption reduction feedback reinforcement sample set as the input of the control model, and performs model parameter back propagation training on the control model using an error function to iteratively update the model parameters.

[0241] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, replacements and changes can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. A method for training a robotic arm control model, characterized by: include, Construct a difference matrix and detect restricted sections. Based on the restricted sections, perform reverse correction on the corresponding parameters after restriction to achieve trajectory buffer reconstruction. The difference matrix is ​​used to measure the difference between the robot arm before and after restriction. After the reverse correction, the robot arm executes the action sequence. By synchronously recording the control input sequence sent by the upper controller to the driver and the execution data fed back by the driver, it is determined whether there are early warning signs of excessive energy accumulation in the current execution action; The action sequence refers to the updated restricted segment; If it exists, the risk level in the current action cycle is identified and the energy anomaly triggering moment is located; Based on the energy anomaly triggering moment, the previous control time index is traversed in reverse to construct a reverse recursive correction function of the control input sequence to perform inertial attenuation compensation for the sharp changes in continuous control input and correct the control input sequence; According to the corrected control input sequence, the relative reduction degree of the local transient energy index before and after the correction is compared to obtain the energy consumption reduction feedback reinforcement sample set; The energy consumption reduction feedback reinforcement sample set is used as the input of the control model, and the error function is used to perform back propagation training on the control model to iteratively update the model parameters.

2. A robot arm control model training method according to claim 1, characterized in that: Construct a difference matrix and detect restricted segments, including: Obtain the behavioral data of the robot arm in the corresponding action cycle in advance, including the pre-limit data set and the post-limit data set; Based on the pre-limited data set, a two-dimensional array is constructed, where the columns of the two-dimensional array are parameter dimensions and the behavior is the sampling time; According to the limiting strategy of each parameter, the difference between the position of each element in the two-dimensional array before and after the limiting is determined to construct a difference matrix; Traverse each column in the difference matrix and identify the continuous non-zero data segments, which are recorded as restricted segments; Based on the restricted section, the corresponding parameters after limitation are reversely corrected to quantify the correction status of progressive buffering of the corresponding parameters while maintaining the limitation strategy of the corresponding parameters, thereby realizing trajectory buffer reconstruction.

3. A method for training a robot arm control model according to claim 2, characterized in that: Based on the restricted section, reverse correction is performed on the corresponding limited parameters to quantify the correction status of the corresponding parameters with progressive buffering while maintaining the limiting strategy of the corresponding parameters, including: For each restricted segment, extract the maximum value within it and mark the row number of the maximum value in the difference matrix and the positive or negative sign of the difference between the maximum value and other parameters in the corresponding restricted segment; Record the row number corresponding to each parameter in the restricted section to obtain a row number set; Take the minimum row number recorded in the row number set as the starting row; S101: For the starting row in each restricted segment, calculate the row number difference based on the distance between the row number corresponding to the starting row and the row number of the maximum value in the column; S102: performing reverse correction on the corresponding parameter after limiting by linearly superimposing the difference before and after limiting with the normalized row number difference and taking the inverse form thereof, and combining the positive and negative signs to obtain a dynamic correction value; S103: Based on the dynamic correction amount, correct the corresponding post-limiting parameters in the post-limiting data set to obtain corrected parameters; Get the next row in the corresponding restricted segment, and repeat S101 to S103 until the last row in the row number set, to update the corresponding restricted segment; The updated restricted section is used as the action sequence that has completed the staged limit correction in the current action cycle.

4. A method for training a robot arm control model according to claim 3, characterized in that: Determine whether there are early warning signs of excessive energy accumulation in the current action, including: S201: During each execution of the motion sequence by the robotic arm, the control input sequence sent from the upper controller to the driver and the execution data fed back by the driver are synchronously recorded. The execution data includes the actual drive current and joint torque. S202: During the execution of the action, in order to dynamically quantify the energy response characteristics of the drive system within a short time window, a local transient energy index is constructed by integrating the product of the actual drive current and the joint torque in the execution data. The local transient energy index is used to reflect the total amount of energy released in the current local stage of the robot arm; If the local transient energy index exceeds 50% of the preset safety threshold, it is judged that the current execution action has early warning signs of excessive energy accumulation, the control time index at this time is recorded, and the input adjustment mechanism is activated, where the control time at this time is recorded as the energy anomaly triggering moment.

5. A robot arm control model training method according to claim 4, characterized in that: After activating the input regulation mechanism, the risk level of the local transient energy indicator in the current action cycle is identified, including: Based on the ratio between the current local transient energy index and the safety threshold, three energy risk levels are divided into low risk range, medium risk range and high risk range; The corresponding inertia dispersion factor value interval is configured according to the corresponding energy risk level, and the inertia dispersion factor corresponding to the current local transient energy index is obtained from the inertia dispersion factor value interval.

6. A method for training a robot arm control model according to claim 5, characterized in that: Based on the located energy anomaly triggering moment, the previous control time index is traversed in reverse to construct a reverse recursive correction function of the control input sequence to perform inertial attenuation compensation for the sharp changes in continuous control input, including: Assume that the control time index corresponding to the energy anomaly triggering moment extracted from the original control input sequence is the local sequence u k,k+Δt }, where u k,k+Δt is the system control quantity at the corresponding time t in the corresponding sliding window, k is the energy anomaly triggering time, and k+Δt is the end time in the corresponding sliding window; The compensation interval [k s ,k], where k s The nearest stable energy boundary point, which refers to the latest moment that satisfies the low-risk interval; In the compensation interval [k s , k], the increments between the original local sequence at the corresponding moment and the original local sequence at the previous moment are fused recursively with the inertia dispersion factor as the weight to generate the modified local sequence; The corrected local sequence is spliced ​​with the undisturbed original control input sequence fragment to obtain a corrected control input sequence; According to the corrected control input sequence, S201 and S202 are re-executed to compare the relative reduction degree of the local transient energy index before and after the correction, and obtain the energy consumption reduction feedback reinforcement sample set.

7. A method for training a robot arm control model according to claim 6, characterized in that: Based on the corrected control input sequence, S201 and S202 are re-executed to compare the relative reduction degree of the local transient energy index before and after the correction, and obtain an energy consumption reduction feedback reinforcement sample set, including: Re-execute S201 and S202 to obtain a corrected local transient energy index; Compare the relative reduction degree of the local transient energy index before and after correction to obtain the relative reduction ratio; If the relative reduction ratio exceeds the preset ratio threshold, the inertia compensation is determined to be effective, and the corresponding corrected control input sequence and the corrected local transient energy index are used as a set of execution mapping samples. Otherwise, it is determined to be invalid and the inertia dispersion factor is readjusted until it is effective. The execution mapping samples corresponding to multiple action cycles are counted to obtain an energy consumption reduction feedback reinforcement sample set.

8. A method for training a robot arm control model according to claim 7, characterized in that: The energy consumption reduction feedback reinforcement sample set is used as the input of the control model, and the target energy index is used as the training output; The error between the predicted energy index of the control model and the target energy index is used as the loss function; The error function is used to perform back-propagation training on the model parameters of the control model. The training process adopts the gradient descent method to iteratively update the model parameters until the control model can generate control input sequences with stable energy-saving characteristics under different action cycles.

9. A robot arm control model training system, used to implement the robot arm control model training method according to any one of claims 1 to 8, characterized in that: include, The Lianchao correction subsystem constructs a difference matrix and detects the restricted section. Based on the restricted section, it performs reverse correction on the corresponding parameters after the restriction to achieve trajectory buffer reconstruction. The difference matrix is ​​used to measure the difference between the robot arm before and after the restriction. The accumulation subsystem, after reverse correction, executes the action sequence of the robot arm and synchronously records the control input sequence sent by the upper controller to the driver and the execution data fed back by the driver to determine whether there are early warning signs of excessive energy accumulation in the current execution action; The action sequence refers to the updated restricted segment; The risk identification subsystem, if present, identifies the risk level within the current action cycle and locates the moment when the energy anomaly is triggered; The correction subsystem, based on the energy anomaly triggering moment, reversely traverses the previous control time index and constructs a reverse recursive correction function of the control input sequence to perform inertial attenuation compensation for the sharp changes in continuous control input and correct the control input sequence; The sample collection subsystem compares the relative reduction degree of the local transient energy index before and after the correction based on the corrected control input sequence, and obtains the energy consumption reduction feedback reinforcement sample set; The training subsystem takes the energy consumption reduction feedback reinforcement sample set as the input of the control model, and uses the error function to perform backpropagation training on the model parameters of the control model to iteratively update the model parameters.

Citation Information

Cited By

  • Linear servo joint reverse driving control method and system

    CN121340307A