Robot path generation system based on teaching and path optimization

CN122584357BActive Publication Date: 2026-09-22XIAN CHENGHE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611063129.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-17
Publication Date
2026-09-22
Estimated Expiration
2046-07-17

AI Technical Summary

Technical Problem

然而,在对复杂动态环境中的机器人进行路径生成时,需要同时协调操作员示教意图与自主避障控制结果;示教意图与自主避障结果之间的一致程度,是衡量路径生成是否稳定、是否安全的重要因素,它直接关系到轨迹平滑性、控制连贯性以及系统振荡风险;为了兼顾操作员控制需求和环境安全约束,通常采用加权方式来定义目标轨迹,例如,上述计算方法虽结合了示教轨迹与避障轨迹,但对权重和的设定高度依赖主观经验,且对动态环境缺乏实时自适应调节能力,存在明显的系统控制滞后性;

Benefits of technology

1)本发明通过计算轨迹偏差值和补偿操作指令的概率分布得出人机控制分歧熵,并结合当前动力学状态数据预测诱发振荡概率;系统依据上述参数动态分配示教顺从权重与自主避障权重,从而生成目标平滑轨迹;该机制克服了固定加权方式的滞后性缺陷,避免了单纯依赖局部避障引发的人机控制对抗,提升了系统对动态扰动的自适应能力;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122584357B_ABST
    Figure CN122584357B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of robot control and intelligent planning, in particular to a robot path generation system based on teaching and path optimization, which comprises a teaching acquisition module, an environment perception module, a state observation module and a control hub; the teaching acquisition module acquires original teaching trajectory data of a remote teaching device, the environment perception module acquires work environment disturbance data and current dynamics state data; the state observation module generates a local obstacle avoidance trajectory, calculates a trajectory deviation value, extracts a high-frequency compensation operation instruction, calculates a man-machine control divergence entropy based on a probability distribution of the high-frequency compensation operation instruction and the trajectory deviation value, and outputs an induced oscillation probability through an oscillation prediction model combined with the current dynamics state data; the control hub dynamically allocates teaching compliance weights and autonomous obstacle avoidance weights according to the above, generates a target smooth trajectory, and converts the target smooth trajectory into trajectory tracking, impedance adjustment and speed attenuation instructions to be sent to the robot, so that remote teaching, safe obstacle avoidance and system stability are coordinated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot control and intelligent planning technology, specifically a robot path generation system based on teaching and path optimization. Background Technology

[0002] Robot path generation technology, as a key component of remote heavy-duty collaborative operations, is widely used in high-risk environments for handling, obstacle avoidance, and precision operations. To ensure the safety and trajectory feasibility of robots in dynamic and disturbed environments, remote teaching input and local path optimization results are usually combined to generate and adjust the robot's target motion trajectory in real time. However, when generating paths for robots in complex dynamic environments, it is necessary to simultaneously coordinate the operator's teaching intentions and the results of autonomous obstacle avoidance control. The consistency between the teaching intentions and the autonomous obstacle avoidance results is a crucial factor in evaluating the stability and safety of path generation, directly affecting trajectory smoothness, control coherence, and the risk of system oscillations. To balance operator control requirements and environmental safety constraints, a weighted approach is typically used to define the target trajectory, for example... Although the above calculation method combines the teaching trajectory and the obstacle avoidance trajectory, it does not adequately address the weighting. and The settings are highly dependent on subjective experience and lack the ability to adapt and adjust in real time to dynamic environments, resulting in obvious system control lag. This means that the robot path generation results are easily affected by factors such as human-machine intention conflicts, communication delays, environmental disturbances, and dynamic abrupt changes, leading to frequent corrections, local jitters, and even control oscillations in the generated trajectory. This uncertainty not only affects the robot's accurate execution of the operator's intentions, but may also have an adverse impact on operational safety and system stability in high-risk scenarios. Summary of the Invention

[0003] The purpose of this invention is to provide a robot path generation system based on teaching and path optimization, and to solve the following technical problems: To avoid human-robot control conflict caused by obstacle avoidance based solely on local shortest paths and forced distance minimization, the robot should retain the interpretability of human intentions while possessing the ability to adapt to dynamic environments, thereby achieving a unified coordination of remote teaching, safe obstacle avoidance, and system stability.

[0004] The objective of this invention can be achieved through the following technical solutions: A robot path generation system based on teaching and path optimization, comprising: The teaching acquisition module communicates with the remote teaching device and is used to collect the original teaching trajectory data from the remote teaching device. The environmental perception module communicates with the robot, which has an end effector; the environmental perception module is used to collect data on environmental disturbances and current dynamic state of the robot. The state observation module is communicatively connected to both the teaching acquisition module and the environment perception module. It generates a local obstacle avoidance trajectory based on the operational environment disturbance data, compares the local obstacle avoidance trajectory with the original teaching trajectory data to obtain a trajectory deviation value, and extracts reverse position adjustment data with frequencies exceeding a preset threshold from the original teaching trajectory data as compensation operation commands. It also determines the human-machine control divergence entropy based on the probability distribution of the trajectory deviation value and the compensation operation commands. The human-machine control divergence entropy and the current dynamic state data are input into an oscillation prediction model, which outputs the induced oscillation probability. The control center is connected to the state observation module. The control center is used to dynamically allocate teaching compliance weights and autonomous obstacle avoidance weights based on human-machine control divergence entropy and induced oscillation probability to generate a target smooth trajectory, and convert the target smooth trajectory into control commands and send them to the robot. The control commands include at least trajectory tracking commands, impedance adjustment commands, and velocity decay commands.

[0005] Preferably, the control center is preset with a safety divergence threshold; used to increase the teaching compliance weight and decrease the autonomous obstacle avoidance weight when the human-machine control divergence entropy is lower than the safety divergence threshold; and to decrease the teaching compliance weight and increase the autonomous obstacle avoidance weight when the human-machine control divergence entropy is higher than or equal to the safety divergence threshold.

[0006] Preferably, the control center is specifically used to: when the human-machine control divergence entropy is higher than or equal to the safety divergence threshold, to perform weighted fusion and filtering on the original teaching trajectory data and the local obstacle avoidance trajectory according to the adjusted weights, to generate a first smooth trajectory as the target smooth trajectory, and to generate the speed decay command accordingly.

[0007] Preferably, the control center is specifically used to: when the human-machine control divergence entropy is lower than the safety divergence threshold, to perform weighted fusion of the original teaching trajectory data and the local obstacle avoidance trajectory according to the adjusted weights, generate a second smooth trajectory as the target smooth trajectory, and generate the trajectory tracking command accordingly.

[0008] Preferably, the control center has an adaptive impedance control algorithm, which is used to dynamically adjust the target stiffness value and target damping value of the robot end effector based on the induced oscillation probability and the current dynamic state data, and generate the impedance adjustment command.

[0009] Preferably, the current dynamic state data includes joint torque data and end-effector acceleration data; the control center is used to generate and send a stop lock command to the robot when the change in the end-effector acceleration data or the joint torque data exceeds the corresponding preset threshold; otherwise, the sending of the control command is maintained.

[0010] Preferably, the operational environment disturbance data includes thermodynamic airflow change data and dynamic obstacle displacement data; the teaching acquisition module is also used to acquire communication status data between the remote teaching device and the system, the communication status data including network latency parameters and signal jitter parameters; the status observation module is also used to correct the human-machine control divergence entropy based on the communication status data.

[0011] The beneficial effects of this invention are: 1) This invention calculates the human-machine control divergence entropy by calculating the trajectory deviation value and the probability distribution of the compensation operation command, and predicts the induced oscillation probability by combining the current dynamic state data; the system dynamically allocates the teaching compliance weight and autonomous obstacle avoidance weight according to the above parameters, thereby generating the target smooth trajectory; this mechanism overcomes the lag defect of the fixed weighting method, avoids the human-machine control confrontation caused by simply relying on local obstacle avoidance, and improves the system's adaptive ability to dynamic disturbances; 2) This invention presets a safety divergence threshold. When the divergence entropy is higher than or equal to the threshold, the autonomous obstacle avoidance weight is increased, high-frequency fine-tuning trajectory segments are filtered out using low-pass filtering, and speed attenuation commands are issued simultaneously. When the divergence entropy is lower than the threshold, the teaching compliance weight is increased to perform high-fidelity trajectory tracking. This mechanism suppresses the risk of control oscillation in high-risk areas while fully ensuring the accurate execution of operational intentions in low-risk areas. 3) Based on the induced oscillation probability and current dynamic state data, the control center of this invention dynamically adjusts the initial stiffness and initial damping values, calculates and outputs the target stiffness and target damping of the end effector; this design enables the robot to actively reduce stiffness and increase damping to absorb impact energy when facing high-risk confrontation, and restore stiffness to support heavy loads under stable working conditions, effectively suppressing mechanical vibration and overshoot after the end effector is disturbed from the mechanical interaction level; 4) This invention monitors the end-effector acceleration change and joint torque data in real time between adjacent sampling periods; when the acceleration change is detected to be higher than the acceleration mutation threshold or the joint torque is higher than the limit threshold, the control center immediately triggers and issues a stop lock command; this mechanism sets strict physical constraint boundaries for heavy-duty operations, preventing the robot from forcibly advancing in transient instability events that exceed the system's flexible adjustment capabilities, thus ensuring safety under extreme working conditions. 5) This invention collects network delay parameters and signal jitter parameters to construct a penalty weight, and uses the penalty weight to correct the human-machine control divergence entropy. This mechanism effectively separates the lag operation input caused by the degraded communication quality link from the actual physical control confrontation, preventing the control center from misjudging the communication degradation as a high-risk human-machine divergence, and significantly improving the reliability of system risk assessment and the accuracy of state observation under complex disturbance environments. Attached Figure Description

[0012] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a schematic diagram of the modules of the robot path generation system based on teaching and path optimization provided in the embodiments of this application. Detailed Implementation

[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0014] Please see Figure 1 A robot path generation system based on teaching and path optimization includes: a teaching acquisition module, which is connected to a remote teaching device; the teaching acquisition module is used to acquire the original teaching trajectory data of the remote teaching device; The environmental perception module communicates with the robot, which has an end effector; the environmental perception module is used to collect data on environmental disturbances and current dynamic state of the robot. The state observation module is communicatively connected to both the teaching acquisition module and the environment perception module. It generates a local obstacle avoidance trajectory based on the operational environment disturbance data, compares the local obstacle avoidance trajectory with the original teaching trajectory data to obtain a trajectory deviation value, and extracts reverse position adjustment data with frequencies exceeding a preset threshold from the original teaching trajectory data as compensation operation commands. It also determines the human-machine control divergence entropy based on the probability distribution of the trajectory deviation value and the compensation operation commands. The human-machine control divergence entropy and the current dynamic state data are input into an oscillation prediction model, which outputs the induced oscillation probability. The control center is connected to the state observation module. The control center is used to dynamically allocate teaching compliance weights and autonomous obstacle avoidance weights based on human-machine control divergence entropy and induced oscillation probability to generate a target smooth trajectory, and convert the target smooth trajectory into control commands and send them to the robot. The control commands include at least trajectory tracking commands, impedance adjustment commands, and velocity decay commands.

[0015] The specific values ​​of the preset thresholds and preset parameters involved in this invention are obtained by those skilled in the art through statistical calibration of historical collision-free safe operation data based on the dynamic parameters of different robot models and the redundancy requirements of specific operating environments; the specific methods for obtaining these values ​​can all be calculated using the following statistical formulas: ,in, This represents a preset threshold value for the target. This represents the statistical mean of the parameter across a historical set of safe operating samples. This represents the corresponding standard deviation. This represents the safety margin coefficient set based on different environmental hazard levels.

[0016] This embodiment provides a robot path generation mechanism based on teaching and path optimization. Specifically, this mechanism is applied to a heavy collaborative robot transfer operation scenario in a nuclear reactor decommissioning area. The operator is located in a shielded control room and guides the robot to move a high-risk reactor vessel through a remote teaching device. The robot runs in the ruins passage where radiation, hot air flow and dynamic debris act together. The system executes a closed-loop control process that includes collecting teaching intentions, sensing external disturbances, observing human-machine control divergences, predicting oscillation risks, and generating smooth trajectories, in order to avoid human-machine control adversarial situations caused by obstacle avoidance based solely on local shortest paths and forced distance minimization. The teaching acquisition module continuously receives raw teaching trajectory data from the remote teaching device; this raw teaching trajectory data can be represented as an end pose sequence arranged according to the sampling period, or it can be further decomposed into a combination of position vector, attitude vector and timestamp. For example, at three consecutive sampling times , , The operator expects the robotic arm's end effector to move to the target pose point. , , ; Meanwhile, the environmental perception module collects environmental disturbance data and current dynamic state data from the robot side. The former may include changes in the direction of hot airflow, changes in the location of falling gravel, and the speed of nearby obstacles, while the latter may include joint torque, end-effector acceleration, joint angular velocity, and contact force estimates. The state observation module performs local path optimization based on the operational environment disturbance data to obtain a local obstacle avoidance trajectory. This local path optimization algorithm can be implemented by obstacle expansion modeling and path cost search in the rolling time domain, and is not limited to a specific solver. Taking the above three teaching trajectories as an example, if in A piece of loose rock rolling to the left was detected nearby. The path obtained after local path optimization may be derived from... , , Composition, in which Compared to It is raised 3 units vertically to temporarily bypass obstacles; The state observation module aligns the local obstacle avoidance trajectory with the original teaching trajectory and calculates the trajectory deviation value; this deviation value can be represented by the average or weighted sum of the Euclidean distances of each sampling point. Taking the above example, the deviation at time t2 is If the deviation is 0 at other times, the average trajectory deviation value within this small window can be recorded as 1. The state observation module also extracts compensation operation instructions with frequencies higher than a preset frequency threshold from the original teaching trajectory data; here, compensation operation instructions refer to the reverse or corrective inputs that the operator quickly applies after sensing that the robot's actions deviate from its expectations. Specifically, the system will The teaching input data within the time window is discretized into five operation states: push forward, pull back, left edit, right edit, and hold. If the statistics within this window show 1 push forward, 4 pulls back, 3 left adjustments, 0 right adjustments, and 2 hold-forwards, then the corresponding probability distribution can be approximated as follows: When certain compensation operation commands appear repeatedly at a high frequency, it indicates that the operator is correcting the system state with high-frequency input. Based on this, the state observation module calculates the information entropy according to the probability distribution of trajectory deviation value and compensation operation command, and uses it as the human-machine control divergence entropy; In specific calculations, to overcome the inherent limitation that the traditional Shannon information entropy converges to 0 and is misjudged as a low-divergence state by the system due to the operator's high-frequency continuous correction in a single direction, such as pulling back with a single input throughout the entire cycle, the state observation module uses the cross-entropy algorithm to solve for the information entropy. Specifically, the system uses the ideal stationary distribution where the operator does not intervene at all, i.e., the probability of operation is kept at 1 and the probability of other compensation operations is 0, as a reference baseline distribution. The system calculates the cross-entropy between the probability distribution of the currently collected compensation operation commands and this ideal stationary distribution, thereby quantifying the overall degree to which the operator's behavior deviates from the stationary state without intervention. The specific mathematical expression for cross-entropy is:

[0017] in, This represents the probability distribution of the currently collected compensation operation commands. This represents an ideal, stable distribution where the operator does not intervene at all. Represents the discrete state space of operation instructions within a preset time window; The normalization process uses the maximum minimum normalization algorithm, taking the maximum reachable distance in the robot's workspace as the maximum reference deviation value, and mapping the actual trajectory deviation value to a dimensionless range of 0 to 1. The trajectory deviation value is normalized and used as a modulation factor, which is then multiplied by the aforementioned cross-entropy to obtain the human-machine control divergence entropy. This design assigns a higher degree of divergence to cases where the trajectory deviation is larger and the deviation from the stationary distribution is more significant. For example, within a window, if the trajectory deviation is normalized to 0.8, and the operator continuously inputs a single-direction forced compensation command multiple times, such as a pull-back operation, although the actions are highly concentrated, the calculated cross-entropy tends to a maximum value because the probability distribution deviates significantly from the baseline state. At this time, the system can correctly output a human-machine control divergence entropy higher than the preset divergence threshold, accurately reflecting the significant one-way control deviation state. If the dispersion of the compensation operation distribution is higher than a preset threshold, the system can also calculate a higher divergence entropy due to deviation from the stationary baseline; conversely, if the trajectory deviation is only... Furthermore, the operator maintains stable input during most sampling periods, the current distribution is close to the ideal distribution, the cross-entropy approaches zero, and the final calculated human-machine control divergence entropy approaches the minimum value. The state observation module inputs the human-machine control divergence entropy and the current dynamic state data into the oscillation prediction model to output the induced oscillation probability; The model can be pre-trained using historical operation data. The oscillation prediction model is constructed using a long short-term memory network or a random forest algorithm. The training samples consist of historical trajectory segments, changes in joint torque and end-effector acceleration at corresponding moments, and labels indicating whether oscillation has occurred. To illustrate the data flow process, we can assume that there are three types of segments in the historical samples: the first type is low divergence entropy + low torque fluctuation, labeled as stable; the second type is high divergence entropy + high acceleration abrupt change, labeled as oscillation precursor; and the third type is medium divergence entropy + high load slow change, labeled as recoverable. When the system is given a state combination at the current moment, such as a divergence entropy of 0.82, a joint torque utilization rate of 0.91, and an end-effector acceleration change rate of 0.75, the model outputs an induced oscillation probability, such as 0.87, indicating a high risk of controlled oscillations occurring within a preset time window in the future. The control center dynamically allocates teaching compliance weights and autonomous obstacle avoidance weights based on divergence entropy and induced oscillation probability. Specifically, when the system determines that the operator's intention is basically consistent with the local optimization result, the teaching compliance weight is increased, and the robot retains more of the inertial direction of the original teaching path. When the system determines that there is a significant control deviation between the local obstacle avoidance trajectory and the human-taught trajectory input, and detects the risk of induced oscillation, the proportion of autonomous obstacle avoidance weight and trajectory smoothing constraint is increased to avoid control command conflicts between operator input and control algorithm output. After the target smooth trajectory is generated, it is converted into control commands by the control center and sent to the robot. The control commands include at least trajectory tracking commands, impedance adjustment commands, and velocity decay commands. Among them, the trajectory tracking command is used to drive the robot to move along the target smooth trajectory, the impedance adjustment command is used to adjust the end-effector compliance characteristics, and the velocity decay command is used to actively reduce the movement speed in high-risk phases. As an abnormal state protection mechanism, if the teaching acquisition module does not receive valid teaching data in a certain sampling period, it can hold the teaching input of the previous valid period for a short time and reduce the teaching compliance weight to prevent the control quantity from abruptly changing. If the environmental perception module temporarily loses some obstacle measurements, the most recent valid obstacle status can be used for short-term extrapolation, while increasing the safety redundancy distance. If the number of compensation operation samples calculated by the state observation module is insufficient and cannot stably form a probability distribution, then the human-machine control divergence entropy within this period can degenerate into a single-factor estimate based on the trajectory deviation value. If the confidence level of the oscillation prediction model is lower than the preset threshold, the control center adopts a conservative strategy, that is, it prioritizes reducing the speed and strengthening the impedance damping, rather than directly executing trajectory switching that exceeds the preset displacement threshold. At the nuclear reactor decommissioning site, robots are carrying reactors containing residual corrosive media through a steel structure corridor with preset clearance tolerances. The operator was initially teaching the robot to move at a constant speed along the center line of the corridor, but the environmental perception module detected falling debris and rising heat flow on the right side. The local path optimization result raised the middle section of the trajectory to the upper left. The operator perceived through the monitoring screen that the robot was not in line with his expectations and input two right correction commands and one pull-back compensation command in succession. Based on this, the state observation module obtains a large trajectory deviation value and a biased high-frequency compensation distribution, and further calculates the increase in divergence entropy; The oscillation prediction model, combined with joint torque data under high load, outputs a higher probability of induced oscillation. The control center thus abandons the constraint of minimizing the spatial distance of obstacles as the sole objective function for local trajectory optimization. Instead, it generates a target smooth trajectory with the velocity parameter decaying to 60% to 80% of the initial velocity and the radius of curvature increasing to more than 1.2 times the radius of curvature of the original trajectory. At the same time, it attaches an impedance adjustment command to make the end effector exhibit higher damping in order to suppress possible oscillations in the human-machine loop. The purpose of this step is to move away from using the minimum trajectory error or the maximum obstacle avoidance distance as the sole objective in high-risk, heavy-load operations. Instead, it incorporates human-machine control discrepancies and oscillation risks into a real-time closed loop, enabling the robot to adapt to dynamic environments while retaining the interpretability of human intentions. This achieves a unified coordination of remote teaching, safe obstacle avoidance, and system stability.

[0018] In a preferred embodiment of the present invention, the control center is preset with a safety divergence threshold; used to increase the teaching compliance weight and decrease the autonomous obstacle avoidance weight when the human-machine control divergence entropy is lower than the safety divergence threshold; and to decrease the teaching compliance weight and increase the autonomous obstacle avoidance weight when the human-machine control divergence entropy is higher than or equal to the safety divergence threshold.

[0019] This embodiment provides a control weight switching mechanism based on a safety divergence threshold. Specifically, in the aforementioned nuclear decommissioning and handling scenario, relying solely on continuous numerical changes to smoothly adjust the control weights can reflect high-precision continuous control state variables. However, when there is network latency and sudden changes in obstacles on-site, the control strategy of the control center may oscillate, causing the robot to be unable to accurately track the taught path or completely output safe obstacle avoidance maneuver commands within a short window, resulting in frequent trajectory planning switching and end effector jitter. Therefore, this embodiment introduces a safety divergence threshold to quickly distinguish the human-machine relationship from two state intervals: basically consistent and significantly contradictory. The control center pre-stores a safety divergence threshold, which can be obtained through historical task data statistics. The specific method of statistics is as follows: collect historical collision-free safe operation data, calculate the mean and standard deviation of human-machine control divergence entropy in all sampling periods, and add three times the standard deviation to the mean as the safety divergence threshold. It can also be obtained by experimental tuning based on multiple sets of different delays and load conditions during the debugging phase; for example, the normalized value range of the divergence entropy is 0 to 1, and 0.6 can be set as the safe divergence threshold. When the human-machine control divergence entropy of the current cycle is less than 0.6, it indicates that the operator still has a high degree of recognition of the robot's current actions. The control center increases the teaching compliance weight and reduces the autonomous obstacle avoidance weight. Specifically, if both weights are 0.5 in the initial state, then when the divergence entropy is 0.3, the teaching compliance weight can be adjusted to 0.7 and the autonomous obstacle avoidance weight to 0.3, so that the target trajectory is closer to the original teaching trajectory. Conversely, when the divergence entropy is higher than or equal to 0.6, it indicates that there is a significant conflict in the understanding of the trajectory between the human and the machine. Continuing to follow the teaching instructions at a high proportion may amplify the repeated corrections caused by delayed input. Therefore, the control center reduces the teaching compliance weight and increases the autonomous obstacle avoidance weight. For example, when the divergence entropy is 0.82, the weights can be adjusted to 0.35 and 0.65, so that the system prioritizes following the safe movement trend after perception and optimization verification; To avoid new control shocks caused by sudden weight changes, a ramp transition can be implemented over two or more sampling periods during actual distribution, for example, gradually changing from 0.5 / 0.5 to 0.45 / 0.55, 0.4 / 0.6, or 0.35 / 0.65. As an abnormal state protection mechanism, if the current divergence entropy is close to the threshold and fluctuates repeatedly, such as swinging between 0.58 and 0.62 for several consecutive cycles, directly switching by a step according to the threshold may cause the weight to flip frequently. To address this, a short hysteresis range can be set, such as 0.58 to 0.62. When entering this range, the control mode of the previous cycle remains unchanged, and the switch is only executed after the system has stably fallen to the same side for several consecutive cycles. If the system is powered on for the first time or the current task segment has not accumulated enough divergence entropy samples, it can start with neutral weights and gradually approach the target weights based on subsequent stable observation results. During the aforementioned steel structure corridor handling process, when the robot first entered the passage, there were few obstacles, and the local obstacle avoidance trajectory only deviated slightly from the taught trajectory. The operator did not make any significant compensation input, and the divergence entropy was 0.28. Based on this, the control center increased the teaching compliance weight, so that the robot maintained a high proportion of teaching follow characteristics. A metal plate on the corridor ceiling tilted and drooped due to aftershocks. In order to avoid the obstacle, the state observation module generated a local obstacle avoidance trajectory that was significantly expanded outward based on the local path optimization algorithm. However, due to the delay in the monitoring screen, the operator still tried to make the robot move along the original center line, and the divergence entropy rose to 0.76. The control center then switches to a high-divergence mode, actively increasing the autonomous obstacle avoidance weight to prevent the robot from being forcibly pulled back to the danger zone under high load conditions by delayed input; The purpose of this mechanism is to establish a clear human-machine control boundary through a safety disagreement threshold, so that the system can retain more of the operator's teaching intentions during the consistency phase and more decisively favor physical safety constraints during the conflict phase, thereby achieving dynamic redistribution of control and suppressing trajectory oscillations.

[0020] In a preferred embodiment of the present invention, the control center is specifically used to: when the human-machine control divergence entropy is higher than or equal to the safety divergence threshold, to perform weighted fusion and filtering on the original teaching trajectory data and the local obstacle avoidance trajectory according to the adjusted weights, to generate a first smooth trajectory as the target smooth trajectory, and to generate the speed decay command accordingly.

[0021] This embodiment provides a flexible smooth attenuation and trajectory frequency reduction mechanism under high divergence conditions. Specifically, in the previous scheme, even if the control weight has been biased towards the autonomous obstacle avoidance side through the safety divergence threshold, if the state observation module is continuously disturbed by high-frequency hot airflow and gravel displacement, its output trajectory may still show high-frequency oscillating trajectory segments. Although such high-frequency fine-tuning is geometrically closer to the local obstacle optimal solution, it appears as high-frequency jittering motion at the end of the remote operator, which can easily further stimulate the compensation input. Therefore, in this embodiment, the original fusion trajectory is not used directly under high divergence conditions. Instead, low-pass filtering and velocity smoothing attenuation are performed to filter out local optimal high-frequency fine-tuning. When the divergence entropy is higher than or equal to the safe divergence threshold, the control center generates an initial trajectory by weighted fusion of the original teaching trajectory and the local obstacle avoidance trajectory based on the reduced teaching compliance weight and the increased autonomous obstacle avoidance weight. For example, in a certain trajectory segment, the sequence of teaching points is: , , , The local obstacle avoidance point sequence is as follows , , , ; If the current weights are set to 0.35 for teaching compliance and 0.65 for autonomous obstacle avoidance, the initial trajectory points obtained after fusion can be C1, C2, C3, and C4, respectively. C2 is closer to B2, and C3 is closer to B3, with the overall trend towards the safe obstacle avoidance side. However, if there are lateral offsets and longitudinal folding changes with frequencies higher than the preset frequency threshold between local obstacle avoidance points, the initial trajectory after fusion may still retain high-frequency oscillations. Therefore, the control center performs low-pass filtering on the initial trajectory; the low-pass filtering is used to extract the main frequency displacement trend of the trajectory and attenuate transient high-frequency fine-tuning components with frequencies higher than the preset filtering frequency threshold. For ease of understanding, it can be assumed that the preset filter frequency threshold corresponds to a frequency exceeding... The direction reversal trajectory segment is considered as high-frequency fine-tuning; If the lateral offsets in five consecutive sampling points are 0, 2, -2, 3, and -1 respectively, it indicates that there is rapid and iterative correction. After low-pass filtering, it can be smoothed to 0, 1, 1, 2, 2, so that the end motion changes from a sawtooth shape to a gradually changing curve. After this processing, the first smooth trajectory is obtained and used as the target smooth trajectory. After generating the first smooth trajectory, the control center further generates a speed decay command based on the target smooth trajectory; the speed decay is not a simple global deceleration, but is executed segment by segment in combination with the local curvature of the current trajectory, load status and oscillation risk; Specifically, if a certain path segment is originally planned to advance at 0.8 m / s, and the segment has both high curvature and high bifurcation characteristics, the control center can reduce the speed of that segment to 0.45 m / s. If the environment stabilizes again in subsequent segments and the divergence entropy decreases, the velocity can gradually recover to 0.6 m / s or higher; this makes the motion state of the system more consistent and predictable. As an anomaly handling mechanism, if the distance between the trajectory obtained after low-pass filtering and the actual obstacle safety boundary is less than the preset safety margin, or even if there is a local crossing of the obstacle expansion zone, the safety constraint should be satisfied first, the obstacle avoidance weight of the local segment should be increased again and secondary smoothing should be performed, instead of directly applying the filtering result; if the distance between the trajectory obtained after low-pass filtering and the actual obstacle safety boundary is greater than or equal to the preset safety margin, the filtering result should be directly output as the first smoothed trajectory. If an unacceptable high curvature abrupt change still exists in the trajectory after multiple filtering, the control center can directly trigger a stronger level of speed decay, and if necessary, order the robot to stop and wait in the current safe zone. If the system still cannot stably track the target's smooth trajectory under the current dynamic constraints after the velocity decays, the trajectory segment can be marked as unexecutable, and the system switches to conservative risk avoidance mode. As the reactor passes through a narrow passage affected by hot airflow, the heat wave at the top pushes the lightweight debris to drift laterally. The state observation module uses a local path optimization algorithm to perform obstacle avoidance correction with a preset offset threshold once every tens of milliseconds on the local obstacle avoidance trajectory. The operator sees the robot end effector frequently changing direction on the monitor and begins to continuously apply correction inputs opposite to those of the system. The divergence entropy quickly crosses the safety threshold. The control center first generates an initial trajectory according to the weight biased towards autonomous obstacle avoidance, and then removes the fine-tuning segment of continuous swing through low-pass filtering to obtain a first smooth trajectory that detours to the left as a whole, but no longer produces high-frequency oscillations exceeding the preset displacement variance threshold locally. At the same time, a speed decay command is issued to control the robot to safely pass through the narrow space at the decayed speed. The operator observes that the end effector becomes predictable and the compensation input is significantly reduced. The purpose of this step is to proactively implement a smooth attenuation strategy that reduces the correction frequency, decreases the operating speed, and maintains the main motion trend when the human-machine divergence has entered the danger zone. This is to filter out the high-frequency local corrections that are most likely to cause human-machine resonance, thereby prioritizing the maintenance of the stability boundary of the entire control loop.

[0022] In a preferred embodiment of the present invention, the control center is specifically used to: when the human-machine control divergence entropy is lower than the safety divergence threshold, to perform weighted fusion of the original teaching trajectory data and the local obstacle avoidance trajectory according to the adjusted weights, generate a second smooth trajectory as the target smooth trajectory, and generate the trajectory tracking command accordingly.

[0023] This embodiment provides an intent-first trajectory following mechanism in a low-divergence state. Specifically, in a high-divergence scenario, strong smoothing and velocity smoothing attenuation are beneficial for suppressing oscillations. However, if this strategy is applied indiscriminately to all operating phases, it will cause the system to exhibit trajectory tracking response lag in operating phases within the safety boundary, significantly reducing the operator's control efficiency over the end-effector's pose accuracy. This is particularly unfavorable for completing high-precision tasks such as docking, critical fitting and crossing, and load attitude fine-tuning. Therefore, this embodiment preserves the dominance of the operator's intent when the divergence entropy is below the safe divergence threshold, in order to obtain higher teaching fidelity and operational efficiency; When the divergence entropy is lower than the safe divergence threshold, the control center performs weighted fusion of the original teaching trajectory and the local obstacle avoidance trajectory based on the increased teaching compliance weight and the decreased autonomous obstacle avoidance weight to generate a second smooth trajectory. Unlike the high divergence state, it is not necessary to focus on suppressing all high-frequency components at this time. Instead, priority should be given to preserving the transitions and posture adjustments that have task significance in human teaching. For example, in order to accurately align the support legs of the reactor with the slots of the transfer platform, the operator inputs a rightward translation and then downward movement that reaches the preset translation step length in the end trajectory. If the current environment is stable and the divergence entropy is low, this movement should be retained in a large proportion. For example, if the teaching trajectory point is , , The local obstacle avoidance trajectory points are , , With a teaching compliance weight of 0.75 and an autonomous obstacle avoidance weight of 0.25, the fusion results are closer to the expected values. The sequence preserves the operator's preset action logic of first making horizontal fine adjustments and then placing the object vertically. After generating the second smooth trajectory, the control center generates trajectory tracking instructions based on the trajectory; the trajectory tracking instructions may include control targets such as desired position, desired velocity and desired attitude, and are sent to the robot controller according to the sampling period; To make tracking more stable, a feedforward term oriented to the current dynamic state can be added to the trajectory tracking command. For example, during the heavy load lifting phase, the driving force compensation in the lifting direction can be appropriately increased, and during the slow landing phase, the upper limit of the end velocity can be reduced. This can ensure tracking response and reduce overshoot caused by inertia. As an anomaly handling mechanism, if the divergence entropy is lower than the safety threshold, but the local environment perceives that the distance between an obstacle and the load outline is less than the minimum safe gap (the minimum safe gap is a dynamic braking distance calculated in real time based on the robot's current running speed, maximum braking acceleration, and system communication network delay parameters), then the teaching compliance ratio cannot be simply increased further. Instead, the autonomous obstacle avoidance weight should be forcibly increased during that local time period. Conversely, if the local environment perceives that the distance between an obstacle and the load outline is greater than or equal to the minimum safe gap, then the currently increased teaching compliance weight should be maintained. If the teaching trajectory itself has discontinuous jump points, such as the operator accidentally touching the wrong point and causing the desired end position to suddenly cross a distance exceeding the preset displacement threshold, then interpolation and smoothing must be performed before the trajectory tracking command is generated to avoid the robot from performing step-change actions that cannot be achieved. If the second smooth trajectory obtained by fusion exceeds the current joint velocity or acceleration constraint, the time parameterization process should be automatically extended to expand the same path over a longer period of time. After the robot completes the passage through the dangerous channel, it enters a relatively stable reactor transfer area with a flat ground and the number of dynamic obstacles is below the preset interference threshold. The operator needs to control the bottom support frame of the reactor to accurately fall into the vibration damping bracket. At this point, the correction amount given by the local path optimizer is lower than the preset correction threshold, the operator input is also relatively consistent, and the divergence entropy remains at around 0.22. The control center then increases the teaching compliance weight, integrates the critical fitting, alignment, and descent actions given by the operator into the second smooth trajectory in a large proportion, and generates fine trajectory tracking instructions. The robot can then perform low-speed precision alignment without deviating from the operator's assembly expectations due to the system excessively increasing its autonomous avoidance response. The purpose of this step is to fully leverage the semantic advantages of remote teaching in a stage where human-machine understanding is consistent and environmental risks are low, so that the robot can maintain the necessary obstacle avoidance capabilities and faithfully execute the operator's fine control intentions as much as possible, thereby improving work efficiency and operational controllability.

[0024] In a preferred embodiment of the present invention, the control center has an adaptive impedance control algorithm for dynamically adjusting the target stiffness and target damping values ​​of the robot end effector based on the induced oscillation probability and the current dynamic state data, and generating the impedance adjustment command.

[0025] This embodiment provides an adaptive impedance adjustment mechanism that incorporates oscillation risk; specifically, the aforementioned trajectory-level weight allocation and smoothing can reduce human-machine control conflicts, but when the robot is carrying a heavy load close to its dynamic limit, simply changing the path is still insufficient to completely suppress mechanical vibration and overshoot after the end effector is disturbed. Especially in narrow spaces, where there is no longer enough space to plan the path to meet the preset redundancy safety distance, if the end stiffness is higher than the preset stiffness upper limit threshold, the operator's compensation input and the local collision avoidance action can easily amplify the system impact together. If the stiffness is lower than the preset lower limit threshold, it may cause load dragging, attitude sag and trajectory tracking error to exceed the preset allowable range. Therefore, this embodiment introduces an adaptive impedance control algorithm to dynamically correct the stiffness and damping of the end effector under different oscillation risk levels. The control center pre-sets a set of initial stiffness and initial damping values; for example, under normal medium-speed transport conditions, the initial stiffness in the end-effector translation direction can be set as follows: The initial damping is The control center receives real-time data on the induced oscillation probability and the current dynamic state, and adjusts the data accordingly. and Perform dynamic compensation; The basic idea is: when the oscillation probability increases, the acceleration changes significantly, or the joint torque approaches its limit, the damping is appropriately increased to absorb high-frequency energy, while the equivalent stiffness is reduced in the necessary direction to reduce the reverse impact caused by rigid tracking; when the oscillation probability is low and precise positioning is required, the stiffness is moderately restored to ensure attitude stability and position accuracy. Specifically, when calculating, the adaptive impedance control algorithm multiplies the induced oscillation probability, the normalized joint torque utilization rate in the current dynamic state data, and the normalized end acceleration change rate by the corresponding risk weight factors, and sums them up to obtain the comprehensive risk coefficient. The target stiffness value is calculated by subtracting the product of the preset stiffness adjustment base and the comprehensive risk coefficient from the initial stiffness value; the target damping value is calculated by adding the product of the preset damping adjustment base and the comprehensive risk coefficient to the initial damping value; wherein, the preset stiffness adjustment base and the preset damping adjustment base are calculated based on the maximum allowable stiffness difference and the critical damping ratio difference of the robot end effector according to a fixed ratio; Assuming the initial stiffness is 100 and the initial damping is 20 at a certain moment; the current oscillation probability is 0.85, the joint torque utilization rate is 0.9, and the end-effector acceleration change rate is 0.7, combined with the corresponding weighting factors, the comprehensive risk coefficient is found to be at a high level. Then, the adaptive impedance algorithm outputs a target stiffness of 70 and a target damping of 35, which indicates that the system avoids the end-effector from generating a high-amplitude mechanical oscillation response when subjected to human-machine control deviation input by dynamically reducing the target stiffness and increasing the target damping. If the oscillation probability at another moment is only 0.15, the torque utilization rate is 0.5, and the end acceleration change rate is 0.2, the calculated comprehensive risk coefficient is low. The algorithm can output a target stiffness of 110 and a target damping of 22, so that the robot can maintain a better load support capacity under stable working conditions. In actual control, this adjustment can be implemented uniformly in all directions, or different compensation amounts can be set for the forward direction, yaw direction and vertical direction respectively; The control center generates an impedance adjustment command based on the target stiffness and target damping values ​​and sends it to the robot's underlying execution controller. This command can be refreshed at a fixed period to keep the impedance parameters updated in sync with the trajectory control. To maintain system stability, parameter changes are typically constrained by amplitude and slope. For example, stiffness changes are limited to no more than 10 units and damping changes to no more than 5 units per sampling period to avoid new instability caused by a step change in impedance parameters. As an anomaly handling mechanism, if there are abnormal missing data in the current dynamic state data, such as short-term distortion of the torque sensor, the impedance compensation amount can be temporarily estimated based only on the oscillation probability and historical effective state, and the compensation amplitude can be limited. If the target stiffness calculated by the algorithm is lower than the lower limit allowed by the equipment, it may cause the load support to become unstable, so it should be truncated to the safe lower limit. If the target damping is higher than the allowable upper limit, it may cause response lag, so saturation processing is also performed; if the oscillation probability suddenly increases but the duration is lower than the preset duration threshold, it is not enough to determine the true risk peak. It can be required that the probability remain high for several consecutive sampling periods before a larger impedance reconstruction is performed. When the robot was moving the reactor close to the temporary vibration damping platform, the front end of the robot had to cross a section of locally heated steel beam. The surface of the steel beam was deformed, which caused the end acceleration data to fluctuate beyond the preset variance. Although the target trajectory had been smoothed, the operator still input a small reverse correction due to the large load inertia. The control center detects that the oscillation probability has increased to 0.78 and the joint torque is at a high level. Based on the increase in the comprehensive risk coefficient, it uses an adaptive impedance control algorithm to reduce the forward stiffness and increase the lateral damping, so that the end effector can absorb the impact energy without significantly deviating from the target trajectory. When the robot enters the platform and waits for precise placement, the oscillation probability decreases, and the system gradually restores its stiffness in order to complete a stable placement. The purpose of this mechanism is not only to alleviate human-machine divergence at the path level, but also to proactively shape execution characteristics with appropriate stiffness and adaptive damping at the mechanical interaction level, so that the robot can maintain a controllable dynamic response when approaching the system dynamic limit, thereby further suppressing induced oscillations.

[0026] In a preferred embodiment of the present invention, the current dynamic state data includes joint torque data and end effector acceleration data; the control center is used to generate and send a stagnation lock command to the robot when the change in the end effector acceleration data or the joint torque data exceeds the corresponding preset threshold; otherwise, the sending of the control command is maintained.

[0027] This embodiment provides a stall lock-in protection mechanism for extreme working conditions; specifically, the aforementioned divergence management, trajectory smoothing and impedance compensation can significantly reduce risks, but in nuclear decommissioning heavy-load transportation tasks, some transient events that exceed normal adjustment capabilities may still occur, such as gravel suddenly getting stuck between the vehicle body and the wall, short-term shift of the reactor center of gravity, and obstruction of a certain joint of the robotic arm near a strange pose. In such transient conditions, if the system maintains the original continuous control strategy based on error correction, it is easy for local transient disturbances to gradually amplify and induce global system instability. Therefore, this embodiment sets a hard boundary, and when the end acceleration changes abruptly or the joint torque exceeds the limit, the stop lock command is directly triggered. The control center reads joint torque data and end-effector acceleration data from the current dynamic state data, and calculates the change in end-effector acceleration between adjacent sampling periods; for example, if the end-effector accelerations sampled in two adjacent periods are 1.2 and 2.0 respectively, then the change is 0.8. If the preset acceleration change threshold is 0.6, then the change has exceeded the threshold. Similarly, if the real-time torque of a critical joint reaches more than 95% of its allowable upper limit, while the preset torque limit threshold corresponds to 90%, then it is considered that the torque is over the limit. As long as either the acceleration change exceeds the threshold or the joint torque exceeds the threshold, the control center will generate a stop lock command. After the stagnation lock command is issued, the robot will no longer continue to perform trajectory tracking of the current propulsion segment, but will enter a controlled stagnation state. This state is not a simple power cut-off, but rather to maintain the stability of the end-load attitude, prioritize locking the current position or fine-tuning it to the nearest safe stationary attitude, and maintain sufficient damping and support force to prevent the load from swinging. If combined with the aforementioned impedance adjustment mechanism, the damping can be increased to a conservative value during the stagnation period, and all degrees of freedom of motion in non-essential directions can be restricted. Conversely, if the change in acceleration at the end of an adjacent sampling period is lower than or equal to the acceleration mutation threshold, and the joint torque is lower than or equal to the torque limit threshold, the control center will continue to issue existing control commands without triggering a stall lock. As an abnormal state protection mechanism, if the change in terminal acceleration exceeds the limit for a short period of time, but the duration is only a single sampling cycle and it immediately returns to normal, two strategies can be selected according to the system settings: instantaneous alarm but not lock or lock first and then manually confirm recovery. If multiple joint torque sensors fail simultaneously, making it impossible to reliably determine whether the limit has been exceeded, the system should handle the situation in the most conservative way, i.e., default to entering a standstill lock and requesting manual review. If the stall lock occurs in a narrow area, and the current position is too close to the obstacle boundary, the control center can perform a safe retraction action with a displacement amount lower than the preset retraction threshold, retracting the end effector to the predefined nearest safe zone before completing the lock. If the environmental disturbance disappears after locking and manual confirmation is passed, the system can replan the recovery segment from the most recent effective smooth trajectory point, instead of directly and seamlessly continuing the original high-risk action; Just before the robot was about to place the reactor onto the vibration damping platform, a loose steel plate on the ground suddenly lifted up due to aftershocks, causing the chassis to shake instantly. The acceleration at the end of the robotic arm jumped rapidly from 0.9 to 1.7, with the change significantly exceeding the threshold. At the same time, the torque supporting the main joint is also approaching its limit; the control center immediately determines that it is close to the physical failure boundary and issues a stop and lock command; The robot stops moving forward, maintains its current posture at the end effector and increases damping to prevent the reactor from continuing to swing forward due to inertia; after the ground stabilizes, the operator and the system jointly confirm whether to start again from a safe position. The purpose of this mechanism is to set a final hard protection boundary for the entire control system. When the dynamic risks exceed the flexible adjustment capabilities, the propulsion stop procedure will be triggered immediately to prevent high-risk loads from becoming unstable, colliding, or falling.

[0028] In a preferred embodiment of the present invention, the working environment disturbance data includes thermodynamic airflow change data and dynamic obstacle displacement data; the teaching acquisition module is also used to acquire communication status data between the remote teaching device and the system, the communication status data including network latency parameters and signal jitter parameters; the status observation module is also used to correct the human-machine control divergence entropy based on the communication status data.

[0029] This embodiment provides a divergence entropy correction mechanism that takes into account communication quality degradation. Specifically, in the aforementioned scheme, the human-machine control divergence entropy is mainly obtained from the trajectory deviation value and the high-frequency compensation operation distribution, which can already reflect the intention conflict between the operator and the control center. However, in the high radiation environment of nuclear decommissioning, remote teaching links are often accompanied by significant network delays and signal jitter. At this point, the compensation instructions issued by the operator may not necessarily fully represent the true intention to resist, but may just be outdated corrections caused by the accumulation of delays; if the difference between pseudo-disagreements caused by communication degradation and actual control disagreements is not distinguished, the system is prone to erroneously amplifying risk assessments and becoming overly conservative. Therefore, this embodiment further incorporates communication status data into the divergence entropy correction process; In addition to conventional obstacle information, the environmental perception module also collects thermal airflow change data and dynamic obstacle displacement data for the disturbance data of the working environment. The former can be represented as the heat flow intensity, direction and fluctuation frequency of a local area, while the latter can be represented as the position, velocity and displacement trend of dynamic obstacles such as gravel, steel plates, and hanging parts. At the same time, the teaching acquisition module also collects communication status data between the remote teaching device and the system, including at least network latency parameters and signal jitter parameters; For example, if the average one-way delays measured in several consecutive sampling windows are 260ms, 320ms, and 410ms, and the jitter amplitudes are 30ms, 70ms, and 120ms, it can be determined that the link quality is deteriorating. The state observation module constructs a penalty weight based on network delay parameters and signal jitter parameters, and uses this penalty weight to correct the original man-machine control divergence entropy. The penalty here is not simply to increase the value, but to reflect whether the current divergence judgment should be amplified or reduced. Suppose that at a certain moment, the original divergence entropy calculated from the trajectory deviation and compensation input is 0.70; If the network latency is only 80ms and the jitter is 10ms, it indicates that the communication link is good. In this case, the penalty weight can be taken as a value close to 1, such as 1.00. After correction, the divergence entropy is still 0.70. If the network latency rises to 400ms and the jitter is 120ms, it indicates that the operator's compensation input is very likely to lag behind the robot's current real state. In this case, a reduction-type correction can be used, such as a penalty weight of 0.85, so that the corrected divergence entropy becomes 0.595, thus preventing the system from misjudging the communication lag as a complete human-machine confrontation. Conversely, if the delay and jitter are moderate, but the direction of the compensation input is consistently opposite to the local obstacle avoidance trend in the long term, it indicates that the discrepancy is more likely to be real. In this case, the penalty weight can be designed to be amplified, such as 1.10, thereby increasing the risk sensitivity. Thermodynamic airflow change data and dynamic obstacle displacement data can also be used together with communication status to determine the construction method of penalty weights; If the environment itself experiences severe high-frequency disturbances and communication is significantly delayed, the monitoring images obtained by the operator will lag behind the actual physical scene. In this case, priority should be given to suppressing misinterpretations of compensation commands. If the environment is relatively stable but the compensation command continues to be high-frequency and reversed, it should be determined that the divergence entropy reflects the actual human-machine control divergence. In this way, the corrected divergence entropy is closer to the real human-machine relationship state, which can provide a more reliable basis for subsequent oscillation prediction and control weight allocation. As an abnormal state protection mechanism, if the network delay parameter or signal jitter parameter is temporarily missing, the penalty weight can be rolled back to 1, that is, the original divergence entropy is not corrected. If the delay reaches the preset delay limit and exceeds the system's tolerable upper limit, such as exceeding the preset limit link threshold, then instead of relying solely on entropy correction, the communication degradation process should be triggered directly, such as freezing high-risk operations, reducing speed, or pausing tasks. If jitter data is affected by transient electromagnetic interference and outliers appear, a sliding window midpoint or quantile can be used to replace single-point observations to avoid excessive disturbance to the divergence entropy correction by a single outlier packet. If the communication status is good within the same window but the environmental disturbance is extremely strong, the proportion of communication factors in the penalty weight can be reduced so that the correction result reflects the on-site physical risk more. When the reactor passes through a high radiation area, the electromagnetic interference between the equipment is enhanced, the remote teaching link delay increases from the original 120ms to 380ms, and is accompanied by signal jitter with an amplitude exceeding the preset jitter threshold. Meanwhile, the hot airflow at the top of the corridor causes lightweight obstacles to drift laterally, and the local path optimizer continuously generates small avoidance trajectories; the monitoring images obtained by the operator lag behind the actual physical scene, and continuously issue compensation inputs opposite to the current optimal obstacle avoidance direction; if calculated only based on the compensation frequency and trajectory deviation, the divergence entropy is higher than the preset limit divergence threshold. After the introduction of communication status correction, the status observation module identifies that a significant portion of the adversarial behavior comes from link lag, and thus reduces and corrects the divergence entropy. Based on this, the control center adopts a more robust and smooth deceleration strategy, rather than over-evaluating the operator's control deviation, thereby avoiding control mode switching that exceeds the preset switching frequency threshold. The purpose of this mechanism is to separate communication non-idealities from human-machine behavior, so that the divergence entropy is no longer just a statistical result of surface action deviations, but becomes a real risk indicator after integrating environmental disturbances, operational compensation, and link status, thereby achieving more reliable state observation of remote heavy-load collaborative tasks.

[0030] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A robot path generation system based on teaching and path optimization, characterized in that, The system includes: The teaching acquisition module is communicatively connected to the remote teaching device and is used to acquire the original teaching trajectory data of the remote teaching device. An environmental perception module is communicatively connected to a robot with an end effector and is used to collect data on environmental disturbances and current dynamic state of the robot. The state observation module is communicatively connected to both the teaching acquisition module and the environment perception module. It generates a local obstacle avoidance trajectory based on the operational environment disturbance data, compares the local obstacle avoidance trajectory with the original teaching trajectory data to obtain a trajectory deviation value, and extracts reverse position adjustment data with frequencies exceeding a preset threshold from the original teaching trajectory data as compensation operation commands. It also determines the human-machine control divergence entropy based on the probability distribution of the trajectory deviation value and the compensation operation commands. The human-machine control divergence entropy and the current dynamic state data are input into an oscillation prediction model, which outputs the induced oscillation probability. The control center, which is communicatively connected to the state observation module, is used to dynamically allocate teaching compliance weights and autonomous obstacle avoidance weights based on the human-machine control divergence entropy and the induced oscillation probability, so as to generate a target smooth trajectory and convert the target smooth trajectory into control commands and send them to the robot; wherein, the control commands include at least trajectory tracking commands, impedance adjustment commands, and velocity decay commands; The state observation module uses the cross-entropy algorithm to solve for information entropy. The system takes the ideal stationary distribution where the operator does not intervene at all, i.e., the operation probability is kept at 1 and the probability of other compensation operations is 0, as the reference distribution. The cross-entropy between the probability distribution of the currently collected compensation operation commands and the ideal stationary distribution is calculated. The trajectory deviation value is normalized and used as a modulation factor. Multiplying it with the above cross-entropy yields the human-machine control divergence entropy.

2. The robot path generation system based on teaching and path optimization according to claim 1, characterized in that: The control center is preset with a safety divergence threshold; when the human-machine control divergence entropy is lower than the safety divergence threshold, the teaching compliance weight is increased and the autonomous obstacle avoidance weight is decreased. And when the human-machine control divergence entropy is higher than or equal to the safety divergence threshold, the teaching compliance weight is reduced and the autonomous obstacle avoidance weight is increased.

3. The robot path generation system based on teaching and path optimization according to claim 2, characterized in that, The control center is specifically used for: When the human-machine control divergence entropy is higher than or equal to the safety divergence threshold, the original teaching trajectory data and the local obstacle avoidance trajectory are weighted, fused and filtered according to the adjusted weights to generate a first smooth trajectory as the target smooth trajectory, and the speed decay command is generated accordingly.

4. The robot path generation system based on teaching and path optimization according to claim 2, characterized in that, The control center is specifically used for: When the human-machine control divergence entropy is lower than the safety divergence threshold, the original teaching trajectory data and the local obstacle avoidance trajectory are weighted and fused according to the adjusted weights to generate a second smooth trajectory as the target smooth trajectory, and the trajectory tracking command is generated accordingly.

5. The robot path generation system based on teaching and path optimization according to claim 1, characterized in that, The control center has an adaptive impedance control algorithm, which is used to dynamically adjust the target stiffness and target damping values ​​of the robot end effector based on the induced oscillation probability and the current dynamic state data, and generate the impedance adjustment command.

6. The robot path generation system based on teaching and path optimization according to claim 1, characterized in that, The current dynamic state data includes joint torque data and end effector acceleration data; the control center is used to generate and send a stop lock command to the robot when the change in the end effector acceleration data or the joint torque data exceeds the corresponding preset threshold; otherwise, the sending of the control command is maintained.

7. The robot path generation system based on teaching and path optimization according to any one of claims 1 to 6, characterized in that, The operational environment disturbance data includes thermodynamic airflow change data and dynamic obstacle displacement data; the teaching acquisition module is also used to acquire communication status data between the remote teaching device and the system, the communication status data including network latency parameters and signal jitter parameters; the status observation module is also used to correct the human-machine control divergence entropy based on the communication status data.

Citation Information

Patent Citations

  • Obstacle avoidance control method based on dynamic obstacle trajectory prediction

    CN121596879A

  • Robot dynamic obstacle avoidance trajectory optimization method and system based on improved DMP

    CN122071117A