Robot cable arrangement risk prevention and control method, device and equipment
By describing the position and risk event probability of the robot's end effector as Markov decision-making process, and combining reinforcement learning and YOLOv8n network, the problem of robots being difficult to predict and control motion trajectory when processing cables is solved, and risk prevention capabilities and collation efficiency are improved.
Patent Information
- Application Number
- CN202411928468.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively predict and control the motion trajectory of a robot when processing cables, especially the dynamic characteristics of the cables and the high-dimensional state space.
By describing the position and risk event probability of the robot end effector as Markov decision-making process, strengthening learning is used for training optimization, SAC risk prevention and control strategies corresponding to different risk situations are obtained, and a risk prediction model is constructed based on the YOLOv8n network to monitor and predict cable status in real time.
It improves the risk prevention capabilities of the robot, achieves efficient and safe cable finishing, and can adapt to different working environments and task requirements.
Smart Images

Figure CN119942041A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot technology, and in particular to a robot cable arrangement risk prevention and control method, device and equipment. Background Art
[0002] Cable arrangement and assembly are common production processes in industrial production, usually completed by workers manually. In order to improve assembly efficiency and reduce labor costs, robots are usually used for cable arrangement and assembly. Cables are deformable linear objects that can significantly change their shape or configuration when subjected to external forces. This dynamic characteristic makes it difficult for robots to predict and control their motion trajectory when handling cables. Moreover, precise control and planning are required during the robot's cable handling process. However, due to the dynamic characteristics and high-dimensional state space of cables, traditional control algorithms and planning methods are often difficult to apply.
[0003] Therefore, the existing technology still needs to be improved and developed. Summary of the invention
[0004] The present invention provides a robot cable arrangement risk prevention control method, device and equipment, which are used to improve the risk prevention capability of the robot.
[0005] The first aspect of the present invention provides a robot cable sorting risk prevention and control method, which includes: describing the robot's motion control as a Markov decision process based on the robot's end effector posture and the probability of risk event occurrence; using reinforcement learning to perform training optimization based on the Markov decision process to obtain SAC risk prevention and control strategies corresponding to different risk situations; building a risk prediction model based on the YOLOv8n network; collecting real-time images of cables, processing the real-time images of cables and inputting them into the risk prediction model to obtain the risk type and risk event probability output by the risk prediction model; determining the risk situation based on the risk type and the probability of risk event occurrence, and calling the corresponding SAC risk prevention and control strategy to control the robot's end effector posture based on the risk situation.
[0006] Preferably, according to the Markov decision process, reinforcement learning is used for training optimization to obtain SAC risk prevention and control strategies corresponding to different risk situations, including: according to the Markov decision process, reinforcement learning is used for training optimization to obtain the SAC risk-based strategy; the robots are initialized one by one to preset risk situations; for each preset risk situation, the robot is trained for risk prevention and control based on the SAC risk-based strategy to obtain SAC risk prevention and control strategies corresponding to different risk situations.
[0007] Preferably, a risk prediction model is constructed based on the YOLOv8n network, including: collecting cable images in different states as sample data, performing data enhancement on the sample data to obtain enhanced data; dividing the enhanced data into a training set and a verification set; constructing a YOLOv8n network structure, and using the training set and the verification set to train and verify the YOLOv8n network structure to obtain a risk prediction model.
[0008] Preferably, real-time images of cables are collected, processed and input into a risk prediction model to obtain risk types and probability of occurrence of risk events output by the risk prediction model, including: collecting real-time images of cables and preprocessing the real-time images of cables; normalizing the preprocessed real-time images of cables to obtain normalized images; and inputting the normalized images into the risk prediction model to obtain risk types and probability of occurrence of risk events output by the risk prediction model.
[0009] Preferably, the normalized image is a three-channel image with a size of 430×430.
[0010] Preferably, the risk situation is determined according to the risk type and the probability of occurrence of the risk event, and the corresponding SAC risk prevention and control strategy is called according to the risk situation to control the position and posture of the robot end effector, specifically including: obtaining a risk threshold corresponding to the risk type, and when the probability of occurrence of the risk event is greater than the risk threshold, the risk situation is determined to be a corresponding initial risk type; if the risk situation includes multiple initial risk types, the risk type with the highest priority is determined as the final risk type according to the priority levels of different risk types; according to the final risk type, the SAC risk prevention and control strategy corresponding to the final risk type is called to control the position and posture of the robot end effector.
[0011] The second aspect of the present invention provides a robot cable sorting risk prevention and control device, including: a design module, which is used to describe the robot's motion control as a Markov decision process according to the robot's end effector posture and the probability of risk event occurrence; a training module, which is used to use reinforcement learning to perform training optimization according to the Markov decision process, and obtain SAC risk prevention and control strategies corresponding to different risk situations; a construction module, which is used to build a risk prediction model based on the YOLOv8n network; a prediction module, which is used to collect real-time images of cables, process the real-time images of cables and input them into the risk prediction model to obtain the risk type and the probability of risk event occurrence output by the risk prediction model; a risk control module, which is used to determine the risk situation according to the risk type and the probability of risk event occurrence, and call the corresponding SAC risk prevention and control strategy to control the robot's end effector posture according to the risk situation.
[0012] The third aspect of the present invention provides a robot cable sorting risk prevention and control device, comprising: a memory and at least one processor, the memory storing computer-readable instructions, the memory and the at least one processor being interconnected via lines; the at least one processor calling the computer-readable instructions in the memory so that the robot cable sorting risk prevention and control device executes each step of the robot cable sorting risk prevention and control method as described above.
[0013] In the technical solution provided by the present invention, the risk prediction model constructed using the YOLOv8n network has a faster processing speed and higher recognition accuracy, can realize real-time monitoring and risk prediction of cable status, and improve the efficiency and safety of cable organization; moreover, according to the Markov decision process, reinforcement learning is used for training optimization to obtain SAC risk prevention and control strategies corresponding to different risk situations, and in the process of the robot organizing cables, the corresponding SAC risk prevention and control strategy is called according to the risk situation to control the position of the robot's end effector, thereby improving the robot's risk prevention ability. In addition, when encountering a new risk type, only targeted fine-tuning needs to be made to the new risk situation, so that the robot can adapt to different working environments and task requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A first flow chart of a robot cable arrangement risk prevention and control method provided by an embodiment of the present invention;
[0015] Figure 2 A schematic diagram of the structure of a robot cable arrangement risk prevention and control device provided in an embodiment of the present invention;
[0016] Figure 3 A schematic structural diagram of a robot cable management risk prevention and control device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0018] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 A first embodiment of a robot cable arrangement risk prevention and control method in an embodiment of the present invention includes:
[0019] S101. Based on the position and posture of the robot end effector and the probability of occurrence of risk events, the robot's motion control is described as a Markov decision process;
[0020] S102. Based on the Markov decision process, reinforcement learning is used for training optimization to obtain SAC risk prevention and control strategies corresponding to different risk situations;
[0021] S103. Build a risk prediction model based on YOLOv8n network;
[0022] S104, collecting real-time images of cables, processing the real-time images of cables and inputting them into the risk prediction model, and obtaining the risk type and probability of occurrence of risk events output by the risk prediction model;
[0023] S105. Determine the risk situation according to the risk type and the probability of occurrence of the risk event, and call the corresponding SAC risk prevention control strategy to control the position and posture of the robot end effector according to the risk situation.
[0024] It is understandable that the execution subject of the present invention may be a robot cable arrangement risk prevention and control device, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.
[0025] In this embodiment, in step S101, the purpose of this step is to transform the robot's motion control problem into a mathematical problem, namely, a Markov decision process. The Markov decision process is a mathematical model used to describe decision-making in an uncertain environment, which takes into account the current state, possible actions, and the results after the action.
[0026] In this embodiment, the pose (i.e., position and posture) of the robot end effector and the probability of risk event occurrence are key elements of the Markov decision process. The pose describes the specific position and direction of the robot end effector in space, while the probability of risk event occurrence reflects the possibility of encountering risk situations such as cable slippage, bending, or entanglement when the robot performs the cable sorting task.
[0027] In this embodiment, in step S102, according to the Markov decision process, reinforcement learning is used for training optimization to obtain SAC risk prevention and control strategies corresponding to different risk situations, specifically including: according to the Markov decision process, reinforcement learning is used for training optimization to obtain the SAC risk-based strategy; the robots are initialized one by one to preset risk situations; for each preset risk situation, the robot is trained for risk prevention and control based on the SAC risk-based strategy to obtain SAC risk prevention and control strategies corresponding to different risk situations.
[0028] In this embodiment, according to the Markov decision process, reinforcement learning is used for training optimization to obtain the SAC risk-based strategy, that is, the reinforcement learning algorithm Soft Actor Critic (SAC) is used to obtain the optimal solution. In reinforcement learning, the robot agent learns how to perform tasks by interacting with the environment. The robot agent observes the current state (state) in each time step, selects an action (action), and then the environment gives a feedback, that is, a reward (reward), and the next state. This process continues until a terminal state is reached, marking the end of a learning cycle (episode). The goal of reinforcement learning is to learn an optimal strategy, that is, a strategy that can bring the maximum expected cumulative reward among all possible strategies, that is, a strategy that can most effectively avoid risks.
[0029] In this embodiment, in a three-dimensional Cartesian coordinate system, the position of any point P can be represented by an ordered triple (x, y, z), where x, y, and z are the projection distances (also called coordinate values) of point P on the x-axis, y-axis, and z-axis, respectively.
[0030] Define the state value S at time t t is a continuous domain state variable, then S i,t =[P t 1 ,P t 2 ,P i ], where P t 1 ,P t 2 are the three-dimensional Cartesian coordinates of the top of the cable and the midpoint between the top of the cable and the robot gripping point, respectively, and p i is the probability of a risk event occurring.
[0031] The robot tool coordinates (TCP coordinates) represent the position and posture of the tool center point in three-dimensional space. In TCP coordinates, x, y, and z represent the positions on the three coordinate axes; rx, ry, and rz represent the robot posture, that is, the angles of rotation around the three coordinate axes x, y, and z in the original coordinate system. Define action variables is the change in TCP coordinates, that is That is, the action variable is the change in the robot end effector position at time t compared to time t-1. In an ideal state, the reinforcement learning algorithm SAC executes action variables Change the robot end effector position and make the state S i,t P t 1 ,P t 2 The coordinates of P tend to the target coordinates and make P i tends to 0 in order to reduce the risk.
[0032] In this embodiment, both the critic and the actor in the SAC algorithm use a three-layer MLP network (256×256×123).
[0033] During training, the SAC algorithm uses stochastic gradient descent (SGD) or its variants (such as the Adam optimizer) to update the parameters of the Critic and Actor networks. SGD is an iterative optimization algorithm that minimizes the loss function by calculating the gradient in each iteration and updating the parameters in the opposite direction of the gradient.
[0034] For the Critic network, the loss function is usually calculated based on the error between the estimated value and the true value of the Q function. This error can be estimated by Temporal Difference Learning (TDL) methods such as TD error.
[0035] For Actor networks, the loss function is usually calculated based on the policy gradient theorem, which involves the gradient of the Q function and the policy distribution. By maximizing this loss function (actually maximizing the weighted sum of expected return and entropy), the Actor network can learn the optimal policy.
[0036] In this embodiment, the risk prevention control reward function r i as follows:
[0037]
[0038] During training, the failure prevention skill strategy is optimized to maximize reward while minimizing risk. To improve the sample efficiency of the preventive control skill, the robot needs to experience risky states more frequently during training.
[0039] In this embodiment, the preset risk situations include bending risk, sliding risk and entanglement risk. The initialization procedure for the bending risk is that the robot moves the cable in a random direction toward the hole wall until the bending risk is triggered; the initialization procedure for the sliding risk is that the robot moves the cable in a plane perpendicular to the axis of the hole when it contacts the hole wall until the sliding risk is triggered; the initialization procedure for the entanglement risk is that the robot rotates the cable to a random angle until the entanglement risk is triggered.
[0040] It should be noted that when a risk situation other than the above three risks is observed, the robot is initialized to the newly discovered risk situation, and the robot is trained for risk prevention and control based on the SAC risk-based strategy to obtain a SAC risk prevention and control strategy corresponding to the newly discovered risk situation.
[0041] In this embodiment, in step S103, a risk prediction model is constructed based on the YOLOv8n network, specifically including: collecting cable images in different states as sample data, performing data enhancement on the sample data to obtain enhanced data; dividing the enhanced data into a training set and a verification set; constructing a YOLOv8n network structure, and using the training set and the verification set to train and verify the YOLOv8n network structure to obtain a risk prediction model.
[0042] Specifically, at least 5 cables of different materials or thicknesses need to be prepared, and the robot is manually controlled to clamp and perform perforation experiments; for each state of each cable (normal, bent, offset, and entangled), 30 image data are collected according to the following steps: 1) the robot end and the cable state are manually changed to make the cable in one of the normal, bent, offset, and entangled conditions; 2) the RGB camera at the robot end is started to take pictures of the cable. Therefore, at least 5x4x30=600 images need to be collected in total, and each image is labeled with a label corresponding to one of the four conditions of normal, bent, offset, and entangled. After the collection is completed, the data set is divided into a training set and a validation set in a ratio of 9:1, that is, 540 images are used as a training set and 60 images are used as a validation set. The online data enhancement option is enabled during YOLOv8n training, and four data enhancement methods, mosaic enhancement, mixed enhancement, random perturbation, and color perturbation, are used to enhance the original images.
[0043] In this embodiment, in step S104, a real-time image of the cable is collected, and the real-time image of the cable is processed and input into the risk prediction model to obtain the risk type and the probability of occurrence of the risk event output by the risk prediction model, which specifically includes: collecting the real-time image of the cable and preprocessing the real-time image of the cable; normalizing the preprocessed real-time image of the cable to obtain a normalized image; inputting the normalized image into the risk prediction model to obtain the risk type and the probability of occurrence of the risk event output by the risk prediction model.
[0044] Optionally, the collected images are preprocessed, such as denoising, contrast enhancement, etc., to improve image quality.
[0045] Optionally, the preprocessed real-time cable image is normalized into a three-channel image with a size of 430×430.
[0046] In this embodiment, the risk prediction model constructed using the YOLOv8n network has a faster processing speed and higher recognition accuracy, and can achieve real-time monitoring and risk prediction of cable status.
[0047] In this embodiment, in step S105, the risk situation is determined according to the risk type and the probability of occurrence of the risk event, and the corresponding SAC risk prevention and control strategy is called according to the risk situation to control the position and posture of the robot end effector, specifically including: obtaining a risk threshold corresponding to the risk type, when the probability of occurrence of the risk event is greater than the risk threshold, the risk situation is determined to be a corresponding initial risk type; if the risk situation includes multiple initial risk types, the risk type with the highest priority is determined as the final risk type according to the priority levels of different risk types; according to the final risk type, the SAC risk prevention and control strategy corresponding to the final risk type is called to control the position and posture of the robot end effector.
[0048] In this implementation, a two-state finite state machine is used to represent a normal and risky situation. For example, when the risk type and the probability of occurrence of a risk event output by the risk prediction model are: the risk type is entanglement risk, and the probability of occurrence of the risk event is p1, the risk threshold p1 of the entanglement risk is obtained. When p1>p1, it indicates that there is an entanglement risk. According to the entanglement risk, the SAC risk prevention control strategy corresponding to the entanglement risk is called to control the robot end effector posture. When p1≤p1, it indicates that the cable state is normal.
[0049] In this implementation, when multiple risks exist simultaneously, the priority of winding risk is the highest, the priority of deviation risk is the second highest, and the priority of bending risk is the lowest.
[0050] Exemplarily, when the risk type and the probability of occurrence of the risk event output by the risk prediction model are: the risk type is entanglement risk and bending risk, the probability of occurrence of the risk event of entanglement risk is p1 and the probability of occurrence of the risk event of bending risk is p3, then the risk threshold p1 of the entanglement risk and the risk threshold p3 of the bending risk are obtained. When p1>p1, p3>p3, it means that there is entanglement risk and bending risk. At this time, the priority of entanglement risk and bending risk needs to be further determined. The priority of entanglement risk is greater than the priority of bending risk. Then the risk type with the highest priority is determined to be entanglement risk. According to the entanglement risk, the SAC risk prevention control strategy corresponding to the entanglement risk is called to control the robot end effector posture. When p1>p1, p3≤p3, it means that there is entanglement risk. According to the entanglement risk, the SAC risk prevention control strategy corresponding to the entanglement risk is called to control the robot end effector posture. When p1≤p)1, p3>p)3, it indicates that there is a bending risk. According to the bending risk, the SAC risk prevention control strategy corresponding to the bending risk is called to control the robot end effector posture. When p1≤p)1, p3≤p)3, it indicates that the cable status is normal.
[0051] In this scheme, the risk prediction model constructed using the YOLOv8n network has a faster processing speed and higher recognition accuracy, and can realize real-time monitoring of cable status and risk prediction, thereby improving the efficiency and safety of cable organization. Moreover, according to the Markov decision process, reinforcement learning is used for training optimization to obtain SAC risk prevention and control strategies corresponding to different risk situations. In the process of robot cable organization, the corresponding SAC risk prevention and control strategy is called according to the risk situation to control the position of the robot's end effector, thereby improving the robot's risk prevention ability. In addition, when encountering a new type of risk, only targeted fine-tuning needs to be made to the new risk situation, so that the robot can adapt to different working environments and task requirements.
[0052] The above describes the robot cable arrangement risk prevention and control method in the embodiment of the present invention. The following describes the device in the embodiment of the present invention. Figure 2 , the implementation of the robot cable arrangement risk prevention and control device in the embodiment of the present invention includes:
[0053] A design module 201 is used to describe the motion control of the robot as a Markov decision process according to the position and posture of the robot end effector and the probability of occurrence of risk events;
[0054] The training module 202 is used to perform training optimization using reinforcement learning according to the Markov decision process to obtain SAC risk prevention and control strategies corresponding to different risk situations;
[0055] A construction module 203 is used to construct a risk prediction model based on the YOLOv8n network;
[0056] The prediction module 204 is used to collect real-time images of cables, process the real-time images of cables and input them into the risk prediction model to obtain the risk type and the probability of occurrence of risk events output by the risk prediction model;
[0057] The risk control module 205 is used to determine the risk situation according to the risk type and the probability of occurrence of the risk event, and call the corresponding SAC risk prevention control strategy to control the position and posture of the robot end effector according to the risk situation.
[0058] In this embodiment, the risk prediction model constructed using the YOLOv8n network has a faster processing speed and higher recognition accuracy, and can realize real-time monitoring and risk prediction of cable status, thereby improving the efficiency and safety of cable organization. Moreover, according to the Markov decision process, reinforcement learning is used for training optimization to obtain SAC risk prevention and control strategies corresponding to different risk situations. In the process of the robot organizing cables, the corresponding SAC risk prevention and control strategy is called according to the risk situation to control the position of the robot's end effector, thereby improving the robot's risk prevention ability. In addition, when encountering a new type of risk, only targeted fine-tuning needs to be made to the new risk situation, so that the robot can adapt to different working environments and task requirements.
[0059] Figure 2 The structure of the robot cable organization risk prevention and control device shown does not constitute a limitation of the robot cable organization risk prevention and control device, and can implement the steps of the robot cable organization risk prevention and control method provided in the above-mentioned method embodiments.
[0060] above Figure 2 The robot cable tidying risk prevention control device in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The robot cable tidying risk prevention control device in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0061] Figure 3It is a schematic diagram of the structure of a robot cable sorting risk prevention and control device provided by an embodiment of the present invention. The device 300 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 (for example, one or more mass storage devices) storing application programs 333 or data 332. Among them, the memory 320 and the storage medium 330 can be short-term storage or permanent storage. The program stored in the storage medium 330 may include one or more modules (not shown), and each module may include a series of instruction operations on the device 300. Furthermore, the processor 310 can be configured to communicate with the storage medium 330 to execute a series of instruction operations in the storage medium on the device 300.
[0062] The device 300 may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows, Mac OS X, Unix, Linux, FreeBSD, etc.
[0063] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A robot cable arrangement risk prevention and control method, characterized in that: The robot cable arrangement risk prevention and control method comprises: According to the end-effector position and the probability of risk events, the robot's motion control is described as a Markov decision process. According to the Markov decision process, reinforcement learning is used for training optimization to obtain SAC risk prevention and control strategies corresponding to different risk situations; Build a risk prediction model based on the YOLOv8n network; Collecting real-time images of cables, processing the real-time images of cables and inputting them into the risk prediction model, and obtaining the risk type and probability of occurrence of risk events output by the risk prediction model; The risk situation is determined according to the risk type and the probability of occurrence of the risk event, and the corresponding SAC risk prevention control strategy is called according to the risk situation to control the position and posture of the robot end effector.
2. The robot cable arrangement risk prevention and control method according to claim 1, characterized in that: According to the Markov decision process, reinforcement learning is used for training optimization to obtain SAC risk prevention and control strategies corresponding to different risk situations, including: According to the Markov decision process, reinforcement learning is used for training optimization to obtain the SAC risk-based strategy; Initialize the robots one by one to the preset risk situations; For each preset risk situation, the robot is trained for risk prevention and control based on the SAC risk-based strategy to obtain the SAC risk prevention and control strategy corresponding to different risk situations.
3. The robot cable arrangement risk prevention and control method according to claim 1, characterized in that: Build a risk prediction model based on the YOLOv8n network, including: Collect cable images in different states as sample data, perform data enhancement on the sample data, and obtain enhanced data; Divide the augmented data into training and validation sets; Construct the YOLOv8n network structure, and use the training set and validation set to train and validate the YOLOv8n network structure to obtain the risk prediction model.
4. The robot cable arrangement risk prevention and control method according to claim 1, characterized in that: Collect real-time images of cables, process them and input them into the risk prediction model to obtain the risk type and probability of occurrence of risk events output by the risk prediction model, including: Collect real-time images of cables and pre-process them; Normalizing the preprocessed real-time cable image to obtain a normalized image; The normalized image is input into the risk prediction model to obtain the risk type and probability of occurrence of risk events output by the risk prediction model.
5. The robot cable arrangement risk prevention and control method according to claim 4, characterized in that: The normalized image is a three-channel image with a size of 430×430.
6. The robot cable arrangement risk prevention and control method according to claim 1, characterized in that: The risk situation is determined according to the risk type and the probability of occurrence of risk events, and the corresponding SAC risk prevention control strategy is called according to the risk situation to control the position and posture of the robot end effector, including: Obtain a risk threshold corresponding to the risk type. When the probability of a risk event occurring is greater than the risk threshold, the risk situation is determined to be the corresponding initial risk type. If the risk situation includes multiple initial risk types, the risk type with the highest priority level is determined as the final risk type based on the priority levels of the different risk types; According to the final risk type, the SAC risk prevention control strategy corresponding to the final risk type is called to control the position and posture of the robot end effector.
7. A robot cable arrangement risk prevention and control device, characterized in that: include: Design a module to describe the robot's motion control as a Markov decision process based on the robot's end-effector pose and the probability of risk events; The training module is used to perform training optimization using reinforcement learning based on the Markov decision process to obtain SAC risk prevention and control strategies corresponding to different risk situations; Building a module for building a risk prediction model based on the YOLOv8n network; A prediction module is used to collect real-time images of cables, process the real-time images of cables and input them into the risk prediction model to obtain the risk type and probability of occurrence of risk events output by the risk prediction model; The risk control module is used to determine the risk situation according to the risk type and the probability of occurrence of the risk event, and call the corresponding SAC risk prevention control strategy to control the position and posture of the robot end effector according to the risk situation.
8. A robot cable arrangement risk prevention and control device, characterized in that: comprising a memory and at least one processor, wherein the memory has computer-readable instructions stored therein; The at least one processor calls the computer-readable instructions in the memory to execute the various steps of the robot cable management risk prevention and control method as described in any one of claims 1-6.