Mechanical arm visual servo control method and device, electronic equipment and medium
By improving the YOLO algorithm for target detection and task decomposition, and combining it with parallel learning, the problem of low efficiency in robotic arm servo systems was solved, achieving efficient servo control and target recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA MOBILE (XIONGAN) ICT CO LTD
- Filing Date
- 2022-03-25
- Publication Date
- 2026-04-21
Smart Images

Figure CN116863262B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of automation control technology, specifically to a visual servo control method, device, electronic device, and medium for a robotic arm. Background Technology
[0002] Currently, the most typical application of multi-joint robots is the robotic arm. However, the robotic arm itself does not have environmental perception capabilities. Therefore, introducing vision sensors into the robotic arm system to form a visual servoing (VS) system not only increases the robotic arm's operating range but also enhances its environmental perception capabilities, thereby meeting the various complex work requirements encountered by the robotic arm in industrial production. However, current methods for controlling robotic arms based on servo systems are inefficient. Summary of the Invention
[0003] This application provides a method, device, electronic device, and medium for visual servo control of a robotic arm, in order to solve the technical problem of low efficiency when controlling a robotic arm to work based on a robotic arm servo system.
[0004] In a first aspect, embodiments of this application provide a visual servo control method for a robotic arm, comprising:
[0005] The initial servo task of the robotic arm is combined with the improved YOLO algorithm to perform target detection and obtain target object information;
[0006] Based on the target object information, the initial servo task is decomposed, and servo control training is performed based on each sub-task obtained from the decomposition to obtain a target servo control strategy that enables the initial servo task to achieve the optimal reward value.
[0007] Based on the target servo control strategy, the robotic arm is controlled to execute a servo task whose target object is the same as the target object in the target object information of the initial servo task.
[0008] In one embodiment, the step of training servo control based on the decomposed subtasks to obtain a target servo control strategy that enables the initial servo task to achieve the optimal reward value includes:
[0009] Servo control training is performed on each of the decomposed subtasks until the servo control strategy that enables the initial servo task to reach the optimal reward value under the preset behavior fusion scenario and its weights is obtained and determined as the target servo control strategy.
[0010] In one embodiment, the loss function of the improved YOLO algorithm includes confidence level and relative error rate.
[0011] In one embodiment, before the step of training servo control based on the sub-tasks obtained from the decomposition, the method further includes:
[0012] The data of each subtask obtained from the decomposition is transformed according to the preset interaction matrix to obtain the transformed subtasks.
[0013] In one embodiment, the subtask includes moving towards the target, avoiding losing target feature points, and maintaining a certain duration after reaching the target.
[0014] In one embodiment, after the step of performing target detection using the improved YOLO algorithm based on the initial servo task of the robotic arm to obtain target object information, the method further includes:
[0015] Determine whether an initial servo control strategy corresponding to the target object information exists;
[0016] If the initial servo control strategy does not exist, then the step of decomposing the initial servo task based on the target object information is executed.
[0017] In one embodiment, after determining whether an initial servo control strategy corresponding to the target object information exists, the method further includes:
[0018] If the initial servo control strategy exists, the robotic arm is controlled to perform the initial servo task based on the initial servo control strategy.
[0019] Secondly, embodiments of this application provide a robotic arm vision servo control device, comprising:
[0020] The detection module is used to perform target detection based on the initial servo task of the robotic arm and the improved YOLO algorithm to obtain target object information;
[0021] The training module is used to decompose the initial servo task based on the target object information, and to train the servo control based on each sub-task obtained from the decomposition, so as to obtain a target servo control strategy that enables the initial servo task to achieve the optimal reward value.
[0022] The control module is used to control the robotic arm to perform a servo task to be executed based on the target servo control strategy, where the target object is the same as the target object in the target object information of the initial servo task.
[0023] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the program to implement the steps of the robotic arm visual servo control method described in the first or second aspect.
[0024] Fourthly, embodiments of this application provide a medium, which is a computer-readable storage medium including a computer program. When the computer program is executed by a processor, it implements the steps of the robotic arm visual servo control method described in the first or second aspect.
[0025] The robotic arm vision servo control method, device, electronic device, and medium provided in this application embodiment use an improved YOLO algorithm for target detection, effectively improving the training effect of target detection; and decompose the initial servo task into multiple sub-tasks based on the target object information obtained from target detection, avoiding the high computational complexity caused by directly solving the source task, and expanding the learning experience of the robotic arm through parallel learning, accelerating the convergence speed of servo control, and quickly obtaining the target servo control strategy that makes the initial servo task reach the optimal reward value, and controlling the robotic arm to execute the servo task to be executed based on the target object in the target object information of the initial servo task according to the target servo control strategy, which can improve the efficiency of controlling the robotic arm to work based on the robotic arm servo system. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is one of the flowcharts of the robotic arm vision servo control method provided in the embodiments of this application;
[0028] Figure 2 This is the second flowchart of the robotic arm vision servo control method provided in the embodiments of this application;
[0029] Figure 3 This is the third flowchart of the robotic arm vision servo control method provided in the embodiments of this application;
[0030] Figure 4 This is a schematic diagram of the functional modules of an embodiment of the robotic arm vision servo control device of this application;
[0031] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0033] Figure 1 This is one of the flowcharts illustrating the robotic arm vision servo control method provided in an embodiment of this application. (Refer to...) Figure 1 This application provides a visual servo control method for a robotic arm, which may include:
[0034] Step S100: Based on the initial servo task of the robotic arm, target detection is performed using the improved YOLO algorithm to obtain target object information;
[0035] In this embodiment, the robotic arm visual servo control method can be applied to electronic devices such as smartphones, tablets, and PCs. The electronic device and the robotic arm containing a camera form a robotic arm servo system. The electronic device can obtain information from the robotic arm and perform motion control on the robotic arm based on the communication connection with the robotic arm. This communication connection can be wired or wireless, and in this embodiment, the robotic arm is specifically a six-axis robotic arm. Robotic arms are generally used in enterprise production activities to replace manual labor in performing some complex and heavy handling tasks, and even to replace manual labor in performing specific tasks on the production line. In this embodiment, a six-axis robotic arm motion control structure is established, in which a camera can be installed to realize the visual servo control of the robotic arm.
[0036] When a user receives a task that the robotic arm needs to perform (e.g., grasping a cuboid object within a specified area), this task is designated as the robotic arm's initial servo task. Understandably, the initial servo task includes a specified object; for example, if the task is to grasp a cuboid object within a specified area, the cuboid object is the target object. Therefore, when controlling the robotic arm to execute the initial servo task, the improved YOLO algorithm, based on images captured by the camera, performs target detection within the specified area to determine information such as the number, position, and shape of the target objects, as well as the target object information itself. YOLO stands for You Only Look Once: Unified, Real-Time Object Detection; "You Only Look Once" refers to requiring only one CNN operation, "Unified" indicates a unified framework providing end-to-end prediction, and "Real-Time" reflects the speed of the YOLO algorithm. In this embodiment, the original YOLO algorithm can specifically be YOLOv2.
[0037] It should be noted that when performing object detection based on the improved YOLO algorithm, the input image is divided into S*S grids, and only objects falling within these grids can be detected. Each grid contains B bounding boxes, typically B is set to 2 by default. The detection result for each bounding box includes five parameters: the center coordinates (x, y) of the bounding box, its width and height (w, h), and a confidence score S. i This embodiment introduces a new loss function based on the traditional YOLO algorithm, resulting in an improved YOLO algorithm. The improved YOLO algorithm's loss function includes confidence score and relative error rate. Specifically, it performs loss calculations on the predicted categories, defining the confidence score as C. i The confidence score of the predicted bounding box is defined as follows: Then we have:
[0038]
[0039] Among them, S 2 B represents the number of grids into which the input image is divided; B represents the number of bounding boxes for each grid, typically B is 2 by default; C represents the number of bounding boxes for each grid. i For reliability score, The confidence score for the predicted bounding box, i = 0, 1, 2...S 2 B = 0, 1, 2; obj is the object, i.e., the target.
[0040] The above indicates that if a target appears in grid i, then the predicted value of the j-th bounding box is valid for that grid prediction. The value is 1; if no target appears in grid i, then A value of 0 indicates that the loss value for this part is not calculated.
[0041] Simultaneously, a loss calculation is performed on the prediction confidence. Specifically, if a target detection process includes K types of objects, then each grid cell predicts the i-th type of object C. i The conditional probability is Pr(C) i If |O), i = 1, 2, ..., K, then we have:
[0042]
[0043] Among them, S 2 B represents the number of grids into which the input image is divided; B represents the number of bounding boxes for each grid, typically B is 2 by default; C represents the number of bounding boxes for each grid. i For reliability score, Let i = 0, 1, 2...S2; B = 0, 1, 2. (p i (c) Predict the i-th class of object C for each grid. i The conditional probability, Predict the i-th class of object C for each grid. i The average of the conditional probabilities; λ noord The coefficients of the classification loss function are denoted by ; classes represent the categories.
[0044] In the above, if the target does not appear in a grid, The value is set to 1, otherwise it is 0.
[0045] Furthermore, by modifying the loss calculation for the width and height portions of the YOLO predicted bounding boxes and introducing a rate of change, objects of different sizes are no longer represented by their dimensions in the loss function, but rather by their relative error rate. This reduces training problems caused by varying target object sizes. Therefore:
[0046]
[0047] Among them, S 2 B represents the number of grids into which the input image is divided; B represents the number of bounding boxes for each grid, typically B is 2 by default; C represents the number of bounding boxes for each grid. i For reliability score, For the predicted bounding box confidence scores, i = 0, 1, 2...S2; B = 0, 1, 2; w i The width of the i-th grid is... h is the average width of the i grid cells. i Let i be the height of the i-th grid. Let be the mean of the heights of the i grid cells.
[0048] In summary, the loss function of the improved YOLO algorithm is shown in the following formula:
[0049]
[0050] Where, λ coord x is the coefficient of the coordinate loss function. i Let x be the x-coordinate of the i-th grid. Let y be the mean of the x-coordinates of the i grid cells. i Let be the ordinate of the i-th grid. The mean of the ordinates of the i grid cells is given. The definitions of the other parameters are the same as above and will not be repeated here.
[0051] After obtaining the aforementioned loss function, the improved YOLO algorithm is trained based on this loss function to achieve optimal detection accuracy. Furthermore, based on the improved and trained YOLO algorithm, target detection is performed on the target objects in the initial servo task of the robotic arm. The number, position, and shape of target objects within a specified area in the initial servo task are determined and included as target object information. By introducing relative error rate and confidence score for loss calculation, the convergence speed of the neural network can be effectively accelerated and fluctuations during the convergence process can be reduced. Even in simulation environments with extremely high Gaussian noise, a higher recognition accuracy can be achieved compared to the original YOLO algorithm, improving the recognition accuracy during target detection and effectively enhancing the training effect of target detection. This also facilitates subsequent task decomposition and servo control training based on target object information.
[0052] Step S200: Based on the target object information, the initial servo task is decomposed into tasks, and servo control training is performed based on each sub-task obtained from the decomposition to obtain a target servo control strategy that enables the initial servo task to achieve the optimal reward value.
[0053] After determining that there is no initial servo control strategy corresponding to the target object information, this embodiment performs task delimitation on the initial servo control based on the target object in the target object information. Specifically, the target object is taken as the target, and the initial servo task is decomposed into three sub-tasks: moving towards the target, avoiding losing target feature points, and maintaining a certain position after reaching the target for a certain period of time. These three sub-tasks allow the robotic arm to complete a series of actions. Servo control training is then performed on the decomposed sub-tasks, such as moving towards the target, avoiding losing target feature points, and maintaining a certain position after reaching the target, until a target servo control strategy is obtained that enables the initial servo task to achieve the optimal reward value under a preset behavior fusion scenario and its weights.
[0054] In this embodiment, the preset behavior fusion scenario includes the process of the robotic arm approaching the target object and the robotic arm reaching the target position. The weights of the robotic arm approaching the target object and the robotic arm reaching the target position can be set according to actual needs. It should be noted that the first sub-task has a higher weight when approaching the target object, while the third sub-task has a higher weight when reaching the target position. For example, in this embodiment, the weights of each sub-task corresponding to the robotic arm approaching the target object can be set to 0.8, 0.15, and 0.05, and the weights of each sub-task corresponding to the robotic arm reaching the target position can be set to 0.05, 0.15, and 0.8.
[0055] It should be further explained that, in this embodiment, when training servo control based on the decomposed subtasks, each subtask is first trained in the scenario of the robotic arm approaching the target object, until each subtask, under this scenario and its weights, enables the initial servo task to reach the optimal reward value. At this point, the parameters in the scenario of the robotic arm approaching the target object no longer change, and convergence is achieved. After convergence in the scenario of the robotic arm approaching the target object, each subtask is trained in the scenario of the robotic arm reaching the target position, until each subtask, under this scenario and its weights, enables the initial servo task to reach the optimal reward value. At this point, the parameters in the scenario of the robotic arm approaching the target object no longer change, and convergence is achieved. Then, the training is complete, and the servo control strategy that enables the initial servo task to reach the optimal reward value is determined as the target servo control strategy. By decomposing the initial servo task into multiple subtasks based on the target object information obtained from target detection, the high computational complexity caused by directly solving the source task is avoided. Furthermore, the learning experience of the robotic arm can be expanded through parallel learning, accelerating the convergence speed of servo control and quickly obtaining the target servo control strategy that makes the initial servo task reach the optimal reward value. This allows the robotic arm to be controlled to execute subsequent servo tasks with the same target object as the target object in the target object information of the initial servo task, thereby improving the efficiency of the robotic arm when controlling the robotic arm to work based on the robotic arm servo system.
[0056] It should be noted that, before the step of training servo control based on the sub-tasks obtained from the decomposition, the following steps are also included:
[0057] Step X: Perform data transformation on each subtask obtained from the decomposition according to the preset interaction matrix, and map the motion of the camera in each subtask to the motion of the robotic arm to obtain the transformed subtasks.
[0058] It should be noted that after decomposing the initial servo task based on the target object information, and before training the servo control based on each sub-task obtained from the decomposition, a speed controller needs to be designed for visual servoing to convert image errors into camera motion speed. In this embodiment, data conversion is also required for each sub-task obtained from the decomposition based on a preset L-interaction matrix, mapping the camera motion in each sub-task to the motion of the robotic arm, resulting in the converted sub-tasks. Specifically, the L-interaction matrix in this embodiment can also be called the image Jacobian matrix, and is shown in the following formula:
[0059] L β(t) =βL e* +(1-β)L e(t)
[0060] Among them, L e(t) L is the current interaction matrix; e* It is called the target interaction matrix, which is obtained from the target feature points and is a constant matrix; β∈[0,1]. When β equals 1, it is equivalent to using the target interaction matrix as the interaction matrix; when β equals 0, it is equivalent to using the current interaction matrix as the interaction matrix.
[0061] Step S300: Based on the target servo control strategy, control the robotic arm to execute the servo task to be executed, where the target object is the same as the target object in the target object information of the initial servo task.
[0062] After training the servo control based on the sub-tasks obtained from the decomposition and obtaining the target servo control strategy that enables the initial servo task to achieve the optimal reward value, if the target object in the servo task to be executed by the robotic arm is the same as the target object in the target object information of the initial servo task, the servo task to be executed can be executed according to the target servo control strategy since the servo control training based on the target object has been carried out and the target servo control strategy that enables the corresponding servo task to achieve the optimal reward value has been obtained.
[0063] For example, if the target object in the initial servo task is a cup, since servo control training has already been performed based on the cup, when the target object in the servo task to be executed is also a cup, the servo task to be executed can be directly executed according to the target servo control strategy learned during training. The robotic arm vision servo control method provided in this application uses an improved YOLO algorithm for target detection, effectively improving the training effect of target detection; and decomposes the initial servo task into multiple sub-tasks based on the target object information obtained from target detection, avoiding the high computational complexity caused by directly solving the source task. Moreover, it can expand the learning experience of the robotic arm through parallel learning, accelerate the convergence speed of servo control, and quickly obtain the target servo control strategy that makes the initial servo task reach the optimal reward value. Based on the target servo control strategy, the robotic arm is controlled to execute the servo task to be executed with the same target object as the target object in the target object information of the initial servo task, which can improve the efficiency of controlling the robotic arm to work based on the robotic arm servo system.
[0064] Figure 2 This is a second schematic flowchart illustrating the robotic arm vision servo control method provided in this application embodiment. (Refer to...) Figure 2 In one embodiment, after obtaining target object information by combining the initial servo task based on the robotic arm with the improved YOLO algorithm for target detection, the method further includes:
[0065] Step A: Determine whether there is an initial servo control strategy corresponding to the target object information;
[0066] Understandably, if servo control training has not been performed on an object identical to the target object in the target object information, then there is currently no servo control strategy that will enable the initial servo task to achieve the optimal reward value during execution. However, if servo control training has been performed on an object identical to the target object in the target object information, then there exists a servo control strategy that will enable the initial servo task to achieve the optimal reward value during execution, and this strategy is determined as the initial servo control strategy. The servo control strategy may include camera speed, angular velocity of each axis, and motion direction, etc.
[0067] Therefore, in this embodiment, the target object in the target object information is compared with the objects corresponding to each existing servo control strategy. By determining whether there is an object identical to the target object among the objects corresponding to each existing servo control strategy, it is determined whether there is an initial servo control strategy corresponding to the target object in the target object information. Specifically, if there is an object identical to the target object among the objects corresponding to each existing servo control strategy, it is determined that an initial servo control strategy corresponding to the target object information exists; if there is no object identical to the target object among the objects corresponding to each existing servo control strategy, it is determined that no initial servo control strategy corresponding to the target object information exists.
[0068] For example: if the target object in the target object information obtained by target detection is a cuboid, and servo control has been trained based on the cuboid before, and the servo control strategy that achieves the optimal reward value when performing a cuboid-based task (such as grasping a cuboid) is learned after training is completed, then it is determined that there is an initial servo control strategy corresponding to the target object information; however, if servo control has not been trained based on the cuboid before, then it is determined that there is no initial servo control strategy corresponding to the target object information.
[0069] Step B: If the initial servo control strategy does not exist, then perform the step of decomposing the initial servo task based on the target object information.
[0070] If, after comparison, it is determined that there is no initial servo control strategy corresponding to the target object information, it means that servo control training has not been performed based on an object that is the same as the target object in the target object information. Then, the step of decomposing the initial servo task based on the target object information is executed. For details, please refer to the description of step S200, which will not be repeated here.
[0071] Furthermore, after the step of determining whether an initial servo control strategy corresponding to the target object information exists, the method further includes:
[0072] Step C: If the initial servo control strategy exists, then control the robotic arm to execute the initial servo task based on the initial servo control strategy.
[0073] If, after comparison, it is determined that an initial servo control strategy corresponding to the target object information exists, it means that servo control training has already been performed based on an object identical to the target object in the target object information, and a servo control strategy that achieves the optimal reward value during the execution of the initial servo task has been obtained. This initial servo control strategy can then be used to control the robotic arm to execute the initial servo task. For example, the initial servo control strategy can be used to grasp a cuboid object in the initial servo task. This embodiment, by determining whether an initial servo control strategy corresponding to the target object information exists, can quickly control the robotic arm to execute the initial servo task if it already exists. If it does not exist, servo control training can be performed based on the target object information to obtain a servo control strategy that achieves the optimal reward value for the initial servo task, which can then be used as the servo control strategy for subsequent servo tasks with the same target object, thus improving the efficiency of controlling the robotic arm using the robotic arm servo system.
[0074] Figure 3 This is the third flowchart illustrating the robotic arm vision servo control method provided in this application embodiment. (Refer to...) Figure 3 In one embodiment, the step of training servo control based on the decomposed subtasks to obtain a target servo control strategy that enables the initial servo task to achieve the optimal reward value includes:
[0075] Step S201: Perform servo control training on each subtask obtained from the decomposition until the target servo control strategy is obtained when each subtask achieves the optimal reward value of the initial servo task under the preset behavior fusion scenario and its weights.
[0076] After decomposing the initial servo task into three subtasks, this embodiment performs servo control training for each subtask. Furthermore, the three subtasks share the same state space and action space during training. Specifically, after decomposing the initial servo task into three subtasks (specifically, {M0, M1, M2}), each subtask is initialized with: Q(i,s,a)←0, where s∈S, a∈A, and Q(i,s,a) represents subtask M. i The Q value obtained after performing action a in state s, where S is the state space, A is the action space, and i = 1, 2, 3. Then, for each initialized subtask, the following steps are performed:
[0077] a. Obtain the current image feature set and target image feature set for the current subtask based on the improved YOLO algorithm;
[0078] b. Calculate the feature error between the current image feature set and the target image feature set based on the loss function of the improved YOLO algorithm;
[0079] c. Obtain the current state based on the feature error and the feature point positions in the current image feature set;
[0080] d. Select the interaction matrix L according to the reinforcement learning strategy β In this step, the reinforcement learning strategy is used to make the robotic arm select the action with the largest Q value in the Q table, while ε-greedy is used as the setting method of the adjustment strategy so that the robotic arm has a 50% exploration probability in the early stage of training.
[0081] e. According to the inverse Jacobian matrix J of the six-axis robotic arm -1 (θ) and the interaction matrix L β Calculate the camera velocity and angular velocities of each axis in a six-axis robotic arm. It should be noted that the six-axis robotic arm has six rotatable joints, each corresponding to one degree of freedom. The first three joints and the last three joints control the changes in position and orientation of the free end of the robotic arm, respectively. According to the robot coordinate system transformation rules, J... n-1 The coordinate system of the joint is transformed to J. n The coordinate system of the joints needs to undergo translation and rotation. The transformation matrix is then used... express;
[0082] f. Based on the calculated camera speed and angular velocity of each axis, rotate each axis of the six-axis robotic arm;
[0083] g. Obtain the next image taken by the camera after rotating each axis, and obtain the new state s. t+1 The reward value for each subtask is determined based on its reward function. It should be noted that in this embodiment, the robotic arm vision servoing task is divided into three subtasks, and a separate reward function is established for each subtask, as follows:
[0084] Subtask 1: If the robotic arm reaches the desired position, it receives a maximum positive reward of 10. Otherwise, it receives a negative reward. This is because we want the robotic arm to choose the shortest path from the camera as much as possible, so we use the camera's negative displacement as the reward. The reward function for subtask 1 is shown in the following formula:
[0085]
[0086] Subtask 2: If the robotic arm loses the target during servoing, it receives a maximum negative reward of -10. Otherwise, it receives a positive reward, using the camera's displacement as the reward. The reward function for subtask 2 is shown in the following formula:
[0087]
[0088] Subtask 3: Rewards are given based on the distance between the robotic arm and the target position. The closer the robotic arm is to the target, the greater the reward value; conversely, if the robotic arm deviates from the target, the greater the deviation, the smaller the reward value. The reward function for subtask 3 is shown in the following formula:
[0089] r = -100D
[0090] Where distance D is the Euclidean distance between the current feature point and the target feature point.
[0091] h. Update Q(i,s,a) based on the reward values of the sub-tasks obtained above, as shown below:
[0092]
[0093] Where Q is the Q-learning process, s is the current state, s' is the new state, a is the current action, a' is the new action, α and γ are adjustment parameters, R is the current reward value, and R' is the new reward value.
[0094] Furthermore, based on the reward values obtained from learning experience during the process of the robotic arm approaching the target object and the scenario of the robotic arm reaching the target position, and the weights in each scenario, the servo control strategy that enables the initial servo task to achieve the optimal reward value is calculated and determined as the target servo control strategy. By decomposing the initial servo task into multiple subtasks based on the target object information obtained from target detection, the high computational complexity caused by directly solving the source task is avoided. Moreover, the learning experience of the robotic arm can be expanded through parallel learning, accelerating the convergence speed of servo control and quickly obtaining the target servo control strategy that enables the initial servo task to achieve the optimal reward value. This allows the subsequent control of the robotic arm based on the target servo control strategy to execute servo tasks whose target objects are the same as those in the target object information of the initial servo task, thereby improving the efficiency of controlling the robotic arm to work based on the robotic arm servo system.
[0095] In one embodiment, the concept of limit testing is used for the experiment, and the experimental steps are as follows:
[0096] ① Input the training set of the improved YOLO algorithm, presented as labeled images. The objects in the training set include four categories: arrows, squares, triangles, and circles. Add a large amount of Gaussian noise to the target image. This abundant Gaussian noise also helps to better simulate the real-world environment.
[0097] ② In this experiment, the network model input format is set to a feature map of fixed size of 448*448*3, and the output format is set to a one-dimensional vector of 7*7*30 containing the target classification result and the target location. After decoding this output vector in a unified way, the target detection result can be displayed in the original image.
[0098] ③ We performed target detection using YOLOv2 and YOLOv2 with an improved loss function, and analyzed the advantages and disadvantages of the two algorithms by recording the accuracy of feature point recognition.
[0099] Simulation experiments show that the FQ-IBVS method, which integrates three sub-tasks, converges in about 30 rounds with minimal fluctuations in the feature error trajectory. By integrating the learning experience of the sub-tasks, the FQ-IBVS method demonstrates excellent servo behavior, indicating that the organic combination of sub-tasks can improve the overall performance of the system.
[0100] Experiments have shown that even in noisy environments, the improved YOLO algorithm can still achieve higher recognition accuracy than the traditional algorithm.
[0101] Furthermore, this application also provides a vision servo control device for a robotic arm.
[0102] Reference Figure 4 , Figure 4 This is a schematic diagram of the functional modules of an embodiment of the robotic arm vision servo control device of this application.
[0103] The robotic arm vision servo control device includes:
[0104] The detection module 100 is used to perform target detection based on the initial servo task of the robotic arm and the improved YOLO algorithm to obtain target object information;
[0105] The training module 200 is used to decompose the initial servo task based on the target object information, and to train the servo control based on each sub-task obtained by the decomposition, so as to obtain a target servo control strategy that enables the initial servo task to achieve the optimal reward value.
[0106] The control module 300 is used to control the robotic arm to perform a servo task to be executed based on the target servo control strategy, where the target object is the same as the target object in the target object information of the initial servo task.
[0107] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 can call the computer program in the memory 830 to execute the steps of the robotic arm visual servo control method, such as including:
[0108] The initial servo task of the robotic arm is combined with the improved YOLO algorithm to perform target detection and obtain target object information;
[0109] Based on the target object information, the initial servo task is decomposed, and servo control training is performed based on each sub-task obtained from the decomposition to obtain a target servo control strategy that enables the initial servo task to achieve the optimal reward value.
[0110] Based on the target servo control strategy, the robotic arm is controlled to execute a servo task whose target object is the same as the target object in the target object information of the initial servo task.
[0111] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0112] On the other hand, embodiments of this application also provide a storage medium, which is a computer-readable storage medium storing a computer program. The computer program is used to cause a processor to execute the steps of the methods provided in the above embodiments, including, for example:
[0113] The initial servo task of the robotic arm is combined with the improved YOLO algorithm to perform target detection and obtain target object information;
[0114] Based on the target object information, the initial servo task is decomposed, and servo control training is performed based on each sub-task obtained from the decomposition to obtain a target servo control strategy that enables the initial servo task to achieve the optimal reward value.
[0115] Based on the target servo control strategy, the robotic arm is controlled to execute a servo task whose target object is the same as the target object in the target object information of the initial servo task.
[0116] The computer-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic storage (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical storage (e.g., CD, DVD, BD, HVD), and semiconductor storage (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).
[0117] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0118] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A visual servo control method for a robotic arm, characterized in that, include: The initial servo task of the robotic arm is combined with the improved YOLO algorithm to perform target detection and obtain target object information; The robotic arm is a six-axis robotic arm; Based on the target object information, the initial servo task is decomposed, and servo control training is performed based on each sub-task obtained from the decomposition to obtain a target servo control strategy that enables the initial servo task to achieve the optimal reward value. Based on the target servo control strategy, the robotic arm is controlled to execute a servo task whose target object is the same as the target object in the target object information of the initial servo task; The step of training servo control based on the sub-tasks obtained from the decomposition to obtain a target servo control strategy that enables the initial servo task to achieve the optimal reward value includes: Servo control training is performed on each subtask obtained from the decomposition until the target servo control strategy is obtained when each subtask achieves the optimal reward value of the initial servo task under the preset behavior fusion scenario and its weights; wherein, the preset behavior fusion scenario includes the process of the robotic arm approaching the target object and the robotic arm reaching the target position. The sub-tasks include moving towards the target, avoiding losing target feature points, and maintaining a certain duration after reaching the target.
2. The robotic arm vision servo control method according to claim 1, characterized in that, The loss function of the improved YOLO algorithm includes confidence level and relative error rate.
3. The robotic arm vision servo control method according to claim 1, characterized in that, Before the step of training servo control based on the sub-tasks obtained from the decomposition, the following is also included: The data of each subtask obtained from the decomposition is transformed according to the preset interaction matrix to obtain the transformed subtasks.
4. The robotic arm vision servo control method according to claim 1, characterized in that, After the initial servo task based on the robotic arm, combined with the improved YOLO algorithm, performs target detection to obtain target object information, the method further includes: Determine whether an initial servo control strategy corresponding to the target object information exists; If the initial servo control strategy does not exist, then the step of decomposing the initial servo task based on the target object information is executed.
5. The robotic arm vision servo control method according to claim 4, characterized in that, After the step of determining whether there is an initial servo control strategy corresponding to the target object information, the method further includes: If the initial servo control strategy exists, the robotic arm is controlled to perform the initial servo task based on the initial servo control strategy.
6. A vision servo control device for a robotic arm, characterized in that, include: The detection module is used to perform target detection based on the initial servo task of the robotic arm and the improved YOLO algorithm to obtain target object information; The robotic arm is a six-axis robotic arm; The training module is used to decompose the initial servo task based on the target object information, and to train the servo control based on each sub-task obtained from the decomposition, so as to obtain a target servo control strategy that enables the initial servo task to achieve the optimal reward value; the sub-tasks include moving towards the target, avoiding losing target feature points, and maintaining a certain time after reaching the target. The control module is used to control the robotic arm to perform a servo task to be executed based on the target servo control strategy, where the target object is the same as the target object in the target object information of the initial servo task. The training module is specifically used to train the servo control of each sub-task obtained by decomposition until the target servo control strategy of each sub-task is obtained so that the initial servo task reaches the optimal reward value under the preset behavior fusion scenario and its weight; wherein, the preset behavior fusion scenario includes the process of the robotic arm approaching the target object and the robotic arm reaching the target position.
7. An electronic device comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the robotic arm vision servo control method according to any one of claims 1 to 5.
8. A medium, said medium being a computer-readable storage medium, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the robotic arm vision servo control method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Manipulator assembly task automatic programming method based on target assembly relationship natural language descriptions
CN106126219A
Dispensing detection method and device and computer readable storage medium
CN110598761A
For hiearchical decomposition deep reinforcement learning for an artificial intelligence model
WO2018236674A1