A method for a robot to learn cell micro - operation skills based on video demonstrations

Through imitation learning methods based on video demonstration, the robot learns cell microoperation skills from the video, uses multi-task observation network and hidden Markov model to obtain space-time trajectories and constraints, solving the problem of collaborative operation of multiple end effectors, improving the safety and accuracy of operations, and reducing the risk of damage to cells.

CN119839849BActive Publication Date: 2025-07-11ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411856879.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-07-11
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

The prior art is difficult to achieve collaborative operation of multi-terminal effectors of biological cells, and there are problems such as large differences in parameters, difficulty in flexibly controlling and may cause damage to cells.

Method used

Using a robot cell micro-operation skills learning method based on video demonstration, the robot learns cell micro-operation skills from the video by imitating the learning framework, uses the multi-task observation network and task parameterized hidden Markov model to obtain space-time trajectory and constraints, and combines the soft actor-criticist network for skill exploration and learning to ensure the safety and accuracy of operations.

Benefits of technology

It improves the safety, flexibility, accuracy and success rate of robot cell microoperation, reduces the requirements for operators and experimental costs, shortens the operation time, and reduces cell deformation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119839849B_ABST
    Figure CN119839849B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for a robot to learn cell micromanipulation skills based on video demonstrations. First, a multi-task observation network is used to identify two end effectors and the object to be manipulated in the robot micromanipulation video under a microscope, and the teaching spatio-temporal trajectory is obtained. Then, a task-parameterized hidden Markov model is used to obtain the spatio-temporal constraints of the robot's actions. In order to simultaneously solve the problems of safety and dexterity in robot micromanipulation, a soft actor-critic network optimized based on imitation learning is proposed, enabling the robot to learn skills through demonstration and exploration. The present invention can perform complex cell manipulation tasks in a real physical environment, and is expected to promote the development of robot skill learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for a robot to learn cell micromanipulation skills in the fields of robot technology and biological cell manipulation, and particularly relates to a method for enabling a robot to learn cell micromanipulation skills through video demonstration. Background Art

[0002] In the field of micro-nano scale automation, robot micromanipulation technology is one of the key technologies, which is widely applied to micro-assembly of mechanical components, biological sample processing, microsurgery, etc. Especially for micro-objects such as biological cells that are fragile and easily deformed, it is particularly challenging to achieve micron-scale cell micromanipulation with multi-end effectors coordinated operation, because cells are easily deformed when stressed and excessive deformation may cause irreversible damage. Traditional micromanipulation methods have problems such as large parameter differences, difficulty in flexibly controlling multiple end effectors, and possible damage to cells.

[0003] Therefore, it is necessary to propose a method for a robot to learn cell micromanipulation skills with a relatively high degree of automation. Summary of the Invention

[0004] The present invention aims to provide a method for a robot to learn cell micromanipulation skills based on video demonstration to overcome the deficiencies of the prior art, enabling the robot to learn cell micromanipulation skills from videos, improving the safety, flexibility, precision, success rate, and degree of automation of the operation, while reducing the requirements for operators and experimental costs. The present invention enables the robot to learn cell micromanipulation skills from demonstration videos through an imitation learning (IL) framework, which is an imitation learning method inspired by the human learning process.

[0005] The technical solution adopted by the present invention is as follows:

[0006] I. A method for a robot to learn cell micromanipulation skills based on video demonstration

[0007] Step 1: Input a robot cell micromanipulation video into a multi-task observation network, and the network outputs a taught spatio-temporal trajectory;

[0008] Step 2: Use a task-parameterized hidden Markov model to obtain the spatio-temporal constraints of the robot actions in the taught spatio-temporal trajectory, and use a behavior cloning modeling strategy to reference the distribution;

[0009] Step 3: Obtain real-time operation image frames, use the multi-task observation network to obtain the micromanipulation state information of the real-time image frames, and then, based on the micromanipulation state information of the real-time image frames, the soft actor-critic network optimized by imitation learning performs skill exploration and learning through the spatio-temporal constraints of the robot actions and the policy reference distribution and outputs the control quantity of the robot, thereby controlling the robot;

[0010] Step 4: Repeat Step 3 to continuously perform the interactive control of the robot until the learning of cell micromanipulation skills is completed.

[0011] In the said Step 1, the multi-task observation network includes a backbone network, an object detection branch, a cell mask segmentation branch, and a key point detection branch. The input of the multi-task observation network serves as the input of the backbone network, and the backbone network is respectively connected to the object detection branch, the cell mask segmentation branch, and the key point detection branch. The object detection branch is used to identify the end effector of the micromanipulation robot, the cell mask segmentation branch is used to generate a cell mask, and the key point detection branch is used to identify the cell pose.

[0012] The total loss of the multi-task observation network during the training process is the weighted sum of the object detection loss, the instance segmentation loss, and the key point detection loss.

[0013] In the soft actor-critic network optimized based on imitation learning, the actual policy π is constructed by combining the divergence loss φ The loss function of the learning process is as follows:

[0014]

[0015] Where is the loss function of the policy function, is the expectation, α is the weight for balancing the policy gradient and the value function estimation, π φ (|) is the probability of the policy function, π z is the policy reference distribution, Q θ () is the action value function, s i and a i are respectively the current state of the agent in the environment and the action selected in state s i , is the weight of the divergence loss, D KL (||) is the divergence value, is the divergence loss.

[0016] In the soft actor-critic network optimized based on imitation learning, the action a of the micromanipulation robot i consists of the moving direction and the micromanipulation speed. At the initial frame, the spatio-temporal constraint of the robot action is used as the constraint state space of the actual policy π φ The threshold parameter λ0 of the exploration ability and the constraint parameter λ s satisfy λ0 < λ s < 1. If an unreasonable state appears in the current state information, a penalty is imposed to avoid potential damage to the cells; during the learning process of subsequent frames, if an unreasonable state appears during the learning process, a penalty is imposed to avoid potential damage to the cells. If the constraint parameter When it is less than the threshold parameter λ0, the unreasonable state is merged with the current constraint state space to update the constraint state space, so that the constraint state space of the policy is gradually and dynamically released. At the same time, the divergence loss gradually decreases during the learning process, so that the soft actor-critic network optimized based on imitation learning gradually learns micro-operation skills through exploration.

[0017] The reward function of the soft actor-critic network optimized based on imitation learning satisfies the following formula:

[0018] r i = r base,i + r task,i

[0019]

[0020] r inject,i = -2||e 2,i - e 3,i || 2

[0021]

[0022]

[0023] where the subscript i represents the current moment, r i is the total reward function, r base,i is the basic reward function, r task,i is the task reward function for different tasks task, task = inject, adjust or composite, e 2,i is the position of the injection needle, m c,i and m c,i-1 are the cell areas at the current moment and the previous moment respectively, e 1,i is the position of the holding needle, std() is the standard deviation, N steps is the number of operation steps, r inject,i is the reward function for the injection needle puncture operation, e 3,i is the position of the cell center, r adjust,i is the reward function for cell pose adjustment, is the current direction of the cell, ω odesire is the target direction of the cell, r composite,i is the reward function for the composite task, c1 is the penalty coefficient for performing tasks in the incorrect order, c2 is the reward coefficient for successful cell pose adjustment, |||| 2 is the square of the vector norm.

[0024] II. A robot cell micro-operation skill learning system based on video demonstration

[0025] An image acquisition unit for acquiring robot cell micro - operation image frames;

[0026] A micro - operation state unit for obtaining state information in the acquired robot cell micro - operation image frames by using a multi - task observation network;

[0027] A teaching spatio - temporal trajectory generation unit for sequentially inputting each image frame in a robot cell micro - operation video into the micro - operation state unit to obtain a teaching spatio - temporal trajectory;

[0028] A robot action spatio - temporal constraint and policy reference distribution generation unit for obtaining robot action spatio - temporal constraints in the teaching spatio - temporal trajectory by using a task - parameterized hidden Markov model and modeling a policy reference distribution by using a behavior cloning strategy;

[0029] A skill exploration and learning unit for, according to the state information corresponding to the real - time operation image frames obtained by the micro - operation state unit, performing skill exploration and learning through the robot action spatio - temporal constraints and the policy reference distribution based on a soft actor - critic network optimized by imitation learning, outputting a control amount of the robot, and controlling the robot.

[0030] III. A computer device

[0031] The device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method for robot cell micro - operation skill learning based on video demonstration are implemented.

[0032] IV. A computer - readable storage medium

[0033] The medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for robot cell micro - operation skill learning based on video demonstration are implemented.

[0034] V. A computer program product

[0035] The product includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method for robot cell micro - operation skill learning based on video demonstration are implemented.

[0036] The beneficial effects of the present invention are as follows:

[0037] The method proposed by the present invention enables a robot to learn cell micro - operation skills from a video. By designing a multi - task observation network, teaching spatio - temporal trajectories of multiple end - effectors and objects to be operated are extracted from an expert video, and the video sequence is converted into an observation sequence containing various information, providing accurate data support for subsequent operations.

[0038] Extracting spatiotemporal constraints in expert demonstrations using a task-parameterized hidden Markov model can effectively identify implicit constraints in micro-operation tasks, ensure the safety and dexterity of robot operations, and avoid excessive damage to cells.

[0039] The soft actor-critic network optimized based on imitation learning combines demonstration and exploration, enabling the robot to perform complex cell operation tasks in an actual physical environment, and improving the success rate and efficiency of operations.

[0040] Compared with existing methods and manual remote operations, the framework of the present invention shortens the operation time, reduces cell deformation, and helps to promote the development of robot skill learning. Brief Description of the Drawings

[0041] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0042] Figure 1 is a schematic flowchart of a method for robot cell micro-operation skill learning based on video demonstration in a preferred embodiment of the present invention;

[0043] Figure 2 is a schematic structural diagram of a multi-task observation network;

[0044] Figure 3 is a schematic diagram of the convergence process of the loss function of the multi-task observation network;

[0045] Figure 4 is a schematic diagram of the recognition result of the multi-task observation network;

[0046] Figure 5 is a schematic diagram of the implicit constraint result identified by compressing, encoding, and decoding the spatiotemporal state trajectory through a task-parameterized hidden Markov model;

[0047] Figure 6 is a schematic diagram of the change process of the reward function of the method of the present invention and other methods in cell operation tasks in a preferred embodiment of the present invention;

[0048] Figure 7 is a schematic diagram of the process of the method of the present invention controlling the robot to complete cell attitude adjustment and cell puncture tasks in a preferred embodiment of the present invention. Detailed Description of the Preferred Embodiments

[0049] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0050] The multi-degree-of-freedom micro-operation robot system used in the experiment includes two robotic arms with three degrees of freedom, an inverted microscope system with automatic switching of optical magnification (such as IXplore Standard, OLYMPUS, Japan), a microscope operation platform equipped with an XY motorized stage, and a CMOS camera for image acquisition. The maximum magnification of the microscope eyepiece and objective lens is 400x, and the motion resolution of the robotic arm is 0.05μm. The system software and control part are integrated into Robot Operating System (ROS) Melodic and run on a desktop computer equipped with an Intel Core i9-10980Xe CPU and an NVIDIA TITAN Xp, with the operating system being Ubuntu 18.04.

[0051] The experimental objects are live zebrafish embryo cells with a diameter of 700 - 1000μm. The diameter of the holding pipette is 800μm, and the diameter of the injection needle is 30μm. 100 experiments are conducted for each method. The evaluation indicators include whether the holding pipette captures the cell, whether the attitude adjustment angle is less than 10 degrees, whether the injection needle penetrates the cell, whether the overall operation time is less than 10 seconds, and whether the cell or the end effector is damaged. Success is defined as the cell being adjusted to the specified attitude and the puncture being completed.

[0052] The present invention proposes a method for a robot to learn cell micro-operation skills based on video demonstrations, such as Figure 1 shown, the method includes the following steps:

[0053] Step 1: Sequentially input several complete robot cell micro-operation videos of the same operation type into a multi-task observation network. The network is used to identify the two end effectors and the manipulated object in the robot micro-operation video under the microscope, so as to output the taught spatio-temporal trajectory;

[0054] Among them, as Figure 2 shown, the multi-task observation network (MTO) includes a backbone network, an object detection branch, a cell mask segmentation branch, and a key point detection branch. The input of the multi-task observation network serves as the input of the backbone network. The backbone network is respectively connected to the object detection branch, the cell mask segmentation branch, and the key point detection branch. The object detection branch is used to identify the end effectors of the two micro-operation robots, the cell mask segmentation branch is used to generate cell masks, and the key point detection branch is used to identify the cell attitude. Finally, the position information e1 and e2 of the end effectors in each frame, the position e3 of the manipulated cell, the cell attitude angle information ω, and the cell mask m c constitute the taught spatio-temporal trajectory Provide a comprehensive basis for state tracking and control of multi-task micro-operations, ensure precise positioning and attitude adjustment during micro-operations, and monitor cell deformation to ensure safe operation. Among them, the position information e1 and e2 of the end effector are used to track its precise position during the operation; the position e3 of the manipulated cell helps to determine the displacement of the cell during the whole operation; the attitude angle information ω of the cell is used to track the attitude change to support the cell attitude adjustment task; and the cell mask m c , which is used for instance segmentation of cells to provide cell shape and deformation information.

[0055] In this embodiment, the backbone network is ResNeXt. The object detection branch is the Region Proposal Network, which generates candidate target regions and identifies the end effectors of two micro-operation robots through bounding box regression and class classification. The cell mask segmentation branch is RoIAlign, which generates instance masks on each region of interest to detect cell deformation. The keypoint detection branch is the Hourglass Network, which generates keypoint heatmaps for identifying cell postures, and finally generates output heatmaps through the keypoint heatmaps, with each channel representing a keypoint.

[0056] The total loss of the multi-task observation network during training is the object detection loss the instance segmentation loss and the keypoint detection loss weighted sum of.

[0057] The object detection loss of the multi-task observation network includes the bounding box loss and the class classification loss. The bounding box loss is The class classification loss is where is the bounding box loss, N is the number of bounding boxes, Smooth L1 () is the smooth L1 loss function, which is used to calculate the difference between the predicted bounding box and the ground truth bounding box, is the i-th ground truth bounding box, is the i-th predicted bounding box; the instance segmentation loss is where H and W are the height and width of the segmentation mask respectively, is the j-th pixel value of the i-th ground truth segmentation mask, is the j-th pixel value of the i-th predicted segmentation mask; the keypoint detection loss is where, is the Gaussian distribution value of the k-th keypoint at the j-th pixel position of the i-th ground truth heatmap, is the Gaussian distribution value of the k-th keypoint at the j-th pixel position of the i-th predicted heatmap, and K is the number of keypoints.

[0058] When training the MTO network, a multi-task observation network is used to process a video dataset containing robotic cell micro-operations. The dataset is divided into an original dataset and a dataset with added noise. During the training process, the convergence of the loss function is recorded, and the changing trends of the object detection loss, instance segmentation loss, and keypoint detection loss with the number of training epochs are observed. As Figure 3 shown, Figure 3 in (a) of Figure 3 (b) of Figure 3 (c) of Figure 3 and (d) of obj are the processes in which the bounding box loss, object classification loss, instance segmentation loss, and keypoint detection loss decrease and converge as the number of training epochs increases. Among them, the three weights λ obj 、λ mask and λ key are set to 0.3, 0.4, and 0.3 respectively.

[0059] To verify the accuracy and robustness of MTO for multi-task observation, tests are conducted on the original dataset and the dataset with added noise. For the object detection task, the mean average precision (mAP) is calculated. When processing the original dataset and the dataset with noise, the object detection mAP (0.5:0.95) is 0.928 and 0.911 respectively, with a change rate of 1.83%. For the instance segmentation task, the mAP is also calculated. The instance segmentation mAP (0.5:0.95) for the original dataset and the dataset with noise is 0.863 and 0.852 respectively, with a change rate of 1.27%. In the pose recognition task, the angle recognition error is measured. The angle recognition errors for the original dataset and the dataset with noise are 1.47° and 1.49° respectively, with a change rate of 1.36%. The experimental results show that MTO has high accuracy and robustness in multi-task observation. The recognition results of the video frames are as Figure 4 shown. After being recognized by the multi-task network, the position information e1 and e2 of the end effector, the position e3 of the manipulated cell, the pose angle information ω of the cell, and the cell mask m c can be obtained.

[0060] Multiple consecutive control tasks are selected for testing, including two end effectors approaching the object to be operated (basic task), cell puncture, cell pose adjustment, and a composite task of puncture after pose adjustment. 10 demonstration videos are prepared for each task.

[0061] Step 2: Use the task-parameterized hidden Markov model (THMM) to obtain the spatio-temporal constraints of the robot actions in the taught spatio-temporal trajectory, and use the behavior cloning modeling strategy to refer to the distribution. The constraint distributions of the end effector operation trajectories e1, e2, the cell puncture trajectory, and the cell pose adjustment trajectory are respectively as Figure 5 shown in (a) of Figure 5 and (b) of

[0062] Specifically:

[0063] The observed sequence is associated with an unknown implicit state sequence, which consists of K discrete hidden states. Different hidden states correspond to different constraints in the robot micro-operation task, and the hidden states follow distribution, a i,j represents the transition probability from state i to state j, and the probability of staying in a state for s consecutive time steps is estimated by a Gaussian distribution The observed feature ξ t comes from a multivariate Gaussian distribution of state j, and the overall parameter set Πi is the probability of the initial state i, a i,m is the probability of transitioning from state i to state m, K is the total number of hidden states, are the mean, covariance of state i, and the mean and covariance of the residence time in state i, respectively. The spatio-temporal state trajectory is compressed, encoded, and decoded by THMM to identify the implicit constraints in the micro-operation task.

[0064] The process of compressing and encoding the spatio-temporal state trajectory is as follows: By using the Expectation-Maximization (EM) method to maximize the expected complete log-likelihood of THMM, the overall parameter set of THMM is estimated.

[0065] The decoding process of the spatio-temporal state trajectory is as follows: Use the Viterbi method to decode the most likely hidden state sequence, calculate the probability of the observed sequence in the hidden state, predict the state sequence in the future time range through deterministic sampling, and obtain the spatio-temporal constraint S THMM ={s1, s2,..., s K}.

[0066] The process of obtaining the policy reference distribution π z using the behavior cloning algorithm is as follows: Take the position information of the end effector at the current moment, the position of the manipulated cell, the attitude angle information of the cell, and the cell mask as the state, and the position information of the end effector at the next moment as the action. Then, after using a multi-layer perceptron (MLP) for behavior cloning modeling, obtain the policy reference distribution π that outputs the corresponding action for any state z .

[0067] Step 3: Obtain real-time operation image frames, and use a multi-task observation network to obtain the micro-operation state information s i of the real-time image frames. Then, based on the micro-operation state information of the real-time image frames, the soft actor-critic network optimized by imitation learning (ILOSAC) explores and learns skills through the spatio-temporal constraints of the robot actions and the policy reference distribution and outputs the control amount of the robot, thereby controlling the robot;

[0068] Step 4: Repeat Step 3 to continuously perform the interactive control of the robot until the learning of cell micromanipulation skills is completed.

[0069] The present invention extends the ordinary Markov decision process (MDP) to a spatio-temporal constrained Markov decision process (SCMPD), including a state space S, an action space A, a transition probability P, a reward function R, and a cost function C. The action a of the micromanipulation robot i consists of a moving direction d i and a micromanipulation speed v i . is the moving direction of the left suction needle, is the moving speed of the left suction needle, is the moving direction of the right suction needle, is the moving speed of the right suction needle, |||| is the modulus of the vector, and the decision is made through the policy π φ (a|s) (parameterized by a Gaussian function, with the mean μ φ (s) and the covariance σ φ (s) given by a deep neural network (DNN)) and the critic function Q θ (s, a) (defined by the DNN and used to estimate the expected return of taking an action in the state). The critic parameter θ is trained by minimizing the loss . Among them, is the expected value of the state-action pair, y i is the target value used to train the critic network and consists of the immediate reward and the expected return of the next state, γ is the reward discount factor, is the expected return under the policy π, exploration rate λ, and temperature coefficient α, s0 is the initial state, a0 is the initial action, y i = r(s i , a i ) + γ(Q θ (s i+1 , a i+1 ) - α log(π φ (a i+1 |s t+1 ))) 2 , and α is the temperature coefficient used to control the balance between exploration and exploitation is the loss function of the temperature coefficient, is the expected value of the action sampled from the policy π i , π i (|) is the action probability distribution under the policy π i , and π i is the policy distribution of the policy in state i. is the minimum entropy objective. The policy π φ is constrained by the reference policy π z and the policy parameters are optimized by minimizing the Kullback-Leibler divergence D KL (π z ||π φ ). Among them, the actual policy π φ is constructed by combining the divergence loss with the loss function of the learning process, and the formula is as follows:

[0070]

[0071] Among them, is the loss function of the policy function, which is used to measure the quality of the current policy parameters φ, is the expectation, which represents the distribution sampling of the state. α is the weight for balancing the policy gradient and the value function estimation. π φ (|) is the probability of the policy function, π z is the policy reference distribution, is the action value function, s i and a i are respectively the current state of the agent in the environment and the action selected in the state s i . is the weight of the divergence loss, λ KL <1. As the interaction time t increases, gradually decreases. D KL (||) is the divergence value, which is used to calculate the difference between two probability distributions, is the divergence loss.

[0072] In the soft actor-critic network optimized based on imitation learning, the action a i of the micromanipulation robot consists of the moving direction and the micromanipulation speed. At the initial frame, the spatio-temporal constraint S THMM of the robot action is used as the constraint state space of the actual policy π φ . The threshold parameter λ0 of the exploration ability and the constraint parameter λ s satisfy λ0<λ s <1. If an unreasonable state appears in the current state information, that is, the state information s i obtained by the multi-task observation network exceeds the current constraint state space, a penalty is imposed to avoid potential damage to the cells; in the subsequent frame learning process, if an unreasonable state appears during the learning process, a penalty is imposed to avoid potential damage to the cells. If the constraint parameter of the exploration ability round t is less than the threshold parameter λ0, the unreasonable state is merged with the current constraint state space to update the constraint state space, so that the constraint state space of the policy is gradually and dynamically released. At the same time, during the learning process Gradually decrease, so the divergence loss gradually decreases. The constraint condition includes the constraint state space and the divergence loss, that is, the constraint condition gradually weakens. Thus, the soft actor-critic network optimized based on imitation learning gradually learns new micro-operation skills through exploration. During the exploration process, the agent's actions are mainly supervised and constrained by the reward function.

[0073] The reward function of the soft actor-critic network optimized based on imitation learning satisfies the following formula:

[0074] r i = r base,i + r task,i

[0075]

[0076] r inject,i = -2||e 2,i - e 3,i || 2

[0077]

[0078]

[0079] where the subscript i represents the current moment, r i is the total reward function, r base,i is the basic reward function, r task,i is the task reward function for different tasks task, task = inject, adjust or composite. Inject, adjust or composite respectively represent the cell puncture, cell pose adjustment, and composite task of first performing cell pose adjustment and then cell puncture. e 2,i is the position of the injection needle, m c,i and m c,i-1 are the cell areas at the current moment and the previous moment respectively, e 1,i is the position of the holding needle, std() is the standard deviation, N steps is the number of operation steps, r inject,i is the reward function for the injection needle puncture operation, e 3,i is the position of the cell center, r adjust,i is the reward function for cell pose adjustment, is the current direction of the cell, ω desire is the target direction of the cell, r composite,i is the reward function for the composite task, c1 is the penalty coefficient for performing tasks in the incorrect order, with a value of 9 in this embodiment, c2 is the reward coefficient for successful cell pose adjustment, with a value of 3 in this embodiment, |||| 2is the square of the vector norm.

[0080] -||e 2,i -m c,i || 2 and ||e 1,i -m c,i || 2 are respectively used to constrain the injection needle and the aspiration needle pipette close to the cell, used to constrain cell deformation, -||m c,i -m c,i-1| | 2 used to prevent large deformation of the cell in a short time, -0.1N steps used to prevent the operation time from being too long, -||e 2,i -e 3,i || 2 used to prompt the injection needle to reach the center of the cell for puncture operation, used to prompt the injection needle to rub the cell to complete the attitude adjustment. r composite,i enable the injection needle to complete the cell attitude adjustment and puncture operation in sequence. When performing the attitude adjustment task, when performing the cell puncture task, r composite,i =-3||e 2,i -e 3,i || 2 . If the system attempts to puncture before adjusting the attitude, then r composite,i =-c1 (c1 = 9), when the cell attitude matches the target, r composite,i =-c2 (c2 = 3).

[0081] Compare the performance of the method of the present invention with other benchmark methods (such as DDPG, SAC, SOIL, DDPGfD, DAPG, etc.), as Figure 6 shown. Among them, Figure 6 (a) is the learning process of the basic end effector approaching the cell, Figure 6 (b) of is the learning process of the puncture task, Figure 6 (c) of is the learning process of the cell attitude adjustment task, Figure 6The learning process of the composite task in (d) is to first adjust the posture and then perform puncture. In the basic task, the learning rates and stabilities of each method are slightly different. The method of the present invention can achieve a higher return. For example, the ILOSAC method reaches -0.32. In the injection task, the ILOSAC method has obvious advantages, with a fast learning speed and the highest average return (-0.81). While for the SAC and DDPG methods, due to the unconstrained policy gradient, the learning curves oscillate greatly. In the adjustment task, guided by imitation learning, the four imitation learning frameworks can all obtain higher returns within a short-time interaction. However, SAC and DDPG start learning from scratch and it is difficult to master effective operation skills in a short time. The average return of the ILOSAC method is significantly better than other methods. For the composite task, due to the high task complexity, the final returns of the DDPG and SAC methods are low and the variances are large. The ILOSAC method can quickly acquire skills by applying spatio-temporal constraints and can learn flexible skills by weakening the constraints to increase the exploration space, obtaining a higher return than the existing imitation learning methods.

[0082] To prove the superiority and task generality of the method of the present invention, verification experiments are carried out on each task in the real environment and the success rates are recorded. The experimental results show that the framework of the present invention has the highest success rate in multiple tasks. For example, the success rates in the basic task, injection task, adjustment task, and composite task reach 97%, 86%, 77%, and 68% respectively, while the success rates of other methods such as DDPG in these tasks are 77%, 56%, 38%, and 19% respectively, SAC is 74%, 61%, 45%, SOIL is 83%, 69%, 58% and 42%, DAPG is 89%, 81%, 63% and 54%, DDPGfD is 85%, 75%, 66% and 47% and at a lower level. Figure 7 Schematic diagram of the cell posture adjustment and cell puncture processes completed by the ILOSAC-controlled injection needle and suction needle respectively.

[0083] The present invention also proposes a robot cell micro-operation skill learning system based on video demonstration. The system includes:

[0084] An image acquisition unit for acquiring robot cell micro-operation image frames;

[0085] A micro-operation state unit for obtaining state information in the acquired robot cell micro-operation image frames by using a multi-task observation network;

[0086] A teaching spatio-temporal trajectory generation unit for sequentially inputting each image frame in the robot cell micro-operation video into the micro-operation state unit to obtain a teaching spatio-temporal trajectory;

[0087] A robot motion spatio-temporal constraint and policy reference distribution generation unit, which is used to obtain the robot motion spatio-temporal constraint and policy reference distribution π in the taught spatio-temporal trajectory by using a task-parameterized hidden Markov model z ;

[0088] A skill exploration and learning unit, which is used to, according to the state information corresponding to the real-time operation image frame obtained by the micro-operation state unit, perform skill exploration and learning based on a soft actor-critic network optimized by imitation learning through the robot motion spatio-temporal constraint and policy reference distribution π z and output the control amount of the robot and control the robot.

[0089] The present invention also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of a robot cell micro-operation skill learning method based on video demonstration are implemented.

[0090] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a robot cell micro-operation skill learning method based on video demonstration are implemented.

[0091] The present invention also provides a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of a robot cell micro-operation skill learning method based on video demonstration are implemented.

[0092] Finally, it should be noted that the above embodiments and descriptions are only used to illustrate the technical solutions of the present invention rather than to limit them. Those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced. Without departing from the spirit and scope of the disclosure of the technical solutions of the present invention, they should all be covered by the protection scope of the claims of the present invention.

Claims

1. A method for a robot to learn cell micromanipulation skills based on video demonstrations, characterized in that, It includes the following steps: Step 1: Input the robot cell micro-operation video into the multi-task observation network, and the network outputs the demonstration spatio-temporal trajectory; Step 2: Use the task-parameterized hidden Markov model to obtain the spatio-temporal constraints of the robot actions in the demonstration spatio-temporal trajectory, and use the behavior cloning algorithm to obtain the policy reference distribution; Step 3: Obtain the real-time operation image frames, use the multi-task observation network to obtain the micro-operation state information of the real-time image frames, and then based on the micro-operation state information of the real-time image frames, the soft actor-critic network optimized by imitation learning explores and learns skills through the spatio-temporal constraints of the robot actions and the policy reference distribution and outputs the control quantity of the robot, so as to control the robot; Step 4: Repeat Step 3 to continuously perform the interactive control of the robot until the learning of the cell micro-operation skills is completed.

2. The method for a robot to learn cell micro-operation skills based on video demonstration according to claim 1, wherein In the said Step 1, the multi-task observation network includes a backbone network, an object detection branch, a cell mask segmentation branch and a key point detection branch. The input of the multi-task observation network is used as the input of the backbone network. The backbone network is respectively connected to the object detection branch, the cell mask segmentation branch and the key point detection branch. The object detection branch is used to identify the end effector of the micro-operation robot, the cell mask segmentation branch is used to generate the cell mask, and the key point detection branch is used to identify the cell pose.

3. A method for a robot to learn cell micromanipulation skills based on video demonstrations according to claim 2, characterized in that, The total loss of the multi-task observation network during the training process is the weighted sum of the object detection loss, the instance segmentation loss and the key point detection loss.

4. A method for a robot to learn cell micromanipulation skills based on video demonstrations according to claim 1, characterized in that, In the soft actor-critic network optimized based on imitation learning, the actual policy π is constructed by combining the divergence loss φ The loss function of the learning process is as follows: Among them, is the loss function of the policy function, is the expectation, α is the weight for balancing the policy gradient and the value function estimation, π φ (|) is the probability of the policy function, π z is the policy reference distribution, is the action value function, s i and a i are respectively the current state of the agent in the environment and the action selected at state s i . is the weight of the divergence loss, D KL (||) is the divergence value, is the divergence loss.

5. A method for a robot to learn cell micromanipulation skills based on video demonstrations according to claim 1, characterized in that, In the soft actor-critic network optimized based on imitation learning, the action a of the micromanipulation robot i consists of the moving direction and the micromanipulation speed. At the initial frame, the spatio-temporal constraint of the robot action is used as the actual policy π φ of the constrained state space, the threshold parameter λ0 of the exploration ability, and the constraint parameter λ s satisfy λ0 < λ s < 1. If an unreasonable state appears in the current state information, a penalty is imposed to avoid potential damage to the cells; in the learning process of subsequent frames, if an unreasonable state appears during the learning process, a penalty is imposed to avoid potential damage to the cells. If the constraint parameter of the exploration ability round t is less than the threshold parameter λ0, the unreasonable state is merged with the current constrained state space to update the constrained state space, so that the constrained state space of the policy is gradually and dynamically released, and at the same time the divergence loss gradually decreases during the learning process. Thus, the soft actor-critic network optimized based on imitation learning gradually learns the micromanipulation skills through exploration.

6. A method for a robot to learn cell micromanipulation skills based on video demonstrations according to claim 1, characterized in that, The reward function of the soft actor-critic network optimized by imitation learning satisfies the following formula: r i =r base,i +r task,i r inject,i = -2||e 2,i -e 3,i || 2 where the subscript i denotes the current moment, r i is the total reward function, r base,i is the basic reward function, r task,i is the task reward function for different tasks task, where task = inject, adjust or composite, e 2,i is the position of the injection needle, m c,i and m c,i-1 are the cell areas at the current moment and the previous moment respectively, e 1,i is the holding needle position, std() is the standard deviation, N steps is the number of operation steps, r inject,i is the reward function for the injection needle puncture operation, e 3,i is the position of the cell center, r adjust,i is the reward function for cell pose adjustment, is the current direction of the cell, ω desire is the target direction of the cell, r composite,i is the reward function for the composite task, c1 is the penalty coefficient for executing tasks in the incorrect order, c2 is the reward coefficient for successful cell pose adjustment, ‖‖ 2 is the square of the vector norm.

7. A robot cell micro-operation skill learning system based on video demonstration, characterized in that, It includes: An image acquisition unit for acquiring the robot cell micro-operation image frames; A micro-operation state unit for using the multi-task observation network to obtain the state information in the acquired robot cell micro-operation image frames; A demonstration spatio-temporal trajectory generation unit for sequentially inputting the image frames in the robot cell micro-operation video into the micro-operation state unit to obtain the demonstration spatio-temporal trajectory; A robot action spatio-temporal constraint and policy reference distribution generation unit for using the task-parameterized hidden Markov model to obtain the spatio-temporal constraints of the robot actions in the demonstration spatio-temporal trajectory and using the behavior cloning algorithm to obtain the policy reference distribution; A skill exploration and learning unit for, according to the state information corresponding to the real-time operation image frames obtained by the micro-operation state unit, exploring and learning skills through the spatio-temporal constraints of the robot actions and the policy reference distribution by the soft actor-critic network optimized by imitation learning and outputting the control quantity of the robot and controlling the robot.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method for learning robot cell micro-operation skills based on video demonstration according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for learning robot cell micro-operation skills based on video demonstration according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, it implements the steps of the method for learning robot cell micro-operation skills based on video demonstration according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Robot sequence task learning method based on visual simulation

    CN111203878A

  • Precise radiotherapy system based on four-dimensional adaptive digital tracking conformal intensity-modulated focusing medical electron accelerator

    CN114796889A