Robot cell micro-operation method based on imitation learning with safety constraint mechanism
By constructing an imitation learning method of the safety constraint mechanism, the potential discrete encoding of the teaching video frame is extracted, and the error accumulation and safety problems in robot cell microoperation are solved, achieving high-precision and safe cell operation.
Patent Information
- Application Number
- CN202510115314.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-01-24
AI Technical Summary
Existing imitation learning algorithms have accumulated errors in robot cell microoperations, difficulty in dealing with soft tissue deformation and uncertainty, and difficulty in ensuring safety in high precision and long-term operations.
By extracting the potential discrete encoding of the teaching video frame, a security action constraint mechanism is built, and the autoregressive Transformer network and security constraint mechanism are used to quantify teaching skills information, predict and execute safety actions, and suppress errors.
It improves the success rate of robot cell microoperation, reduces operation errors and cumulative errors, ensures the safety and accuracy of operations, and adapts to changes in complex environments.
Smart Images

Figure CN119858159B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for robotic cell micromanipulation, involving the fields of robotics technology and biological cell manipulation, and particularly to a method for robotic cell micromanipulation based on imitation learning with a safety constraint mechanism. Background Art
[0002] In the field of robotics technology, especially in cell micromanipulation tasks such as delicate operations like stripping cell membranes, there is a need for long-term and high-precision operation of robots. These tasks not only require the robot to have a high degree of dexterity and stability but also ensure safety during the operation process to avoid damage to the fragile cell structure. Although imitation learning provides an effective method to control the robot by imitating taught behaviors, existing algorithms are prone to accumulating errors when dealing with long-term tasks, resulting in operation failures. In addition, these algorithms have limitations in adapting to the deformation and uncertainty of soft tissues such as cells and face challenges in precise coordination and extended action planning in microscale operations. Therefore, it is necessary to develop a new imitation learning algorithm that can improve the success rate of cell micromanipulation tasks, especially in scenarios requiring long-term operation and high-precision control, by accurately representing taught skill information and implementing safety action constraints. Summary of the Invention
[0003] To solve the problems in the background art, the present invention provides a method for robotic cell micromanipulation based on imitation learning with a safety constraint mechanism. The present invention aims to quantify the representation of taught skill information by extracting the latent discrete encoding of the taught video frames and modeling the distribution of these encodings, and extract safety action constraints. Finally, by fusing the actions of the previous time step to predict actions and execute actions that meet the safety constraints, the compound error is effectively suppressed.
[0004] The technical solution adopted by the present invention is as follows:
[0005] The method for robotic cell micromanipulation based on imitation learning with a safety constraint mechanism of the present invention includes: [[ID=2l]]
[0006] Step 1: Obtain the taught video of robotic cell micromanipulation and extract the video frames as a training set to train the image generation network and the attention network including the reconstruction and skill information distribution loss function, obtain the encoders of the trained image generation network and attention network, and extract the discrete latent encoding of each taught video frame.
[0007] Step 2: Jointly construct the trained image generation network, the encoder of the attention network, and the multi-layer perceptron into a robotic cell operation model. Construct the safety action constraints of the robotic cell operation model based on each discrete latent code, quantify the representation of the demonstration skill information by calculating the log-likelihood of the latent discrete code, so as to extract the safety action constraints. Train the robotic cell operation model through the video frames of the demonstration video and the action trajectories during the robotic cell micro-operations in the demonstration video until the loss function of the robotic cell operation model converges, and obtain the trained robotic cell operation model.
[0008] Step 3: When the robot performs autonomous cell micro-operations, obtain the video frames during the robot's autonomous cell micro-operations in real time and input them together with the robot's actions at the previous moment into the trained robotic cell operation model. Output the robot's actions at the next moment under the condition of satisfying the safety action constraints, and then control the robot to execute the actions to achieve robotic cell micro-operations.
[0009] In the above-mentioned Step 1, the demonstration video of the robotic cell micro-operations is a video of the robot performing micro-operations on cells according to preset actions. Specifically, it is a video of controlling the manipulator of a remote robot to perform cell micro-operations according to preset actions through manual teleoperation. The image generation network and the attention network respectively adopt vector quantization and the generative adversarial network VQ-GAN (Vector Quantized Generative Adversarial Network) and the autoregressive Transformer network. Input the training set into the vector quantization and generative adversarial network VQ-GAN for training to obtain the trained vector quantization and generative adversarial network VQ-GAN. At the same time, the encoder of the vector quantization and generative adversarial network VQ-GAN compresses each demonstration video frame to obtain the discrete latent code z of each demonstration video frame. Input each discrete latent code z into the autoregressive Transformer network for training, use the autoregressive Transformer network to model the distribution of the discrete latent code z, maximize the likelihood value of each discrete latent code z until the reconstruction and skill information distribution loss function converges, and obtain the trained autoregressive Transformer network.
[0010] The above-mentioned autoregressive Transformer network optimizes the distribution of the discrete latent code based on the maximum likelihood estimation. The optimization objective function of the autoregressive Transformer network is as follows:
[0011]
[0012] Among them, represents the model parameters of the autoregressive Transformer network; Indicates the time step of the demonstration video; Indicates the number of dimensions of the discrete latent code; Indicates the log-likelihood; Indicates the encoding value at the th position in the discrete latent code at time Indicates the encoding values at the 1st to the th positions in the discrete latent code at time Indicates the latent discrete code from to subsequence, Indicates the frame length in the subsequence.
[0013] The reconstruction and skill information distribution loss function of the autoregressive Transformer network described above is as follows:
[0014]
[0015] where, Indicates the cross-entropy loss between the predicted value and the true value, i.e., the reconstruction loss; and respectively indicate the true value and the predicted value of the discrete latent code at time ; Indicates the predicted value of the encoding value at the th position in the discrete latent code at time ; Indicates the predicted values of the encoding values at the 1st to the th positions in the discrete latent code at time .
[0016] In the second step described above, the safety action constraint is specifically as follows:
[0017]
[0018]
[0019] where, Indicates the log-likelihood; Indicates the predicted value of the discrete latent code at time ; Indicates the discrete latent code from time to time ; Indicates the frame length in the subsequence; Indicates the threshold parameter of the safety constraint; Indicates during the training process Safety constraint value at a moment Indicates the total number of samples of the potential discrete coding of the teaching video Indicates the log-likelihood value of the potential discrete coding
[0020] In the second step described above, when training the robot cell operation model, the video frames of the teaching video and the action trajectories during the robot cell micro-operation in the teaching video are input into the robot cell operation model. The action trajectories during the robot cell micro-operation in the teaching video are the action trajectories of the end of the robot's manipulator and the microgripper. The video frames of the teaching video are sequentially input into the encoders of the trained image generation network and the attention network. After processing, the results are fused with the action trajectories during the robot cell micro-operation in the teaching video, and then the fused results are input into a multi-layer perceptron for training. During training, the network parameters of the encoder of the attention network are frozen, and the image generation network and the multi-layer perceptron are trained
[0021] In the second step described above, the loss function of the robot cell operation model Specifically as follows
[0022]
[0023] Wherein Indicates the mean square error And Respectively indicate The real action in the action trajectory during the robot cell micro-operation at a moment and the predicted action output by the robot cell operation model And Respectively indicate the reconstruction and skill information distribution loss function of the autoregressive Transformer network and the balance weight parameter of its skill information loss
[0024] In the third step described above, the robot action at the next moment currently output by the trained robot cell operation model , is input into the controller of the robot. The controller smooths the robot action at the next moment Using a trajectory smoothing method, and then controls the robot to execute the smoothed action
[0025] The electronic device of the present invention includes: a memory and a processor coupled to each other. Wherein, the memory stores program data, and the processor calls the program data to execute the method as described above
[0026] The computer-readable storage medium of the present invention stores program data thereon, and when the program data is executed by a processor, the method as described above is implemented
[0027] The beneficial effects of the present invention are as follows:
[0028] 1. The method proposed by the present invention can extract potential discrete codes by imitating the operation video frames of teaching, and use the autoregressive transformer model for modeling, which can effectively imitate teaching skills, improve the accuracy and safety of the robot in cell micro-operation tasks, and reduce the risk of damage to cells.
[0029] 2. The present invention quantifies the teaching skill information by calculating the log-likelihood of the potential discrete codes and extracts safety action constraints. The present invention can effectively suppress the compound errors common in long-term tasks when performing actions, and improve the success rate of the tasks.
[0030] 3. The safety constraint mechanism integrated in the method of the present invention can be dynamically adjusted during the operation to ensure that the operation of the robot remains within a safe and effective range even when facing unknown or changing environmental conditions.
[0031] In summary, by introducing the skill information representation technology and the safety constraint mechanism, the problems of multi-task coupling, long-time operation accuracy, and safety in micro-operation tasks are solved. The algorithm shows a high success rate in the embryo cell membrane stripping task, and effectively reduces operation errors and cumulative errors. The present invention can perform complex cell manipulation tasks in a real physical environment, and is expected to promote the development of robot skill learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a schematic flow chart of the method of the present invention;
[0033] Figure 2 is a schematic overall structure diagram of the robot cell operation model of the present invention, where Figure 2 in (a) is a schematic diagram of the encoder and decoder architectures of the vector quantization and generative adversarial network VQ-GAN, Figure 2 in (b) is a schematic diagram of the model of the autoregressive Transformer network, Figure 2 in (c) is a schematic network structure diagram of the robot cell operation model including a multilayer perceptron for controlling the left and right robotic arms and the microgripper during specific implementation;
[0034] Figure 3 is a schematic mechanical structure diagram of the microgripper used during specific implementation of the present invention;
[0035] Figure 4 is a schematic diagram of the change of the loss function during the training process, where Figure 4 in (a) is a schematic diagram of the loss function during the training of the vector quantization and generative adversarial network VQ-GAN, Figure 4(b) is a schematic diagram of the loss during the training process of the autoregressive Transformer network, Figure 4 (c) is a schematic diagram of the loss function during the behavior cloning training process, Figure 4 (b) is a schematic diagram of the normalized log-likelihood value of the training dataset after behavior cloning;
[0036] Figure 5 is a schematic diagram of controlling the robot system to perform the cell membrane tearing process during the specific implementation of the present invention;
[0037] In the figure: 1, micro-gripper; 2, Bragg grating; 3, optical fiber; 4, moving slider; 5, rotating shaft; 6, first DC servo motor; 7, second DC servo motor; 8, optoelectronic switch; 9, cylindrical propulsion cam; 10, rotating gear; 11, slider push rod; 12, rotating shaft rotating gear; 13, base. Detailed implementation manners
[0038] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0039] As Figure 1 shown, the robot cell micro-operation method based on imitation learning with a safety constraint mechanism of the present invention is specifically as follows:
[0040] First, obtain the teaching video of the robot cell micro-operation and extract video frames as the training set to train the image generation network and the attention network including the reconstruction and skill information distribution loss function, obtain the trained encoders of the image generation network and the attention network, and extract the discrete latent encodings of each teaching video frame; the teaching video of the robot cell micro-operation is a video of the robot performing micro-operations on cells according to preset actions, specifically a video of controlling the manipulator of a remote robot to perform cell micro-operations through manual teleoperation according to preset actions; the image generation network and the attention network respectively adopt vector quantization and the generative adversarial network VQ-GAN and the autoregressive Transformer network. After inputting the training set into the vector quantization and the generative adversarial network VQ-GAN for training, the trained vector quantization and the generative adversarial network VQ-GAN are obtained. At the same time, the encoder of the vector quantization and the generative adversarial network VQ-GAN compresses each teaching video frame to obtain the 8×8 discrete latent encoding z of each teaching video frame; input each discrete latent encoding z into the autoregressive Transformer network for training, use the autoregressive Transformer network to model the distribution of the discrete latent encoding z, maximize the likelihood value of each discrete latent encoding z until the reconstruction and skill information distribution loss function converges, and obtain the trained autoregressive Transformer network.
[0041] The autoregressive Transformer network optimizes the discrete latent encoding based on maximum likelihood estimation For the distribution of the autoregressive Transformer network, the optimization objective function is as follows:
[0042]
[0043] where represents the model parameters of the autoregressive Transformer network; represents the number of time steps of the demonstration video; represents the number of dimensions of the discrete latent encoding; represents the log-likelihood; represents the encoding value at the -th position in the 8×8 discrete latent encoding at time , represents the encoding values at the 1st to the -th positions in the discrete latent encoding at time ; represents the latent discrete encoding of the subsequence from to , represents the frame length in the subsequence; represents and the log-likelihood given .
[0044] The reconstruction and skill information distribution loss function of the autoregressive Transformer network is as follows:
[0045]
[0046] where represents the cross-entropy loss between the predicted value and the true value, i.e., the reconstruction loss; and respectively represent the true value and the predicted value of the discrete latent encoding at time ; represents the predicted value of the encoding value at the -th position in the 8×8 discrete latent encoding at time , represents the predicted values of the encoding values at the 1st to the -th positions in the discrete latent encoding at time ; is used for the maximum likelihood estimation optimization of the latent encoding and is the skill information distribution loss.
[0047] As shown in Figure 2 (a), Figure 2in (b) and Figure 2 As shown in (c), the encoder of the trained image generation network, the attention network, and the multi-layer perceptron are jointly constructed into a robotic cell operation model. Safety action constraints for the robotic cell operation model are constructed based on each discrete latent code. The representation of the taught skill information is quantified by calculating the log-likelihood of the latent discrete code, thereby extracting the safety action constraints. The robotic cell operation model is trained through the video frames of the teaching video and the action trajectories during the robotic cell micro-operations in the teaching video until the loss function of the robotic cell operation model converges, obtaining the trained robotic cell operation model. The safety action constraints are specifically as follows:
[0048]
[0049]
[0050] where, represents the log-likelihood; represents the predicted value of the discrete latent code at time , represents the discrete latent codes from time to time , represents the frame length in the subsequence; represents the threshold parameter of the safety constraint; represents the safety constraint value at time during the training process; represents the total number of samples of the latent discrete code of the teaching video; represents the log-likelihood value of the latent discrete code, represents the log-likelihood value after the corresponding latent discrete code of the subsequence is given, represents the log-likelihood value after the corresponding latent discrete code of the subsequence is given.
[0051] When training the robotic cell operation model, the video frames of the teaching video and the action trajectories during the robotic cell micro-operations in the teaching video are input into the robotic cell operation model. The action trajectories during the robotic cell micro-operations in the teaching video are the action trajectories of the end of the robotic arm and the microgripper of the robot. The video frames of the teaching video are sequentially input into the encoders of the trained image generation network and the attention network. After processing, the results are fused with the action trajectories during the robotic cell micro-operations in the teaching video, and then the fused results are input into the multi-layer perceptron for training. During training, the network parameters of the encoder of the attention network are frozen, and the image generation network and the multi-layer perceptron are trained.
[0052] Train the robotic cell operation model, i.e., perform behavior cloning. Use the demonstration video frames and motion trajectories as inputs. Compress the demonstration video frames into discrete latent encodings through the encoder of VQ-GAN. Use the discrete latent encoding sequence as the input, extract features using the autoregressive Transformer, fuse the features with the action trajectory at the previous time step, and predict the action at the next moment through a multi-layer perceptron. The network parameters of the VQ-GAN encoder are frozen, and the autoregressive Transformer network and the multi-layer perceptron network are trained, and the network parameters are saved.
[0053] Loss function of the robotic cell operation model Specifically as follows:
[0054]
[0055] Among them, represents the mean square error; and respectively represent the real action in the action trajectory during the robotic cell micro-operation at time and the predicted action output by the robotic cell operation model;
[0056] Finally, when the robot performs autonomous cell micro-operation, the video frames during the robot's autonomous cell micro-operation are obtained in real time and jointly input into the trained robotic cell operation model together with the robot's action at the previous moment. Under the condition of satisfying the safe action constraints, the action of the robot at the next moment is output. The action of the robot at the next moment currently output by the trained robotic cell operation model , is input into the controller of the robot. The controller smooths the action of the robot at the next moment using the trajectory smoothing method, and then controls the robot to execute the smoothed action to achieve robotic cell micro-operation.
[0057] The trajectory smoothing method used when executing the action is as follows:
[0058]
[0059] Among them, and respectively represent the smoothed actions of the robot at the current and next moments; represents the smoothing parameter.
[0060] For example, Figure 3As shown in the figure, the hardware structure of the robot system used in the specific implementation of the present invention mainly consists of an inverted microscope, a microscope operating robotic arm, a camera, a force feedback device, and a micro gripper, etc. Among them, the inverted microscope is an IXplore Standard, equipped with an XY electric stage, and its maximum magnification can reach 400x through the eyepiece and objective lens. The robotic arm selected is TransferMan, which has three degrees of freedom and a motion resolution of 0.05 µm. The camera W160 is installed on the inverted microscope and is used to capture image information at a rate of 60 frames per second. Two Phantom Touch devices are used to achieve the force feedback of the end effector. Two micro grippers with symmetrical structures are designed for cell dexterous operation and are integrated onto the left and right robotic arms. As Figure 3 shown, the micro gripper includes a micro claw 1, two fiber Bragg gratings 2, an optical fiber 3, a motion slider 4, a rotating shaft 5, a first DC servo motor 6, a second DC servo motor 7, a photoelectric gate 8, a cylindrical push cam 9, a rotating gear 10, a slider push rod 11, a rotating shaft rotating gear 12, and a base 13. The photoelectric gate 8, the cylindrical push cam 9, the rotating gear 10, the slider push rod 11, and the rotating shaft rotating gear 12 are installed on one side of the base 13. The rotating shaft 5, the first DC servo motor 6, and the second DC servo motor 7 are installed on the other side of the base 13. The output shaft of the first DC servo motor 6 is synchronously connected to the central shaft of the rotating gear 10. One end of the rotating shaft 5 is synchronously connected to one central end of the rotating shaft rotating gear 12. The slider push rod 11 is synchronously connected to the other central end of the rotating shaft rotating gear 12. The rotating gear 10 and the rotating shaft rotating gear 12 are meshed. The cylindrical push cam 9 is pressed against the slider push rod 11. The central shaft of the cylindrical push cam 9 is synchronously connected to the output shaft of the second DC servo motor 7. The photoelectric gate 8 is installed on one side close to the cylindrical push cam 9 and is used for the reset operation of the cylindrical push cam. The cylindrical push cam 9, the rotating gear 10, and the rotating shaft rotating gear 12 are arranged in sequence. One end of the motion slider 4 is synchronously connected and installed at the center of the other end of the rotating shaft 5. The two fiber Bragg gratings 2 and the optical fiber 3 are arranged on the motion slider 4. The micro claw 1 is installed at the other end of the motion slider 4. The micro gripper has two degrees of freedom, namely the grasping action and the wrist-like rotational motion. The axial rotation is driven by a set of rotating gears 10 of the first DC servo motor 6. The grasping action is completed by the second DC servo motor 7 pushing the motion slider 4 through the cylindrical push cam 9. Three fiber Bragg gratings FBG (Fiber Bragg Grating) 2 are integrated on the motion slider 4 to sense the bending deformation and provide force sensing and collision detection to prevent damage to the micro gripper. In terms of software control, the software and control part of the system are integrated into the Robot Operating System (ROS) Melodic running on the Ubuntu 18.04 operating system. The computer is configured with an Intel Core i9 - 10980Xe CPU and an NVIDIA TITAN Xp.
[0061] In the specific implementation of the present invention, zebrafish embryo cells are used as experimental objects. These cells are placed in a culture dish added with nutrient solution to maintain their activity, so as to ensure that the cells can maintain a normal physiological state during the experimental operation, facilitating precise cell micromanipulation experimental research.
[0062] The entire cell micromanipulation task mainly includes three subtasks: pushing cells, grasping cells, and cell membrane tearing. In the subtask of pushing cells, a single end effector pushes the cell, requiring relatively low dexterity and cooperation. The main purpose is to move the cell to a suitable operation position, such as the center of the field of view. The subtask of grasping cells requires using a double end effector to grasp the cell membrane. This process requires high dexterity and cooperation, and it is necessary to ensure that only the cell membrane is grasped to avoid damaging the cell interior and ensure the activity of the cell. The subtask of cell membrane tearing is to slowly separate and tear the cell membrane with left and right microgrippers to release the embryo. In this process, the tearing and squeezing of the membrane require the highest dexterity and cooperation, and the interaction process between the robot and the cell is also the most complex.
[0063] In the stage of teaching operation and teaching data collection, cell operation teaching controls the robot to perform operation demonstrations through a teleoperation system. According to the cell images under the microscope, the manipulator and microgrippers are precisely manipulated by using a force feedback device to complete the above three subtasks. During the whole process, various data including video information, end effector trajectory, axial rotation angle, grasping action, and multi-joint manipulator trajectory are recorded through the ROS (Robot Operating System) platform of the robot operating system. Finally, 175 effective demonstration samples are collected, with a total duration of 186.1 minutes.
[0064] The change of the loss function during the training process of the network is as Figure 4 shown in (a) of Figure 4 shown in (b) of Figure 4 shown in (c) of Figure 4 and shown in (d) of Figure 4 As shown in (a) of Figure 4 the loss function of the VQ-GAN network converges, indicating that the network can effectively compress image data into discrete latent code z. As shown in (b) of Figure 4 the reconstruction and skill information distribution loss function of the Transformer network Figure 4 converges. As shown in (d) of Effectively converge, proving that the network can predict the correct cell manipulation actions.
[0065] To verify the effectiveness of the VQ-GAN encoder, it was replaced with the encoders of the Residual Network ResNet18, the MobileNetV2 (Inverted Residuals and Linear Bottlenecks) convolutional neural network, and the VGG16 (Visual Geometry Group 16-layer network) convolutional neural network respectively, and the normalized log-likelihood values were calculated for the subtasks (pushing the cell, grasping the cell, cell membrane tearing) during the cell membrane tearing process. The experimental results show that the VQ-GAN encoder demonstrates significant superiority in all subtasks. Specifically, the normalized log-likelihood value of VQ-GAN in the cell pushing task is 0.74 ± 0.07, significantly higher than 0.39 ± 0.11 of ResNet18, 0.27 ± 0.06 of MobileNetV2, and 0.44 ± 0.15 of VGG16; in the cell grasping task, the normalized log-likelihood value of VQ-GAN is 0.86 ± 0.10, far exceeding other encoders (ResNet18: 0.21 ± 0.08, MobileNetV2: 0.19 ± 0.13, VGG16: 0.28 ± 0.09); in the cell membrane tearing task, the normalized log-likelihood value of VQ-GAN is 0.82 ± 0.12, also significantly higher than other encoders (ResNet18: 0.33 ± 0.16, MobileNetV2: 0.25 ± 0.07, VGG16: 0.26 ± 0.12). These data fully prove the excellent performance of the VQ-GAN encoder in the cell membrane tearing task, and its efficient feature extraction and representation capabilities give it significant advantages in dealing with complex tasks.
[0066] Finally, during the autonomous operation of the robot, video frames are obtained in real time and the above processes of encoding, feature extraction, action prediction, and constraint verification are repeated, and operations are performed according to the safe action constraint conditions to ensure the safety and accuracy of the robot operation. During the experiment, the safety constraint threshold is set to 0.8, and the smoothing parameter is set to 0.2. As Figure 5 shown, it shows that the robot-controlled end effector completed the cell membrane tearing process of zebrafish embryos four times. The robot successfully completed 4 tests of cell membrane tearing operations on zebrafish embryos, and the embryos after membrane tearing were intact, proving the effectiveness of the robot cell micromanipulation method based on the imitation learning with a safety constraint mechanism.
[0067] To verify the superiority of the proposed method, it was compared with existing imitation learning methods, including Behavioral Cloning (BC), Action Chunking with Transformers (ACT), Visual Imitation Learning with Neural Networks (VINN), and Diffusion Models (Diffusion). Each method was tested 30 times, and the success rates of sub-tasks, average success rates, and final success rates of the cell membrane tearing task were statistically analyzed. The specific data are as follows: The success rate of the BC method in the cell pushing task was 93.3% (28 / 30), the success rate of grasping the cell was 32.1% (9 / 28), the success rate of cell membrane tearing was 66.7% (6 / 9), the average success rate was 64.0%, and the final success rate was 20.0% (6 / 30); The success rate of the ACT method in the cell pushing task was 96.7% (29 / 30), the success rate of grasping the cell was 58.6% (17 / 29), the success rate of cell membrane tearing was 76.5% (13 / 17), the average success rate was 77.2%, and the final success rate was 43.3% (13 / 30); The success rate of the VINN method in the cell pushing task was 96.7% (29 / 30), the success rate of grasping the cell was 31.0% (9 / 29), the success rate of cell membrane tearing was 22.2% (2 / 9), the average success rate was 50.0%, and the final success rate was 6.67% (2 / 30); The success rate of the Diffusion method in the cell pushing task was 76.7% (23 / 30), the success rate of grasping the cell was 26.1% (6 / 23), the success rate of cell membrane tearing was 0 (0 / 6), the average success rate was 34.3%, and the final success rate was 0 (0 / 30); The success rate of the proposed method in the cell pushing task was 96.7% (29 / 30), the success rate of grasping the cell was 82.8% (24 / 29), the success rate of cell membrane tearing was 79.2% (19 / 24), the average success rate was 86.2%, and the final success rate was 63.3% (19 / 30). These data indicate that the proposed method is significantly superior to existing methods in all key indicators. Its success rate in the cell pushing task is comparable to that of ACT and VINN, but its success rates in the cell grasping and cell membrane tearing tasks are significantly higher than those of other methods, and both the average success rate and the final success rate are much higher than those of other methods, fully demonstrating its excellent performance, high efficiency, and robustness in the cell membrane tearing task, and showing its significant technical advantages in complex task processing.
[0068] Finally, it should be noted that the above embodiments and descriptions are only used to illustrate the technical solutions of the present invention and not to limit them. Those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced. Without departing from the spirit and scope of the disclosure of the technical solutions of the present invention, they should all be covered by the protection scope of the claims of the present invention.
Claims
1. A robot cell micromanipulation method based on imitation learning with a safety constraint mechanism, characterized in that Including: Step 1: Obtain the teaching video of robot cell micromanipulation and extract video frames as the training set to train the image generation network and the attention network including the reconstruction and skill information distribution loss function, obtain the encoders of the trained image generation network and attention network, and extract the discrete latent codes of each teaching video frame; Step 2: Jointly construct the robot cell operation model with the encoders of the trained image generation network, attention network and multi-layer perceptron, and construct the safety action constraints of the robot cell operation model according to each discrete latent code; Train the robot cell operation model through the video frames of the teaching video and the action trajectories during the robot cell micromanipulation in the teaching video until the loss function of the robot cell operation model converges to obtain the trained robot cell operation model; Step 3: When the robot performs autonomous cell micromanipulation, obtain the video frames during the robot's autonomous cell micromanipulation in real time and input them together with the robot's actions at the previous moment into the trained robot cell operation model, and output the robot's actions at the next moment under the condition of satisfying the safety action constraints, and then control the robot to execute the actions to achieve robot cell micromanipulation; In the said Step 2, the safety action constraints are specifically as follows: Among them, represents the log-likelihood; represents the predicted value of the discrete latent encoding at time , represents the discrete latent encoding from time to time represents the frame length; represents the threshold parameter of the safety constraint; represents the safety constraint value at time during the training process; represents the total number of samples of the latent discrete encoding of the demonstration video; represents the log-likelihood value of the latent discrete encoding.
2. The method for robotic cell micromanipulation based on imitation learning with a safety constraint mechanism according to claim 1, characterized in that: In the said Step 1, the teaching video of robot cell micromanipulation is the video of the robot performing micromanipulation on cells according to preset actions; The image generation network and the attention network respectively adopt vector quantization and generative adversarial network VQ-GAN and autoregressive Transformer network. After inputting the training set into the vector quantization and generative adversarial network VQ-GAN for training, the trained vector quantization and generative adversarial network VQ-GAN is obtained. At the same time, the encoder of the vector quantization and generative adversarial network VQ-GAN compresses each teaching video frame to obtain the discrete latent code z of each teaching video frame; Input each discrete latent code z into the autoregressive Transformer network for training, maximize the likelihood value of each discrete latent code z until the reconstruction and skill information distribution loss function converges to obtain the trained autoregressive Transformer network.
3. The robot cell micro-operation method based on imitation learning with a safety constraint mechanism according to claim 2, characterized in that: The autoregressive Transformer network optimizes the distribution of discrete latent codes based on maximum likelihood estimation, and the optimization objective function of the autoregressive Transformer network is as follows: Among them, represents the model parameters of the autoregressive Transformer network; represents the number of time steps of the demonstration video; represents the number of dimensions of the discrete latent encoding; represents the log-likelihood; represents the encoding value at the -th position in the discrete latent encoding at the -th moment, represents the encoding values at the 1st to the -th positions in the discrete latent encoding at the -th moment; represents the latent discrete encoding of the subsequence from to , represents the frame length in the subsequence.
4. The method for robotic cell micromanipulation based on imitation learning with a safety constraint mechanism according to claim 3, wherein: The reconstruction and skill information distribution loss function of the autoregressive Transformer network is as follows: Among them, represents the cross-entropy loss between the predicted value and the true value, that is, the reconstruction loss; and respectively represent the true value and the predicted value of the discrete latent code at the th moment; represents the predicted value of the coding value at the th moment at the th position in the discrete latent code, represents the predicted value of the coding values at the 1st to the th moments at the th positions in the discrete latent code.
5. The method for robotic cell micromanipulation based on imitation learning with a safety constraint mechanism according to claim 1, characterized in that: In the said Step 2, when training the robot cell operation model, input the video frames of the teaching video and the action trajectories during the robot cell micromanipulation in the teaching video into the robot cell operation model. The video frames of the teaching video are sequentially input into the encoders of the trained image generation network and attention network. After processing, the results are fused with the action trajectories during the robot cell micromanipulation in the teaching video, and then the fused results are input into the multi-layer perceptron for training. During training, the network parameters of the encoder of the attention network are frozen, and the image generation network and the multi-layer perceptron are trained.
6. The method for robotic cell micromanipulation based on imitation learning with a safety constraint mechanism according to claim 1, wherein: In the second step described above, the loss function of the robotic cell operation model is specifically as follows: Among them, represents the mean squared error; and respectively represent the true action in the action trajectory during robot cell micromanipulation at time and the reconstruction of the autoregressive Transformer network and the balance weight parameter of the skill information distribution loss function and its skill information loss, respectively.
7. The method for robot cell micromanipulation based on imitation learning with a safety constraint mechanism according to claim 1, characterized in that: In the third step described above, the robot action at the next moment currently output by the trained robot cell operation model is input into the controller of the robot, and the controller smooths the robot action at the next moment using a trajectory smoothing method, and then controls the robot to execute the smoothed action.
8. An electronic device, characterized in that, Including: A memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method according to any one of claims 1-7.
9. A computer-readable storage medium having program data stored thereon, characterized in that, When the described program data is executed by a processor, it implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Robot active learning method based on image input
CN109800864A
Microscopic visual servo control method based on deep learning
CN111239085A