Non-invasive retrieval method of semi-transparent flexible ultrathin brain slices based on multi-modal perception

CN122378717BActive Publication Date: 2026-09-11FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610673761.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-15
Publication Date
2026-09-11
Estimated Expiration
2046-05-15

AI Technical Summary

Technical Problem

一旦操作失败,机器无法自救或判断状态,导致整体流程的鲁棒性和成功率低下

Benefits of technology

本发明提供的基于多模态感知的半透明柔性超薄脑片无损捞取方法,将力控数据作为模型输入,使得动作生成算法打破了纯视觉的局限性,模型能够根据接触培养液或组织边缘的微小力反馈,自适应地调整下压深度和捞取速度,像人类手部一样展现出柔顺特性,从根本上防止了组织的机械性损伤。对于形态多变、难以用数学公式精确描述的半透明软组织,本发明利用模仿学习直接从数据中映射控制策略,降低了底层解析算法的开发难度,提高了系统的泛化能力。本发明利用端到端模型处理复杂的捞取动作,利用显式视觉模型进行结果验收,既保证了动作的流畅自然,又确保了流程的绝对可靠。针对半透明超薄脑片在液体中漂浮、易受流体扰动及表面张力影响的极端物理特性,本发明结合高频本体感知与力矩反馈,在视觉特征因液面折射而失效的临界接触阶段,依靠力觉变化自适应调整刚度,复刻了人类专家在克服表面张力时的柔顺拖曳手法,实现了超薄脑片提取的零破损率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122378717B_ABST
    Figure CN122378717B_ABST
Patent Text Reader

Abstract

The application discloses a translucent flexible ultrathin brain slice nondestructive fishing method based on multi-modal perception, and comprises the following steps: taking a camera image acquisition moment as a benchmark, aligning a global visual angle image, a wrist follow-up visual angle image, a mechanical arm joint position vector and an end contact force vector into a time-stamped unified multi-modal observation frame, and constructing a training data set with an expert demonstration future action sequence; training an action generation model, sampling a latent variable representing an operation strategy from the demonstration action through a first encoder, extracting multi-modal features of the current observation through a second encoder, and predicting a future action instruction sequence by a decoder combining the latent variable and the multi-modal features; in the running stage, inputting the aligned current observation frame into the model to output an action sequence, which is interpolated and smoothed to drive the mechanical arm to perform fishing; after the action sequence is executed, a target detection model is used to determine whether there is target tissue in the original container area, and if there is no residue, the operation is determined to be successful.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot intelligent control and embodied micromanipulation technology, and particularly relates to a non-destructive retrieval method for semi-transparent flexible ultrathin brain slices based on multimodal perception. Background Technology

[0002] In biomedical and neuroscience experiments, the acquisition and transfer of translucent, flexible biological tissues are typically performed manually by experienced researchers. Taking ultrathin brain slices in neuroscience experiments as an example, their thickness is usually on the order of 200 micrometers. They are extremely fragile and translucent, exhibiting very low contrast with the background when suspended in culture medium. Furthermore, surface fluctuations cause severe light refraction, posing a significant challenge to automated operations. While visual servoing-based automated transfer systems have emerged in recent years, existing technologies still have the following limitations.

[0003] First, modeling flexible contact is extremely difficult. Traditional analytical control methods struggle to establish accurate dynamic contact models for highly deformable, floating flexible biological tissues in liquids. At the moment the manipulator penetrates the liquid surface and retrieves the brain slice, it is subjected to nonlinear fluid disturbances and surface tension adsorption. Pure visual servoing algorithms are prone to losing target at critical contact points due to liquid surface refraction. Traditional rigid trajectory planning or robotic arms lacking high-frequency force feedback are highly susceptible to generating destructive contact forces even with minute positional errors, leading to brain slice tearing or cellular structural damage.

[0004] Secondly, pure visual imitation learning lacks tactile perception. Most existing end-to-end robot imitation learning models rely solely on visual images and the position of the robotic arm joints as input. When contacting flexible tissues, the complete lack of force perception makes them highly susceptible to generating excessive interaction forces at the moment of contact, leading to tissue rupture or severe damage. Because they cannot perceive the minute contact forces with liquid surfaces and tissues, robots cannot exhibit the compliant characteristics of a human hand.

[0005] Furthermore, a closed-loop verification mechanism is lacking. Traditional imitation learning often involves open-loop execution, where the model completes the task after outputting a sequence of actions once, lacking explicit verification of the execution results. Once an operation fails, the machine cannot recover or determine its state, resulting in low robustness and success rate of the overall process. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention proposes a non-destructive retrieval method for semi-transparent, flexible, ultrathin brain slices based on multimodal sensing, thereby resolving the issues present in the prior art.

[0007] In a first aspect, to achieve the above objectives, the present invention provides a method for non-destructive retrieval of semi-transparent, flexible, ultrathin brain slices based on multimodal sensing, comprising: Based on the time of image acquisition by the camera, the global view image, wrist follow-up view image, robotic arm joint position vector and end contact force vector are aligned to construct a multimodal observation frame with unified timestamps, and associated with the future action sequence taught by experts to form a training dataset; The action generation model is trained using the training dataset. The model samples latent variables representing the operation strategy from the taught actions through a first encoder, extracts features from the current multimodal observation frame through a second encoder, and combines the latent variables with the extracted multimodal features by a decoder to predict the action command sequence for multiple future time steps. During the operation phase, whenever the camera generates a new image frame, the aligned current multimodal observation frame is acquired, the trained action generation model is input, the predicted action sequence is output, and after interpolation and smoothing, the robotic arm is driven to perform the retrieval operation. After the robotic arm completes a sequence of actions, it acquires the current global view image and uses a target detection model to determine whether there is target tissue in the original container area. If there is no residue, the operation is considered successful.

[0008] Optionally, the process of constructing multimodal observation frames includes: Maintain a global reference clock. When a new frame image is generated by the global view camera or the wrist servo camera, record the millisecond-level timestamp of the frame. Extract the joint position vector and end contact force vector with the timestamp closest to that moment from the robot state circular buffer and package them into a data frame with a unified timestamp.

[0009] Optionally, the training process for the action generation model includes: The global view image and the wrist-following view image contained in the current multimodal observation frame in the dataset are respectively input into the pre-trained residual network to extract spatial feature maps, which are then flattened and dimension-upgraded through a linear projection layer to generate two sets of visual word sequences. The joint position vector and end contact force vector contained in the current multimodal observation frame are mapped to generate state lexes through a linear projection layer.

[0010] Optionally, the training process for the action generation model may also include: The future action sequence is mapped through a linear projection layer and then injected with sinusoidal absolute position encoding to generate a temporal action lexical sequence. The state lexical, a learnable classification marker lexical, and the temporal action lexical sequence are concatenated along the sequence length dimension and input into the first encoder; The output features of the classification marker lexical are extracted, the mean and logarithmic variance of the action latent space distribution are calculated, and the latent variables are obtained by sampling through reparameterization techniques.

[0011] Optionally, the training process for the action generation model may also include: The latent variables, the state lexical units, and the two sets of visual lexical unit sequences are concatenated along the sequence dimension to form a global multimodal input sequence, which is then input into the second encoder. The decoder combines the hidden state matrix output by the second encoder with a cross-attention mechanism to calculate and output the action instruction sequence for the next multiple time steps in parallel.

[0012] Optionally, the interpolation smoothing and execution process of the action instruction sequence includes: The inference frequency of the action generation model is synchronized with the frequency at which the camera generates new image frames. The predicted action sequence output by the model each time is weighted and averaged using time ensemble technology over the overlapping time steps of adjacent inference cycles to generate the target instruction at the current moment. The target command is sent to the underlying controller, which, through a high-frequency trajectory planner, interpolates and reconstructs the low-frequency target command into a high-frequency physical command stream under the condition of satisfying speed and acceleration constraints, thereby driving the robotic arm to move.

[0013] Optionally, the state closure verification process includes: After the action generation model has completed a sequence of actions, it sends an execution completion signal to the control state machine. The system then actively pauses the underlying motion and activates the global visual verification module. The pre-trained object detection model is called to process the current global view image, filter the bounding boxes that are classified as flexible tissue and have a confidence score greater than a preset threshold, and extract their geometric center pixel coordinates. Perform Boolean logic judgment. If the coordinates of the geometric center pixel are not within the preset original container region of interest, the operation is considered successful, and the next stage task instruction is issued.

[0014] Optionally, the state closure verification process also includes: After the placement task is performed, if the detected geometric center pixel coordinates are located within the preset target container region of interest, the tissue transfer process is considered to have successfully closed the loop. If the target tissue is still detected within the original container's region of interest, the operation is deemed a failure, the robotic arm's posture is frozen, and it returns to a safe point, triggering a retry mechanism.

[0015] Secondly, the present invention also provides a computer terminal device, comprising: One or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the non-destructive retrieval method for semi-transparent flexible ultrathin brain slices based on multimodal perception in the first aspect described above.

[0016] Thirdly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the steps of the non-destructive retrieval method for semi-transparent flexible ultrathin brain slices based on multimodal perception in the first aspect described above.

[0017] Compared with the prior art, the present invention has the following advantages and technical effects: This invention provides a non-destructive retrieval method for semi-transparent, flexible, ultra-thin brain slices based on multimodal perception. By using force control data as model input, the action generation algorithm overcomes the limitations of pure vision. The model can adaptively adjust the pressing depth and retrieval speed based on minute force feedback from contact with the culture medium or tissue edges, exhibiting compliant characteristics like a human hand and fundamentally preventing mechanical damage to the tissue. For semi-transparent soft tissues with varied morphologies that are difficult to describe precisely using mathematical formulas, this invention utilizes imitation learning to directly map control strategies from the data, reducing the development difficulty of the underlying analytical algorithm and improving the system's generalization ability. This invention uses an end-to-end model to handle complex retrieval actions and an explicit visual model for result acceptance, ensuring both smooth and natural actions and absolute reliability of the process. Addressing the extreme physical characteristics of translucent ultrathin brain slices—floating in liquids and susceptible to fluid disturbances and surface tension—this invention combines high-frequency proprioception and torque feedback. At the critical contact stage where visual features fail due to liquid surface refraction, stiffness is adaptively adjusted by changes in force perception. This replicates the compliant dragging technique used by human experts to overcome surface tension, achieving zero breakage rate in the extraction of ultrathin brain slices. Attached Figure Description

[0018] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a flowchart illustrating the non-destructive retrieval method for semi-transparent, flexible, ultrathin brain slices based on multimodal sensing, according to an embodiment of the present invention. Detailed Implementation

[0019] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0021] Example 1 like Figure 1 As shown, this embodiment provides a non-destructive retrieval method for semi-transparent, flexible, ultrathin brain slices based on multimodal sensing, including: Based on the time of image acquisition by the camera, the global view image, wrist follow-up view image, robotic arm joint position vector and end contact force vector are aligned to construct a multimodal observation frame with unified timestamps, and associated with the future action sequence taught by experts to form a training dataset; The action generation model is trained using the training dataset. The model samples latent variables representing the operation strategy from the taught actions through a first encoder, extracts features from the current multimodal observation frame through a second encoder, and combines the latent variables with the extracted multimodal features by a decoder to predict the action command sequence for multiple future time steps. During the operation phase, whenever the camera generates a new image frame, the aligned current multimodal observation frame is acquired, the trained action generation model is input, the predicted action sequence is output, and after interpolation and smoothing, the robotic arm is driven to perform the retrieval operation. After the robotic arm completes a sequence of actions, it acquires the current global view image and uses a target detection model to determine whether there is target tissue in the original container area. If there is no residue, the operation is considered successful.

[0022] Furthermore, the process of constructing multimodal observation frames includes: Maintain a global reference clock. When a new frame image is generated by the global view camera or the wrist servo camera, record the millisecond-level timestamp of the frame. Extract the joint position vector and end contact force vector with the timestamp closest to that moment from the robot state circular buffer and package them into a data frame with a unified timestamp.

[0023] Specifically, the implementation process of this embodiment includes: Hardware triggering and timestamp alignment mechanism: A multimodal perception module is constructed, consisting of a Realsense depth camera mounted on the wrist of a robotic arm with a fixed global viewpoint, and a control interface based on libfranka. The system maintains a global reference clock, recording a precise millisecond-level timestamp when the camera generates a new frame (30Hz). And immediately extract the timestamp closest from the robot's state circular buffer. High-frequency joint data and force data (1000Hz).

[0024] Furthermore, the training process for the action generation model includes: The global view image and the wrist-following view image contained in the current multimodal observation frame in the dataset are respectively input into the pre-trained residual network to extract spatial feature maps, which are then flattened and dimension-upgraded through a linear projection layer to generate two sets of visual word sequences. The joint position vector and end contact force vector contained in the current multimodal observation frame are mapped to generate state lexes through a linear projection layer.

[0025] Specifically, the implementation process of this embodiment includes: Visual feature space unfolding: Color images from two perspectives in the dataset are independently input into a pre-trained ResNet (e.g., ResNet-18) residual backbone network to extract two-dimensional spatial feature maps. Subsequently, they must be flattened and dimensionality increased by the corresponding linear projection layers to uniformly map them into two sets of independent 512-dimensional high-dimensional visual token sequences.

[0026] Joint state feature alignment: Reusing the joint state vector at the current time step The input is then fed into an independent linear projection layer to map and generate a joint state token with the same dimension (512 dimensions) as the visual token.

[0027] Furthermore, the training process for the action generation model also includes: The future action sequence is mapped through a linear projection layer and then injected with sinusoidal absolute position encoding to generate a temporal action lexical sequence. The state lexical, a learnable classification marker lexical, and the temporal action lexical sequence are concatenated along the sequence length dimension and input into the first encoder; The output features of the classification marker lexical are extracted, the mean and logarithmic variance of the action latent space distribution are calculated, and the latent variables are obtained by sampling through reparameterization techniques.

[0028] Specifically, the implementation process of this embodiment includes: Action frequency domain isometric sampling and feature dimensionality reduction: The expert teaching trajectory is a continuous temporal signal at the physical level. The system sets a fixed control reference frequency (e.g., 30Hz, corresponding to the sampling period). This involves performing equidistant divergence sampling on the continuous joint trajectory. A segment is captured from the current physical moment. To the Future The discrete action target points within the window are composed of dimensions. Timing action block Subsequently, the discrete sequence is input into a linear projection layer in the form of a multilayer perceptron (MLP), which maps and elevates the feature dimension to the unified hidden layer dimension of the model. (e.g., 512-dimensional), to obtain the basic action feature matrix.

[0029] Sinusoidal Position Embedding: Because the Transformer architecture itself lacks the ability to handle inductive biases related to temporal sequence, temporal dimension information needs to be injected into the basic action feature matrix. For each time step position in the action sequence... ( ) and specific feature dimension index The fixed absolute position code is calculated using the following injection formula: ; ; The basic action feature matrix is ​​added element-wise to the positional encoding matrix to generate a sequence of temporal action tokens (denoted as...). It internalizes the precise temporal dependencies of each action frame.

[0030] Cross-modal feature mapping and global sequence concatenation constitute: the current time step Joint observation state vector (Including multimodal physical feedback after filtering and downsampling) is mapped to the same dimension through another set of independent linear projection layers. State token (denoted as Simultaneously, a classification token (CLS Token, denoted as CLS Token) of the same dimension and with learnable parameters is initialized in the feature space. ), which is specifically used for subsequent extraction of global distribution features.

[0031] Finally, before being fed into the first Transformer encoder, the three tensors are concatenated along the sequence length dimension. The concatenated tensor forms the global input sequence. The total sequence length is .

[0032] Distribution sampling calculation: After concatenating action features and state features, the result is input into a Transformer encoder containing four layers of self-attention blocks. The output features of the CLS token are extracted, and the mean vector of the action latent space distribution is calculated through a linear projection layer. Sum of logarithmic variance : ; ; Subsequently, the reparameterization trick was applied to perform sampling calculations to obtain latent variables. : ; This latent variable It internalized the fine-tuning style and compliance strategies of experts when facing force feedback from flexible tissues.

[0033] Furthermore, the training process for the action generation model also includes: The latent variables, the state lexical units, and the two sets of visual lexical unit sequences are concatenated along the sequence dimension to form a global multimodal input sequence, which is then input into the second encoder. The decoder combines the hidden state matrix output by the second encoder with a cross-attention mechanism to calculate and output the action instruction sequence for the next multiple time steps in parallel.

[0034] Specifically, the implementation process of this embodiment includes: Action block prediction based on Transformer decoder: latent variables The joint state lexical units and two sets of visual lexical sequences are concatenated along the sequence dimension to form a global multimodal input sequence. This sequence is then input into a second Transformer encoder containing four layers of self-attention blocks.

[0035] Matrix operations through multi-head self-attention mechanism: ; The model implicitly calculates the coupling mapping relationship between "multi-view visual features", "high-frequency force-feedback", and "action style", and outputs a global multimodal latent state feature matrix.

[0036] Action block prediction based on Transformer decoder: A set of fixed-position embedding vectors representing the query time sequence is constructed and input into a Transformer decoder containing four layers of self-attention blocks. The decoder combines the hidden state matrix output by the second Transformer encoder with a cross-attention mechanism to compute the output of future continuous... The predicted action sequence for each time step serves as the actual block of action instructions issued.

[0037] Furthermore, the interpolation smoothing and execution process of the action instruction sequence includes: The inference frequency of the action generation model is synchronized with the frequency at which the camera generates new image frames. The predicted action sequence output by the model each time is weighted and averaged using time ensemble technology over the overlapping time steps of adjacent inference cycles to generate the target instruction at the current moment. The target command is sent to the underlying controller, which, through a high-frequency trajectory planner, interpolates and reconstructs the low-frequency target command into a high-frequency physical command stream under the condition of satisfying speed and acceleration constraints, thereby driving the robotic arm to move.

[0038] Specifically, the implementation process of this embodiment includes: The trained network is deployed on a physical system to establish an interactive relationship between "macroscopic low-frequency decision-making" and "microscopic high-frequency force control": Visual hardware frame-driven inference computation: The system's operating frequency is strictly synchronized with the output frame rate of the multi-view camera (e.g., 30fps). Whenever the camera hardware outputs a new frame, the system uses it as a trigger signal to instantly extract the corresponding underlying joint and force state data, stitch them together, and feed them into the ACT model. At this point, latent variables... By forcing the value to the mean, the model computes forward at the same frequency as the camera (i.e., 30Hz) and outputs the future. Predicted joint sequences for each step.

[0039] Low-level high-frequency interpolation execution: Low-frequency commands (30Hz) driven by the camera frame rate are sent to the second computing node (low-level cerebellum). The low-level node uses a high-frequency trajectory generation algorithm (such as the Rukcig library) to interpolate and reconstruct the low-frequency decision points of the camera into a continuous physical command stream of 1000Hz (1ms cycle) under the physical conditions that meet the maximum speed and acceleration constraints of the robotic arm, driving the robotic arm to complete compliant retrieval.

[0040] Real-time end-to-end inference and physical smoothing: During the automatic retrieval process, the higher brain inputs the current multimodal observation frame into the ACT model every 33ms (30Hz) to obtain the predicted action block.

[0041] The temporal ensemble technique is used to perform a weighted average of the predicted actions that overlap in adjacent time steps, generate the current target joint angle command, and send it to the underlying cerebellum.

[0042] After receiving the instruction, the lower cerebellum uses a high-frequency trajectory planner (such as Ruckig) to limit the speed, acceleration, and jerk (Jerk), resampling the coarse 30Hz instruction into an absolutely smooth 1000Hz trajectory, driving the robotic arm to complete a highly compliant retrieval action.

[0043] Furthermore, the state closure verification process includes: After the action generation model has completed a sequence of actions, it sends an execution completion signal to the control state machine. The system then actively pauses the underlying motion and activates the global visual verification module. The pre-trained object detection model is called to process the current global view image, filter the bounding boxes that are classified as flexible tissue and have a confidence score greater than a preset threshold, and extract their geometric center pixel coordinates. Perform Boolean logic judgment. If the coordinates of the geometric center pixel are not within the preset original container region of interest, the operation is considered successful, and the next stage task instruction is issued.

[0044] Furthermore, the state closure verification process also includes: After the placement task is performed, if the detected geometric center pixel coordinates are located within the preset target container region of interest, the tissue transfer process is considered to have successfully closed the loop. If the target tissue is still detected within the original container's region of interest, the operation is deemed a failure, the robotic arm's posture is frozen, and it returns to a safe point, triggering a retry mechanism.

[0045] Specifically, the implementation process of this embodiment includes: To verify the results of end-to-end black-box execution, the following white-box closed-loop verification is designed: Verification trigger conditions and coordinate definitions: After the ACT model completes a long sequence of actions, it sends a completion signal to the control state machine. The system actively pauses the underlying motion and activates the global vision verification module. The transformation matrix from the camera to the robotic arm base is pre-observed through hand-eye calibration, and the region of interest in the original petri dish is quantized and defined in pixel coordinates. Region of interest in the target culture dish .

[0046] Object detection and state machine decision calculation: The pre-trained YOLO object detection model is used to process the current visual image. Bounding boxes with the category "flexible tissue" and a confidence score of >0.5 are selected, and their geometric center pixel coordinates are extracted. .

[0047] After performing the PICK task: Perform a Boolean logic check; if detected... If there is no residue in the original container, the current action block is considered to have been executed successfully, and the next stage task instruction is issued; otherwise, it is considered to have failed, the robotic arm posture is frozen, it returns to the safe point, and a retry is triggered.

[0048] After performing the PLACE task: if it is determined This confirms that the organizational transfer process has been successfully closed.

[0049] Example 2 In this embodiment, a computer terminal device is provided, including: One or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the above-described method for non-destructive retrieval of semi-transparent flexible ultrathin brain slices based on multimodal perception.

[0050] In this embodiment, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-described method for non-destructive retrieval of semi-transparent flexible ultrathin brain slices based on multimodal perception.

[0051] An application example of this invention: Hardware platform and communication deployment: Physical equipment: A robotic arm with 7 degrees of freedom (such as a robotic arm equipped with a torque sensor) is used, with a manipulator (such as a brush) installed at the end for retrieving flexible tissue. Visual acquisition is performed using a Real-Sense depth camera (resolution set at 640x480, frame rate 30fps).

[0052] Computing architecture: A dual-machine asynchronous architecture is adopted. The first computing node (high-level brain) is responsible for running the ACT model and YOLO vision model, with a control cycle of 30Hz; the second computing node (lower cerebellum) communicates with the robotic arm controller through a real-time kernel, with a control cycle of 1000Hz (1ms), and is responsible for high-frequency smoothing of instructions and reporting of low-level sensor data.

[0053] Acquisition and alignment of multimodal data: During the expert teaching (such as remotely retrieving brain slices) and real-time reasoning phases, the system simultaneously collects and aligns data from the following three modalities: Visual image data: Extract the color image stream from the camera.

[0054] Body perception data: Read the current angles and angular velocities of the robotic arm's seven joints.

[0055] Force feedback data (core feature): Directly read the end-effector force and torque data estimated by the robot arm's underlying controller, i.e., the six-dimensional end-effector force / torque vector (e.g., O_F_ext_hat_K).

[0056] Timestamp alignment: Since the refresh rate of the underlying torque and the body data (1000Hz) is much higher than the camera frame rate (30fps), the system uses the camera's acquisition timestamp as a reference to downsample or perform moving average filtering on the underlying high-frequency torque data, and strictly aligns and packages the three into a multimodal observation frame.

[0057] Construction and training of an ACT model integrating force perception: Feature Encoder: Visual images are processed by the ResNet-18 backbone network to extract high-dimensional spatial features.

[0058] The proprioceptive data and force feedback data are mapped into one-dimensional latent feature vectors by a multilayer perceptron (MLP).

[0059] The three features mentioned above are concatenated with positional encoding and used as a token input to the Transformer encoder. At this point, the network not only learns spatial location, but also establishes an implicit mapping relationship between "specific visual location" and "specific contact torque" (i.e., learns to remain compliant when in contact with liquid surfaces or tissues).

[0060] Action prediction (Decoder): The Transformer decoder predicts action chunks for the next dozens of steps based on the current observation token, such as predicting the absolute joint angle target sequence for the next 50 steps (lasting 1 second).

[0061] Real-time end-to-end inference and physical smoothing: During the automatic retrieval process, the higher brain inputs the current multimodal observation frame into the ACT model every 33ms (30Hz) to obtain the predicted action block.

[0062] The temporal ensemble technique is used to perform a weighted average of the predicted actions that overlap in adjacent time steps, generate the current target joint angle command, and send it to the underlying cerebellum.

[0063] After receiving the instruction, the lower cerebellum uses a high-frequency trajectory planner (such as Ruckig) to limit the speed, acceleration, and jerk (Jerk), resampling the coarse 30Hz instruction into an absolutely smooth 1000Hz trajectory, driving the robotic arm to complete a highly compliant retrieval action.

[0064] Introducing YOLO's state closure verification: Once the ACT model has completed the "Pick" action primitive, the system pauses end-to-end action output.

[0065] The pre-trained YOLO object detection model is invoked to infer the region of interest (ROI) of the original petri dish in the current camera image.

[0066] Judgment logic: If no target bounding box of category "flexible tissue (Sample)" with a confidence level greater than 0.5 is detected within the ROI area, the tissue is determined to have been successfully retrieved, and the subsequent transfer task continues; if the target is still detected, the retrieval is determined to have failed, the robotic arm returns to the safe observation point, and the retry mechanism is triggered.

[0067] This invention provides a non-destructive retrieval method for semi-transparent, flexible, ultra-thin brain slices based on multimodal perception. By using force control data as model input, the action generation algorithm overcomes the limitations of pure vision. The model can adaptively adjust the pressing depth and retrieval speed based on minute force feedback from contact with the culture medium or tissue edges, exhibiting compliant characteristics like a human hand and fundamentally preventing mechanical damage to the tissue. For semi-transparent soft tissues with varied morphologies that are difficult to describe precisely using mathematical formulas, this invention utilizes imitation learning to directly map control strategies from the data, reducing the development difficulty of the underlying analytical algorithm and improving the system's generalization ability. This invention uses an end-to-end model to handle complex retrieval actions and an explicit visual model for result acceptance, ensuring both smooth and natural actions and absolute reliability of the process. Addressing the extreme physical characteristics of translucent ultrathin brain slices—floating in liquids and susceptible to fluid disturbances and surface tension—this invention combines high-frequency proprioception and torque feedback. At the critical contact stage where visual features fail due to liquid surface refraction, stiffness is adaptively adjusted by changes in force perception. This replicates the compliant dragging technique used by human experts to overcome surface tension, achieving zero breakage rate in the extraction of ultrathin brain slices.

[0068] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for non-destructive retrieval of a translucent flexible ultrathin brain slice based on multi-modal perception, characterized in that, include: Based on the time of image acquisition by the camera, the global view image, wrist follow-up view image, robotic arm joint position vector and end contact force vector are aligned to construct a multimodal observation frame with unified timestamps, and associated with the future action sequence taught by experts to form a training dataset; The action generation model is trained using the training dataset. The model samples latent variables representing the operation strategy from the taught actions through a first encoder, extracts features from the current multimodal observation frame through a second encoder, and combines the latent variables with the extracted multimodal features by a decoder to predict the action command sequence for multiple future time steps. During the operation phase, whenever the camera generates a new image frame, the aligned current multimodal observation frame is acquired, the trained action generation model is input, the predicted action sequence is output, and after interpolation and smoothing, the robotic arm is driven to perform the retrieval operation. After the robotic arm completes a sequence of actions, it acquires the current global view image and uses a target detection model to determine whether there is target tissue in the original container area. If there is no residue, the operation is considered successful.

2. The method according to claim 1, characterized in that, The process of constructing multimodal observation frames includes: Maintain a global reference clock. When a new frame image is generated by the global view camera or the wrist servo camera, record the millisecond-level timestamp of the frame. Extract the joint position vector and end contact force vector with the timestamp closest to that moment from the robot state circular buffer and package them into a data frame with a unified timestamp.

3. The method according to claim 1, characterized in that, The training process of the action generation model includes: The global view image and the wrist-following view image contained in the current multimodal observation frame in the dataset are respectively input into the pre-trained residual network to extract spatial feature maps, which are then flattened and dimension-upgraded through a linear projection layer to generate two sets of visual word sequences. The joint position vector and end contact force vector contained in the current multimodal observation frame are mapped to generate state lexes through a linear projection layer.

4. The method according to claim 3, characterized in that, The training process for the action generation model also includes: The future action sequence is mapped through a linear projection layer and then injected with sinusoidal absolute position encoding to generate a temporal action lexical sequence. The state lexical, a learnable classification marker lexical, and the temporal action lexical sequence are concatenated along the sequence length dimension and input into the first encoder; The output features of the classification marker lexical are extracted, the mean and logarithmic variance of the action latent space distribution are calculated, and the latent variables are obtained by sampling through reparameterization techniques.

5. The method according to claim 4, characterized in that, The training process for the action generation model also includes: The latent variables, the state lexical units, and the two sets of visual lexical unit sequences are concatenated along the sequence dimension to form a global multimodal input sequence, which is then input into the second encoder. The decoder combines the hidden state matrix output by the second encoder with a cross-attention mechanism to calculate and output the action instruction sequence for the next multiple time steps in parallel.

6. The method according to claim 1, characterized in that, The interpolation smoothing and execution process of the action command sequence includes: The inference frequency of the action generation model is synchronized with the frequency at which the camera generates new image frames. The predicted action sequence output by the model each time is weighted and averaged using time ensemble technology over the overlapping time steps of adjacent inference cycles to generate the target instruction at the current moment. The target command is sent to the underlying controller, which, through a high-frequency trajectory planner, interpolates and reconstructs the low-frequency target command into a high-frequency physical command stream under the condition of satisfying speed and acceleration constraints, thereby driving the robotic arm to move.

7. The method according to claim 1, characterized in that, The process of state closure verification includes: After the action generation model has completed a sequence of actions, it sends an execution completion signal to the control state machine. The system then actively pauses the underlying motion and activates the global visual verification module. The pre-trained object detection model is called to process the current global view image, filter the bounding boxes that are classified as flexible tissue and have a confidence score greater than a preset threshold, and extract their geometric center pixel coordinates. Perform Boolean logic judgment. If the coordinates of the geometric center pixel are not within the preset original container region of interest, the operation is considered successful, and the next stage task instruction is issued.

8. The method according to claim 7, characterized in that, The process of state closure verification also includes: After the placement task is performed, if the detected geometric center pixel coordinates are located within the preset target container region of interest, the tissue transfer process is considered to have successfully closed the loop. If the target tissue is still detected within the original container's region of interest, the operation is deemed a failure, the robotic arm's posture is frozen, and it returns to a safe point, triggering a retry mechanism.

9. A computer terminal device, characterized by include: One or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors perform the steps of the method as described in any one of claims 1-8.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Soft manipulator and use method thereof

    CN116330250A

  • Mechanical arm motion control method based on multi-agent cooperation

    CN120620234A