High-precision body intelligent mechanical arm combined black light laboratory sample pretreatment system

CN122775889APending Publication Date: 2026-09-18SICHUAN ACADEMY OF MEDICAL SCI SICHUAN PROVINCIAL PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610624691.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-08
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0003]上述现有技术在完全无光且样本位姿存在偏移的场景中,红外对射传感器仅能提供一维的到位开关量,无法获取样本容器在三维空间中的具体位姿偏差,机械臂按固定坐标下压或夹取时容易与发生偏移的样本发生刚性碰撞;微光夜视方案在无可见光环境下严重依赖图像增强算法引入大量噪点并产生处理延迟,无法满足高精度对准的实时要求;同时,传统位置与速度控制律缺乏对末端接触力与环境刚度的感知能力,在抓取脆弱样本管或执行开盖微操作时,一旦存在微小坐标偏差,机械臂刚性执行动作会直接导致样本容器破裂

Benefits of technology

[0017] 1. This solution uses an active infrared structured light projector to project coded light spots and combines them with a binocular infrared camera for stereo matching to reconstruct the 3D point cloud contour of the sample container in a no-visible-light environment. A six-dimensional torque sensor and an array of tendon-type tactile sensors deployed at the end of the robotic arm acquire the contact area, local normal stress, and pure six-dimensional interactive torque. The embodied intelligent control platform's embedded physics engine, a large-scale embodied model, fuses the 3D point cloud contour with the temporal data input to a graph neural network, outputting the desired impedance parameters and corrected trajectory. After coarse positioning, the robotic arm switches to an impedance control mode based on temporal data feedforward. The physics engine calculates the angular velocity sequence that satisfies the contact force threshold based on the penalty function contact mechanics constraints and outputs the diagonal stiffness matrix and diagonal damping matrix. This solves the problem of fragile sample damage caused by the lack of accurate 3D spatial perception and adaptive contact mechanics control under no-light conditions, achieving accurate 3D perception and high-precision non-destructive pre-processing micro-operations for pose-shifted samples in a no-light environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122775889A_ABST
    Figure CN122775889A_ABST
Patent Text Reader

Abstract

This invention relates to the field of robotic arms and materials testing and analysis technology, specifically to a light-out laboratory sample preprocessing system combining a high-precision embodied intelligent robotic arm. The system includes a light-out sensing cluster, an embodied intelligent control platform, and a high-precision robotic arm actuator. An active infrared structured light projector and an infrared camera reconstruct the 3D point cloud contour of the sample container in a light-out environment. A six-dimensional torque sensor and an array of tendon-type tactile sensors collect contact mechanics time-series data. The embodied large model fuses the 3D point cloud contour and time-series data using a graph neural network, outputting the desired impedance parameters and corrected trajectory. After coarse positioning, the robotic arm switches to an impedance control mode based on time-series data feedforward to perform preprocessing operations. This solution achieves accurate 3D sensing and high-precision non-destructive micro-manipulation of samples with pose displacement in a light-out environment, avoiding sample damage caused by rigid collisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robotic arms and materials testing and analysis technology, specifically to a light-out laboratory sample pretreatment system that incorporates a high-precision, embodied intelligent robotic arm. Background Technology

[0002] In a darkened laboratory, ambient light sources are switched off to prevent degradation of photosensitive samples or light pollution. In this environment, automated sample pretreatment primarily relies on a robotic arm with a pre-defined trajectory, coupled with an infrared beam sensor, or guided by a low-light night vision camera. The infrared beam sensor has transmitters and receivers positioned around the robotic arm's path and the sample station. When a sample bottle is placed in position and blocks the light path, a switching signal is generated. The robotic arm then uses this signal to perform the grasping action according to a pre-stored fixed coordinate sequence. The low-light night vision solution acquires images under low-light conditions, enhances contrast using image enhancement algorithms, extracts sample edge contours, and calculates coordinates to guide the robotic arm's movement. The robotic arm's underlying layer typically employs position and velocity control laws, generating drive commands based on calculated coordinate deviations.

[0003] In scenarios with complete darkness and sample pose deviation, the aforementioned existing technologies can only provide one-dimensional position switching signals from infrared photoelectric sensors, failing to acquire the specific pose deviation of the sample container in three-dimensional space. When the robotic arm presses down or grips at fixed coordinates, it is prone to rigid collisions with the shifted sample. Low-light night vision solutions heavily rely on image enhancement algorithms in the absence of visible light, introducing significant noise and processing delays, failing to meet the real-time requirements of high-precision alignment. Simultaneously, traditional position and velocity control laws lack the ability to sense end-effector contact forces and environmental stiffness. When grasping fragile sample tubes or performing micro-operations like opening lids, even a slight coordinate deviation can cause the robotic arm's rigid execution to directly lead to sample container breakage. The core technical problem arising from this is that in a dark laboratory environment, conventional solutions, lacking precise three-dimensional spatial perception and adaptive contact mechanics control mechanisms, prevent the robotic arm from performing non-destructive, high-precision pre-processing micro-operations on fragile samples with pose deviations. Summary of the Invention

[0004] The purpose of this invention is to provide a light-out laboratory sample pretreatment system that incorporates a high-precision, embodied intelligent robotic arm, which can solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A light-less laboratory sample preprocessing system incorporating a high-precision embodied intelligent robotic arm includes a light-free sensing cluster, an embodied intelligent control platform, and a high-precision robotic arm actuator. The light-free sensing cluster comprises an active infrared structured light projector, an infrared camera, a six-dimensional torque sensor deployed at the end joint of the robotic arm, and an array of tendon-type tactile sensors deployed on the gripper base. The active infrared structured light projector projects coded light spots onto the sample area, and the infrared camera reconstructs the three-dimensional point cloud contour of the sample container based on the coded light spots. The embodied intelligent control platform deploys an embodied large-scale model with an embedded physics engine. The embodied large model receives sample pose coordinates provided by the 3D point cloud contour and time-series data collected by the six-dimensional torque sensor and the array-type tendon-type tactile sensor. It inputs the sample pose coordinates and the time-series data into a graph neural network for feature alignment and fusion, and outputs the expected impedance parameters and corrected trajectory of the robotic arm in the joint space. The high-precision robotic arm actuator performs coarse positioning based on the sample pose coordinates. After the gripper contacts the sample, it closes the position control mode and switches to the impedance control mode based on the time-series data feedforward. It performs preprocessing operations based on the expected impedance parameters and the corrected trajectory.

[0007] Preferably, the active infrared structured light projector in the lightless sensing cluster uses diffractive optical elements to modulate the infrared laser beam into a pseudo-random coded light spot matrix with a specific spatial arrangement structure; the infrared camera is a binocular infrared camera, which simultaneously acquires an infrared left view and an infrared right view containing the pseudo-random coded light spot matrix; the embodied intelligent control platform performs stereo matching on the infrared left view and the infrared right view based on epipolar constraints, extracts the disparity value of the matching pixels, calculates the depth information by combining the intrinsic and extrinsic parameter matrices of the binocular infrared camera, and generates a sparse three-dimensional point cloud containing only the sample container and the surrounding preset operation area after filtering out background depth noise, and reconstructs the three-dimensional point cloud contour of the sample container by surface fitting of the sparse three-dimensional point cloud.

[0008] Preferably, the array-type tendon-type tactile sensor includes a flexible substrate attached to the inner side of the gripper and a microelectromechanical system (MEMS) strain gauge array embedded in the flexible substrate in a grid-like distribution. The MEMS strain gauge array uses an analog switch matrix to time-division select each strain gauge node to obtain the resistance change of each node when the inner side of the gripper contacts the sample container. The embodied intelligent control platform converts the resistance change into a strain value matrix for each node. Based on the pre-calibrated elastic modulus and Poisson's ratio of the flexible substrate, the strain value matrix is ​​mapped to a contact area distribution matrix and a local normal stress distribution matrix of the contact area between the inner side of the gripper and the sample container. The contact area distribution matrix and the local normal stress distribution matrix are combined into a tactile feedback data stream in the time-series data.

[0009] Preferably, the six-dimensional torque sensor is rigidly coaxially connected between the output shaft of the robotic arm's end effector joint and the gripper base via a flange; the six-dimensional torque sensor contains three sets of orthogonally arranged Wheatstone full-bridge strain gauge circuits; the embodied intelligent control platform differentially amplifies and converts the analog voltage signals output by the three sets of orthogonally arranged Wheatstone full-bridge strain gauge circuits to obtain raw six-dimensional torque data containing three-dimensional force components and three-dimensional torque components; the embodied intelligent control platform performs gravity compensation and end effector inertial force decoupling calculations on the raw six-dimensional torque data based on the forward kinematics Jacobian matrix of the robotic arm's end effector joint to generate pure six-dimensional interactive torque data free from static interference, and uses the pure six-dimensional interactive torque data as the torque feedback data stream in the time series data.

[0010] Preferably, the graph neural network in the embodied large model constructs the sample pose coordinates into spatial location nodes, and aligns the tactile feedback data stream and torque feedback data stream in the temporal data according to timestamps to construct a first temporal feature node and a second temporal feature node, respectively; the graph neural network establishes undirected edge connections based on the spatiotemporal correlation of the spatial location nodes, the first temporal feature node, and the second temporal feature node within a preset time step; the graph neural network performs message passing on the undirected edge connections through graph convolutional layers, maps the point cloud geometric features of the spatial location nodes, the contact mechanical features of the first temporal feature node, and the six-dimensional torque features of the second temporal feature node to a latent space feature vector of the same dimension, and performs a concatenation operation on the latent space feature vector to generate a fused multimodal feature tensor.

[0011] Preferably, the physics engine embedded in the embodied large model stores a three-dimensional rigid body model of the sample container and a flexible body model of the gripper; the physics engine receives the contact area distribution information in the fused multimodal feature tensor and establishes a penalty function contact mechanical constraint between the flexible body model and the three-dimensional rigid body model; the physics engine performs forward dynamic integration calculation based on the penalty function contact mechanical constraint and the real-time angle of each joint of the current robotic arm to solve for the angular velocity sequence of each joint that satisfies the preset contact force threshold; the embodied large model inputs the angular velocity sequence of each joint to the policy network, and the policy network outputs the diagonal stiffness matrix and diagonal damping matrix in the desired impedance parameters, and generates the corrected trajectory by combining the angular velocity sequence of each joint.

[0012] Preferably, before performing stereo matching of the infrared left view and the infrared right view based on epipolar constraints, the embodied intelligent control platform performs adaptive local histogram equalization on the infrared left view and the infrared right view to eliminate grayscale shifts caused by differences in infrared thermal radiation. During stereo matching, the embodied intelligent control platform uses a cost aggregation algorithm based on semi-global matching to calculate the initial disparity value of the matching pixels, performs a left-right consistency check on the initial disparity value to remove mismatched pixels in occluded areas, performs median filtering on the retained initial disparity value to remove isolated noise points, generates the depth information based on the filtered disparity value, and downsamples the depth information based on a preset spatial voxel grid to fix the point cloud density of the sparse 3D point cloud.

[0013] Preferably, each strain gauge node in the microelectromechanical system strain gauge array is connected to the input terminal of the analog switch matrix via row and column scan lines; the output terminal of the analog switch matrix is ​​connected to the differential input terminal of the instrumentation amplifier, and the reference voltage terminal of the instrumentation amplifier is connected to the output terminal of the digital-to-analog converter; when the embodied intelligent control platform selects each strain gauge node in a time-division manner, it synchronously controls the digital-to-analog converter to output a common-mode bias voltage corresponding to the currently selected node to the reference voltage terminal of the instrumentation amplifier; the instrumentation amplifier differentially amplifies the weak voltage signal corresponding to the resistance change of the strain gauge node and the common-mode bias voltage to eliminate the common-mode interference voltage generated by the flexible substrate in the bending state, and outputs a differentially amplified voltage signal that only reflects the normal deformation of the strain gauge node.

[0014] Preferably, when performing gravity compensation, the embodied intelligent control platform retrieves the mass parameters and center-of-gravity position parameters of each link of the robotic arm from the storage, and calculates the static gravitational torque vector generated by each link on the end joint under the gravitational field using the recursive Newton-Euler algorithm based on the real-time angle of each joint. When performing inertial force decoupling calculation, the embodied intelligent control platform obtains the joint angular velocity and joint angular acceleration fed back by the encoders installed at each joint motor, constructs the joint space mass matrix of the robotic arm by combining the mass parameters and center-of-gravity position parameters, calculates the inertial force vector by multiplying the joint space mass matrix by the joint angular acceleration, and generates the pure six-dimensional interactive torque data by subtracting the static gravitational torque vector and the inertial force vector from the original six-dimensional torque data.

[0015] Preferably, the gripper flexible body model in the physics engine is represented by a spring-mass mesh model. The spring-mass mesh model includes multiple mass points distributed on the geometric surface of the gripper and structural springs, shear springs, and bending springs connecting adjacent mass points. When establishing the penalty function contact mechanical constraints, the physics engine detects the penetration depth between the mass points in the spring-mass mesh model and the surface of the three-dimensional rigid body model of the sample container in real time. When a penetration depth exists, the normal rebound force and tangential friction force on the mass point are calculated based on the penetration depth and a preset penalty stiffness coefficient. The normal rebound force and the tangential friction force are applied as external loads to the spring-mass mesh model. The spatial position coordinates and velocity vectors of each mass point in the spring-mass mesh model are updated according to the explicit Euler integral method.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0017] 1. This solution uses an active infrared structured light projector to project coded light spots and combines them with a binocular infrared camera for stereo matching to reconstruct the 3D point cloud contour of the sample container in a no-visible-light environment. A six-dimensional torque sensor and an array of tendon-type tactile sensors deployed at the end of the robotic arm acquire the contact area, local normal stress, and pure six-dimensional interactive torque. The embodied intelligent control platform's embedded physics engine, a large-scale embodied model, fuses the 3D point cloud contour with the temporal data input to a graph neural network, outputting the desired impedance parameters and corrected trajectory. After coarse positioning, the robotic arm switches to an impedance control mode based on temporal data feedforward. The physics engine calculates the angular velocity sequence that satisfies the contact force threshold based on the penalty function contact mechanics constraints and outputs the diagonal stiffness matrix and diagonal damping matrix. This solves the problem of fragile sample damage caused by the lack of accurate 3D spatial perception and adaptive contact mechanics control under no-light conditions, achieving accurate 3D perception and high-precision non-destructive pre-processing micro-operations for pose-shifted samples in a no-light environment.

[0018] 2. In the graph neural network processing, this scheme constructs spatial location nodes from sample pose coordinates and establishes temporal feature nodes from tactile feedback data streams and torque feedback data streams based on timestamps, creating undirected edge connections. Graph convolutional layers are used for message passing to generate fused multimodal feature tensors, improving the feature alignment accuracy of multimodal heterogeneous data. In the infrared sensing stage, adaptive local histogram equalization and semi-global matching cost aggregation combined with left-right consistency checks and median filtering fix the point cloud density of the sparse 3D point cloud, eliminating matching interference caused by infrared thermal radiation and occlusion. In the tactile... In the torque acquisition stage, a common-mode bias voltage is provided to the reference terminal of the instrumentation amplifier through a digital-to-analog converter to eliminate common-mode interference from the bending of the flexible substrate. Static gravity and inertial force are decoupled by a recursive Newton-Euler algorithm and a joint space mass matrix, eliminating the coupling interference of the robot arm's motion state on the contact mechanics perception. In the flexible body modeling of the physics engine, a spring-mass mesh model including structural springs, shear springs, and bending springs is adopted. The normal rebound force and tangential friction force are calculated based on the penetration depth and penalty stiffness coefficient, which improves the physical fidelity of the calculation of local deformation during the contact between the gripper and the sample container. Attached Figure Description

[0019] Figure 1 A flowchart illustrating the overall system workflow provided in this embodiment of the invention;

[0020] Figure 2 A flowchart illustrating the 3D point cloud reconstruction process of a sample container provided in an embodiment of the present invention;

[0021] Figure 3 This is a flowchart of the multimodal sensing signal processing provided in an embodiment of the present invention;

[0022] Figure 4 A flowchart of the graph neural network multimodal feature fusion process provided in an embodiment of the present invention;

[0023] Figure 5 A flowchart of the physical engine contact mechanics calculation process provided in an embodiment of the present invention;

[0024] Figure 6 A flowchart illustrating the switching and execution process of the robotic arm control mode provided in an embodiment of the present invention. Detailed Implementation

[0025] refer to Figure 1In one embodiment, a light-less laboratory sample preprocessing system incorporating a high-precision embodied intelligent robotic arm includes a light-free sensing cluster, an embodied intelligent control platform, and a high-precision robotic arm actuator. The light-free sensing cluster includes an active infrared structured light projector, an infrared camera, a six-dimensional torque sensor deployed at the end joint of the robotic arm, and an array of tendon-type tactile sensors deployed on the gripper base. The active infrared structured light projector projects coded light spots onto the sample area, and the infrared camera reconstructs the three-dimensional point cloud contour of the sample container based on the coded light spots. The embodied intelligent control platform deploys an embodied large model with an embedded physics engine. The embodied large model receives the sample pose coordinates provided by the three-dimensional point cloud contour and the temporal data collected by the six-dimensional torque sensor and the array of tendon-type tactile sensors. It inputs the sample pose coordinates and temporal data into a graph neural network for feature alignment and fusion, outputting the desired impedance parameters and corrected trajectory of the robotic arm in joint space. The high-precision robotic arm actuator performs coarse positioning based on the sample's pose coordinates. After the gripper contacts the sample, the position control mode is turned off and switched to the impedance control mode based on time-series data feedforward. Pre-processing operations are performed based on the desired impedance parameters and the corrected trajectory.

[0026] Specifically, the projection light path of the active infrared structured light projector and the imaging light path of the infrared camera are arranged along the same optical axis, with the projection field of view and the imaging field of view completely overlapping. The sample area is within the overlapping working distance range of the projection field of view and the imaging field of view. The active infrared structured light projector continuously projects coded light spots onto the sample area, covering the placement position of the sample container and the surrounding preset operating area. The infrared camera synchronously acquires infrared images of the sample area with projected coded light spots at a fixed frame rate. The acquired infrared images are decoded and stereo matching calculations are performed to extract the three-dimensional spatial coordinate set of the sample container surface, generating a three-dimensional point cloud contour of the sample container. The three-dimensional point cloud contour includes the geometric shape features of the sample container and its pose coordinates in the world coordinate system. The pose coordinates include the centroid spatial coordinates and attitude angles of the sample container, providing spatial positioning basis for motion planning of the high-precision robotic arm actuator.

[0027] Furthermore, a six-dimensional torque sensor is deployed between the output shaft of the robotic arm's end-effector joint and the gripper base, maintaining a rigid coaxial connection with both. This sensor acquires six-dimensional mechanical signals in real-time during the interaction between the robotic arm's end-effector and the environment. An array-type tendon-shaped tactile sensor is fitted onto the inner gripping surface of the gripper, moving synchronously with the opening and closing motion of the gripper. When the gripper contacts the sample container, it acquires real-time deformation and stress distribution signals in the contact area. The six-dimensional torque sensor and the array-type tendon-shaped tactile sensor output signals at the same sampling frequency. The integrated intelligent control platform marks each set of acquired signals with a corresponding timestamp, combining the six-dimensional mechanical signal and tactile stress signal at the same timestamp into time-series data. This time-series data includes tactile feedback data streams and torque feedback data streams within continuous time steps.

[0028] The embodied large model is deployed in the edge computing unit of the embodied intelligent control platform, embedding a physics engine and a graph neural network. The embodied large model receives sample pose coordinates obtained from 3D point cloud contour analysis, as well as temporal data within continuous time steps, and inputs these coordinates into the graph neural network. The graph neural network performs spatiotemporal feature alignment on the input multimodal heterogeneous data, mapping spatial pose features and temporal mechanical features to the same feature space, and performs feature fusion to generate a fused multimodal feature tensor. The embodied large model inputs the fused multimodal feature tensor into the embedded physics engine. Based on the contact mechanics information and sample pose information in the fused multimodal feature tensor, the physics engine constructs a dynamic constraint model of the contact scene, performs forward dynamics calculations, and solves for the motion parameter sequence of each joint of the robotic arm. The embodied large model inputs the solved motion parameter sequence into the policy network. Based on the dynamics calculation results, the policy network outputs the desired impedance parameters of the robotic arm joint space, and simultaneously generates a corrected trajectory for the robotic arm's end effector by combining the sample pose coordinates and contact mechanics information. The desired impedance parameters include the diagonal stiffness matrix and diagonal damping matrix of the joint space, and the corrected trajectory includes the desired angle sequence of each joint of the robotic arm in a continuous time step.

[0029] refer to Figure 6 The high-precision robotic arm actuator comprises a multi-degree-of-freedom serial robotic arm body, an end effector gripper, and a joint drive controller. The joint drive controller receives sample pose coordinates output from the embodied intelligent control platform, plans a coarse positioning trajectory for the robotic arm end effector based on these coordinates, and drives the robotic arm body to move the end effector gripper to a preset position above the sample container, completing the coarse positioning operation. During coarse positioning, the joint drive controller operates in position control mode, calculating the position deviation and generating joint drive current based on the planned trajectory and the real-time angle feedback from the joint encoder, controlling the robotic arm end effector to move along the planned trajectory. When the gripper contacts the sample container, the embodied intelligent control platform determines the establishment of the contact state based on the tactile feedback data stream and torque feedback data stream in the timing data, and sends a mode switching command to the joint drive controller. Upon receiving the mode switching command, the joint drive controller deactivates the position control mode and switches to an impedance control mode based on timing data feedforward. In impedance control mode, the joint drive controller receives the desired impedance parameters and corrected trajectory output by the embodied intelligent control platform. Combining the torque feedback data stream and tactile feedback data stream collected in real time from the time series data, it performs impedance control law calculation, generates joint drive commands, and controls the end gripper of the robotic arm to perform pre-processing operations such as sample grasping, capping, liquid transfer, and centrifuge tube sorting according to the corrected trajectory.

[0030] In this embodiment, the impedance control dynamic equation of the robotic arm joint space is:

[0031]

[0032] in, This represents the real-time angle vectors of each joint of the robotic arm. This represents the number of joint degrees of freedom of the robotic arm; This is the real-time angular velocity vector of each joint of the robotic arm; This represents the real-time angular acceleration vectors of each joint of the robotic arm; Let be the joint space inertia matrix of the robotic arm; The matrix of Coriolis force and centrifugal force of the robotic arm; Let be the gravitational torque vector of the robotic arm; This represents the output driving torque vector of the motors at each joint of the robotic arm; The interactive torque vector exerted by the external environment on the end effector of the robotic arm is obtained by mapping the pure six-dimensional interactive torque data collected by the six-dimensional torque sensor through the Jacobian matrix.

[0033] The tensor generation formula for multimodal feature fusion is:

[0034]

[0035] in, To fuse multimodal feature tensors, The dimension of the fusion feature; The point cloud geometric feature vector extracted from the sample pose coordinates. The dimension of the geometric feature; The contact mechanics feature vector extracted from the haptic feedback data stream; The six-dimensional torque feature vector extracted from the torque feedback data stream; This is a feature vector concatenation operation that concatenates three feature vectors with the same dimensions along the feature dimensions to form a fused multimodal feature tensor.

[0036] Table 1. Input / output parameters of each functional module in the sample pretreatment system for a dark laboratory.

[0037] Lightless sensing cluster Active infrared structured light projection unit Projection trigger signal, coded spot modulation parameters Pseudo-random coding of light spot projection in sample region Hard synchronization trigger with infrared camera frame rate Lightless sensing cluster Binocular infrared imaging and point cloud reconstruction unit Infrared left view, infrared right view, camera intrinsic and extrinsic parameter matrix 3D point cloud contour of sample container, sample pose coordinates Aligned by image frame timestamps Lightless sensing cluster Six-dimensional torque sensing unit Six-dimensional mechanical simulation signal of the end joint Raw six-dimensional torque data, pure six-dimensional interactive torque data Hard synchronization with the sampling frequency of the tactile sensing unit Lightless sensing cluster Array-type tendon-type tactile sensing unit Analog signal of strain gauge nodal resistance change Strain matrix, contact area distribution matrix, local normal stress distribution matrix Hard synchronization with the sampling frequency of the six-dimensional torque sensing unit Embossed Intelligent Control Platform Graph Neural Network Feature Fusion Unit Sample pose coordinates, haptic feedback data stream, torque feedback data stream Fusion of multimodal feature tensors Align multimodal data by timestamp Embossed Intelligent Control Platform Dynamics calculation unit with embedded physics engine Fusion of multimodal feature tensors and real-time joint angles of the robotic arm Joint angular velocity sequence and contact mechanical constraint parameters Synchronized with the joint controller sampling period Embossed Intelligent Control Platform Policy Network Control Unit Joint angular velocity sequences and fused multimodal feature tensors Joint space desired impedance parameters, end-effector correction trajectory Update output according to control cycle High-precision robotic arm actuator Joint drive and control unit Sample pose coordinates, desired impedance parameters, and corrected trajectory Joint drive current, robotic arm end effector motion control commands Dual-mode switching, closed-loop update according to control cycle

[0038] Specifically, the table clarifies the correspondence between input and output parameters and the data synchronization method of each functional module of the system, ensuring the consistency of multimodal data in the spatiotemporal dimension, providing a timing benchmark for feature alignment and fusion of multi-source data, and clarifying the data interaction logic between modules, thus ensuring the closed-loop operation of the system control link.

[0039] Furthermore, during the coarse positioning process, the embodied intelligent control platform continuously receives time-series data collected by the light-free sensing cluster. It monitors the tactile feedback and torque feedback data streams within this time-series data in real time. When the proportion of non-zero elements in the contact area distribution matrix exceeds a preset threshold, or when the normal force component in the pure six-dimensional interactive torque data exceeds a preset contact force threshold, it determines that the gripper and sample container have made effective contact and immediately triggers a control mode switching operation. In impedance control mode, the joint drive controller introduces real-time torque and tactile feedback from the time-series data via a feedforward method, dynamically adjusting the desired impedance parameters to keep the contact force at the end of the robotic arm within a preset safe range, avoiding rigid collisions with the sample container. Simultaneously, it performs high-precision micro-operations based on the corrected trajectory.

[0040] This embodiment constructs a complete technical chain for three-dimensional perception, contact mechanics perception, intelligent decision-making, and adaptive control in a dark environment through the above technical solution. It realizes the fully automated execution of sample preprocessing operations in a dark laboratory environment, and provides accurate spatial positioning basis and adaptive contact mechanics control capability for samples with pose deviations.

[0041] In one optional embodiment, the active infrared structured light projector in the lightless sensing cluster uses diffractive optical elements to modulate the infrared laser beam into a pseudo-random coded light spot matrix with a specific spatial arrangement structure; the infrared camera is a binocular infrared camera, which simultaneously acquires an infrared left view and an infrared right view containing the pseudo-random coded light spot matrix; the embodied intelligent control platform performs stereo matching on the infrared left view and the infrared right view based on epipolar constraints, extracts the disparity value of the matching pixels, calculates the depth information by combining the intrinsic and extrinsic parameter matrices of the binocular infrared camera, and generates a sparse three-dimensional point cloud containing only the sample container and the surrounding preset operation area after filtering out background depth noise; and reconstructs the three-dimensional point cloud contour of the sample container by surface fitting of the sparse three-dimensional point cloud.

[0042] refer to Figure 2 Specifically, the active infrared structured light projector outputs a single-mode infrared laser beam from its infrared laser source. This beam is incident on the incident surface of a diffractive optical element, which performs phase modulation on the beam, splitting the single-mode Gaussian beam and modulating it into a pseudo-random coded spot matrix with a specific spatial arrangement. Each spot unit in the pseudo-random coded spot matrix has a unique coding feature, and the coding features of adjacent spot units are distinguishable. Within the imaging field of view of the binocular infrared camera, the combination of spot codes in any local area is unique, providing a unique feature identifier for stereo matching.

[0043] Furthermore, the binocular infrared camera includes a left infrared camera and a right infrared camera, which are fixedly arranged along a horizontal baseline with a fixed length. The intrinsic parameter matrix, distortion coefficient, and extrinsic parameter matrix of the left and right infrared cameras are all pre-calibrated and stored. The left and right infrared cameras synchronously trigger acquisition with the same frame rate and the same exposure parameters, respectively acquiring an infrared left view and an infrared right view containing a pseudo-random coded spot matrix. Each set of image frames of the infrared left view and the infrared right view is marked with the same timestamp to ensure the spatiotemporal consistency of the left and right views.

[0044] Before performing stereo matching on the infrared left and right views based on epipolar constraints, the embodied intelligent control platform performs adaptive local histogram equalization on the infrared left and right views to eliminate grayscale shifts caused by differences in infrared thermal radiation. Specifically, the adaptive local histogram equalization process divides the infrared left and right views into multiple non-overlapping local sub-blocks, performs histogram equalization on each sub-block, adjusts the grayscale mapping relationship based on the grayscale distribution within the sub-block, enhances the grayscale contrast between the coded light spot and the background in the local area, and eliminates global grayscale shifts caused by differences in infrared thermal radiation from the sample container and the operating platform, thus avoiding the impact of uneven grayscale on the accuracy of subsequent stereo matching.

[0045] During stereo matching, the embodied intelligent control platform uses a cost aggregation algorithm based on semi-global matching to calculate the initial disparity value of the matching pixels. The corresponding global energy function is:

[0046]

[0047] in, For disparity map The global energy function; For pixels in an image; For pixels disparity value; For pixels In disparity value The matching cost below; For pixels The set of neighboring pixels; For each pixel in the neighborhood pixel set; This is the penalty coefficient when the disparity change of neighboring pixels is 1. This is the penalty coefficient when the disparity change of neighboring pixels is greater than 1. ; This is an indicator function that takes the value 1 when the condition within the parentheses is true, and 0 otherwise.

[0048] Specifically, firstly, the matching cost value of each pixel in the infrared left view within a preset disparity range is calculated. The matching cost value is calculated based on the grayscale features of the coded light spot, generating a cost space for each pixel. Then, one-dimensional cost aggregation is performed on the cost space along multiple directions, and the aggregated cost values ​​of each pixel in each direction are summed to obtain the global aggregated cost value for that pixel. For each pixel, the disparity value corresponding to the minimum global aggregated cost value is selected as the initial disparity value for that pixel, generating an initial disparity map.

[0049] Furthermore, the embodied intelligent control platform performs a left-right consistency check on the initial disparity values ​​to eliminate mismatched pixels in occluded areas. Specifically, an initial disparity map generated based on the infrared left view is the left disparity map, and a corresponding right disparity map is generated based on the infrared right view. For each pixel in the left disparity map, its corresponding pixel in the right disparity map is calculated based on its disparity value. The disparity values ​​of the two pixels are compared. When the difference in disparity values ​​exceeds a preset threshold, the pixel is determined to be a mismatched pixel, and its disparity value is set to invalid. After completing the left-right consistency check, median filtering is performed on the retained initial disparity values ​​to remove isolated noise points. A median filter kernel of a preset size is used to traverse the disparity map, and the median of the neighborhood of each valid pixel's disparity value is taken to eliminate isolated noise points in the disparity map, resulting in an optimized disparity map.

[0050] The embodied intelligent control platform calculates depth information based on the optimized disparity map and the intrinsic and extrinsic parameter matrices of the binocular infrared camera. The corresponding depth calculation formula is as follows:

[0051]

[0052] in, This represents the depth value of the spatial point corresponding to the pixel in the camera coordinate system. The equivalent focal length of the binocular infrared camera is obtained by calibration using the camera's intrinsic parameter matrix. The horizontal baseline length of the binocular infrared camera; This is the optimized disparity value corresponding to this pixel.

[0053] Specifically, the depth information corresponding to each valid pixel is calculated using the above formula to generate a depth map. The depth information is downsampled based on a preset spatial voxel grid to fix the point cloud density of the sparse 3D point cloud. Specifically, the sample operation area in the world coordinate system is divided into a 3D voxel grid of equal size. The size of each voxel grid is fixed. For all depth points falling within the same voxel grid, only the depth point at the centroid of the voxel grid is retained as a valid point, and the remaining redundant points are discarded. Through voxel downsampling, the point cloud density of the generated sparse 3D point cloud remains fixed, unaffected by sample distance or field of view coverage.

[0054] Furthermore, the embodied intelligent control platform performs background filtering on the sparse 3D point cloud. Based on the preset operating area spatial range, background depth points outside the operating area are removed, retaining only the effective sparse 3D point cloud containing only the sample container and the surrounding preset operating area. The effective sparse 3D point cloud is then reconstructed using surface fitting to recreate the 3D point cloud contour of the sample container. A least-squares surface fitting algorithm is used to perform piecewise surface fitting on the sparse 3D point cloud. The corresponding least-squares solution formula is:

[0055]

[0056] in, For the first sparse 3D point cloud The spatial coordinates of the points; The effective number of points in a sparse 3D point cloud; For the first One basis function; The number of basis functions; For the first The fitting coefficients of each basis function; The fitting coefficient vector is obtained by solving the least squares problem, and the expression of the fitting surface is determined.

[0057] Specifically, a continuous sample container surface mesh model is generated through the above fitting process, and the surface mesh model is uniformly sampled to generate a high-density three-dimensional point cloud contour of the sample container.

[0058] Table 2. Core parameter configuration table for pseudo-random encoded spot matrix

[0059] Optical parameters of light spot Wavelength of light spot unit The center wavelength of the infrared laser beam Within the spectral response range of the infrared camera, avoiding the peak wavelengths of ambient thermal radiation. Optical parameters of light spot beam unit divergence angle beam divergence half-angle of a single spot unit The light spot is less than a preset threshold, ensuring that there is no significant diffusion of the light spot unit within the working distance range. Encoding structure parameters Number of rows and columns of a matrix Number of rows and columns of the pseudo-random coded spot matrix The product of the number of rows and columns is greater than the preset minimum number of codes to ensure the uniqueness of local region codes. Encoding structure parameters Minimum spot spacing Minimum distance between the centers of adjacent light spot units Larger than twice the diameter of the light spot unit to avoid overlapping interference between adjacent light spots. Encoding structure parameters Encoding window size Pixel size of the local coding window used for stereo matching The number of light spot units contained within the window is greater than the preset minimum number to ensure the uniqueness of the encoding. Projection parameters Projection frame rate Projection update frequency of the encoded spot matrix To achieve hard synchronization triggering, the frame rate of the data acquisition is kept consistent with that of the binocular infrared camera. Projection parameters Working distance range Effective projection range of the coded light spot Covering the entire height variation range of the sample container placement station

[0060] Specifically, the table clarifies the core parameter configuration rules of the pseudo-random coded spot matrix, providing clear constraints for the modulation design of diffractive optical elements, ensuring that the projected coded spot has stable optical characteristics and unique coded features within the working distance range, providing highly reliable feature basis for binocular stereo matching, and improving the accuracy and robustness of 3D point cloud reconstruction in the dark.

[0061] This embodiment achieves high-precision reconstruction of the 3D point cloud contour of a sample container in a dark environment through the above technical solution. By modulating the pseudo-random encoded light spot, preprocessing the adaptive local histogram equalization, aggregating the cost of semi-global matching, eliminating mismatches and voxel downsampling operations, the interference of infrared thermal radiation differences, occlusion and noise on stereo matching is eliminated, and a 3D point cloud contour of the sample container with fixed density and accurate pose is generated, providing a stable and high-precision spatial coordinate basis for the coarse positioning of the robotic arm.

[0062] In another optional embodiment, the array-type tendon-type tactile sensor includes a flexible substrate attached to the inner side of the gripper and a microelectromechanical system (MEMS) strain gauge array embedded in the flexible substrate in a grid-like distribution. The MEMS strain gauge array uses an analog switch matrix to time-division select each strain gauge node to obtain the resistance change of each node when the inner side of the gripper contacts the sample container. The embodied intelligent control platform converts the resistance change into a strain value matrix of each node. Based on the pre-calibrated elastic modulus and Poisson's ratio of the flexible substrate, the strain value matrix is ​​mapped to a contact area distribution matrix and a local normal stress distribution matrix of the contact area between the inner side of the gripper and the sample container. The contact area distribution matrix and the local normal stress distribution matrix are combined into a tactile feedback data stream in the time-series data. The six-dimensional torque sensor is rigidly coaxially connected between the output shaft of the robotic arm's end effector joint and the gripper base via a flange. The sensor contains three sets of orthogonally arranged Wheatstone full-bridge strain gauge circuits. The embodied intelligent control platform differentially amplifies and converts the analog voltage signals output by the three sets of orthogonally arranged Wheatstone full-bridge strain gauge circuits to obtain raw six-dimensional torque data containing three-dimensional force and three-dimensional torque components. Based on the Jacobian matrix of the kinematic forward solution of the robotic arm's end effector joint, the platform performs gravity compensation and end effector inertial force decoupling calculations on the raw six-dimensional torque data to generate pure six-dimensional interactive torque data free from static interference. This pure six-dimensional interactive torque data is then used as the torque feedback data stream in the time-series data.

[0063] refer to Figure 3 Specifically, the flexible substrate of the array-type tendon-type tactile sensor is made of flexible polymer material and has a shape structure that completely conforms to the inner clamping surface of the gripper, allowing it to deform synchronously with the bending of the gripper's clamping surface. An array of microelectromechanical system (MEMS) strain gauges is embedded in the neutral layer inside the flexible substrate, uniformly distributed in a grid pattern. Each strain gauge node corresponds to a local contact area on the inner side of the gripper. The resistance value of the strain gauge node changes linearly with the deformation of the flexible substrate. When the inner side of the gripper comes into contact with the sample container, the flexible substrate in the contact area undergoes normal compressive deformation, and the resistance value of the corresponding strain gauge node changes accordingly.

[0064] Furthermore, each strain gauge node in the microelectromechanical system (MEMS) strain gauge array is connected to the input of an analog switch matrix via row and column scan lines; the output of the analog switch matrix is ​​connected to the differential input of an instrumentation amplifier, and the reference voltage of the instrumentation amplifier is connected to the output of a digital-to-analog converter (DAC). When the embodied intelligent control platform selects each strain gauge node in a time-division multiplexing manner, it synchronously controls the DAC to output a common-mode bias voltage corresponding to the currently selected node to the reference voltage of the instrumentation amplifier. The instrumentation amplifier differentially amplifies the weak voltage signal corresponding to the resistance change of the strain gauge node and the common-mode bias voltage, eliminating the common-mode interference voltage generated by the flexible substrate under bending conditions, and outputting a differentially amplified voltage signal that only reflects the normal deformation of the strain gauge node. Specifically, the analog switch matrix uses row and column addressing to sequentially select each strain gauge node in the strain gauge array. When a single strain gauge node is selected, the remaining strain gauge nodes are in an off state to avoid signal crosstalk between adjacent nodes. The common-mode bias voltage output by the digital-to-analog converter is the same as the static output voltage of the currently selected strain gauge node in the non-contact state. The instrumentation amplifier performs differential calculation on the real-time output voltage of the strain gauge node and the common-mode bias voltage to eliminate the common-mode voltage interference caused by the overall bending of the flexible substrate, and only amplifies the differential voltage signal corresponding to the normal deformation caused by contact, thereby improving the signal-to-noise ratio of the strain signal.

[0065] The embodied intelligent control platform performs analog-to-digital conversion on the differential amplified voltage signal output from the instrumentation amplifier to obtain the digital voltage value corresponding to each strain gauge node. Based on the pre-calibrated voltage-resistance mapping relationship, the digital voltage value is converted into the resistance change of each strain gauge node, and the strain value of each strain gauge node is calculated. The corresponding strain calculation formula is as follows:

[0066]

[0067] in, The first strain gauge in the strain gauge array Line 1 Strain values ​​of column nodes; The resistance change at the strain gauge node is the difference between the resistance value in the contact state and the static resistance value in the non-contact state. The strain sensitivity coefficient of the strain gauge node is obtained through pre-calibration; This represents the static resistance value of the strain gauge node in a non-contact state.

[0068] Specifically, the strain values ​​of all nodes in the strain gauge array are calculated using the above formula, generating a strain value matrix. The number of rows and columns of the strain value matrix is ​​consistent with the number of rows and columns of the microelectromechanical system strain gauge array, and each element in the matrix corresponds to the strain value of a strain gauge node.

[0069] The embodied intelligent control platform maps the strain value matrix into a local normal stress distribution matrix based on the pre-calibrated elastic modulus and Poisson's ratio of the flexible substrate. The corresponding stress calculation formula is as follows:

[0070]

[0071] in, The first strain gauge in the strain gauge array Line 1 Local normal stress values ​​corresponding to column nodes; Pre-calibrated elastic modulus for flexible substrates; Pre-calibration of Poisson's ratio for flexible substrates; This represents the strain value at that node.

[0072] Specifically, using the above formula, each strain value in the strain value matrix is ​​converted into a corresponding local normal stress value, generating a local normal stress distribution matrix for the current sampling time. Simultaneously, the strain value matrix is ​​binarized, setting nodes with strain values ​​exceeding a preset strain threshold to 1 and all other nodes to 0, generating a contact area distribution matrix. The area corresponding to the element with a value of 1 in the contact area distribution matrix is ​​the effective contact area between the inner side of the gripper and the sample container. The embodied intelligent control platform combines the contact area distribution matrix and the local normal stress distribution matrix at the same sampling time, marks the corresponding timestamp, and generates a tactile feedback data stream in the time-series data. The tactile feedback data stream is a sequence set of contact area distribution matrices and local normal stress distribution matrices within consecutive time steps.

[0073] Furthermore, the six-dimensional torque sensor has flange connection structures at both ends. One end is rigidly connected coaxially to the output shaft of the robotic arm's end joint via a flange, and the other end is rigidly connected coaxially to the base of the gripper via a flange, ensuring that the force and torque generated by the interaction between the robotic arm's end and the environment are completely transmitted to the sensitive element of the six-dimensional torque sensor. Inside the six-dimensional torque sensor are three sets of orthogonally arranged Wheatstone full-bridge strain gauge circuits, corresponding to the X, Y, and Z orthogonal axes in three-dimensional space. One set of Wheatstone full-bridge strain gauge circuits corresponds to the force and torque components in the X and Y axes, one set corresponds to the force and torque components in the X and Y axes in the Z axis direction, and one set corresponds to the torque component in the Z axis direction. All three sets of Wheatstone full-bridge strain gauge circuits output an analog voltage signal containing three-dimensional force and torque components.

[0074] The embodied intelligent control platform differentially amplifies and converts the analog voltage signals output from three orthogonally arranged Wheatstone full-bridge strain gauge circuits to obtain raw six-dimensional torque data. This raw six-dimensional torque data is a combination of three-dimensional force vectors and three-dimensional torque vectors in the world coordinate system. Because the gravity of the robotic arm's links, the gravity of the end effector, and the inertial forces during the robotic arm's movement can couple and interfere with the measurement results of the six-dimensional torque sensor, it is necessary to perform gravity compensation and end effector inertial force decoupling calculations on the raw six-dimensional torque data to generate pure six-dimensional interactive torque data free from static interference.

[0075] When performing gravity compensation, the embodied intelligent control platform retrieves the stored mass parameters and center-of-mass position parameters of each link of the robotic arm. Based on the real-time angles of each joint, it calculates the static gravitational torque vector generated by each link on the end joint under the gravitational field using a recursive Newton-Euler algorithm. The corresponding formula for calculating the gravitational torque is as follows:

[0076]

[0077] in, For the first The gravitational torque vector generated by each link under the gravitational field; For the first The rotation matrix of the coordinate system of each link relative to the world coordinate system is calculated from the real-time angles of each joint of the robotic arm using forward kinematics. For the first The mass parameters of each link are pre-calibrated and stored. This is the gravitational acceleration vector in the world coordinate system; This represents the number of joint degrees of freedom of the robotic arm; For the first From the origin of the first link coordinate system to the first The position vector of the origin of the coordinate system of each link; This refers to the cross product operation of vectors.

[0078] Specifically, by summing the static gravitational torque vectors of all the links, the total static gravitational torque vector generated by the robotic arm body and the end effector on the six-dimensional torque sensor is obtained.

[0079] When performing inertial force decoupling calculations, the embodied intelligent control platform obtains the joint angular velocity and joint angular acceleration fed back from the encoders installed at each joint motor. It then combines these with mass parameters and center-of-mass position parameters to construct the joint space mass matrix of the robotic arm and calculates the inertial torque vector. The corresponding inertial force decoupling formula is as follows:

[0080]

[0081] in, This is the torque vector corresponding to the inertial force, Coriolis force, and centrifugal force generated during the movement of the robotic arm; The joint space mass matrix of the robotic arm is calculated from the mass parameters of each link, the position parameters of the center of mass, and the real-time joint angles. This represents the real-time angular acceleration vectors of each joint of the robotic arm; This is the real-time angular velocity vector of each joint of the robotic arm; Let be the matrix of Coriolis force and centrifugal force of the robotic arm.

[0082] Specifically, the static gravitational torque vector and inertial force vector are subtracted from the original six-dimensional torque data to generate pure six-dimensional interactive torque data. This pure six-dimensional interactive torque data only contains the external forces and torques generated by the interaction between the robotic arm's end effector and the environment, eliminating the coupling interference between the robotic arm's own gravity and motion inertia. The embodied intelligent control platform marks the pure six-dimensional interactive torque data with corresponding timestamps, generating a torque feedback data stream in the time-series data. The torque feedback data stream is a sequence set of pure six-dimensional interactive torque data within continuous time steps.

[0083] Table 3. Calibration parameters and performance indicators of arrayed MEMS strain gauge nodes.

[0084] Basic electrical parameters static resistance value Resistance value of strain gauge nodes in non-contact, non-deformed state The relative deviation of the static resistance values ​​of nodes within the same array is less than a preset threshold. Basic electrical parameters Strain sensitivity coefficient The ratio of the relative change in the nodal resistance of the strain gauge to the strain value The relative deviation of the strain sensitivity coefficient of nodes within the same array is less than a preset threshold. Mechanical calibration parameters elastic modulus Elastic modulus of flexible substrate materials The measurement uncertainty of the calibration value is less than the preset threshold and covers the operating temperature range. Mechanical calibration parameters Poisson's ratio Poisson's ratio of flexible substrate materials The measurement uncertainty of the calibration value is less than the preset threshold and covers the operating temperature range. Performance indicators Strain measurement range The range of strain values ​​that can be accurately measured at strain gauge nodes The strain value corresponding to the maximum deformation during the gripper clamping operation Performance indicators Strain resolution The smallest resolvable strain value at strain gauge nodes Less than the preset threshold, meeting the requirements for micro-contact force detection. Performance indicators linearity Linear fitting error between strain value and resistance change The linearity error is less than a preset threshold across the entire measurement range. Performance indicators Crosstalk rejection ratio Signal crosstalk suppression capability between adjacent strain gauge nodes The signal strength exceeds a preset threshold to avoid signal interference from adjacent nodes.

[0085] Specifically, the table clarifies the calibration parameters and performance requirements of array-type MEMS strain gauge nodes, providing a clear basis for the factory calibration and field calibration of tactile sensors. It ensures that each node of the strain gauge array has stable and consistent electrical and mechanical properties, guarantees the detection accuracy of contact area distribution and local normal stress distribution, and provides reliable feedback data for accurate judgment of contact state and precise control of contact mechanics.

[0086] This embodiment achieves high-precision tactile perception of the contact state between the gripper and the sample container, as well as high signal-to-noise ratio detection of the interaction torque at the end of the robotic arm through the above technical solutions. The interference of flexible substrate bending is eliminated by time-division gating and common-mode voltage suppression, and the coupling interference between the weight of the robotic arm and the motion inertia is eliminated by decoupling the recursive Newton-Euler algorithm and inertial force. High-fidelity tactile feedback data stream and torque feedback data stream are generated, providing accurate contact mechanics feedback basis for feature fusion and adaptive control of the embodied large model.

[0087] In another optional embodiment, the graph neural network in the embodied large model constructs spatial location nodes from sample pose coordinates, and constructs first temporal feature nodes and second temporal feature nodes from haptic feedback data streams and torque feedback data streams in temporal data after aligning them according to timestamps, respectively. The graph neural network establishes undirected edge connections based on the spatiotemporal correlation between spatial location nodes, first temporal feature nodes, and second temporal feature nodes within a preset time step. The graph neural network uses graph convolutional layers to pass messages through the undirected edge connections, mapping the point cloud geometric features of spatial location nodes, the contact mechanical features of first temporal feature nodes, and the six-dimensional torque features of second temporal feature nodes to latent space feature vectors of the same dimension, and performing a concatenation operation on the latent space feature vectors to generate a fused multimodal feature tensor. The embodied large model's embedded physics engine stores a 3D rigid body model of the sample container and a flexible body model of the gripper. The physics engine receives contact area distribution information from the fused multimodal feature tensor and establishes penalty function contact mechanical constraints between the flexible body model and the 3D rigid body model. Based on the penalty function contact mechanical constraints and the real-time angles of each joint of the current robotic arm, the physics engine performs forward dynamic integration to solve for the angular velocity sequence of each joint that satisfies the preset contact force threshold. The embodied large model inputs the angular velocity sequence of each joint into the policy network, which outputs the diagonal stiffness matrix and diagonal damping matrix from the desired impedance parameters, and combines them with the angular velocity sequence of each joint to generate a corrected trajectory. The gripper flexible body model in the physics engine is represented by a spring-mass mesh model. The spring-mass mesh model includes multiple mass points distributed on the geometric surface of the gripper, as well as structural springs, shear springs, and bending springs connecting adjacent mass points. When establishing the penalty function contact mechanical constraints, the physics engine detects the penetration depth between the mass points in the spring-mass mesh model and the surface of the three-dimensional rigid body model of the sample container in real time. When a penetration depth exists, the normal rebound force and tangential friction force on the mass point are calculated based on the penetration depth and the preset penalty stiffness coefficient. The normal rebound force and tangential friction force are applied as external loads to the spring-mass mesh model, and the spatial position coordinates and velocity vectors of each mass point in the spring-mass mesh model are updated according to the explicit Euler integral method.

[0088] refer to Figure 4Specifically, the graph neural network is a spatiotemporal graph convolutional network, including an input layer, multiple graph convolutional layers, fully connected layers, and an output layer. The embodied large model first extracts features from the input sample pose coordinates, encoding the centroid spatial coordinates, pose angles, and geometric features of the 3D point cloud contour of the sample container to construct spatial position nodes. The feature vector of each spatial position node contains all pose and geometric shape information of the sample container in the world coordinate system. Simultaneously, the embodied large model timestamps the haptic feedback data stream and torque feedback data stream in the temporal data. It flattens and encodes the contact area distribution matrix and local normal stress distribution matrix within the same time step, constructing the first temporal feature node; it also encodes the pure six-dimensional interactive torque data within the same time step, constructing the second temporal feature node. Both the first and second temporal feature nodes are temporal feature sequences within consecutive time steps, with one feature node corresponding to each time step.

[0089] Furthermore, the graph neural network establishes undirected edge connections based on the spatiotemporal correlation of spatial location nodes, first temporal feature nodes, and second temporal feature nodes within a preset time step. Specifically, for spatial location nodes, first temporal feature nodes, and second temporal feature nodes within the same time step, undirected edge connections are established between each pair to represent the correlation between spatial pose information and contact mechanics information at the same moment; for first temporal feature nodes and second temporal feature nodes within adjacent time steps, undirected edge connections are established to represent the correlation between contact mechanics information in the time dimension; for spatial location nodes and first temporal feature nodes and second temporal feature nodes within adjacent time steps, undirected edge connections are established to represent the correlation between changes in spatial pose and changes in contact mechanics. Through the above undirected edge connections, a spatiotemporal graph structure containing information related to both spatial and temporal dimensions is constructed.

[0090] Graph neural networks use graph convolutional layers to perform message passing on undirected edge connections. The corresponding graph convolutional message passing formula is:

[0091]

[0092] in, For the first Nodes in a layered graph convolutional layer eigenvectors; For the first Nodes in a layered graph convolutional layer The output feature vector; For nodes The set of neighboring nodes; and For the first Trainable weight matrix of a layered graph convolutional layer; It is a non-linear activation function; For nodes The number of neighboring nodes; For neighboring nodes The number of neighboring nodes.

[0093] Specifically, for each node in the graph, the feature information of its neighboring nodes is aggregated using the above formula, and combined with its own feature information, the output feature vector of that node is updated. Through iterative calculations of multiple graph convolutional layers, the point cloud geometric features of spatial location nodes, the contact mechanical features of the first temporal feature nodes, and the six-dimensional moment features of the second temporal feature nodes are mapped to the latent space feature vector of the same dimension. This ensures that features of different modalities have consistent dimensions and comparable feature distributions in the same latent space, thus completing the feature alignment of multimodal heterogeneous data.

[0094] Furthermore, the graph neural network concatenates the latent space feature vectors of the aligned spatial location nodes, the latent space feature vectors of the first temporal feature node, and the latent space feature vectors of the second temporal feature node along the feature dimension to generate a fused multimodal feature tensor. This fused multimodal feature tensor contains the spatial pose information, contact area distribution information, local normal stress distribution information, and six-dimensional interaction torque information of the sample, while also incorporating spatiotemporal correlation features, providing comprehensive feature basis for subsequent physics engine dynamics calculations and policy network control decisions.

[0095] refer to Figure 5 The physics engine embedded in the large-scale model is a rigid-flexible body coupled dynamics engine, which pre-stores 3D rigid body models of sample containers of various sizes, as well as flexible body models of the grippers. The 3D rigid body model of the sample container includes rigid body dynamic parameters such as the sample container's geometry, mass, moment of inertia, and surface friction coefficient. The flexible body model of the gripper is represented by a spring-mass mesh model, which includes multiple masses distributed on the geometric surface of the gripper, as well as structural springs, shear springs, and bending springs connecting adjacent masses. Structural springs are arranged along the edge direction of the mass mesh to constrain the axial deformation between masses; shear springs are arranged along the diagonal direction of the mass mesh to constrain the shear deformation between masses; bending springs connect adjacent nodes separated by one mass to constrain the bending deformation of the mass mesh. Through the combination of these three types of springs, the mechanical properties of the flexible gripping surface of the gripper are fully characterized.

[0096] The dynamic equilibrium equations of the spring-mass mesh model are:

[0097]

[0098] in, Let be the mass matrix of the spring-mass mesh model, and let be a diagonal matrix, with the diagonal elements representing the mass of each mass point. Let be the acceleration vector of each particle; Let be the velocity vector of each particle; Let be the displacement vector of each particle; Here is the damping matrix; The stiffness matrix is ​​obtained by assembling the stiffness coefficients of structural springs, shear springs, and bending springs. The external load vector applied to the particle; This is the contact force vector, which includes the normal rebound force and the tangential friction force.

[0099] The physics engine receives contact area distribution information from the fused multimodal feature tensor to determine the contact area between the gripper and the sample container, establishing penalty function contact mechanical constraints between the flexible body model and the 3D rigid body model. Specifically, the physics engine detects the distance between each particle in the spring-mass mesh model and the surface of the 3D rigid body model of the sample container in real time. When a particle is located inside the surface of the 3D rigid body model, penetration is determined, and the penetration depth between the particle and the surface of the rigid body model is calculated. The corresponding contact force calculation formula is as follows:

[0100]

[0101] in, Let be the contact force vector acting on the particle; The default penalty stiffness coefficient; This represents the penetration depth between the particle and the surface of the rigid body model; it is a positive value when penetration occurs. This is the unit normal vector of the rigid body model's surface at the contact point, pointing outwards from the rigid body. This is the preset contact damping coefficient; The velocity vector of the particle; The preset dynamic friction coefficient between the rigid body model and the flexible body model; This represents the magnitude of the normal contact force. It is the unit tangent vector at the point of contact, which is opposite to the direction of the relative tangential velocity of the particle.

[0102] Specifically, when penetration depth exists, the physics engine calculates the normal rebound force and tangential friction force acting on the particle according to the above formula, combines the normal rebound force and tangential friction force into a contact force vector, and applies it as an external load to the corresponding particle in the spring-particle mesh model. The physics engine then numerically solves the dynamic equilibrium equations of the spring-particle mesh model using the explicit Euler integral method, updating the spatial coordinates and velocity vectors of each particle in the spring-particle mesh model, simulating the local deformation of the flexible gripping surface of the gripper during the contact process.

[0103] Furthermore, based on the penalty function contact mechanics constraints and the real-time angles of each joint of the current robotic arm, the physics engine performs forward dynamics integration to solve for the joint angular velocity sequence that satisfies the preset contact force threshold. Specifically, the physics engine uses the real-time angles and angular velocities of each joint of the current robotic arm as the initial state, and the preset contact force threshold as the constraint condition. Within the preset prediction time domain, it performs multiple forward dynamics integrations, calculates the joint angle, angular velocity, and end-effector contact force at each step, and selects the joint angular velocity sequences whose end-effector contact forces are always within the preset contact force threshold range, and outputs them to the policy network.

[0104] The mapping formula for the desired output impedance parameter of the strategy network is:

[0105]

[0106] in, Let be the desired diagonal stiffness matrix in the joint space; Let be the desired diagonal damping matrix in the joint space; This is the sequence of joint angular velocities output by the physics engine; To fuse multimodal feature tensors; and These are two multilayer perceptrons in the policy network, used to map and obtain the desired stiffness matrix and desired damping matrix, respectively.

[0107] Specifically, the embodied large model inputs the joint angular velocity sequences output by the physics engine and the fused multimodal feature tensor into the policy network. The policy network uses two independent multilayer perceptrons to map the expected diagonal stiffness matrix and diagonal damping matrix in the joint space, respectively, forming the expected impedance parameters. Simultaneously, the policy network combines the joint angular velocity sequences and sample pose coordinates to correct the original trajectory of the robotic arm's end effector, generating a corrected trajectory. The corrected trajectory considers contact mechanics constraints and sample pose offsets, ensuring that the motion of the robotic arm's end effector meets contact safety requirements.

[0108] Furthermore, the high-precision robotic arm's actuator performs coarse positioning based on the sample's pose coordinates. After the gripper contacts the sample, the position control mode is deactivated and switched to an impedance control mode based on time-series data feedforward. Pre-processing operations are performed based on the desired impedance parameters and the corrected trajectory. In impedance control mode, the joint drive controller uses the desired impedance parameters as the control target and the torque feedback data stream acquired in real time from the time-series data as feedforward compensation to perform impedance control law calculations. This adjusts the driving torque of each joint in real time, ensuring that the dynamic characteristics of the robotic arm's end effector conform to the desired impedance model. This achieves adaptive control of the contact force while maintaining positioning accuracy, preventing damage to the fragile sample container.

[0109] Table 4. Spring parameter configuration table for the gripper flexible body spring-mass mesh model.

[0110] Structural springs Axial stiffness coefficient Stiffness coefficient for constraining axial deformation of a mass The elastic modulus and cross-sectional dimensions of the flexible substrate material of the gripper are calculated as follows: Structural springs Axial damping coefficient Damping coefficient of constrained axial vibration of a mass Calculated based on the preset damping ratio and the natural frequency of the structural spring. shear spring Shear stiffness coefficient Stiffness coefficient of constrained mass shear deformation The axial stiffness coefficient is smaller than that of springs in the same region, and the shear modulus matches that of the flexible substrate. shear spring Shear damping coefficient Damping coefficient of constrained mass shear vibration The axial damping coefficient is consistent with that of the springs in the same region. Bending spring Bending stiffness coefficient Stiffness coefficient of constrained mass mesh bending deformation The value is calculated based on the flexural modulus and thickness of the flexible substrate, and is less than the shear stiffness coefficient. Bending spring Bending damping coefficient Damping coefficient of bending vibration of constrained mass grid Calculated based on the preset bending damping ratio and the natural frequency of the bending spring. All types of springs Maximum deformation threshold Maximum elastic deformation that a spring can withstand The deformation is less than the yield deformation corresponding to the flexible substrate material, thus avoiding plastic deformation.

[0111] Specifically, the table clarifies the core parameter configuration rules for the three types of springs in the spring-mass mesh model, providing clear parameter basis for the construction of the gripper flexible body model. This ensures that the spring-mass mesh model can accurately simulate the axial, shear, and bending mechanical properties of the gripper flexible clamping surface, improve the accuracy and physical fidelity of contact mechanics calculations in the physics engine, and provide reliable dynamic model support for the accurate solution of the desired impedance parameters.

[0112] This embodiment achieves spatiotemporal feature alignment and fusion of multimodal heterogeneous data through the above technical solution. By using rigid-flexible body coupled dynamics calculation with an embedded physics engine, it accurately simulates the contact process between the gripper and the sample container, solves for the joint motion sequence and desired impedance parameters that meet the contact safety constraints, realizes the adaptive impedance control of the robotic arm in a dark environment, and completes high-precision non-destructive preprocessing micro-operations on the pose offset sample.

Claims

1. A light-out laboratory sample pretreatment system incorporating a high-precision, embodied intelligent robotic arm, characterized in that: This includes a light-free sensing cluster, an embodied intelligent control platform, and a high-precision robotic arm actuator; The light-free sensing cluster includes an active infrared structured light projector, an infrared camera, a six-dimensional torque sensor deployed at the end joint of the robotic arm, and an array of tendon-type tactile sensors deployed at the gripper base. The active infrared structured light projector projects coded light spots onto the sample area, and the infrared camera reconstructs the three-dimensional point cloud contour of the sample container based on the coded light spots. The embodied intelligent control platform deploys an embodied large model with an embedded physics engine. The embodied large model receives sample pose coordinates provided by the three-dimensional point cloud contour and time-series data collected by the six-dimensional torque sensor and the array-type tendon-type tactile sensor. The sample pose coordinates and the time-series data are input into a graph neural network for feature alignment and fusion, and the desired impedance parameters and corrected trajectory of the robotic arm in the joint space are output. The high-precision robotic arm actuator performs coarse positioning based on the sample pose coordinates. After the gripper contacts the sample, it closes the position control mode and switches to the impedance control mode based on the time-series data feedforward. It then performs preprocessing operations based on the desired impedance parameters and the corrected trajectory.

2. The light-out laboratory sample pretreatment system combined with a high-precision embodied intelligent robotic arm according to claim 1, characterized in that, The active infrared structured light projector in the lightless sensing cluster uses diffractive optical elements to modulate the infrared laser beam into a pseudo-random coded light spot matrix with a specific spatial arrangement structure. The infrared camera is a binocular infrared camera, which simultaneously acquires an infrared left view and an infrared right view containing the pseudo-random coded spot matrix. The embodied intelligent control platform performs stereo matching of the infrared left view and the infrared right view based on epipolar constraints, extracts the disparity value of the matching pixels, calculates the depth information by combining the intrinsic and extrinsic parameter matrices of the binocular infrared camera, and generates a sparse three-dimensional point cloud containing only the sample container and the surrounding preset operation area after filtering out background depth noise. The sparse three-dimensional point cloud is then surface-fitted to reconstruct the three-dimensional point cloud contour of the sample container.

3. The light-out laboratory sample pretreatment system combined with a high-precision embodied intelligent robotic arm according to claim 1, characterized in that, The array-type tendon-type tactile sensor includes a flexible substrate attached to the inside of the gripper and a microelectromechanical system strain gauge array embedded in the flexible substrate in a grid-like distribution. The microelectromechanical system strain gauge array uses an analog switch matrix to select each strain gauge node in a time-division manner, thereby obtaining the resistance change of each node when the inner side of the gripper contacts the sample container. The embodied intelligent control platform converts the resistance change into a strain value matrix for each node. Based on the pre-calibrated elastic modulus and Poisson's ratio of the flexible substrate, the strain value matrix is ​​mapped to the contact area distribution matrix and the local normal stress distribution matrix of the contact area between the inner side of the gripper and the sample container. The contact area distribution matrix and the local normal stress distribution matrix are combined into the tactile feedback data stream in the time series data.

4. The light-out laboratory sample pretreatment system combined with a high-precision embodied intelligent robotic arm according to claim 1, characterized in that, The six-dimensional torque sensor is rigidly connected coaxially between the output shaft of the end joint of the robotic arm and the gripper base via a flange. The six-dimensional torque sensor contains three sets of orthogonally arranged Wheatstone full-bridge strain gauge circuits. The embodied intelligent control platform differentially amplifies and converts the analog voltage signals output by the three sets of orthogonally arranged Wheatstone full-bridge strain gauge circuits to obtain six-dimensional torque raw data containing three-dimensional force components and three-dimensional torque components. The embodied intelligent control platform performs gravity compensation and end-effector inertial force decoupling calculations on the raw six-dimensional torque data based on the forward kinematics Jacobian matrix of the end-effector joints of the robotic arm, generating pure six-dimensional interactive torque data that eliminates static interference, and using the pure six-dimensional interactive torque data as the torque feedback data stream in the time-series data.

5. The light-out laboratory sample pretreatment system combined with a high-precision embodied intelligent robotic arm according to claim 1, characterized in that, The graph neural network in the embodied large model constructs the sample pose coordinates as spatial location nodes, and constructs the tactile feedback data stream and torque feedback data stream in the temporal data as the first temporal feature node and the second temporal feature node after aligning them according to the timestamps respectively. The graph neural network establishes undirected edge connections based on the spatiotemporal correlation between the spatial location nodes, the first temporal feature nodes, and the second temporal feature nodes within a preset time step. The graph neural network uses graph convolutional layers to pass messages through the undirected edge connections, mapping the point cloud geometric features of the spatial location nodes, the contact mechanical features of the first temporal feature nodes, and the six-dimensional torque features of the second temporal feature nodes to a latent space feature vector of the same dimension, and performing a concatenation operation on the latent space feature vector to generate a fused multimodal feature tensor.

6. The light-out laboratory sample pretreatment system combined with a high-precision embodied intelligent robotic arm according to claim 1, characterized in that, The physical engine embedded in the embodied large model stores a three-dimensional rigid body model of the sample container and a flexible body model of the gripper. The physics engine receives the contact area distribution information from the fused multimodal feature tensor and establishes a penalty function contact mechanical constraint between the flexible body model and the three-dimensional rigid body model. The physics engine performs forward dynamic integration calculations based on the penalty function contact mechanics constraints and the real-time angles of each joint of the current robotic arm, and solves for the angular velocity sequence of each joint that satisfies the preset contact force threshold. The embodied large model inputs the angular velocity sequences of each joint into the policy network, and the policy network outputs the diagonal stiffness matrix and diagonal damping matrix in the desired impedance parameters, and generates the corrected trajectory by combining the angular velocity sequences of each joint.

7. The light-out laboratory sample pretreatment system combined with a high-precision embodied intelligent robotic arm according to claim 2, characterized in that, Before performing stereo matching of the infrared left view and the infrared right view based on epipolar constraints, the embodied intelligent control platform performs adaptive local histogram equalization processing on the infrared left view and the infrared right view to eliminate grayscale shift caused by differences in infrared thermal radiation. During stereo matching, the embodied intelligent control platform uses a cost aggregation algorithm based on semi-global matching to calculate the initial disparity value of the matching pixels. It performs a left-right consistency check on the initial disparity value to remove mismatched pixels in occluded areas. It performs median filtering on the retained initial disparity value to remove isolated noise points. It generates the depth information based on the filtered disparity value and downsamples the depth information according to a preset spatial voxel grid to fix the point cloud density of the sparse 3D point cloud.

8. The light-out laboratory sample pretreatment system combined with a high-precision embodied intelligent robotic arm according to claim 3, characterized in that, Each strain gauge node in the microelectromechanical system strain gauge array is connected to the input terminal of the analog switch matrix via row and column scan lines; The output of the analog switch matrix is ​​connected to the differential input of the instrumentation amplifier, and the reference voltage of the instrumentation amplifier is connected to the output of the digital-to-analog converter. When the embodied intelligent control platform selects each strain gauge node in a time-division manner, it synchronously controls the digital-to-analog converter to output the common-mode bias voltage corresponding to the currently selected node to the reference voltage terminal of the instrumentation amplifier. The instrumentation amplifier differentially amplifies the weak voltage signal corresponding to the resistance change of the strain gauge node and the common-mode bias voltage to eliminate the common-mode interference voltage generated by the flexible substrate in the bending state, and outputs a differentially amplified voltage signal that only reflects the normal deformation of the strain gauge node.

9. The light-out laboratory sample pretreatment system combined with a high-precision embodied intelligent robotic arm according to claim 4, characterized in that, When performing gravity compensation, the embodied intelligent control platform retrieves the mass parameters and center of mass position parameters of each link of the robotic arm stored in the database, and calculates the static gravitational torque vector generated by each link on the end joint under the gravitational field using the recursive Newton-Euler algorithm based on the real-time angle of each joint. When performing inertial force decoupling calculation, the embodied intelligent control platform acquires the joint angular velocity and joint angular acceleration fed back by the encoders installed at each joint motor. It then constructs the joint space mass matrix of the robotic arm by combining the mass parameters and the center of mass position parameters. The inertial force vector is calculated by multiplying the joint space mass matrix by the joint angular acceleration. Finally, the static gravity torque vector and the inertial force vector are subtracted from the original six-dimensional torque data to generate the pure six-dimensional interactive torque data.

10. The light-out laboratory sample pretreatment system combining a high-precision embodied intelligent robotic arm according to claim 6, characterized in that, The flexible gripper model in the physics engine is represented by a spring-mass mesh model, which includes multiple mass points distributed on the geometric surface of the gripper and structural springs, shear springs, and bending springs connecting adjacent mass points. When establishing the penalty function contact mechanical constraints, the physics engine detects the penetration depth between the particles in the spring-mass mesh model and the surface of the three-dimensional rigid body model of the sample container in real time. When a penetration depth exists, the engine calculates the normal rebound force and tangential friction force on the particles based on the penetration depth and the preset penalty stiffness coefficient. The normal rebound force and the tangential friction force are applied as external loads to the spring-mass mesh model. The engine then updates the spatial position coordinates and velocity vectors of each particle in the spring-mass mesh model according to the explicit Euler integral method.