A microscopic positioning control method and system for adaptive visual and force sense fusion

Through the micropositioning control method of adaptive visual force fusion, high-precision, stable and safe micro-operation in complex microenvironment is achieved using a high-frame rate microcamera and MEMS force sensor array, combined with dynamic timestamp alignment and nonlinear coordinate mapping.

CN120259435BActive Publication Date: 2025-08-01NANCHANG INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510742094.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-08-01
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

When facing a micro-operating environment with dynamic complexity and extreme sensitivity of the operating object, existing micro-positioning control technologies are difficult to perceive environmental changes in real time and adjust control parameters adaptively, resulting in accumulated position estimation errors, positioning offsets and operational insecurity problems.

Method used

Adaptive visual force fusion control method is adopted to capture data in real time through high-frame-rate microcamera and MEMS force sensing sensor arrays. Combined with dynamic timestamp alignment, nonlinear coordinate mapping, multimodal data fusion and reinforcement learning, control parameters are dynamically adjusted to realize synchronization and adaptive control between vision and force sensation.

Benefits of technology

It achieves submicron-level fusion positioning accuracy, improves operation stability and safety, can effectively resist liquid fluctuations and mechanical vibrations, meets the real-time needs of minimally invasive surgery, and avoids overload damage to brittle samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259435B_ABST
    Figure CN120259435B_ABST
Patent Text Reader

Abstract

The present invention proposes a microscopic positioning control method and system for adaptive visual-force fusion, wherein: the control method includes the following steps: image capture; dynamic timestamp alignment; non-linear mapping of the spatial coordinate system; visual data enhancement and feature extraction; dynamic calibration of the force signal; spatio-temporal registration of multi-modal data; feature fusion based on the attention mechanism; fusion pose generation and error correction; real-time optimization of reinforcement learning parameters; generation of fuzzy sliding mode disturbance rejection control instructions; dynamic impedance protection and real-time retraction; fault recovery; performance calibration. The present invention eliminates the spatio-temporal asynchrony and dimensional heterogeneity of visual and force data through bilinear interpolation timestamp alignment and non-linear coordinate mapping based on deep learning, achieving a fusion positioning accuracy of sub-micron level; based on the attention mechanism and force gradient compensation, dynamically allocates the weights of vision and force.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of optical measurement, and specifically to a microscopic positioning control method and system for adaptive visual and force sense fusion. Background Art

[0002] Microscopic positioning control technology is the core support in fields such as minimally invasive surgery, biological cell manipulation, and micro-nano device assembly. The accuracy and robustness of its control strategy directly determine the success, failure, and safety of microscopic operations. In biomedical scenarios, it is necessary to perform non-destructive manipulation of fragile biological samples (such as cells, tissues) with nanometer-level accuracy; in the field of micro-nano manufacturing, high-stability assembly of sub-micron-level devices is required. The common challenges in such operating environments lie in dynamic complexity (such as liquid disturbance, target deformation) and the extreme sensitivity of the operating object. There is an urgent need for an intelligent control method that can sense environmental changes in real time, adaptively adjust control parameters, and ensure operation safety.

[0003] Existing microscopic positioning control technologies mostly rely on single-modal feedback mechanisms. For example, the control method based on visual servo extracts the target pose through microscopic images, but it is vulnerable to optical noise, changes in sample transparency, and dynamic occlusion interference, resulting in the accumulation of pose estimation errors; while the control method based on force sense feedback can sense the contact force in real time, but it cannot provide effective displacement information in the non-contact stage, and is prone to misjudgment due to environmental vibration. In addition, traditional multi-modal control strategies often adopt simple serial feedback or fixed-weight data fusion, ignoring the spatio-temporal asynchrony and dimensional heterogeneity between visual and force sense data, resulting in fusion lag or noise amplification, seriously affecting the dynamic response accuracy of the control system. More critically, existing control methods mostly adopt static parameter configuration, unable to adapt to sudden changes in the mechanical properties of the operation target or environmental disturbances, easily causing positioning deviation, contact force exceeding the limit, and even sample damage.

[0004] Although some researches have attempted to introduce multi-modal fusion control in recent years, they still face the following bottleneck problems: First, the lack of a dynamic spatio-temporal registration mechanism makes it difficult to solve the spatio-temporal misalignment and heterogeneous feature alignment problems of visual and force sense data; second, the static optimization mode of control parameters cannot respond to the dynamic changes of the microscopic environment in real time, resulting in insufficient robustness; third, the computational delay of traditional control architectures is relatively high, making it difficult to meet the millisecond-level real-time requirements in scenarios such as minimally invasive surgery. These problems seriously restrict the reliability and universality of microscopic positioning control in complex microenvironments. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the existing defects and provide a microscopic positioning control method and system for adaptive visual and force sense fusion, effectively solving the problems in the background art.

[0006] To achieve the above object, the present invention provides the following technical solution: An adaptive visual and force sense fusion microscopic positioning control method, comprising the following steps:

[0007] S1), Image capture: Real-time capture of the operation area image by a high-frame-rate microscopic camera, using an adaptive exposure algorithm to eliminate optical noise, and extracting the target contour through an edge detection operator, and outputting the target pose

[0008] Real-time acquisition of six-dimensional force and torque signals at the operation end through a MEMS force sense sensor array , and suppressing environmental vibration interference through a band-pass filter;

[0009] S2), Dynamic timestamp alignment: Based on the visual frame timestamp and the force sense sampling timestamp , asynchronous data alignment is achieved through a bilinear interpolation function:

[0010] ;

[0011] wherein is the calibration compensation term, dynamically optimized through a sliding window variance minimization algorithm;

[0012] is the sampling frequency of the visual system;

[0013] is the sampling frequency of the force sense system;

[0014] represents the unified timestamp after dynamic timestamp alignment;

[0015] S3), Nonlinear mapping of the spatial coordinate system: Establish a mapping model from the visual pixel coordinate system to the force sense Cartesian coordinate system :

[0016] ;

[0017] wherein is the calibration parameter, is the nonlinear error compensated by a convolutional neural network, is the Z-axis prediction function based on deep learning;

[0018] S4), Visual data enhancement and feature extraction: Perform Gaussian filtering and Laplacian sharpening on the original image to enhance the target edge features:

[0019] ;

[0020] where:

[0021] : The original microscopic image data;

[0022] is a Gaussian kernel used to smooth the image; where σ = 1.5 pixels, representing the standard deviation of the Gaussian kernel, which controls the smoothness of the filtering;

[0023] is the convolution operation symbol, indicating Gaussian filtering of the original image;

[0024] is the Laplacian operator of the original image, used to extract the high-frequency edge features of the image;

[0025] is the sharpening coefficient, used to control the intensity of Laplacian sharpening;

[0026] is the enhanced image, combining the effects of Gaussian filtering and Laplacian sharpening;

[0027] S5), Force sense signal dynamic calibration: According to the ambient temperature and the pose of the robotic arm , compensate the zero drift of the MEMS force sense sensor in real time:

[0028] ;

[0029] Where: : The original uncompensated contact force data, directly output by the MEMS force sense sensor;

[0030] : The temperature drift coefficient, used to quantify the influence of ambient temperature change on the zero offset of the MEMS force sense sensor;

[0031] : The current ambient temperature, collected by a thermosensitive device or a thermocouple;

[0032] : The ambient temperature during calibration, used as the reference for temperature compensation;

[0033] : The pose coupling weight matrix, a 2×1 or 2×2 vector or matrix, reflecting the interference weight of the pose on the force offset;

[0034] : The two-dimensional displacement vector of the current operation end in the visual coordinate system, with the unit of pixel or mm;

[0035] : The finally calibrated force signal, eliminating the pose interference, used as the true feedback value for subsequent control calculations;

[0036] S6), Spatiotemporal registration of multimodal data: Perform spatiotemporal registration on the coordinate mapping result output by S3 and the calibrated force sense signal of S5 to generate a synchronized data set , and fuse the temporal noise through a Kalman filter;

[0037] : Represents a structured data set after time alignment, coordinate correction, and filtering fusion, used for subsequent training of the depth control model or state estimation;

[0038] S7): Feature fusion based on the attention mechanism: Dynamically allocate visual and force sense weights through a hybrid deep learning model:

[0039] ;

[0040] =1- ;

[0041] The above structure realizes dynamic weighting of the two channels of vision and force sense during the fusion process through the softmax normalization mechanism;

[0042] Among them, the channel scoring factors 、 Are respectively defined as:

[0043] , ;

[0044] Variable definitions are as follows:

[0045] ∇I, ∇F: Feature gradient vectors of the image and the force sense signal respectively;

[0046] 、 : Channel weight matrices for feature mapping, parameters of a 1×N fully connected layer, obtained through end-to-end training;

[0047] ReLU(x)=max(0,x): Rectified linear unit activation function, introducing non-linearity to enhance the sparsity of feature expression;

[0048] : Visual modality attention weight, reflecting the importance of visual features in the current fusion process;

[0049] , Is the visual gradient and force gradient feature vector;

[0050] S8), Fusion pose generation and error correction: Output the fusion pose :

[0051] ;

[0052] ;

[0053] wherein , is the force gradient compensation term;

[0054] In the formula, is the target pose coordinate output by the vision system, from the image capture and edge detection results of step S1, belonging to the vision pixel coordinate system;

[0055] is the target coordinate in the MEMS force sensor coordinate system, obtained through the non-linear mapping of the space coordinate system in step S3, with the unit of physical length;

[0056] is the weight parameter, the weight coefficient of the vision data, dynamically calculated through the attention mechanism in step S7, with the value range of [0,1];

[0057] is the weight coefficient of the force sense data, also generated by the attention mechanism, and is related to ;

[0058] S9), real-time optimization of reinforcement learning parameters: Define the state space (pose error, force error, disturbance level), and output the control parameter increment through the ActorCritic network:

[0059] ;

[0060] : The control gain adjustment amount of the vision channel, used to dynamically adjust the visual servo response intensity;

[0061] : The control gain adjustment amount of the force sense channel, used to dynamically adjust the force control stiffness or impedance;

[0062] : The learning rate, which controls the speed of network weight update;

[0063] : The sliding mode surface variable, representing the error measure between the current state and the target state;

[0064] 、 : The network parameter matrices, corresponding to the weight matrices of the vision gain regulator and the force sense gain regulator respectively;

[0065] 、 : Bias term, adjusting the intermediate offset of the network's non-linear output;

[0066] The above structure constitutes a single-layer or double-layer neural network mapping for mapping the sliding mode error to the control gain adjustment amount, realizing the adaptive control adjustment under different disturbance levels;

[0067] S10), generating the fuzzy sliding mode disturbance rejection control command: Design the sliding mode surface

[0068] ;

[0069] Among them Fuzzy adjustment according to the disturbance level ;

[0070] In the formula: is the pose error; The sliding mode surface gain coefficient; The non-linear exponential parameter (μ ∈ (0, 1); The sign function ( outputs +1 when > 0, otherwise -1);

[0071] Generate the disturbance rejection control command:

[0072] ;

[0073] : The saturation function limiting parameter;

[0074] : The visual channel sliding mode control gain coefficient, which controls the sensitivity of the control system to the visual error sliding mode surface. The larger the value, the faster the system responds to the change of visual error;

[0075] : The force sense channel disturbance rejection gain coefficient, which controls the action intensity of the sat function in the force channel and is often used to suppress the force sense disturbance or external micro-vibration;

[0076] : The disturbance compensation gain, which adjusts the influence of the force sense channel compensation on the controller;

[0077] : The change rate of the contact force, which characterizes the trend of external disturbance;

[0078] Among them , ;

[0079] The function sat(x) is used to suppress the chattering of the sliding mode controller, and the output value is limited in the interval [-1, 1];

[0080] S11), dynamic impedance protection and real-time withdrawal: according to the environment stiffness Dynamically adjust impedance parameters :

[0081] ;

[0082] in: : Basic quality parameter, which represents the reference quality of the system under an ideal rigid environment;

[0083] : Environmental stiffness, a dynamic estimation value, reflecting the stiffness of the current contact object;

[0084] : Equivalent mass, which is used to convert the foundation mass based on the environmental flexibility;

[0085] : Equivalent damping, used for speed control / energy dissipation regulation in the withdrawal process control;

[0086] : Damping adjustment proportional factor;

[0087] : controller equivalent stiffness coefficient, which models the controller's ability to respond compliantly to terminal disturbances;

[0088] ;

[0089] : Indicates the end position fine-tuning vector when the protection is triggered, in nanometers (nm);

[0090] : concession scale factor;

[0091] : The current contact force collected in real time by the MEMS force sensor;

[0092] :Preset contact safety threshold, when When the retracement is triggered;

[0093] ∇x: is the unit velocity direction vector of the current position.

[0094] Preferably, in step S2, during the dynamic timestamp alignment, the compensation item is calibrated Optimization via sliding window variance minimization algorithm:

[0095] ;

[0096] in: : The visual end pose vector of the k-th sliding window sample point;

[0097] : kth force sensing end sampling point, time lag The corresponding pose vector;

[0098] : Euclidean norm squared, used to minimize the temporal alignment error between vision and force perception;

[0099] : represents the time compensation amount to minimize the objective function;

[0100] in is the number of data frames in the sliding window;

[0101] In the nonlinear mapping of the spatial coordinate system in step S3, the Z-axis prediction function It is implemented through pre-trained deep neural network (DNN), and the input is the local feature block of the image , the output is The predicted value of:

[0102] ;

[0103] Wherein, the input and output parameters include:

[0104] : The input image local feature block is a local image block containing the target area extracted from the microscopic image;

[0105] : The output Z-axis coordinate value is part of the force Cartesian coordinate system and represents the position of the target in the depth direction;

[0106] Neural network parameters include:

[0107] W h : The weight matrix of the hidden layer;

[0108] : bias vector of the hidden layer;

[0109] : rectified linear unit activation function;

[0110] The neural network parameters output layer are as follows:

[0111] : The weight matrix of the output layer of the deep neural network;

[0112] : Bias scalar for the output layer.

[0113] Preferably, in step S4 of visual data enhancement, the size of the Gaussian kernel G(σ) is 5×5, and the sharpening coefficient λ is dynamically adjusted according to the signal-to-noise ratio SNR of the image (the signal-to-noise ratio of the local image block, estimated from the image mean and gradient noise):

[0114] ;

[0115] In the dynamic calibration of the force sense signal, the pose coupling weight matrix W is determined through an offline calibration experiment and satisfies:

[0116] ;

[0117] : is a two-dimensional coupling weight matrix used to model the systematic perturbation of the contact force signal caused by the displacement of the image coordinates;

[0118] : represents the position variable (such as the image coordinates ) changes to the force signal partial derivative, reflecting the sensitivity of the force sense channel to image displacement;

[0119] All parameters are obtained through linear fitting during the calibration stage (calibration platform + known pose changes) and are used for position compensation during actual operation.

[0120] Preferably, the step S6 further includes: during the spatio-temporal registration process, an extended Kalman filter (EKF) is used to fuse the temporal noise, and the state equation and the observation equation are respectively:

[0121] ;

[0122] : is the system state variable at the current moment, including the pose information estimated by visual-force sense fusion;

[0123] : is the observable quantity, such as the center coordinates and direction of the end image obtained by microscopic image processing;

[0124] where A is the state transition matrix and H is the observation matrix, , is Gaussian noise;

[0125] The fused synchronous data set contains the visual-force sense covariance matrix Cvf with timestamp alignment.

[0126] Preferably, the step S7 further includes: in the attention mechanism, the visual feature confidence is calculated based on the normalized variance of the image gradient magnitude :

[0127] ;

[0128] : The variance of the image gradient distribution, reflecting the overall texture or edge complexity;

[0129] : The maximum gradient value in the image, corresponding to the most prominent edge, used for normalization;

[0130] : The minimum value constant, used to avoid division-by-zero errors during the image contrast normalization process;

[0131] ∇I: The image gray gradient, calculated by the image Sobel operator, used to represent the image edge intensity;

[0132] Force gradient compensation term , Calculated by real-time difference:

[0133] ;

[0134] : Represents the time difference gradient of the contact force, used to detect the rapid change trend of the force;

[0135] : Are the contact force values collected at the current moment and the previous moment respectively, collected by the MEMS force sensor at a high frequency;

[0136] : Is the time interval between two adjacent sampling moments, set by the system sampling frequency;

[0137] : Is the time calibration coefficient, used to scale the time base of the difference result to uniformly correspond to the withdrawal response rate (set to 0.1ms here).

[0138] Preferably, the step S9 further includes: In the ActorCritic network, the Critic network outputs the state value function , and updates the weights through the temporal difference (TD) error:

[0139] ;

[0140] : Temporal difference error (TDerror), representing the gap between the current predicted value and the target return, is the core basis for updating the policy and the Critic network;

[0141] : Immediate reward value, generated by the environment feedback, used to measure the effect of the current action;

[0142] : Discount factor, which controls the influence weight of future rewards;

[0143] : Learning rate, which determines the step size of each update of network parameters;

[0144] : The gradient of the Critic network with respect to its parameters ;

[0145] Gaussian noise is added to the action exploration strategy , and the noise variance decays exponentially with the number of training rounds:

[0146] ;

[0147] Where: : The standard deviation of the Gaussian policy perturbation at the current moment, which controls the randomness in the action output;

[0148] : The initial perturbation intensity, which is adjusted according to the environmental complexity;

[0149] : The annealing time constant, which affects the noise decay rate;

[0150] This formula implements exponential annealing, enabling the agent to explore sufficiently in the initial stage of training and tend to be stable in the later stage of training.

[0151] Preferably, the step S10 further includes: The perturbation level is divided into three levels: low, medium, and high, The value of , The value of ;

[0152] A high-frequency perturbation suppression term is superimposed, and its frequency is estimated in real time through FFT, and the suppression gain is adaptively adjusted:

[0153] ;

[0154] Where: : High-frequency perturbation force vector, which is obtained by band-pass filtering or fast Fourier transform (FFT) of the original force sense signal and reflects the perturbation intensity;

[0155] : Original sequence of contact force, which is continuously collected by a MEMS force sensor;

[0156] : Represents the two - norm (magnitude) of a vector, used to calculate the perturbation size;

[0157] : The maximum value of the force - sense signal within the current sliding window, used for normalization.

[0158] Preferably, the S11 further includes: environmental stiffness Estimated online through the slope of the force - displacement curve:

[0159] ;

[0160] : The change in contact force within the time window, measured by the MEMS force - sense sensor;

[0161] : The change in displacement of the end - effector during this time period, calculated by the vision system;

[0162] →0: Represents approximate differentiation in an extremely short time, used to calculate the response slope of stiffness in real - time;

[0163] In nanoscale retraction, the pose gradient Calculated by the visual feature optical flow method:

[0164] ;

[0165] : The two - dimensional pixel coordinates of the target in the current frame image;

[0166] : Represents the velocity vector component of the image target, calculated by the optical flow method;

[0167] : The estimated result of the final pose change rate;

[0168] : The time interval between two adjacent frames, in milliseconds (ms), determined by the vision frame rate.

[0169] Preferably, it further includes S12, fault recovery: When it is detected that the signal mutation of the MEMS force - sense sensor ΔF>5mN or the communication delay Td>2ms, a three - level emergency response is triggered:

[0170] a. Pause the movement of the robotic arm and switch to the safe impedance mode;

[0171] b. Reconstruct the target pose based on the weighted average of historical data, and the weight w i Satisfies ∑w i =1;

[0172] c. If the continuous fault timeout Tfault > 5 s, start the self-checking protocol and report the error code.

[0173] Preferably, it further includes S13, performance calibration: verifying that the pose tracking error ep ≤ 0.1 μm through a standard micro-nano reference device;

[0174] When switching between multiple platforms, compensate for the coordinate system offset through an adaptive calibration algorithm:

[0175] ; : Represents the position offset error of a single calibration point in visual measurement (such as pixel error or millimeter deviation), which is obtained by manual or algorithm extraction;

[0176] : Represents the number of target points participating in the average calibration, where ≥10 is the number of calibration points;

[0177] : Is the translation offset correction amount during system initialization, used to eliminate systematic deviations in overall image recognition and uniformly compensate in all subsequent coordinate conversions.

[0178] The present invention also provides a system for implementing the above-mentioned microscopic positioning control method for adaptive visual force sense fusion, including:

[0179] 1) Multimodal data acquisition module:

[0180] A visual acquisition unit for capturing the image of the operation area in real time, extracting the target contour through an edge detection algorithm, and outputting the target pose data;

[0181] A force sense acquisition unit for obtaining the six-dimensional force / torque signal at the operation end in real time;

[0182] A temperature monitoring unit for collecting the ambient temperature in real time through a thermal sensor;

[0183] (2) Asynchronous data synchronization module:

[0184] A data buffer for caching visual frame data and force sense sampling data respectively;

[0185] An interpolation alignment unit for realizing data synchronization based on the visual frame timestamp and the force sense sampling timestamp through a bilinear interpolation function, and generating a unified timestamp;

[0186] A calibration compensation unit for storing preset compensation parameters;

[0187] (3) Spatial coordinate mapping module:

[0188] A calibration parameter database for storing calibration parameters;

[0189] A non - linear error compensation unit that runs a convolutional neural network model for compensation;

[0190] A prediction unit that implements deep - learning - based prediction;

[0191] (4) Data processing and enhancement module:

[0192] A visual enhancement unit that performs Gaussian filtering and Laplacian sharpening on the original image to generate an enhanced image;

[0193] A force - sense calibration unit that compensates for the zero - point drift of the sensor in real - time according to the environmental temperature and the pose of the robotic arm and outputs a calibrated force signal;

[0194] 5) Multimodal data fusion module:

[0195] A spatio - temporal registration unit that spatio - temporally aligns the mapped coordinates and the calibrated force signal to generate a synchronized data set and reduces noise through a Kalman filter;

[0196] An attention - weighted unit that uses a hybrid deep - learning model to dynamically allocate visual and force - sense weights;

[0197] A pose synthesis unit that generates a fused pose;

[0198] (6) Adaptive control module:

[0199] A reinforcement - learning optimization unit that adjusts control parameters in real - time through an Actor - Critic network;

[0200] A sliding - mode control unit that generates control commands;

[0201] An impedance - adjustment unit that dynamically calculates the equivalent mass and damping;

[0202] (7) Safety protection module:

[0203] A safety - retraction unit that triggers nano - level displacement compensation;

[0204] A fault - emergency unit that performs exception handling;

[0205] A self - inspection and recovery unit that runs a historical - data reconstruction algorithm;

[0206] (8) Calibration verification single module:

[0207] An error - analysis unit that verifies the pose - tracking error;

[0208] An adaptive - compensation unit that calculates the offset correction amount.

[0209] Preferably, it further includes: a time alignment module that optimizes the calibration compensation term through a sliding window variance minimization algorithm and a spatial mapping module of a pre-trained deep neural network;

[0210] A force-sense coupling calibration module for online compensating the systematic offset in the contact force signal.

[0211] Preferably, it further includes:

[0212] A multi-source fusion processing module that implements spatio-temporal registration using an extended Kalman filter. The multi-source fusion processing module includes: a state prediction module burned in the kinematic model ROM of the DSP and an observation update module that executes the matrix mapping visual pose extraction module output to the state space; a covariance calculation module that generates a synchronous data set matrix and stores it in the dual-port RAM for the decision-making module to call;

[0213] A confidence evaluation module that calculates the gradient magnitude by the image processor and outputs it through a variance normalization unit;

[0214] A force gradient compensation module that includes a high-speed ADC and a differential calculation circuit.

[0215] Preferably, it further includes:

[0216] A high-frequency disturbance suppressor composed of a coprocessor and an adaptive gain controller and an environmental stiffness estimator integrated in the force control feedback loop, a fault recovery module composed of a three-stage emergency trigger circuit, a historical data reconstruction unit, and a self-check protocol executor; a micro-nano reference component positioning platform for verifying the tracking error and a calibration module for coordinate system offset compensation.

[0217] Compared with the prior art, the beneficial effects of the present invention are:

[0218] 1. Through bilinear interpolation timestamp alignment and deep learning-based non-linear coordinate mapping, this application eliminates the spatio-temporal asynchrony and dimensional heterogeneity of visual and force data, and achieves a fusion positioning accuracy of sub-micron level.

[0219] 2. Based on the attention mechanism and force gradient compensation, dynamically allocate visual and force weights, suppress noise interference and enhance key features, and improve the data reliability in complex environments; online adjust the visual servo gain and force control stiffness through the Actor-Critic network, adapt to the sudden change of the mechanical characteristics of the operation target, and reduce the positioning offset and contact force overrun.

[0220] 3. Combine the fuzzy rule base with the high-frequency disturbance suppression term to effectively resist dynamic disturbances such as liquid fluctuations and mechanical vibrations, and improve operation stability; compress the data processing delay to the millisecond level through hardware-accelerated spatio-temporal registration and feature extraction to meet the real-time requirements of minimally invasive surgery; adjust the impedance parameters in real time based on the environmental stiffness, and combine the nanoscale retraction mechanism to avoid overloading damage to brittle samples, and the operation safety is improved significantly. BRIEF DESCRIPTION OF THE DRAWINGS

[0221] Figure 1 It is a control flow chart of a microscopic positioning control method for adaptive visual and force sense fusion of the present invention;

[0222] Figure 2 It is a system framework diagram of a microscopic positioning control method for adaptive visual and force sense fusion of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0223] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0224] As Figure 1 shown, a microscopic positioning control method for adaptive visual and force sense fusion includes the following steps:

[0225] S1), Image capture: Real-time capture the image of the operation area through a high-frame-rate microscopic camera, use an adaptive exposure algorithm to eliminate optical noise, and extract the target contour through an edge detection operator, and output the target pose

[0226] Real-time obtain the six-dimensional force and torque signals of the operation end through the MEMS sensor array , and suppress the environmental vibration interference through a band-pass filter;

[0227] For force feedback control: Real-time detect the contact force to avoid damage to brittle samples (such as cells) due to overload.

[0228] Dynamically adjust the pose: Combine the visual data to correct the movement trajectory of the robotic arm and improve the positioning accuracy.

[0229] Safety protection: Trigger a protection mechanism (such as nanoscale retraction) when the contact force exceeds the limit to ensure operation safety.

[0230] S2), Dynamic timestamp alignment: Based on the visual frame timestamp and the force sense sampling timestamp , asynchronous data alignment is achieved through a bilinear interpolation function:

[0231] Bilinear interpolation utilizes the haptic data points at adjacent timestamps and calculates the haptic signal value corresponding to the visual frame timestamp by weighted calculation according to the time ratio.

[0232] Asynchronous data: refers to the misaligned data caused by the asynchronous acquisition times of visual (high frame rate) and haptic (different sampling rates) data in step S1.

[0233] Alignment purpose:

[0234] Eliminate the temporal differences between visual and haptic data, ensure strict synchronization between the two in time, and provide a consistent basis for multimodal fusion.

[0235] ;

[0236] Among them is the calibration compensation term, which is dynamically optimized through the sliding window variance minimization algorithm;

[0237] is the sampling frequency of the visual system;

[0238] is the sampling frequency of the haptic system;

[0239] represents the unified timestamp after dynamic timestamp alignment.

[0240] S3), non-linear mapping of the spatial coordinate system: Establish a mapping model from the visual pixel coordinate system to the haptic Cartesian coordinate system :

[0241] Solve spatial heterogeneity: Visual data is in two-dimensional pixel coordinates, and haptic data is in three-dimensional Cartesian coordinates. The mapping model unifies the two into the same spatial framework.

[0242] Improve the positioning accuracy: Compensate for non-linear errors (such as lens distortion) through deep learning to achieve sub-micron level coordinate alignment.

[0243] Support multimodal fusion: Provide the basic data for spatial alignment for subsequent spatio-temporal registration (S6).

[0244] ;

[0245] Among them is the calibration parameter, is the non-linear error compensated by the convolutional neural network, is the Z-axis prediction function based on deep learning.

[0246] S4), Visual Data Enhancement and Feature Extraction: For the original image Perform Gaussian filtering and Laplacian sharpening to enhance the target edge features:

[0247] Gaussian filtering with σ = 1.5: Smooth the image, suppress optical noise, and improve the signal-to-noise ratio.

[0248] λ = 0.8 is the sharpening coefficient, which is used to control the intensity of Laplacian sharpening, enhance the high-frequency components of the target edge, highlight the contour features, and facilitate subsequent edge detection and pose estimation.

[0249] Comprehensive effect: Improve the image quality and ensure the robustness and accuracy of feature extraction.

[0250] ;

[0251] Where: : The original microscopic image data;

[0252] is the Gaussian kernel, which is used to smooth the image; where σ = 1.5 pixels, representing the standard deviation of the Gaussian kernel, controlling the smoothness of the filtering;

[0253] The convolution operation symbol indicates performing Gaussian filtering on the original image;

[0254] is the Laplacian operator (second derivative) of the original image, which is used to extract the high-frequency edge features of the image;

[0255] is the enhanced image, which combines the effects of Gaussian filtering (noise suppression) and Laplacian sharpening (edge highlighting);

[0256] S5), Dynamic Calibration of Force Sensation Signals: According to the environmental temperature and the pose of the robotic arm , compensate the zero-point drift of the MEMS force sensation sensor in real time:

[0257] Compensate for zero-point drift: The MEMS force sensation sensor is vulnerable to the influence of environmental temperature T and the pose of the robotic arm (x, y), resulting in zero-point offset.

[0258] Correct the error in real time: Dynamically calibrate the force sensation signal through the temperature coefficient α_T and the pose coupling weight matrix W,

[0259] Improve data reliability: Ensure that the force sensation signal accurately reflects the real contact force and reduce misjudgment. <s

[0260] ; Where is the temperature coefficient, is the pose coupling weight matrix.

[0261] S6), Spatiotemporal Registration of Multimodal Data: Perform spatiotemporal registration on the coordinate mapping result output by S3 and the calibrated force sense signal of S5 to generate a synchronized dataset , and fuse the temporal noise through a Kalman filter;

[0262] Spatial and Temporal Synchronization: Align the visual coordinate mapping result output by S3 with the calibrated force sense signal of S5 in space and time to generate a synchronized dataset .

[0263] Noise Reduction Processing: Fuse the temporal noise through an Extended Kalman Filter (EKF) to improve the data quality.

[0264] Data Fusion Preparation: Provide clean and consistent input for the feature fusion based on the attention mechanism (S7).

[0265] S7): Feature Fusion Based on Attention Mechanism: Dynamically allocate visual and force sense weights through a hybrid deep learning model:

[0266] ;

[0267] where , , , are the visual gradient and force gradient feature vectors.

[0268] S8), Fusion Pose Generation and Error Correction: Output the fusion pose :

[0269] ;

[0270] ;

[0271] where , is the force gradient compensation term;

[0272] In the formula, is the target pose coordinate output by the visual system, from the image capture and edge detection results of step S1, belonging to the visual pixel coordinate system (two-dimensional plane coordinate, unit is pixel or calibrated physical unit);

[0273] The function is to provide the position information of the target in the microscopic image for initial pose estimation.

[0274] is the target coordinate in the MEMS force sense sensor coordinate system, obtained through the nonlinear mapping of the spatial coordinate system in step S3 (from the visual pixel coordinate system to the force sense Cartesian coordinate system), unit is physical length;

[0275] is the weight parameter, The weight coefficient of the visual data is dynamically calculated by the attention mechanism in step S7 and has a value range of [0, 1];

[0276] is the weight coefficient of force data, which is also generated by the attention mechanism. .

[0277] S9), real-time optimization of reinforcement learning parameters: defining the state space (pose error, force error, disturbance level), the control parameter increments are output by the ActorCritic network:

[0278] ;

[0279] in is the learning rate, and the network weights are updated through TD error back propagation;

[0280] S10), fuzzy sliding mode anti-disturbance control command generation: design sliding surface

[0281] ;

[0282] in According to the disturbance level Blur adjustment;

[0283] Where: s is the sliding surface variable, for Pose error (the difference between the current pose and the target pose); Sliding surface gain coefficient; Nonlinear exponential parameter (μ∈(0,1); Symbolic function ( >0, output +1, otherwise -1);

[0284] Generate disturbance rejection control instructions:

[0285] ;

[0286] in , ;

[0287] S11), dynamic impedance protection and real-time withdrawal: according to the environment stiffness Dynamically adjust impedance parameters :

[0288] ;

[0289] When the contact force exceeds the limit When triggered, a nanoscale retraction occurs:

[0290] .

[0291] Preferably, step S2 further includes:

[0292] In dynamic timestamp alignment, the calibration compensation term is optimized by the sliding window variance minimization algorithm:

[0293] ;

[0294] where is the number of data frames within the sliding window;

[0295] In the non - linear mapping of the spatial coordinate system, the Z - axis prediction function is implemented by a pre - trained deep neural network (DNN), with the input being the local feature block of the image , and the output being the predicted value of:

[0296] ;

[0297] In the formula, the input and output parameters include:

[0298] : The input local feature block of the image is a local image block containing the target area extracted from the microscopic image.

[0299] As the input of the deep neural network (DNN), it is used to extract visual features related to the Z - axis coordinate.

[0300] : The output Z - axis coordinate value, which is part of the force - sense Cartesian coordinate system, represents the position of the target in the depth direction (Z - axis).

[0301] Its function is the Z - axis coordinate predicted by the DNN model, which is used to supplement the non - linear mapping from the visual pixel coordinate system to the force - sense coordinate system (to solve the problem of insufficient accuracy of the traditional linear mapping in the Z - axis direction).

[0302] The neural network parameters (hidden layer) include:

[0303] W h : The weight matrix of the hidden layer (Weight matrix of the hidden layer).

[0304] Type: Matrix, and its dimension is determined by the size of the input feature block and the number of neurons in the hidden layer.

[0305] Function: For the input feature block Perform a linear transformation to map pixel-level features to the abstract feature space of the hidden layer.

[0306] : Bias vector of the hidden layer.

[0307] Type: Vector, with the dimension consistent with the number of neurons in the hidden layer.

[0308] Function: Add a bias term to the linear transformation of the hidden layer to improve the model fitting ability.

[0309] : Rectified Linear Unit activation function.

[0310] Function: For the result of the linear transformation of the hidden layer Introduce non-linear activation to prevent the network from becoming a simple linear model and enhance the ability to model complex non-linear relationships.

[0311] The neural network parameters (output layer) are as follows:

[0312] : Weight matrix of the output layer of the deep neural network (DNN).

[0313] Type: Matrix, and its dimension is determined by the number of neurons in the hidden layer and the output dimension (in the formula, the output is a single Z-axis coordinate, so the dimension is 1 X the number of neurons in the hidden layer). <=

[0314] Function: Map the abstract features of the hidden layer to the predicted value of the final Z-axis coordinate, and complete the non-linear mapping from visual features to the Z-axis of the force perception coordinate system.

[0315] : Bias scalar of the output layer.

[0316] Type: Scalar.

[0317] Function: Add a bias term to the linear transformation of the output layer to adjust the offset of the final predicted value.

[0318] This formula realizes the non-linear mapping from the visual pixel coordinate system to the Z-axis of the force perception Cartesian coordinate system through a deep neural network, solves the mapping error problem of traditional linear methods in the depth direction, and improves the three-dimensional positioning accuracy (especially the sub-micron level Z-axis positioning).

[0319] Preferably, the S4 further includes: In visual data augmentation, the size of the Gaussian kernel G(σ) is 5×5, and the sharpening coefficient λ is dynamically adjusted according to the signal-to-noise ratio SNR of the image:

[0320] ;

[0321] In the dynamic calibration of force signals, the posture coupling weight matrix W is determined through offline calibration experiments and satisfies:

[0322] .

[0323] Preferably, the S6 further includes: in the spatiotemporal registration, using an extended Kalman filter (EKF) to fuse the time series noise, and the state equation and the observation equation are respectively:

[0324] ;

[0325] Where A is the state transfer matrix, H is the observation matrix, , is Gaussian noise;

[0326] Synchronized dataset after fusion Contains the timestamp-aligned visual force covariance matrix Cvf.

[0327] Preferably, the S7 further includes: in the attention mechanism, the visual feature confidence The calculation is based on the image gradient magnitude The normalized variance of :

[0328] ;

[0329] Force gradient compensation term , By real-time difference calculation:

[0330] .

[0331] Preferably, the S9 further includes: in the ActorCritic network, the Critic network outputs a state value function , and update the weights by the temporal difference (TD) error:

[0332] ;

[0333] Action Exploration Strategy Adding Gaussian Noise , noise variance Exponential decay with training epochs:

[0334] .

[0335] Preferably, the S10 further includes: disturbance level It is divided into three levels: low, medium and high, corresponding to The value of , The value of ;

[0336] Superimpose a high-frequency disturbance suppression term , whose frequency is estimated in real time by FFT, and the suppression gain is adaptively adjusted:

[0337] ;

[0338] Preferably, the S11 further includes: environmental stiffness is estimated online through the slope of the force-displacement curve:

[0339] ;

[0340] In the nanoscale retraction, the pose gradient is calculated by the visual feature optical flow method:

[0341]

[0342] Preferably, it further includes S12 fault recovery: when a sensor signal mutation ΔF>5 mN or a communication delay Td>2 ms is detected, a three-level emergency response is triggered:

[0343] a. Pause the movement of the robotic arm and switch to the safe impedance mode;

[0344] b. Reconstruct the target pose based on the weighted average of historical data, and the weight w i satisfies ∑w i = 1;

[0345] c. If the continuous fault timeout Tfault>5 s, start the self-checking protocol and report the error code.

[0346] Preferably, it further includes S13, performance calibration: verify that the pose tracking error ep≤0.1 μm through a standard micro-nano reference device;

[0347] When switching between multiple platforms, compensate for the coordinate system offset through an adaptive calibration algorithm:

[0348] ;

[0349] where N≥10 is the number of calibration points.

[0350] The 1-13 steps proposed by the present invention are not executed in isolation, but are closely associated through data flow and control flow to form a closed-loop feedback system:

[0351] S1 (data acquisition): Provide the original visual image and force sense signal, and provide input for the subsequent steps.

[0352] S2 (Timestamp Alignment) and S3 (Coordinate System Mapping): Solve the spatio-temporal asynchrony and dimensional heterogeneity of visual and haptic data, unify heterogeneous data into the same spatio-temporal framework (S2 aligns timestamps, S3 aligns spatial coordinates), and lay the foundation for multi-modal fusion.

[0353] S4 (Visual Enhancement) and S5 (Haptic Calibration): Improve the feature robustness of visual data (such as edge sharpening) and the accuracy of haptic data (such as temperature drift compensation) respectively, and ensure the reliability of subsequent fused data.

[0354] S6 (Spatio-temporal Registration): Synchronize the mapping result of S3 and the calibration signal of S5 in space and time, generate a synchronized dataset for fusion, and denoise it through Kalman filtering.

[0355] S7 (Feature Fusion) and S8 (Pose Generation): Dynamically allocate visual and haptic weights based on the attention mechanism (S7), correct errors by combining force gradient compensation, and generate a fused pose (S8).

[0356] S9 (Parameter Optimization) and S10 (Disturbance Rejection Control): Optimize control parameters in real time through reinforcement learning (S9), combine fuzzy sliding mode control to generate disturbance rejection commands (S10), and respond to dynamic environmental changes.

[0357] S11 (Impedance Protection): Adjust impedance parameters according to environmental stiffness, and trigger nanoscale retraction when the contact force exceeds the limit to protect the sample safety.

[0358] S12 (Fault Recovery) and S13 (Performance Calibration): As the safety guarantee (S12) and accuracy verification (S13) of the system, supplement the core processes of the previous 11 steps.

[0359] Solution to the lack of dynamic spatio-temporal registration mechanism:

[0360] Bilinear interpolation timestamp alignment (S2): Compensate the sampling time difference between visual and haptic through an interpolation function to eliminate asynchrony.

[0361] Deep learning non-linear coordinate mapping (S3): Use a pre-trained DNN model to establish a non-linear mapping from pixel coordinates to Cartesian coordinates to solve dimensional heterogeneity.

[0362] Extended Kalman filter (S6): Fuse temporal noise to further improve spatio-temporal registration accuracy.

[0363] Solution to the insufficient static optimization of control parameters:

[0364] Reinforcement learning parameter real-time optimization (S9): Online adjust visual servo gain and force control stiffness through an Actor-Critic network to respond to sudden changes in mechanical properties.

[0365] Dynamic weight allocation (S7): Based on the attention mechanism and force gradient compensation, visually and haptically weighted values are adaptively allocated to suppress noise interference.

[0366] Solutions for high computing latency:

[0367] Hardware acceleration (S6, S13): Using the FPGA processing unit to accelerate spatio-temporal registration and feature extraction, compressing the latency to the millisecond level.

[0368] High-frequency disturbance suppression (S10): By estimating the disturbance frequency in real time through FFT and superimposing a suppression term to reduce the computational load.

[0369] Realization of sub-micron-level fusion positioning accuracy technology:

[0370] Bilinear interpolation (S2) to align timestamps, eliminating the sampling latency between vision and haptics.

[0371] DNN non-linear mapping (S3) to compensate for the error from pixel coordinates to Cartesian coordinates (such as the Z-axis prediction function), reducing the spatial mapping deviation.

[0372] Kalman filtering (S6) further fuses noise to ensure spatio-temporal consistency of data.

[0373] Realization of dynamic weight allocation and robustness improvement technology:

[0374] Attention mechanism (S7): Calculate the visual confidence based on the image gradient magnitude (such as the normalized variance) and dynamically adjust the weights.

[0375] Force gradient compensation (S8): Calculate the force gradient term through real-time differentiation to correct the pose error.

[0376] Actor-Critic network (S9): Online optimize control parameters to adapt to mechanical mutations (such as stiffness changes during cell puncture).

[0377] Realization of anti-interference ability and real-time performance technology:

[0378] Fuzzy sliding mode control (S10): Fuzzily adjust the sliding mode surface parameters according to the disturbance level to resist disturbances such as liquid fluctuations.

[0379] Hardware acceleration (FPGA): Improve the data processing speed to meet the real-time requirements at the millisecond level.

[0380] Dynamic impedance protection (S11): Adjust the impedance parameters in real time according to the environmental stiffness, and combine with the nano-level retraction mechanism to avoid overload damage.

[0381] Steps S1 to S11 constitute the core technical solution, covering the entire process from data acquisition, fusion to control instruction generation.

[0382] Step S12 (Fault Recovery): As a safety guarantee extension, an emergency response is triggered in case of system anomalies (such as signal mutations or communication delays) to ensure operational safety.

[0383] Step S13 (Performance Calibration): As a verification extension, the pose error is tested through standard devices and the coordinate system offset for multi-platform switching is compensated to ensure system accuracy and compatibility.

[0384] Steps S12 and S13 are supplements to the core processes (S1 - S11), respectively targeting system reliability and accuracy verification, forming a complete "core function + safety redundancy + performance guarantee" system.

[0385] The present invention provides the following technical solutions:

[0386] Biological cell nanoscale manipulation experiment:

[0387] 1. Implementation environment and equipment

[0388] 1.1 Hardware platform:

[0389] High - frame - rate microscopic camera (Basler acA2000 - 340km, frame rate );

[0390] MEMS sensor array (Futek LSB200, range 0.1 μN - 10 mN, sampling rate ); <E

[0391] Piezoelectric ceramic - driven nanorobot arm (Physik Instrumente P - 611.3S, displacement resolution 1 nm);

[0392] FPGA processing unit (Xilinx Zynq UltraScale +, main frequency 500 MHz).

[0393] 1.2 Software environment:

[0394] Image processing: OpenCV4.5 (edge detection, Kalman filtering);

[0395] Deep learning framework: PyTorch1.9 (CNN, DNN model training);

[0396] Control algorithm: MATLAB / Simulink real - time control module.

[0397] 1.3 Experimental object:

[0398] Human umbilical vein endothelial cells (HUVEC), cultured in a PDMS microfluidic chip;

[0399] Standard silicon microcantilever (size 10μm × 100μm × 1μm).

[0400] The specific implementation steps are as follows:

[0401] Step S1: Multi-modal data synchronous acquisition - Visual data: The microscopic camera captures cell images at 1200Hz, with an adaptive exposure time of 0.1ms. Canny edge detection is used to extract the cell contour, and the pose (x v , y v , θv) is output.

[0402] Force sense data: The MEMS force sense sensor collects cell contact force signals (Fx, Fy), and a band-pass filter (10Hz~1kHz) suppresses the culture medium fluctuation noise.

[0403] Step S2: Dynamic timestamp alignment - Align the visual (t v ) and force sense (t f ) timestamps through bilinear interpolation. The sliding window N = 50, = 0.05ms.

[0404] Step S3: Nonlinear mapping of the spatial coordinate system - Pre-trained DNN model (input: 32×32 image block; output: ), error <0.05μm.

[0405] Step S4: Visual data enhancement Gaussian filtering (σ = 1.5pixel) and Laplacian sharpening (λ = 0.8), with the signal-to-noise ratio increased by 30%.

[0406] Step S5: Force sense signal calibration Temperature compensation coefficient , pose coupling matrix Determined by offline calibration.

[0407] Step S6: Spatiotemporal registration and Kalman filtering EKF parameters: State transition matrix , observation matrix , noise covariance , .

[0408] Steps S7 - S8: Attention mechanism fusion and pose generation Visual weight , force sense weight , force gradient compensation .

[0409] Step S9: Reinforcement learning parameter optimization Actor-Critic network structure: Input layer with 3 nodes, hidden layer with 64 nodes, output layer with 2 nodes; Learning rate , discount factor , exploration noise 。

[0410] Step S10: Fuzzy Sliding Mode Disturbance Rejection Control - Disturbance Level In , ,high - frequency suppression gain 。

[0411] Step S11: Dynamic Impedance Protects the Environmental Stiffness ,impedance parameter , 。

[0412] Step S12: Fault Recovery: Simulate Communication Delay , trigger a three - level response, and reconstruct the pose error m.

[0413] Step S13: Performance Calibration: Calibrate the Pose Error of the Silicon Cantilever m, after multi - platform switching offset compensation 。

[0414] Verification of Implementation Effect

[0415] 。

[0416] Comparative Experiment

[0417] Experimental Group: Use the method of the present invention to perform puncture operations on HUVEC (target depth 5μm);

[0418] Control Group: Traditional Visual Servo + Fixed Impedance Control.

[0419] 。

[0420] Thus, this embodiment verifies the significant advantages of the method in biological cell nanoscale operations:

[0421] High Precision: The fused pose error is low, meeting the requirements of cell non - destructive operation;

[0422] Strong Robustness: The liquid disturbance suppression rate is relatively high, adapting to the dynamic micro - environment;

[0423] High Safety: The response time for contact force exceeding the limit is short, and the cell survival rate is increased significantly;

[0424] High - efficiency Compatibility: The calibration time for multi - platform switching is significantly shortened, supporting cross - scenario applications.

[0425] As Figure 2 shown, another embodiment of the present invention provides a system for implementing the above - mentioned microscopic positioning control method for adaptive visual - force fusion, including:

[0426] (1) Multi-modal data acquisition module:

[0427] A visual acquisition unit for capturing the image of the operation area in real time, extracting the target contour through an edge detection algorithm, and outputting the target pose data;

[0428] A force perception acquisition unit for obtaining the six-dimensional force / torque signal at the end of the operation in real time;

[0429] A temperature monitoring unit for a thermal sensor to collect the ambient temperature in real time;

[0430] (2) Asynchronous data synchronization module:

[0431] A data buffer for caching visual frame data and force perception sampling data respectively;

[0432] An interpolation alignment unit for data synchronization based on the visual frame timestamp and the force perception sampling timestamp, and generating a unified timestamp through a bilinear interpolation function;

[0433] A calibration compensation unit for storing preset compensation parameters;

[0434] (3) Spatial coordinate mapping module:

[0435] A calibration parameter database for storing calibration parameters;

[0436] A non-linear error compensation unit for running a convolutional neural network model for compensation;

[0437] A prediction unit based on deep learning;

[0438] (4) Data processing and enhancement module:

[0439] A visual enhancement unit for performing Gaussian filtering and Laplacian sharpening on the original image to generate an enhanced image;

[0440] A force perception calibration unit for compensating the sensor zero drift in real time according to the ambient temperature and the manipulator pose, and outputting a calibrated force signal;

[0441] (5) Multi-modal data fusion module:

[0442] A spatio-temporal registration unit for spatio-temporally aligning the mapped coordinates and the calibrated force signal, generating a synchronized data set, and denoising through a Kalman filter;

[0443] An attention weighting unit for dynamically allocating visual and force perception weights using a hybrid deep learning model;

[0444] A pose synthesis unit for generating a fused pose;

[0445] (6) Adaptive control module:

[0446] An Actor-Critic network-based reinforcement learning optimization unit for real-time adjustment of control parameters;

[0447] A sliding mode control unit for generating control commands;

[0448] An impedance adjustment unit for dynamically calculating equivalent mass and damping;

[0449] 7) Safety protection module:

[0450] A safety retraction unit for triggering nanoscale displacement compensation;

[0451] A fault emergency unit for performing exception handling;

[0452] A self-check and recovery unit for running historical data reconstruction algorithms;

[0453] (8) Calibration verification single module:

[0454] An error analysis unit for verifying pose tracking errors;

[0455] An adaptive compensation unit for calculating offset correction amounts.

[0456] Preferably, it further includes: a time alignment module for optimizing calibration compensation terms through a sliding window variance minimization algorithm and a spatial mapping module of a pre-trained deep neural network;

[0457] A force perception coupling calibration module for online compensating systematic offsets in contact force signals.

[0458] Preferably, it further includes:

[0459] A multi-source fusion processing module that uses an extended Kalman filter to achieve spatio-temporal registration. The multi-source fusion processing module includes: a state prediction module burned in the ROM of the kinematic model in the DSP and an observation update module that executes matrix mapping to extract visual poses and outputs them to the state space; a covariance calculation module that generates a synchronous dataset matrix and stores it in the dual-port RAM for the decision-making module to call;

[0460] A confidence evaluation module that calculates the gradient magnitude by the image processor and outputs it through a variance normalization unit;

[0461] A force gradient compensation module that includes a high-speed ADC and a differential calculation circuit.

[0462] Preferably, it further includes:

[0463] A high-frequency disturbance suppressor composed of a coprocessor and an adaptive gain controller, and an environmental stiffness estimator integrated in a force control feedback loop; a fault recovery module composed of a three-stage emergency trigger circuit, a historical data reconstruction unit, and a self-check protocol executor; a micro-nano reference positioning platform for verifying tracking errors and a calibration module for coordinate system offset compensation.

[0464] The data flow of the micro-positioning control system for adaptive vision-force fusion of the present invention is as follows:

[0465] 1. After the sensing module obtains the original data, it aligns the timestamps through the spatio-temporal synchronization module;

[0466] 2. The preprocessed data enters the coordinate mapping module to complete the spatial transformation;

[0467] 3. The fusion module integrates multi-modal information to generate control parameters;

[0468] 4. The intelligent control module receives the feedback from the safety module while outputting control instructions;

[0469] 5. The calibration verification module provides closed-loop calibration support.

[0470] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An adaptive microscopic positioning control method for visual and force sense fusion, characterized in that: Including the following steps: S1), Image capture: The image of the operation area is captured in real time by a high-frame-rate microscopic camera, and the target contour is extracted by an edge detection operator to output the target pose The six-dimensional force and torque signals at the operation end are obtained in real time through the MEMS force sensor array ; S2), Dynamic timestamp alignment: Based on the visual frame timestamp and the haptic sampling timestamp , Asynchronous data alignment is achieved through a bilinear interpolation function; S3), Nonlinear mapping of the spatial coordinate system: Establish a mapping model from the visual pixel coordinate system to the haptic Cartesian coordinate system ; S4), Visual data enhancement and feature extraction: For the original image perform Gaussian filtering and Laplacian sharpening to enhance the target edge features; S5), Dynamic calibration of force sense signals: According to the ambient temperature and the pose of the robotic arm , compensate for the zero drift of the MEMS force sense sensor in real time; S6), spatio-temporal registration of multi-modal data: Spatially and temporally register the coordinate mapping result output in step S3 with the calibrated force sense signal in step S5 to generate a synchronized data set; S7): Feature fusion based on the attention mechanism: Dynamically allocate visual and force sense weights through a hybrid deep learning model; S8), integrating pose generation and error correction: output the integrated pose ; S9), Real-time optimization of reinforcement learning parameters: Define the state space , and output the control parameter increment through the Actor-Critic network; S10), Generate a sliding mode surface for fuzzy sliding mode disturbance rejection control instruction design; S11), Dynamic impedance protection and real-time withdrawal: Dynamically adjust the impedance parameters according to the environmental stiffness ( ).

2. The microscopic positioning control method with adaptive visual and force sense fusion according to claim 1, characterized in that: S2) Dynamic Timestamp Alignment: Based on the visual frame timestamps and the haptic sampling timestamps , asynchronous data alignment is achieved through a bilinear interpolation function: ; wherein is a calibration compensation term; is the sampling frequency of the vision system; is the sampling frequency of the force sense system; Represents the unified timestamp after dynamic timestamp alignment; Calibration compensation term Optimized by the sliding window variance minimization algorithm: ; : The visual end - pose vector of the k - th sliding window sample point; : The k-th force perception end sampling point, the pose vector corresponding to the time lag after that; : The Euclidean norm squared, which is used to minimize the temporal alignment error between vision and force perception; : represents solving for the time compensation amount that minimizes the objective function; Among them is the number of data frames within the sliding window; S3) Nonlinear mapping of the spatial coordinate system: Establish a visual pixel coordinate system to the haptic Cartesian coordinate system mapping model: ; wherein is a calibration parameter, is the non - linear error compensated by the convolutional neural network, is the Z - axis prediction function based on deep learning; Z-axis prediction function Implemented by a pre-trained deep neural network (DNN), with the input being the local feature block of the image ( ), and the output being the predicted value of the Z-axis coordinate value ( ): ; In the formula, the input and output parameters include: : The input local feature block of the image is a local image block containing the target area extracted from the microscopic image; : The output Z-axis coordinate value, which is part of the force-sensing Cartesian coordinate system, represents the position of the target in the depth direction; The neural network parameters include: W h : The weight matrix of the hidden layer; : Bias vector of the hidden layer; : Rectified linear unit activation function; The neural network parameter output layer is as follows: : The weight matrix of the output layer of the deep neural network; : The bias scalar of the output layer; S4) Visual data enhancement and feature extraction: For the original image perform Gaussian filtering and Laplacian sharpening to enhance the target edge features: ; Wherein: : Original microscopic image data; is a Gaussian kernel; where σ = 1.5 pixels, representing the standard deviation of the Gaussian kernel; The convolution operation symbol indicates Gaussian filtering of the original image; is the Laplacian operator of the original image, which is used to extract the high-frequency edge features of the image; is the sharpening coefficient, which is used to control the intensity of Laplacian sharpening; For the enhanced image; S5) Dynamic calibration of force sense signal: According to the ambient temperature and the pose of the robotic arm , compensate the zero drift of the MEMS force sense sensor in real time: ; Among them : The original uncompensated contact force data, directly output by the MEMS force sensor : Temperature drift coefficient, which is used to quantify the influence of environmental temperature change on the zero-point offset of the MEMS force sensor; : The current ambient temperature, collected by a thermosensitive device or a thermocouple; : Ambient temperature during calibration, serving as the reference for temperature compensation; : Pose coupling weight matrix, which is a 2×1 or 2×2 vector or matrix; : The two-dimensional displacement vector of the current operation end in the visual coordinate system, with the unit of pixel or mm; : The finally calibrated force signal, after removing the pose interference, is used as the true feedback value for subsequent control calculations; S6) Spatiotemporal registration of multimodal data: Perform spatiotemporal registration on the coordinate mapping result output in step S3 and the calibrated force sense signal in step S5 to generate a synchronized data set: , and fuse the temporal noise through a Kalman filter; : Represents the structured data set after time alignment, coordinate calibration, and filtering fusion are completed; S7) Feature fusion based on the attention mechanism: Dynamically allocate visual and force sense weights through a hybrid deep learning model: ; =1- ; The above structure realizes dynamic weighting of the visual and force sense channels during the fusion process through the softmax normalization mechanism; Among them, the channel scoring factors , are respectively defined as: , ; The variable definitions are as follows: ∇I, ∇F: Feature gradient vectors of the image and the force sense signal respectively; , : It is the channel weight matrix for feature mapping, which is the parameter of a 1×N fully connected layer and is obtained through end-to-end training; ReLU(x)=max(0,x): Is the rectified linear unit activation function, introducing non-linearity to enhance the sparsity of feature expression; : is the visual modality attention weight, reflecting the importance of visual features in the current fusion process; , are the visual gradient and force gradient feature vectors; S8) Fusion of pose generation and error correction: Output the fused pose : ; ; Among them , is the force gradient compensation term; wherein, is the target pose coordinate output by the vision system, which comes from the image capture and edge detection results of step S1 and belongs to the vision pixel coordinate system; is the target coordinate in the coordinate system of the MEMS force sensor, obtained through the non-linear mapping of the space coordinate system in step S3, with the unit of physical length; is a weight parameter, the weight coefficient of visual data, which is dynamically calculated by the attention mechanism in step S7 and has a value range of [0, 1]; is the weight coefficient of the force perception data, which is also generated by the attention mechanism and is related to ; S9) Real-time optimization of reinforcement learning parameters: Define the state space , and output the control parameter increment through the Actor-Critic network: ; : The control gain adjustment amount of the visual channel, which is used to dynamically adjust the visual servo response intensity; : The control gain adjustment amount of the force sensing channel, which is used to dynamically adjust the force control stiffness or impedance; : Learning rate, which controls the speed of network weight update; : The sliding mode surface variable, representing the error metric between the current state and the target state; , : The network parameter matrix, corresponding to the weight matrices of the visual gain regulator and the force sense gain regulator respectively; , : Bias term, adjusting the intermediate offset of the network's non-linear output; The above structure constitutes a single-layer or double-layer neural network mapping for mapping the sliding mode error ( ) to the control gain adjustment amount to achieve adaptive control adjustment under different disturbance levels; S10) Generate a sliding mode surface for fuzzy sliding mode disturbance rejection control instruction design: ; Among them According to the perturbation level Fuzzy adjustment; In the formula: is pose error; sliding mode surface gain coefficient; nonlinear exponential parameter, μ ∈ (0, 1); sign function; Generate a disturbance rejection control instruction: ; : is the saturation function limiting parameter; : Visual channel sliding mode control gain coefficient, which controls the sensitivity of the system to the visual error sliding surface. The larger the value, the faster the system responds to changes in visual error; : Haptic channel anti-disturbance gain coefficient, which controls the action intensity of the sat function in the force channel and is commonly used to suppress haptic disturbances or external micro-vibrations; : Disturbance compensation gain, which adjusts the influence of force sense channel compensation on the controller; : is the change rate of the contact force, characterizing the trend of external disturbance; Among them , ; The sat(x) function is used to suppress the chattering of the sliding mode controller, and the output value is limited to the interval [-1, 1]; S11) Dynamic impedance protection and real-time withdrawal: Dynamically adjust the impedance parameters according to the environmental stiffness ( ) ; : The basic quality parameter, representing the reference quality of the system in an ideal rigid environment; : Environmental stiffness, dynamic estimated value, reflecting the rigidity of the current contact object; : Equivalent mass, which is used to convert the basic mass by combining with the environmental flexibility; : Equivalent damping, used for speed control / energy dissipation regulation in the retraction process; : Damping adjustment proportional factor; : Controller equivalent stiffness coefficient, modeling the compliant response ability of the controller to end disturbances; ; : Represents the fine-tuning vector of the end position when the protection is triggered, with the unit of nanometer (nm); : Retreating ratio factor; : The current contact force collected in real time by the MEMS force sensor; : Preset contact safety threshold, which triggers a retraction when occurs; ∇x: Is the unit velocity direction vector of the current position.

3. The microscopic positioning control method for adaptive visual and force sense fusion according to claim 2, characterized in that: The size of the Gaussian kernel G(σ) is 5×5, and the sharpening coefficient λ is dynamically adjusted according to the image signal-to-noise ratio SNR: ; In the dynamic calibration of the force sense signal, the pose coupling weight matrix (W) is determined through an off-line calibration experiment, satisfying: ; : It is a two-dimensional coupling weight matrix used to model the systematic perturbation of the contact force signal caused by the displacement of the image coordinates; : represents the position variable The partial derivative of the change with respect to the force signal , which reflects the sensitivity of the force sensing channel to image displacement; All parameters are obtained through linear fitting during the calibration stage and are used for position compensation during actual operation.

4. The microscopic positioning control method for adaptive visual and force sense fusion according to claim 2, wherein: Fuse the time series noise through a Kalman filter, and the state equation and the observation equation are respectively: ; : The system state variable at the current moment, including the pose information estimated by visual-force fusion; : An observed quantity, such as the center coordinates and direction of the end image obtained by microscopic image processing; where A is the state transition matrix and H is the observation matrix, , is Gaussian noise; Fused synchronous dataset Contains the visual-force covariance matrix Cvf with timestamp alignment; In the attention mechanism, the visual feature confidence is calculated based on the normalized variance of the image gradient magnitude as follows: ; : Variance of the image gradient distribution, reflecting the overall texture or edge complexity; : The maximum gradient value in the image, corresponding to the most prominent edge, is used for normalization; : A minimum constant used to avoid division-by-zero errors during the image contrast normalization process; ∇I: Image gray gradient, calculated by the image Sobel operator, used to represent the image edge intensity; Force gradient compensation term , by real-time differential calculation: ; : represents the time-difference gradient of the contact force, which is used to detect the rapid change trend of the force; : The contact force values collected at the current moment and the previous moment respectively, which are collected by the MEMS force sensor at a high frequency; : The time interval between two adjacent sampling instants, which is set by the system sampling frequency; : It is the time calibration coefficient, which is used to scale the time reference of the differential result so that it uniformly corresponds to the drawdown response rate; In the Actor-Critic network, the Critic network outputs the state value function , and updates the weights through the temporal difference (TD) error: ; : Temporal Difference error (TDerror), which represents the gap between the current predicted value and the target return, is the core basis for updating the policy and the Critic network; : The immediate reward value, generated by the environmental feedback, is used to measure the effect of the current action; : Discount factor, which controls the influence weight of future rewards; : Learning rate, which determines the step size for each update of the network parameters; : The gradient of the Critic network with respect to its parameters ; The action exploration strategy adds Gaussian noise , and the noise variance exponentially decays with the number of training rounds: ; : The standard deviation of the Gaussian policy perturbation at the current moment, which controls the magnitude of randomness in the action output; : is the initial perturbation intensity, adjusted according to the environmental complexity; : It is the annealing time constant and affects the noise attenuation rate.

5. The microscopic positioning control method with adaptive visual and force sense fusion according to claim 2, characterized in that: Disturbance level It is divided into three levels: low, medium, and high. The value of is The value of ; Superimposed high-frequency disturbance suppression term ( ), whose frequency ( ) is estimated in real time by FFT, and the suppression gain is adaptively adjusted: ; : High-frequency disturbance force vector, obtained by band-pass filtering or fast Fourier transform of the original force sense signal, reflecting the disturbance intensity; : The original sequence of contact force, continuously collected by the MEMS force sensor; : represents the two-norm of a vector and is used to calculate the perturbation magnitude; : The maximum value of the force perception signal within the current sliding window, used for normalization; Environmental stiffness Online estimation through the slope of the force-displacement curve: ; : The change in contact force within the time window, measured by the MEMS force sensor; : The displacement change of the end during this time period, which is calculated by the vision system; →0: Represents approximate differentiation within an extremely short time, used for calculating the response slope of stiffness in real time; Pose gradient Calculated by the optical flow method of visual features: ; : The two-dimensional pixel coordinates of the target in the current frame image; : Represents the velocity vector component of the image target, calculated by the optical flow method; : is the estimated result of the final pose change rate; : The time interval between two adjacent frames, with the unit of millisecond (ms), is determined by the visual frame rate.

6. The microscopic positioning control method with adaptive visual and force sense fusion according to claim 1, characterized in that: It further includes step S12, fault recovery: When it is detected that the MEMS force sense sensor signal mutation ΔF>5 mN or the communication delay Td>2 ms, trigger a three-level emergency response: a. Pause the movement of the robotic arm and switch to the safe impedance mode; b. Reconstruct the target pose based on the weighted average of historical data, with weight w i satisfying ∑w i = 1; c. If the continuous fault timeout Tfault>5 s, start the self-check protocol and report the error code; It further includes step S13, performance calibration: Verify that the pose tracking error ep≤0.1 μm through a standard micro-nano reference device; When switching between multiple platforms, compensate for the coordinate system offset through an adaptive calibration algorithm: ; : Represents the position offset error of a single calibration point in visual measurement, which is obtained by manual or algorithm extraction; : It represents the number of target points participating in the calibration average, where ≥10 is the number of calibration points; : The translation offset correction amount during system initialization, which is used to eliminate systematic deviations in overall image recognition and is uniformly compensated in all subsequent coordinate conversions.

7. A system for implementing the microscopic positioning control method of adaptive visual and force fusion according to any one of claims 1-6, characterized in that, Including (1) Multi-modal data acquisition module: A visual acquisition unit used to capture images of the operation area in real time and extract the target contour through an edge detection algorithm; output the target pose data; A force sense acquisition unit used to obtain the six-axis force / torque signal of the operation end in real time; A temperature monitoring unit used to collect the ambient temperature in real time through a thermal sensor; (2) Asynchronous data synchronization module: A data buffer used to cache the visual frame data and the force sense sampling data respectively; Based on the visual frame timestamp and the force sensing sampling timestamp, a data synchronization is achieved through a bilinear interpolation function to generate an interpolation alignment unit with a unified timestamp; A calibration compensation unit for storing preset compensation parameters; (3) Spatial coordinate mapping module: A calibration parameter database for storing calibration parameters; A non-linear error compensation unit that runs a convolutional neural network model for compensation; A prediction unit that implements deep learning-based prediction; (4) Data processing and enhancement module: A visual enhancement unit for performing Gaussian filtering and Laplacian sharpening on the original image to generate an enhanced image; A force sensing calibration unit that compensates for the zero drift of the sensor in real time according to the environmental temperature and the pose of the robotic arm and outputs a calibrated force signal; (5) Multimodal data fusion module: A spatio-temporal registration unit for spatio-temporally aligning the mapped coordinates and the calibrated force signal to generate a synchronized data set and denoising it through a Kalman filter; An attention weighting unit that uses a hybrid deep learning model to dynamically allocate visual and force sensing weights; A pose synthesis unit for generating a fused pose; (6) Adaptive control module: A reinforcement learning optimization unit that adjusts control parameters in real time through an Actor-Critic network; A sliding mode control unit for generating control commands; An impedance adjustment unit that dynamically calculates the equivalent mass and damping; (7) Safety protection module: A safety retraction unit for triggering nanoscale displacement compensation; A fault emergency unit for performing exception handling; A self-check recovery unit for running a historical data reconstruction algorithm; (8) Calibration verification single module: An error analysis unit for verifying the pose tracking error; An adaptive compensation unit for calculating the offset correction amount.

8. The system of the microscopic positioning control method for adaptive visual and force sense fusion according to claim 7, wherein, It also includes: A time alignment module that optimizes the calibration compensation term through a sliding window variance minimization algorithm and a spatial mapping module of a pre-trained deep neural network; A force sensing coupling calibration module for online compensating for systematic offsets in the contact force signal.

9. The system of the microscopic positioning control method for adaptive visual and force sense fusion according to claim 7, wherein It further includes: A multi-source fusion processing module that uses an extended Kalman filter to achieve spatio-temporal registration. The multi-source fusion processing module includes: a state prediction module burned in the kinematic model ROM of the DSP and an observation update module that executes the matrix mapping visual pose extraction module output to the state space; a covariance calculation module that generates a synchronized data set matrix and stores it in the dual-port RAM for the decision-making module to call; a confidence evaluation module that calculates the gradient magnitude by the image processor and outputs it through a variance normalization unit; a force gradient compensation module that includes a high-speed ADC and a differential calculation circuit.

10. The system of the microscopic positioning control method for adaptive visual and force sense fusion according to claim 7, characterized in that, It further includes: A high-frequency disturbance suppressor composed of a coprocessor and an adaptive gain controller and an environmental stiffness estimator integrated in the force control feedback loop, a fault recovery module composed of a three-stage emergency trigger circuit, a historical data reconstruction unit and a self-check protocol executor, and a calibration module for verifying the tracking error and compensating for the coordinate system offset of the micro-nano reference component positioning platform.

Citation Information

Patent Citations

  • Virtual force feedback remote nano operation platform based on scanning electron microscope and method for realizing virtual force sensing interacting

    CN103092346A

  • Robot pose estimation method based on laser point cloud and visual SLAM

    CN119596328A