Microscopic positioning control method and system for self-adaptive vision and force sense fusion

Through the adaptive visual force fusion method, the spatiotemporal asynchronousness and dimensional heterogeneity of visual and force sensation data in micropositioning control are solved, and the submicron-level fusion positioning accuracy and operation stability are achieved, meeting the real-time requirements of minimally invasive surgery.

CN120259435AActive Publication Date: 2025-07-04NANCHANG INST OF TECH

Patent Information

Application Number
CN202510742094.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The existing micropositioning control technology has problems with spatiotemporal asynchronousness and dimension heterogeneity of visual and force-sensing data in complex microenvironments, resulting in fusion lag and noise amplification, and it is impossible to adapt to the mechanical properties of the operating target and environmental disturbances, making it difficult to meet the real-time needs of scenarios such as minimally invasive surgery.

Method used

Adaptive visual force-conscious fusion method is adopted, and the synchronization and dynamic weight allocation of visual and force-conscious data are achieved through bilinear interpolation timestamp alignment, nonlinear coordinate mapping based on deep learning, visual data augmentation and feature extraction, force-conscious signal calibration, multimodal data spatiotemporal registration, attention mechanism-based feature fusion, reinforcement learning parameter optimization, fuzzy slip mode immunity control and dynamic impedance protection.

Benefits of technology

It achieves submicron-level fusion positioning accuracy, improves data reliability and operation stability in complex environments, meets the real-time requirements of minimally invasive surgery, and avoids overload damage to brittle samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259435A_ABST
    Figure CN120259435A_ABST
Patent Text Reader

Abstract

The invention provides a self-adaptive visual sense and force sense fusion microscopic positioning control method and system. The control method comprises the following steps: capturing an image; aligning the dynamic timestamps; carrying out non-linear mapping on a space coordinate system; visual data enhancement and feature extraction; dynamically calibrating the force sense signal; carrying out space-time registration on the multi-modal data; feature fusion based on an attention mechanism; carrying out fusion pose generation and error correction; real-time optimization of reinforcement learning parameters; generating a fuzzy sliding mode anti-interference control instruction; performing dynamic impedance protection and real-time withdrawing; fault recovery; and performance calibration. According to the method, through bilinear interpolation timestamp alignment and nonlinear coordinate mapping based on deep learning, space-time asynchronism and dimension isomerism of vision and force sense data are eliminated, and submicron-level fusion positioning precision is achieved; based on an attention mechanism and force gradient compensation, visual and haptic weights are dynamically distributed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of optical measurement, and particularly to a microscopic positioning control method and system for adaptive visual-force fusion. Background Art

[0002] Microscopic positioning control technology is the core support in fields such as minimally invasive surgery, biological cell manipulation, and micro-nano device assembly. The accuracy and robustness of its control strategy directly determine the success, failure, and safety of microscopic operations. In biomedical scenarios, it is necessary to perform non-destructive manipulation of fragile biological samples (such as cells, tissues) with nanometer-level accuracy; in the field of micro-nano manufacturing, high-stability assembly of sub-micron-level devices is required. The common challenges in such operating environments lie in dynamic complexity (such as liquid disturbance, target deformation) and the extreme sensitivity of the operating object. There is an urgent need for an intelligent control method that can sense environmental changes in real time, adaptively adjust control parameters, and ensure operation safety.

[0003] Existing microscopic positioning control technologies mostly rely on single-modal feedback mechanisms. For example, the vision servo-based control method extracts the target pose through microscopic images, but it is vulnerable to optical noise, changes in sample transparency, and dynamic occlusion interference, resulting in the accumulation of pose estimation errors; while the force feedback-based control method can sense the contact force in real time, but it cannot provide effective displacement information in the non-contact stage and is prone to misjudgment due to environmental vibration. In addition, traditional multi-modal control strategies often adopt simple serial feedback or fixed-weight data fusion, ignoring the spatio-temporal asynchrony and dimensional heterogeneity between visual and force data, resulting in fusion lag or noise amplification, seriously affecting the dynamic response accuracy of the control system. More critically, existing control methods mostly adopt static parameter configuration, unable to adapt to sudden changes in the mechanical properties of the operation target or environmental disturbances, easily causing positioning deviation, excessive contact force, or even sample damage.

[0004] Although some studies have tried to introduce multi-modal fusion control in recent years, they still face the following bottleneck problems: First, the lack of a dynamic spatio-temporal registration mechanism makes it difficult to solve the spatio-temporal misalignment and heterogeneous feature alignment problems between visual and force data; second, the static optimization mode of control parameters cannot respond to the dynamic changes of the microscopic environment in real time, resulting in insufficient robustness; third, the computational delay of traditional control architectures is relatively high, making it difficult to meet the millisecond-level real-time requirements of scenarios such as minimally invasive surgery. These problems seriously restrict the reliability and universality of microscopic positioning control in complex microenvironments. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the existing defects and provide a microscopic positioning control method and system for adaptive visual-force fusion, effectively solving the problems in the background art.

[0006] To achieve the above object, the present invention proposes: The present invention adopts the following technical solutions to achieve: A microscopic positioning control method for adaptive visual and force sense fusion, comprising the following steps: S1), Image capture: The operation area image is captured in real time by a high-frame-rate microscopic camera, the adaptive exposure algorithm is used to eliminate optical noise, and the target contour is extracted by an edge detection operator, and the target pose is output

[0007] The six-dimensional force and torque signals at the operation end are obtained in real time through the MEMS force sense sensor array , and the environmental vibration interference is suppressed by a band-pass filter; S2), Dynamic timestamp alignment: Based on the visual frame timestamp and the force sense sampling timestamp , asynchronous data alignment is achieved through a bilinear interpolation function: ; wherein is the calibration compensation term, which is dynamically optimized by the sliding window variance minimization algorithm; is the sampling frequency of the visual system; is the sampling frequency of the force sense system; represents the unified timestamp after dynamic timestamp alignment; S3), Nonlinear mapping of the spatial coordinate system: Establish a mapping model from the visual pixel coordinate system to the force sense Cartesian coordinate system : ; wherein is the calibration parameter, is the nonlinear error compensated by the convolutional neural network, is the Z-axis prediction function based on deep learning; S4), Visual data enhancement and feature extraction: Gaussian filtering and Laplacian sharpening are performed on the original image to enhance the target edge features: ; wherein: : Original microscopic image data; is the Gaussian kernel, which is used to smooth the image; where σ = 1.5 pixels, representing the standard deviation of the Gaussian kernel, controlling the smoothness of the filtering; is the convolution operation symbol, indicating that Gaussian filtering is performed on the original image; is the Laplacian operator of the original image, which is used to extract the high-frequency edge features of the image; is the sharpening coefficient, which is used to control the intensity of Laplacian sharpening; is the enhanced image, which combines the effects of Gaussian filtering and Laplacian sharpening; S5), force sense signal dynamic calibration: According to the environmental temperature and the pose of the robotic arm , compensate the zero-point drift of the MEMS force sense sensor in real time: ; Among them: : The original uncompensated contact force data, which is directly output by the MEMS force sense sensor; : The temperature drift coefficient, which is used to quantify the influence of environmental temperature change on the zero-point offset of the MEMS force sense sensor; : The current environmental temperature, which is collected by a thermosensitive device or a thermocouple; : The environmental temperature during calibration, which is used as the reference for temperature compensation; : The pose coupling weight matrix, which is a 2×1 or 2×2 vector or matrix, reflecting the interference weight of the pose on the force offset; : The two-dimensional displacement vector of the current operation end in the visual coordinate system, with the unit of pixel or mm; : The finally calibrated force signal, which eliminates the pose interference and is used as the true feedback value for subsequent control calculations; S6), multi-modal data spatio-temporal registration: Perform spatio-temporal registration on the coordinate mapping result output by S3 and the force sense signal calibrated by S5 to generate a synchronized data set , and fuse the temporal noise through a Kalman filter; : Represents the structured data set after time alignment, coordinate correction, and filtering fusion, which is used for subsequent deep control model training or state estimation; S7): Feature fusion based on the attention mechanism: Dynamically allocate the visual and force sense weights through a hybrid deep learning model: ; = 1 - ; The above structure realizes the dynamic weighting of the visual and force sense channels during the fusion process through the softmax normalization mechanism; Among them, the channel scoring factor , are respectively defined as: , ; The variables are defined as follows: ∇I, ∇F: Feature gradient vectors of the image and the force sense signal respectively; , : Channel weight matrix for feature mapping, parameters of a 1×N fully connected layer, obtained through end-to-end training; ReLU(x) = max(0, x): Rectified linear unit activation function, introducing non-linearity to enhance the sparsity of feature expression; : Visual modality attention weight, reflecting the importance of visual features in the current fusion process; , is the visual gradient and force gradient feature vector; S8), Fusion pose generation and error correction: Output the fusion pose : ; ; where , is the force gradient compensation term; In the formula, is the target pose coordinate output by the visual system, from the image capture and edge detection results of step S1, belonging to the visual pixel coordinate system; is the target coordinate in the MEMS force sense sensor coordinate system, obtained through the non-linear mapping of the space coordinate system in step S3, with the unit of physical length; is the weight parameter, the weight coefficient of visual data, dynamically calculated through the attention mechanism in step S7, with the value range of [0, 1]; is the weight coefficient of force sense data, also generated by the attention mechanism, and ; S9), Real-time optimization of reinforcement learning parameters: Define the state space (pose error, force error, disturbance level), and output the control parameter increment through the Actor-Critic network: ; : Control gain adjustment amount of the visual channel, used to dynamically adjust the visual servo response intensity; : The control gain adjustment amount of the force perception channel, which is used to dynamically adjust the force control stiffness or impedance; : The learning rate, which controls the speed of network weight update; : The sliding mode surface variable, which represents the error metric between the current state and the target state; 、 : The network parameter matrix, which respectively corresponds to the weight matrices of the visual gain regulator and the force perception gain regulator; 、 : The bias term, which adjusts the middle offset of the network's non-linear output; The above structure constitutes a single-layer or double-layer neural network mapping, which is used to map the sliding mode error to the control gain adjustment amount, and realize the adaptive control adjustment under different disturbance levels; S10), generating the fuzzy sliding mode disturbance rejection control command: designing the sliding mode surface ; where fuzzy adjustment according to the disturbance level ; In the formula: is the pose error; the sliding mode surface gain coefficient; the non-linear exponential parameter (μ ∈ (0, 1); the sign function ( outputs +1 when >0, otherwise -1); generating the disturbance rejection control command: ; : the saturation function limiting parameter; : the visual channel sliding mode control gain coefficient, which controls the sensitivity of the control system to the visual error sliding mode surface. The larger the value, the faster the system responds to the change of visual error; : the force perception channel disturbance rejection gain coefficient, which controls the action intensity of the sat function in the force channel and is often used to suppress the force perception disturbance or external micro-vibration; : the disturbance compensation gain, which adjusts the influence of the force perception channel compensation on the controller; : the change rate of the contact force, which characterizes the trend of external disturbance; where , ; sat(x) This function is used to suppress the chattering of the sliding mode controller, and the output value is limited to the interval [-1,1]; S11), dynamic impedance protection and real-time withdrawal: according to the environmental stiffness Dynamically adjust impedance parameters : ; in: : Basic quality parameter, which indicates the reference quality of the system under an ideal rigid environment; : Environmental stiffness, a dynamic estimated value, reflecting the stiffness of the current contact object; : Equivalent mass, which is used to convert the foundation mass in combination with the environmental flexibility; : Equivalent damping, used for speed control / energy dissipation regulation in the withdrawal process control; : Damping adjustment proportional factor; : Controller equivalent stiffness coefficient, modeling the controller's compliant response capability to terminal disturbances; ; : Indicates the end position fine-tuning vector when the protection is triggered, in nanometers (nm); : concession scale factor; : The current contact force collected in real time by the MEMS force sensor; :Preset contact safety threshold, when When the retracement is triggered; ∇x: is the unit velocity direction vector of the current position.

[0008] Preferably, in step S2, in the dynamic timestamp alignment, the compensation item is calibrated Optimization via sliding window variance minimization algorithm: ; in: : The visual end pose vector of the k-th sliding window sample point; : kth force sensing end sampling point, time lag The corresponding pose vector; : Euclidean norm squared, used to minimize the temporal alignment error between vision and force perception; : It represents solving for the time compensation amount that minimizes the objective function; wherein is the number of data frames within the sliding window; In step S3 of the non - linear mapping of the spatial coordinate system, the Z - axis prediction function is implemented through a pre - trained deep neural network (DNN). The input is the local feature block of the image , and the output is the predicted value of: ; In the formula, the input and output parameters include: : The input local feature block of the image is a local image block containing the target area extracted from the microscopic image; : The output Z - axis coordinate value, which is part of the force - perception Cartesian coordinate system, represents the position of the target in the depth direction; The neural network parameters include: W h : The weight matrix of the hidden layer; : The bias vector of the hidden layer; : The rectified linear unit activation function; The output layer of the neural network parameters is as follows: : The weight matrix of the output layer of the deep neural network; : The bias scalar of the output layer.

[0009] Preferably, in step S4 of visual data enhancement, the size of the Gaussian kernel G(σ) is 5×5, and the sharpening coefficient λ is dynamically adjusted according to the signal - to - noise ratio SNR of the image (the signal - to - noise ratio of the local image block, estimated from the image mean and gradient noise): ; In the dynamic calibration of the force - perception signal, the pose - coupling weight matrix W is determined through an offline calibration experiment and satisfies: ; : The two - dimensional coupling weight matrix is used to model the systematic perturbation of the contact - force signal caused by the displacement of the image coordinates; : Represents the position variable (such as the image coordinates ) the partial derivative of the change in the force signal , reflecting the sensitivity of the force - perception channel to image displacement; All parameters are obtained by linear fitting during the calibration phase (calibration platform + known pose changes) and are used for position compensation during actual operation.

[0010] Preferably, the step S6 further includes: during the spatio-temporal registration process, an extended Kalman filter (EKF) is used to fuse the temporal noise, and the state equation and the observation equation are respectively: ; : The system state variable at the current moment, including the pose information estimated by visual-force fusion; : The observed quantity, such as the center coordinates and direction of the end image obtained by microscopic image processing; where A is the state transition matrix and H is the observation matrix, , is Gaussian noise; The fused synchronous data set contains the visual-force covariance matrix Cvf with aligned timestamps.

[0011] Preferably, the step S7 further includes: in the attention mechanism, the visual feature confidence is calculated based on the normalized variance of the image gradient magnitude : ; : The variance of the image gradient distribution, reflecting the overall texture or edge complexity; : The maximum value of the gradient in the image, corresponding to the most significant edge, used for normalization; : A minimum constant, used to avoid division-by-zero errors during the image contrast normalization process; ∇I: The image gray gradient, calculated by the image Sobel operator, used to represent the image edge intensity; The force gradient compensation term , is calculated by real-time difference: ; : Represents the time difference gradient of the contact force, used to detect the rapid change trend of the force; : Are respectively the contact force values collected at the current moment and the previous moment, collected by the MEMS force sensor at a high frequency; : Is the time interval between two consecutive sampling moments, set by the system sampling frequency; : is the time calibration coefficient, which is used to scale the time reference of the differential result so that it uniformly corresponds to the drawdown response rate (set to 0.1 ms here).

[0012] Preferably, the step S9 further includes: in the ActorCritic network, the Critic network outputs the state value function , and updates the weights through the temporal difference (TD) error: ; : Temporal difference error (TDerror), which represents the gap between the current predicted value and the target return, and is the core basis for updating the policy and the Critic network; : Immediate reward value, which is generated by the environment feedback and is used to measure the effect of the current action; : Discount factor, which controls the influence weight of future rewards; : Learning rate, which determines the step size of each update of the network parameters; : Gradient of the Critic network with respect to its parameter ; Gaussian noise is added to the action exploration strategy , and the noise variance decays exponentially with the number of training rounds: ; Where: : Standard deviation of the Gaussian policy perturbation at the current moment, which controls the randomness size in the action output; : Initial perturbation intensity, which is adjusted according to the environmental complexity; : Annealing time constant, which affects the noise decay speed; This formula realizes exponential annealing, enabling the agent to explore sufficiently in the initial stage of training and tending to be stable in the later stage of training.

[0013] Preferably, the step S10 further includes: The perturbation level is divided into three levels: low, medium, and high, The value of , The value of ; A high-frequency perturbation suppression term is superimposed, and its frequency is estimated in real time by FFT, and the suppression gain is adaptively adjusted: ; Wherein: : The high-frequency disturbance force vector is obtained by band-pass filtering or fast Fourier transform (FFT) of the original force sense signal, reflecting the disturbance intensity; : The original sequence of contact forces is continuously collected by the MEMS force sense sensor; : Represents the two-norm (magnitude) of the vector, used to calculate the disturbance magnitude; : The maximum value of the force sense signal within the current sliding window, used for normalization.

[0014] Preferably, the S11 further includes: environmental stiffness Estimated online through the slope of the force-displacement curve: ; : The change in contact force within the time window, measured by the MEMS force sense sensor; : The change in displacement of the end within this time period, calculated by the vision system; →0: Represents approximate differentiation in an extremely short time, used to calculate the response slope of the stiffness in real time; In the nanoscale retraction, the pose gradient Calculated by the vision feature optical flow method: ; : The two-dimensional pixel coordinates of the target in the current frame image; : Represents the velocity vector component of the image target, calculated by the optical flow method; : The final estimated result of the pose change rate; : The time interval between two adjacent frames, in milliseconds (ms), determined by the vision frame rate.

[0015] Preferably, it further includes S12, fault recovery: When it is detected that the signal mutation of the MEMS force sense sensor ΔF>5mN or the communication delay Td>2ms, a three-level emergency response is triggered: a. Pause the movement of the robotic arm and switch to the safe impedance mode; b. Reconstruct the target pose based on the weighted average of historical data, and the weight w i Satisfies ∑w i =1; c. If the continuous fault timeout Tfault>5s, start the self-check protocol and report the error code.

[0016] Preferably, it further includes S13, performance calibration: verifying that the pose tracking error ep ≤ 0.1 μm through a standard micro-nano reference device; When switching between multiple platforms, compensating for the coordinate system offset through an adaptive calibration algorithm: ; : Represents the position offset error of a single calibration point in visual measurement (such as pixel error or millimeter deviation), obtained through manual or algorithm extraction; : Represents the number of target points participating in the average of calibration, where ≥10 is the number of calibration points; : Is the translational offset correction amount during system initialization, used to eliminate systematic deviations in overall image recognition, and uniformly compensated in all subsequent coordinate conversions.

[0017] The present invention also provides a system for implementing the above-mentioned microscopic positioning control method for adaptive visual force sense fusion, including: 1) Multimodal data acquisition module: A visual acquisition unit for capturing images of the operation area in real time, extracting the target contour through an edge detection algorithm, and outputting target pose data; A force sense acquisition unit for obtaining six-dimensional force / torque signals at the operation end in real time; A temperature monitoring unit for collecting the ambient temperature in real time through a thermal sensor; (2) Asynchronous data synchronization module: A data buffer for caching visual frame data and force sense sampling data respectively; An interpolation alignment unit for realizing data synchronization based on the visual frame timestamp and the force sense sampling timestamp through a bilinear interpolation function, and generating a unified timestamp; A calibration compensation unit for storing preset compensation parameters; (3) Spatial coordinate mapping module: A calibration parameter database for storing calibration parameters; A non-linear error compensation unit for running a convolutional neural network model for compensation; A prediction unit for realizing deep learning-based prediction; (4) Data processing and enhancement module: A visual enhancement unit for performing Gaussian filtering and Laplacian sharpening on the original image to generate an enhanced image; A force sense calibration unit for compensating the zero drift of the sensor in real time according to the ambient temperature and the pose of the robotic arm, and outputting a calibrated force signal; 5) Multimodal data fusion module: A spatio-temporal registration unit for spatio-temporally aligning the mapped coordinates with the calibrated force signal, generating a synchronized data set, and denoising through a Kalman filter; An attention weighting unit that dynamically assigns visual and force weights using a hybrid deep learning model; A pose synthesis unit for generating a fused pose; (6) Adaptive control module: A reinforcement learning optimization unit that adjusts control parameters in real time through an Actor-Critic network; A sliding mode control unit for generating control commands; An impedance adjustment unit that dynamically calculates equivalent mass and damping; (7) Safety protection module: A safety retraction unit for triggering nano-level displacement compensation; A fault emergency unit for performing exception handling; A self-check recovery unit for running a historical data reconstruction algorithm; (8) Calibration verification single module: An error analysis unit for verifying pose tracking errors; An adaptive compensation unit for calculating offset correction amounts.

[0018] Preferably, it further includes: a time alignment module that optimizes calibration compensation terms through a sliding window variance minimization algorithm and a spatial mapping module of a pre-trained deep neural network; A force sense coupling calibration module for online compensating systematic offsets in contact force signals.

[0019] Preferably, it further includes: A multi-source fusion processing module that implements spatio-temporal registration using an extended Kalman filter. The multi-source fusion processing module includes: a state prediction module burned in the kinematic model ROM of the DSP and an observation update module that executes matrix mapping to extract visual poses and outputs them to the state space; a covariance calculation module that generates a synchronized data set matrix and stores it in a dual-port RAM for the decision-making module to call; A confidence evaluation module that calculates the gradient magnitude by an image processor and outputs it through a variance normalization unit; A force gradient compensation module that includes a high-speed ADC and a differential calculation circuit.

[0020] Preferably, it further includes: A high-frequency disturbance suppressor composed of a coprocessor and an adaptive gain controller and an environmental stiffness estimator integrated in the force control feedback loop, a fault recovery module composed of a three-stage emergency trigger circuit, a historical data reconstruction unit, and a self-check protocol executor; a micro-nano reference component positioning platform for verifying tracking errors and a calibration module for coordinate system offset compensation.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Through bilinear interpolation timestamp alignment and deep learning-based non-linear coordinate mapping, this application eliminates the spatio-temporal asynchrony and dimensional heterogeneity of visual and haptic data, and achieves sub-micron fusion positioning accuracy.

[0022] 2. Based on the attention mechanism and force gradient compensation, the visual and haptic weights are dynamically allocated to suppress noise interference and enhance key features, improving data reliability in complex environments; through the Actor-Critic network, the visual servo gain and force control stiffness are adjusted online to adapt to sudden changes in the mechanical characteristics of the operation target, reducing positioning deviation and contact force overrun.

[0023] 3. Combining the fuzzy rule base with the high-frequency disturbance suppression term, it effectively resists dynamic interferences such as liquid fluctuations and mechanical vibrations, improving operation stability; through hardware-accelerated spatio-temporal registration and feature extraction, the data processing delay is compressed to the millisecond level to meet the real-time requirements of minimally invasive surgery; based on the environmental stiffness, the impedance parameters are adjusted in real time, combined with the nano-level retraction mechanism, to avoid overloading damage to brittle samples, and the operation safety is highly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is the control flow chart of a microscopic positioning control method for adaptive visual-haptic fusion of the present invention; Figure 2 is the system framework diagram of a microscopic positioning control method for adaptive visual-haptic fusion of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] As Figure 1 shown, a microscopic positioning control method for adaptive visual-haptic fusion includes the following steps: S1), Image capture: The operation area image is captured in real time by a high-frame-rate microscopic camera, the optical noise is eliminated by an adaptive exposure algorithm, and the target contour is extracted by an edge detection operator, and the target pose is output

[0027] The six-dimensional force and torque signals at the operation end are obtained in real time through the MEMS sensor array , and the environmental vibration interference is suppressed by a band-pass filter; For force feedback control: Detect the contact force in real time to avoid damaging brittle samples (such as cells) due to overload.

[0028] Dynamically adjust the pose: Combine visual data to correct the motion trajectory of the robotic arm and improve the positioning accuracy.

[0029] Safety protection: Trigger a protection mechanism (such as nanoscale retraction) when the contact force exceeds the limit to ensure the safety of the operation.

[0030] S2), Dynamic timestamp alignment: Based on the visual frame timestamp and the force sensing sampling timestamp , Achieve asynchronous data alignment through a bilinear interpolation function: Bilinear interpolation uses the force sensing data points at adjacent timestamps and calculates the force sensing signal value corresponding to the visual frame timestamp by weighted calculation according to the time ratio.

[0031] Asynchronous data: It is the misaligned data caused by the asynchronous acquisition times of visual (high frame rate) and force sensing (different sampling rates) data in step S1.

[0032] Alignment purpose: Eliminate the timing difference between visual and force sensing data, ensure that the two are strictly synchronized in time, and provide a consistent basis for multi-modal fusion.

[0033] ; where is the calibration compensation term, which is dynamically optimized by the sliding window variance minimization algorithm; is the sampling frequency of the visual system; is the sampling frequency of the force sensing system; represents the unified timestamp after dynamic timestamp alignment.

[0034] S3), Nonlinear mapping of the spatial coordinate system: Establish a mapping model from the visual pixel coordinate system to the force sensing Cartesian coordinate system : Solve the spatial heterogeneity: Visual data is in two-dimensional pixel coordinates, and force sensing data is in three-dimensional Cartesian coordinates. The mapping model unifies the two into the same spatial framework.

[0035] Improve the positioning accuracy: Compensate for the nonlinear error (such as lens distortion) through deep learning to achieve sub-micron coordinate alignment.

[0036] Support multi-modal fusion: Provide the basic data for spatial alignment for subsequent spatio-temporal registration (S6).

[0037] ; wherein is the calibration parameter, is the non - linear error compensated by the convolutional neural network, is the Z - axis prediction function based on deep learning.

[0038] S4), Visual data enhancement and feature extraction: For the original image perform Gaussian filtering and Laplacian sharpening to enhance the target edge features: Gaussian filtering σ = 1.5: Smooth the image, suppress optical noise, and improve the signal - to - noise ratio.

[0039] λ = 0.8 is the sharpening coefficient, which is used to control the intensity of Laplacian sharpening, enhance the high - frequency components of the target edge, highlight the contour features, and facilitate subsequent edge detection and pose estimation.

[0040] Comprehensive effect: Improve the image quality and ensure the robustness and accuracy of feature extraction.

[0041] ; wherein: : The original microscopic image data; is the Gaussian kernel, which is used to smooth the image; where σ = 1.5 pixels, representing the standard deviation of the Gaussian kernel, controlling the smoothness of filtering; The convolution operation symbol, indicating performing Gaussian filtering on the original image; is the Laplacian operator (second - order derivative) of the original image, which is used to extract the high - frequency edge features of the image; is the enhanced image, combining the effects of Gaussian filtering (noise suppression) and Laplacian sharpening (edge highlighting); S5), Force - sense signal dynamic calibration: According to the environmental temperature and the robotic arm pose , compensate the zero - point drift of the MEMS force - sense sensor in real - time: Compensate the zero - point drift: The MEMS force - sense sensor is vulnerable to the influence of environmental temperature T and the robotic arm pose (x, y), resulting in zero - point offset.

[0042] Correct the error in real - time: Through the temperature coefficient α_T and the pose coupling weight matrix W, dynamically calibrate the force - sense signal, Improve data reliability: Ensure that the force - sense signal accurately reflects the real contact force and reduce misjudgment.

[0043] ; wherein is the temperature coefficient, is the pose coupling weight matrix.

[0044] S6), Multimodal data spatio-temporal registration: Spatially and temporally register the coordinate mapping result output by S3 with the calibrated force sense signal of S5 to generate a synchronized data set , and fuse the temporal noise through a Kalman filter; Spatio-temporal synchronization: Spatially and temporally align the visual coordinate mapping result output by S3 with the calibrated force sense signal of S5 to generate a synchronized data set .

[0045] Noise reduction processing: Fuse the temporal noise through an Extended Kalman Filter (EKF) to improve the data quality.

[0046] Data fusion preparation: Provide clean and consistent input for the attention mechanism-based feature fusion (S7).

[0047] S7): Attention mechanism-based feature fusion: Dynamically allocate visual and force sense weights through a hybrid deep learning model: ; where , , , are the visual gradient and force gradient feature vectors.

[0048] S8), Fusion pose generation and error correction: Output the fusion pose : ; ; where , is the force gradient compensation term; In the formula, is the target pose coordinate output by the visual system, from the image capture and edge detection results of step S1, belonging to the visual pixel coordinate system (two-dimensional plane coordinate, unit is pixel or calibrated physical unit); The role is to provide the position information of the target in the microscopic image for initial pose estimation.

[0049] is the target coordinate in the MEMS force sense sensor coordinate system, obtained through the non-linear mapping of the spatial coordinate system in step S3 (from the visual pixel coordinate system to the force sense Cartesian coordinate system), unit is physical length; is the weight parameter, the weight coefficient of the visual data, dynamically calculated through the attention mechanism of step S7, and the value range is [0,1]; is the weight coefficient of the force perception data, which is also generated by the attention mechanism and is related to .

[0050] S9), Real-time optimization of reinforcement learning parameters: Define the state space (pose error, force error, disturbance level), and output the control parameter increment through the Actor-Critic network: ; where is the learning rate, and the network weights are updated by backpropagation of the TD error; S10), Generation of fuzzy sliding mode disturbance rejection control command: Design the sliding mode surface ; where is fuzzily adjusted according to the disturbance level ; In the formula: s is the sliding mode surface variable, is the pose error (the difference between the current pose and the target pose); the sliding mode surface gain coefficient; the non-linear exponential parameter (μ ∈ (0, 1); the sign function ( outputs +1 when > 0, otherwise -1); Generate the disturbance rejection control command: ; where , ; S11), Dynamic impedance protection and real-time retraction: Dynamically adjust the impedance parameter according to the environmental stiffness : ; When the contact force exceeds the limit , trigger nanoscale retraction: .

[0051] Preferably, step S2 further includes: In dynamic timestamp alignment, calibrate the compensation term by the sliding window variance minimization algorithm: ; where is the number of data frames in the sliding window; In the non-linear mapping of the spatial coordinate system, the Z-axis prediction function is implemented by a pre-trained deep neural network (DNN), and the input is the local feature block of the image , the output is Predicted value of: ; The input and output parameters include: : The input image local feature block is a local image block containing the target area extracted from the microscopic image.

[0052] As the input of the deep neural network (DNN), it is used to extract visual features related to the Z-axis coordinate.

[0053] : The output Z-axis coordinate value is part of the force Cartesian coordinate system and represents the position of the target in the depth direction (Z-axis).

[0054] The function is to obtain the Z-axis coordinate predicted by the DNN model, which is used to supplement the nonlinear mapping from the visual pixel coordinate system to the force coordinate system (to solve the problem of insufficient accuracy of the traditional linear mapping in the Z-axis direction).

[0055] The neural network parameters (hidden layers) include: W h : Weight matrix of the hidden layer.

[0056] Type: Matrix, the dimension is determined by the input feature block size and the number of hidden layer neurons.

[0057] Function: Input feature block Perform a linear transformation to map pixel-level features to the abstract feature space of the hidden layer.

[0058] : Bias vector of the hidden layer.

[0059] Type: vector, with the same dimension as the number of neurons in the hidden layer.

[0060] Function: Add bias terms to the linear transformation of the hidden layer to improve the model fitting ability.

[0061] : RectifiedLinearUnit activation function.

[0062] Function: Linear transformation result of hidden layer Introducing nonlinear activation prevents the network from becoming a simple linear model and enhances the ability to model complex nonlinear relationships.

[0063] The neural network parameters (output layer) are as follows: : The weight matrix of the output layer of a deep neural network (DNN).

[0064] Type: Matrix, the dimension is determined by the number of neurons in the hidden layer and the output dimension (in the formula, the output is a single Z-axis coordinate, so the dimension is 1 X the number of neurons in the hidden layer).

[0065] Function: Map the abstract features of the hidden layer to the final predicted value of the Z-axis coordinate, and complete the non-linear mapping from visual features to the Z-axis of the force perception coordinate system.

[0066] : The bias scalar of the output layer (Bias scalar of the output layer).

[0067] Type: Scalar.

[0068] Function: Add a bias term to the linear transformation of the output layer and adjust the offset of the final predicted value.

[0069] This formula realizes the non-linear mapping from the visual pixel coordinate system to the Z-axis of the force perception Cartesian coordinate system through a deep neural network, solves the mapping error problem of traditional linear methods in the depth direction, and improves the three-dimensional positioning accuracy (especially the sub-micron-level Z-axis positioning).

[0070] Preferably, the S4 further includes: In visual data augmentation, the size of the Gaussian kernel G(σ) is 5×5, and the sharpening coefficient λ is dynamically adjusted according to the signal-to-noise ratio SNR of the image: ; In the dynamic calibration of the force perception signal, the pose coupling weight matrix W is determined through an offline calibration experiment and satisfies: .

[0071] Preferably, the S6 further includes: In spatio-temporal registration, an extended Kalman filter (EKF) is used to fuse the temporal noise, and the state equation and the observation equation are respectively: ; where A is the state transition matrix and H is the observation matrix, , is Gaussian noise; The fused synchronous data set contains the visual-force covariance matrix Cvf with timestamp alignment.

[0072] Preferably, the S7 further includes: In the attention mechanism, the visual feature confidence is calculated based on the normalized variance of the image gradient magnitude : ; Force gradient compensation term , Through real-time differential calculation: .

[0073] Preferably, the S9 further includes: In the ActorCritic network, the Critic network outputs the state value function , and updates the weights through the temporal difference (TD) error: ; The action exploration strategy adds Gaussian noise , and the noise variance exponentially decays with the number of training rounds: .

[0074] Preferably, the S10 further includes: The perturbation level is divided into three levels: low, medium, and high, corresponding to the value of , the value of ; Superimpose a high-frequency perturbation suppression term , and its frequency is estimated in real time by FFT, and the suppression gain is adaptively adjusted: ; Preferably, the S11 further includes: The environmental stiffness is estimated online through the slope of the force-displacement curve: ; During nanoscale retraction, the pose gradient is calculated by the visual feature optical flow method:

[0075] Preferably, it further includes S12 fault recovery: When a sudden change in the sensor signal ΔF > 5 mN or a communication delay Td > 2 ms is detected, a three-level emergency response is triggered: a. Pause the movement of the robotic arm and switch to the safe impedance mode; b. Reconstruct the target pose based on the weighted average of historical data, and the weight w i satisfies ∑w i = 1; c. If the continuous fault timeout Tfault > 5 s, start the self-check protocol and report the error code.

[0076] Preferably, it further includes S13, performance calibration: verifying that the pose tracking error ep ≤ 0.1 μm through a standard micro-nano reference device; When switching between multiple platforms, compensating for coordinate system offsets through an adaptive calibration algorithm: ; where N ≥ 10 is the number of calibration points.

[0077] The 1 - 13 steps proposed by the present invention are not executed in isolation, but are closely associated through data flow and control flow to form a closed-loop feedback system: S1 (data acquisition): providing the original visual image and force sense signal, and providing input for subsequent steps.

[0078] S2 (timestamp alignment) and S3 (coordinate system mapping): solving the spatio-temporal asynchrony and dimensional heterogeneity of visual and force sense data, unifying the heterogeneous data into the same spatio-temporal framework (S2 aligns timestamps, S3 aligns spatial coordinates), and laying a foundation for multi-modal fusion.

[0079] S4 (visual enhancement) and S5 (force sense calibration): respectively enhancing the feature robustness of visual data (such as edge sharpening) and the accuracy of force sense data (such as temperature drift compensation), and ensuring the reliability of subsequent fused data.

[0080] S6 (spatio-temporal registration): synchronizing the mapping result of S3 and the calibration signal of S5 in space and time, generating a synchronous data set for fusion, and denoising through Kalman filtering.

[0081] S7 (feature fusion) and S8 (pose generation): dynamically allocating visual and force sense weights based on the attention mechanism (S7), combining force gradient compensation to correct errors, and generating a fused pose (S8).

[0082] S9 (parameter optimization) and S10 (disturbance rejection control): real-time optimizing control parameters through reinforcement learning (S9), and generating disturbance rejection instructions in combination with fuzzy sliding mode control (S10) to respond to dynamic changes in the environment.

[0083] S11 (impedance protection): adjusting impedance parameters according to the environmental stiffness, and triggering a nano-level retraction when the contact force exceeds the limit to protect the safety of the sample.

[0084] S12 (fault recovery) and S13 (performance calibration): serving as the safety guarantee (S12) and accuracy verification (S13) of the system, and supplementing the core processes of the previous 11 steps.

[0085] Solution to the lack of dynamic spatio-temporal registration mechanism: Bilinear interpolation timestamp alignment (S2): compensating for the sampling time difference between vision and force sense through an interpolation function to eliminate asynchrony.

[0086] Deep learning non-linear coordinate mapping (S3): Establish a non-linear mapping from pixel coordinates to Cartesian coordinates using a pre-trained DNN model to address dimensional heterogeneity.

[0087] Extended Kalman filter (S6): Incorporate temporal noise to further improve spatio-temporal registration accuracy.

[0088] Solution to the problem of insufficient static optimization of control parameters: Reinforcement learning parameter real-time optimization (S9): Online adjust visual servo gain and force control stiffness through an Actor-Critic network to respond to sudden changes in mechanical characteristics.

[0089] Dynamic weight allocation (S7): Based on the attention mechanism and force gradient compensation, adaptively allocate visual and force perception weights to suppress noise interference.

[0090] Solution to the problem of high computational latency: Hardware acceleration (S6, S13): Utilize FPGA processing units to accelerate spatio-temporal registration and feature extraction, compressing the latency to the millisecond level.

[0091] High-frequency perturbation suppression (S10): Real-time estimate the perturbation frequency through FFT and superimpose suppression terms to reduce the computational load.

[0092] Realization of sub-micron-level fusion positioning accuracy technology: Bilinear interpolation (S2) to align timestamps and eliminate sampling delays between vision and force perception.

[0093] DNN non-linear mapping (S3) compensates for the error from pixel coordinates to Cartesian coordinates (such as the Z-axis prediction function) to reduce spatial mapping deviation.

[0094] Kalman filter (S6) further incorporates noise to ensure spatio-temporal consistency of data.

[0095] Realization of dynamic weight allocation and robustness improvement technology: Attention mechanism (S7): Calculate visual confidence based on the image gradient magnitude (such as normalized variance) and dynamically adjust the weights.

[0096] Force gradient compensation (S8): Calculate the force gradient term through real-time differentiation to correct the pose error.

[0097] Actor-Critic network (S9): Online optimize control parameters to adapt to mechanical mutations (such as stiffness changes during cell puncture).

[0098] Realization of anti-interference ability and real-time performance technology: Fuzzy sliding mode control (S10): Fuzzily adjust the sliding mode surface parameters according to the perturbation level to resist interference such as liquid fluctuations.

[0099] Hardware Acceleration (FPGA): Enhances data processing speed to meet millisecond-level real-time requirements.

[0100] Dynamic Impedance Protection (S11): Adjusts impedance parameters in real time according to environmental stiffness, and combines with a nanoscale retraction mechanism to avoid overload damage.

[0101] Steps S1 to S11 constitute the core technical solution, covering the entire process from data acquisition, fusion to control instruction generation.

[0102] Step S12 (Fault Recovery): As a safety guarantee extension, triggers an emergency response when the system is abnormal (such as signal mutation or communication delay) to ensure operational safety.

[0103] Step S13 (Performance Calibration): As a verification extension, tests the pose error through standard devices and compensates for the coordinate system offset during multi-platform switching to ensure system accuracy and compatibility.

[0104] Steps S12 and S13 are supplements to the core process (S1 - S11), respectively targeting system reliability and accuracy verification, forming a complete "core function + safety redundancy + performance guarantee" system.

[0105] The present invention provides the following technical solutions: Biological Cell Nanoscale Manipulation Experiment: 1. Implementation Environment and Equipment 1.1 Hardware Platform: High-frame-rate Microscopic Camera (Basler acA2000 - 340km, frame rate ); MEMS Sensor Array (Futek LSB200, range 0.1 μN ~ 10 mN, sampling rate ); Piezoelectric Ceramic-driven Nanorobot Arm (Physik Instrumente P-611.3S, displacement resolution 1 nm); FPGA Processing Unit (Xilinx Zynq UltraScale +, main frequency 500 MHz).

[0106] 1.2 Software Environment: Image Processing: OpenCV4.5 (edge detection, Kalman filtering); Deep Learning Framework: PyTorch1.9 (CNN, DNN model training); Control Algorithm: MATLAB / Simulink Real-time Control Module.

[0107] 1.3 Experimental Object: Human Umbilical Vein Endothelial Cells (HUVEC), cultured in a PDMS microfluidic chip; Standard silicon microcantilever (size 10μm × 100μm × 1μm).

[0108] The specific implementation steps are as follows: Step S1: Multimodal data synchronous acquisition - Visual data: The microscopic camera captures cell images at 1200Hz, with an adaptive exposure time of 0.1ms. Canny edge detection is used to extract the cell contour, and the pose (x v , y v , θv) is output.

[0109] Force sense data: The MEMS force sense sensor collects cell contact force signals (Fx, Fy), and a band-pass filter (10Hz ~ 1kHz) suppresses the culture medium fluctuation noise.

[0110] Step S2: Dynamic timestamp alignment - Align the visual (t v ) and force sense (t f ) timestamps through bilinear interpolation. The sliding window N = 50, = 0.05ms.

[0111] Step S3: Nonlinear mapping of the spatial coordinate system - Pre-trained DNN model (input: 32×32 image block; output: ), error <0.05μm.

[0112] Step S4: Visual data enhancement Gaussian filtering (σ = 1.5pixel) and Laplacian sharpening (λ = 0.8), with the signal-to-noise ratio increased by 30%.

[0113] Step S5: Calibration of the force sense signal Temperature compensation coefficient , pose coupling matrix Determined by offline calibration.

[0114] Step S6: Spatiotemporal registration and Kalman filtering EKF parameters: State transition matrix , observation matrix , noise covariance , .

[0115] Step S7 - S8: Attention mechanism fusion and pose generation Visual weight , force sense weight , force gradient compensation .

[0116] Step S9: Reinforcement learning parameter optimization Actor-Critic network structure: Input layer with 3 nodes, hidden layer with 64 nodes, output layer with 2 nodes; Learning rate , discount factor , exploration noise 。

[0117] Step S10: Fuzzy sliding mode disturbance rejection control - disturbance level In , ,high - frequency suppression gain 。

[0118] Step S11: Dynamic impedance to protect environmental stiffness ,impedance parameter , 。

[0119] Step S12: Fault recovery: Simulate communication delay , trigger a three - level response, and reconstruct the pose error m.

[0120] Step S13: Performance calibration: Calibrate the pose error of the silicon cantilever beam m, after multi - platform switching offset compensation 。 Verification of implementation effect

[0121] 。

[0122] Comparative experiment Experimental group: Use the method of the present invention to perform puncture operations on HUVEC (target depth 5μm); Control group: Traditional visual servo + fixed impedance control.

[0123] 。

[0124] Thus, this embodiment verifies the significant advantages of the method in biological cell nano - level operations: High precision: The fused pose error is low, meeting the requirements of cell non - destructive operation; Strong robustness: The liquid disturbance suppression rate is relatively high, adapting to the dynamic micro - environment; High safety: The response time for contact force exceeding the limit is short, and the cell survival rate is increased significantly; High - efficiency compatibility: The calibration time for multi - platform switching is greatly shortened, supporting cross - scenario applications.

[0125] As Figure 2 shown, another embodiment of the present invention provides a system for implementing the above - mentioned microscopic positioning control method for adaptive visual - force fusion, including: (1) Multi - modal data acquisition module: A visual acquisition unit for capturing images of the operation area in real time, extracting the target contour through an edge detection algorithm, and outputting target pose data; A force perception acquisition unit for real-time acquisition of six-dimensional force / torque signals at the end of an operation; A temperature monitoring unit that uses a thermal sensor to collect the ambient temperature in real time; 2) Asynchronous data synchronization module: A data buffer for caching visual frame data and force perception sampling data respectively; An interpolation alignment unit that realizes data synchronization based on the visual frame timestamp and the force perception sampling timestamp through a bilinear interpolation function to generate a unified timestamp; A calibration compensation unit for storing preset compensation parameters; 3) Spatial coordinate mapping module: A calibration parameter database for storing calibration parameters; A non-linear error compensation unit that runs a convolutional neural network model for compensation; A prediction unit that realizes deep learning-based prediction; (4) Data processing and enhancement module: A visual enhancement unit for performing Gaussian filtering and Laplacian sharpening on the original image to generate an enhanced image; A force perception calibration unit that compensates for the zero-point drift of the sensor in real time according to the ambient temperature and the pose of the robotic arm and outputs a calibrated force signal; (5) Multimodal data fusion module: A spatio-temporal registration unit for spatio-temporally aligning the mapped coordinates and the calibrated force signal to generate a synchronized data set and reducing noise through a Kalman filter; An attention weighting unit that uses a hybrid deep learning model to dynamically allocate visual and force perception weights; A pose synthesis unit for generating a fused pose; (6) Adaptive control module: A reinforcement learning optimization unit that adjusts control parameters in real time through an Actor-Critic network; A sliding mode control unit for generating control instructions; An impedance adjustment unit that dynamically calculates the equivalent mass and damping; 7) Safety protection module: A safety retraction unit for triggering nanoscale displacement compensation; A fault emergency unit for performing exception handling; A self-check recovery unit for running a historical data reconstruction algorithm; (8) Calibration verification single module: An error analysis unit for verifying the pose tracking error; An adaptive compensation unit for calculating the offset correction amount.

[0126] Preferably, it further includes: a time alignment module for optimizing the calibration compensation term through a sliding window variance minimization algorithm and a spatial mapping module of a pre-trained deep neural network; A force perception coupling calibration module for online compensating the systematic offset in the contact force signal.

[0127] Preferably, it further includes: A multi-source fusion processing module that uses an extended Kalman filter to achieve spatio-temporal registration, and the multi-source fusion processing module includes: a state prediction module burned in the kinematic model ROM of the DSP and an observation update module that executes the matrix mapping visual pose extraction module output to the state space; a covariance calculation module that generates a synchronous data set matrix and stores it in the dual-port RAM for the decision-making module to call; A confidence evaluation module that calculates the gradient magnitude by the image processor and outputs it through a variance normalization unit; A force gradient compensation module that includes a high-speed ADC and a differential calculation circuit.

[0128] Preferably, it further includes: A high-frequency disturbance suppressor composed of a coprocessor and an adaptive gain controller and an environmental stiffness estimator integrated in the force control feedback loop, a fault recovery module composed of a three-stage emergency trigger circuit, a historical data reconstruction unit, and a self-check protocol executor; a micro-nano reference component positioning platform for verifying the tracking error and a calibration module for coordinate system offset compensation.

[0129] The data flow of the micro-positioning control system for adaptive visual force perception fusion of the present invention is as follows: 1. After the sensing module obtains the original data, it aligns the timestamps through the spatio-temporal synchronization module; 2. The preprocessed data enters the coordinate mapping module to complete the spatial transformation; 3. The fusion module integrates multi-modal information to generate control parameters; 4. The intelligent control module receives the feedback from the safety module while outputting control instructions; 5. The calibration verification module provides closed-loop calibration support.

[0130] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An adaptive microscopic positioning control method for visual and force sense fusion, characterized in that: It includes the following steps: S1), Image capture: The image of the operation area is captured in real time by a high-frame-rate microscopic camera, and the target contour is extracted by an edge detection operator, and the target pose is output The six-dimensional force and torque signals at the operation end are obtained in real time through the MEMS force sensor array ; S2), Dynamic timestamp alignment: Based on the visual frame timestamp and the haptic sampling timestamp , Asynchronous data alignment is achieved through a bilinear interpolation function; S3), Nonlinear mapping of the spatial coordinate system: Establish a mapping model from the visual pixel coordinate system to the haptic Cartesian coordinate system ; S4), Visual data enhancement and feature extraction: For the original image Perform Gaussian filtering and Laplacian sharpening to enhance the target edge features; S5), Dynamic calibration of force sense signal: According to the ambient temperature and the pose of the robotic arm , compensate the zero drift of the MEMS force sense sensor in real time; S6), spatio-temporal registration of multi-modal data: Spatially and temporally register the coordinate mapping result output in step S3 and the calibrated force sense signal in step S5 to generate a synchronized data set; S7): Feature fusion based on the attention mechanism: Dynamically allocate visual and force sense weights through a hybrid deep learning model; S8), integrating pose generation and error correction: output the integrated pose ; S9), Real-time optimization of reinforcement learning parameters: Define the state space , and output the control parameter increment through the Actor-Critic network; S10), generating a sliding mode surface for fuzzy sliding mode disturbance rejection control instruction design; S11), Dynamic impedance protection and real-time withdrawal: Dynamically adjust the impedance parameters according to the environmental stiffness ( ).

2. The microscopic positioning control method for adaptive visual and force sense fusion according to claim 1, characterized in that: S2) Dynamic Timestamp Alignment: Based on the visual frame timestamp and the haptic sampling timestamp , asynchronous data alignment is achieved through a bilinear interpolation function: ; wherein is a calibration compensation term; is the sampling frequency of the vision system; is the sampling frequency of the force sensing system; Represents the unified timestamp after dynamic timestamp alignment; Calibration compensation item Optimized by the sliding window variance minimization algorithm: ; : The visual end - pose vector of the k - th sliding window sample point; : The k-th force sensing end sampling point, the pose vector corresponding to the time lag after that; : The Euclidean norm squared, which is used to minimize the temporal alignment error between vision and force sensing; : represents finding the time compensation amount that minimizes the objective function; Among them is the number of data frames within the sliding window; S3) Nonlinear mapping of the spatial coordinate system: Establish a visual pixel coordinate system to the Cartesian coordinate system of the sense of force mapping model: ; wherein is a calibration parameter, is the non-linear error compensated by the convolutional neural network, is the Z-axis prediction function based on deep learning; Z-axis prediction function Implemented by a pre-trained deep neural network (DNN), with the input being the local feature block of the image ( ), and the output being the predicted value of the Z-axis coordinate value ( ): ; In the formula, the input and output parameters include: : The input local image feature block is a local image block containing the target area extracted from the microscopic image; : The output Z-axis coordinate value, which is part of the force Cartesian coordinate system, represents the position of the target in the depth direction; The neural network parameters include: W h : The weight matrix of the hidden layer; : Bias vector of the hidden layer; : Rectified Linear Unit Activation Function; The output layer of the neural network parameters is as follows: : The weight matrix of the output layer of the deep neural network; : The bias scalar of the output layer; S4) Visual data enhancement and feature extraction: For the original image perform Gaussian filtering and Laplacian sharpening to enhance the target edge features: ; Wherein: : The original microscopic image data; is a Gaussian kernel; where σ = 1.5 pixels, representing the standard deviation of the Gaussian kernel; The convolution operation symbol indicates Gaussian filtering of the original image; is the Laplacian operator of the original image, which is used to extract the high-frequency edge features of the image; is the sharpening coefficient, which is used to control the intensity of Laplacian sharpening; is the enhanced image; S5) Dynamic calibration of force sense signal: According to the ambient temperature and the pose of the robotic arm , compensate the zero drift of the MEMS force sense sensor in real time: ; Among them : The original uncompensated contact force data is directly output by the MEMS force sensor; : Temperature drift coefficient, which is used to quantify the influence of ambient temperature change on the zero-point offset of the MEMS force sensor; : The current ambient temperature, collected by a thermosensitive device or a thermocouple; : Ambient temperature during calibration, serving as the benchmark for temperature compensation; : Pose coupling weight matrix, which is a 2×1 or 2×2 vector or matrix; : The two-dimensional displacement vector of the current operation end in the visual coordinate system, with the unit of pixel or mm; : The finally calibrated force signal, after removing the pose interference, is used as the true feedback value for subsequent control calculations; S6) Spatiotemporal registration of multimodal data: Perform spatiotemporal registration on the coordinate mapping result output in step S3 and the calibrated force sense signal in step S5 to generate a synchronized dataset: , and fuse the temporal noise through a Kalman filter; : Represents the structured data set after time alignment, coordinate calibration, and filtering fusion are completed; S7) Feature fusion based on the attention mechanism: Dynamically allocate visual and force sense weights through a hybrid deep learning model: ; =1- ; The above structure realizes dynamic weighting of the visual and force sense channels during the fusion process through the softmax normalization mechanism; Among them, the channel scoring factor , are respectively defined as: , ; The variable definitions are as follows: ∇I, ∇F: Feature gradient vectors of the image and the force sense signal respectively; , : The channel weight matrix for feature mapping, which is the parameter of a 1×N fully connected layer and is obtained through end-to-end training; ReLU(x)=max(0,x): Is the rectified linear unit activation function, introducing non-linearity to enhance the sparsity of feature expression; : is the visual modality attention weight, reflecting the importance of visual features in the current fusion process; , are the visual gradient and force gradient feature vectors; S8) Fusion of pose generation and error correction: Output the fused pose : ; ; Among them , is the force gradient compensation term; Wherein, is the target pose coordinate output by the vision system, from the image capture and edge detection results of step S1, belonging to the vision pixel coordinate system; is the target coordinate in the coordinate system of the MEMS force sensor, obtained through the non-linear mapping of the spatial coordinate system in step S3, with the unit of physical length; is a weight parameter, the weight coefficient of visual data, which is dynamically calculated by the attention mechanism in step S7 and ranges from [0, 1]; is the weight coefficient of the force sense data, which is also generated by the attention mechanism and is related to ; S9) Real-time optimization of reinforcement learning parameters: Define the state space , and output the control parameter increment through the Actor-Critic network: ; : The control gain adjustment amount of the visual channel, which is used to dynamically adjust the visual servo response intensity; : The control gain adjustment amount of the force sensing channel, which is used to dynamically adjust the force control stiffness or impedance; : Learning rate, which controls the speed of network weight update; : The sliding mode surface variable, representing the error metric between the current state and the target state; , : The network parameter matrix, corresponding to the weight matrices of the visual gain regulator and the haptic gain regulator respectively; , : Bias term, adjusting the intermediate offset of the non-linear output of the network; The above structure constitutes a single-layer or double-layer neural network mapping for mapping the sliding mode error ( ) to the control gain adjustment amount to achieve adaptive control adjustment under different disturbance levels; S10) Generating a sliding mode surface for fuzzy sliding mode disturbance rejection control instruction design: ; Among them According to the disturbance level Fuzzy adjustment; In the formula: is pose error; sliding mode surface gain coefficient; nonlinear exponential parameter, μ ∈ (0, 1); sign function; Generating a disturbance rejection control instruction: ; : is the saturation function limiting parameter; : Visual channel sliding mode control gain coefficient, which controls the sensitivity of the system to the visual error sliding surface. The larger the value, the faster the system responds to the change of visual error; : Force perception channel anti-disturbance gain coefficient, which controls the action intensity of the sat function in the force channel and is commonly used to suppress force perception disturbances or external micro-vibrations; : Disturbance compensation gain, which adjusts the influence of the force sense channel compensation on the controller; : is the change rate of the contact force, representing the trend of external disturbance; Among them , ; The sat(x) function is used to suppress the chattering of the sliding mode controller, and the output value is limited to the interval [-1,1]; S11) Dynamic impedance protection and real-time withdrawal: Dynamically adjust the impedance parameters according to the environmental stiffness ( ) ; : Base quality parameter, representing the reference quality of the system in an ideal rigid environment; : Environmental stiffness, a dynamically estimated value that reflects the rigidity of the current contact object; : Equivalent mass, which is used to convert the foundation mass in combination with the environmental flexibility; : Equivalent damping, used for speed control / energy dissipation regulation in the retraction process; : Damping adjustment proportional factor; : Controller equivalent stiffness coefficient, which models the compliant response ability of the controller to end disturbances; ; : Represents the fine-tuning vector of the end position when the protection is triggered, with the unit of nanometer (nm); : Yield ratio factor; : The current contact force collected in real time by the MEMS force sensor; : A preset contact safety threshold that triggers a retraction when occurs; ∇x: Is the unit velocity direction vector of the current position.

3. The microscopic positioning control method for adaptive visual and force sense fusion according to claim 2, wherein: The size of the Gaussian kernel G(σ) is 5×5, and the sharpening coefficient λ is dynamically adjusted according to the signal-to-noise ratio SNR of the image: ; In the dynamic calibration of the force sense signal, the pose coupling weight matrix (W) is determined through an offline calibration experiment, satisfying: ; : is a two-dimensional coupling weight matrix used to model the systematic perturbation of the contact force signal caused by the displacement of the image coordinates; : represents the position variable The partial derivative of the change with respect to the force signal reflects the sensitivity of the force perception channel to image displacement; All parameters are obtained through linear fitting during the calibration stage and are used for position compensation during actual operation.

4. The microscopic positioning control method for adaptive visual and force sense fusion according to claim 2, wherein: Fusing the temporal noise through a Kalman filter, the state equation and the observation equation are respectively: ; : The system state variable at the current moment, including the pose information estimated by visual-force fusion; : Observation quantities, such as the center coordinates and direction of the end image obtained by microscopic image processing; where A is the state transition matrix and H is the observation matrix, , is Gaussian noise; Fused synchronized dataset Contains the visual-force covariance matrix Cvf with timestamp alignment; In the attention mechanism, the visual feature confidence is calculated based on the normalized variance of the image gradient magnitude as follows: ; : The variance of the image gradient distribution, which reflects the overall texture or edge complexity; : The maximum gradient value in the image, corresponding to the most prominent edge, is used for normalization; : A minimum value constant used to avoid division-by-zero errors during the image contrast normalization process; ∇I: Image gray gradient, calculated by the image Sobel operator, used to represent the edge intensity of the image; Force gradient compensation term , by real-time differential calculation: ; : Represents the time difference gradient of the contact force, used to detect the rapid change trend of the force; : are the contact force values collected at the current moment and the previous moment respectively, and are collected by the MEMS force sensor at a high frequency; : The time interval between two consecutive sampling instants, which is set by the system sampling frequency; : It is the time calibration coefficient, which is used to scale the time reference of the differential result so that it uniformly corresponds to the drawdown response rate; In the Actor-Critic network, the Critic network outputs the state value function , and updates the weights through the temporal difference (TD) error: ; : Temporal Difference error (TDerror), which represents the gap between the current predicted value and the target return, is the core basis for updating the policy and the Critic network; : The immediate reward value, generated by the environmental feedback, is used to measure the effect of the current action; : Discount factor, which controls the influence weight of future rewards; : Learning rate, which determines the step size for each update of the network parameters; : The gradient of the Critic network with respect to its parameters ; The action exploration strategy adds Gaussian noise , and the noise variance exponentially decays with the number of training rounds: ; : It is the standard deviation of the Gaussian policy perturbation at the current moment, which controls the magnitude of randomness in the action output; : is the initial perturbation intensity, adjusted according to the environmental complexity; : is the annealing time constant, which affects the noise decay rate.

5. The microscopic positioning control method for adaptive visual and force sense fusion according to claim 2, characterized in that: Disturbance level It is divided into three levels: low, medium, and high. The value of is The value of ; Superimposed high-frequency disturbance suppression term ( ), whose frequency ( ) is estimated in real time by FFT, and the suppression gain is adaptively adjusted: ; : The high-frequency disturbance force vector is obtained by band-pass filtering or fast Fourier transform of the original force sense signal, reflecting the disturbance intensity; : The original sequence of contact forces, continuously collected by the MEMS force sensor; : Represents the two-norm of a vector and is used to calculate the perturbation magnitude; : The maximum value of the force sense signal within the current sliding window, used for normalization; Environmental stiffness Online estimation through the slope of the force-displacement curve: ; : The change in contact force within the time window, measured by the MEMS force sensor; : The displacement change amount at the end during this time period, which is calculated by the vision system; →0: Represents approximate differentiation in an extremely short time, used for calculating the response slope of stiffness in real time; Pose gradient Calculated by the visual feature optical flow method: ; : The two-dimensional pixel coordinates of the target in the current frame image; : Represents the velocity vector component of the image target, calculated by the optical flow method; : is the estimated result of the final pose change rate; : The time interval between two adjacent frames, with the unit of millisecond (ms), is determined by the visual frame rate.

6. The microscopic positioning control method for adaptive visual and force sense fusion according to claim 1, characterized in that: It further includes step S12, fault recovery: When it is detected that the signal mutation of the MEMS force sense sensor ΔF>5mN or the communication delay Td>2ms, trigger a three-level emergency response: a. Pause the movement of the robotic arm and switch to the safe impedance mode; b. Reconstruct the target pose based on the weighted average of historical data, with weight w i satisfying ∑w i = 1; c. If the continuous fault timeout Tfault>5s, start the self-check protocol and report the error code; It further includes step S13, performance calibration: Verify that the pose tracking error ep≤0.1μm through a standard micro-nano reference device; When switching between multiple platforms, compensate for the coordinate system offset through an adaptive calibration algorithm: ; : Represents the position offset error of a single calibration point in visual measurement, which is obtained by manual or algorithm extraction; : Represents the number of target points participating in the calibration average, where ≥10 is the number of calibration points; : The translation offset correction amount during system initialization, which is used to eliminate systematic deviations in overall image recognition and is uniformly compensated in all subsequent coordinate conversions.

7. A system for implementing the microscopic positioning control method of adaptive visual and force sense fusion according to any one of claims 1-6, characterized in that, It includes (1) Multi-modal data acquisition module: A visual acquisition unit for capturing images of the operation area in real time and extracting the target contour through an edge detection algorithm; outputting the target pose data; A force sense acquisition unit for obtaining the six-dimensional force / torque signal at the operation end in real time; A temperature monitoring unit for collecting the ambient temperature in real time through a thermal sensor; (2) Asynchronous data synchronization module: A data buffer for caching visual frame data and force sense sampling data respectively; Based on the visual frame timestamp and the force sensing sampling timestamp, a data synchronization is achieved through a bilinear interpolation function to generate an interpolation alignment unit with a unified timestamp; A calibration compensation unit for storing preset compensation parameters; (3) Spatial coordinate mapping module: A calibration parameter database for storing calibration parameters; A non-linear error compensation unit that runs a convolutional neural network model for compensation; A prediction unit that implements deep learning-based prediction; (4) Data processing and enhancement module: A visual enhancement unit for performing Gaussian filtering and Laplacian sharpening on the original image to generate an enhanced image; A force sensing calibration unit that compensates for the zero drift of the sensor in real time according to the environmental temperature and the pose of the robotic arm and outputs a calibrated force signal; (5) Multimodal data fusion module: A spatio-temporal registration unit for spatio-temporally aligning the mapped coordinates and the calibrated force signal to generate a synchronized data set and denoising it through a Kalman filter; An attention weighting unit that uses a hybrid deep learning model to dynamically allocate visual and force sensing weights; A pose synthesis unit for generating a fused pose; (6) Adaptive control module: A reinforcement learning optimization unit that adjusts control parameters in real time through an Actor-Critic network; A sliding mode control unit for generating control commands; An impedance adjustment unit that dynamically calculates the equivalent mass and damping; (7) Safety protection module: A safety retraction unit for triggering nanoscale displacement compensation; A fault emergency unit for performing exception handling; A self-check recovery unit for running a historical data reconstruction algorithm; (8) Calibration verification single module: An error analysis unit for verifying the pose tracking error; An adaptive compensation unit for calculating the offset correction amount.

8. The system of the microscopic positioning control method for adaptive visual and force sense fusion according to claim 7, wherein It also includes: A time alignment module that optimizes the calibration compensation term through a sliding window variance minimization algorithm and a spatial mapping module of a pre-trained deep neural network; A force sensing coupling calibration module for online compensating the systematic offset in the contact force signal.

9. The system of the microscopic positioning control method for adaptive visual and force sense fusion according to claim 7, wherein It further includes: A multi-source fusion processing module that uses an extended Kalman filter to achieve spatio-temporal registration. The multi-source fusion processing module includes: a state prediction module burned in the kinematic model ROM of the DSP and an observation update module that executes the matrix mapping visual pose extraction module output to the state space; a covariance calculation module that generates a synchronized data set matrix and stores it in a dual-port RAM for the decision-making module to call; a confidence evaluation module that calculates the gradient magnitude by the image processor and outputs it through a variance normalization unit; a force gradient compensation module that includes a high-speed ADC and a differential calculation circuit.

10. The system of the microscopic positioning control method for adaptive visual and force sense fusion according to claim 7, wherein It further includes: A high-frequency disturbance suppressor composed of a coprocessor and an adaptive gain controller and an environmental stiffness estimator integrated in the force control feedback loop, a fault recovery module composed of a three-stage emergency trigger circuit, a historical data reconstruction unit and a self-check protocol executor, and a calibration module for verifying the tracking error and the micro-nano reference component positioning platform and coordinate system offset compensation.

Citation Information

Patent Citations

  • Virtual force feedback remote nano operation platform based on scanning electron microscope and method for realizing virtual force sensing interacting

    CN103092346A

  • Virtual interaction system and method of scanning electron microscope

    CN118732857A

  • Robot pose estimation method based on laser point cloud and visual SLAM

    CN119596328A

  • Bridge pier measuring method and system with laser gyroscope measuring rod and unmanned aerial vehicle vision fused

    CN119935095A

  • Self-adaptive gravity center adjusting method and system for transfer robot

    CN119974028A

Cited By

  • Lightweight container correction method and system based on machine vision

    CN120472324A

  • Visual touch double-channel control hardware method and system for robot impedance self-adaption

    CN120663335A

  • Air cooler defrosting management control method and system based on machine vision and AI algorithm

    CN120760395A

  • Visual guidance method and system for intelligent auxiliary assembly

    CN120779889A

  • Robot precision compensation method based on space-time Kriging model

    CN120791778A