Multi-modal fusion puncture tissue characteristic identification system

The multimodal fusion puncture tissue characteristic recognition system utilizes cross-modal contrastive learning and causal discovery algorithms to generate a high-precision tissue recognition model, solving the problems of insufficient tissue recognition accuracy and data scarcity in existing technologies, and realizing real-time and safe ophthalmic puncture surgery recognition.

CN122005085AInactive Publication Date: 2026-05-12SHANGHAI EAST HOSPITAL EAST HOSPITAL TONGJI UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI EAST HOSPITAL EAST HOSPITAL TONGJI UNIV SCHOOL OF MEDICINE
Filing Date
2026-04-13
Publication Date
2026-05-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In ophthalmic puncture surgery, existing technologies cannot reflect the dynamic deformation of tissues during surgery due to the inability of preoperative images. Reliance on one-dimensional force signals results in insufficient accuracy, and the lack of high-quality training data prevents the effective application of artificial intelligence models.

Method used

A multimodal fusion puncture tissue feature recognition system is adopted. The encoder network is trained through cross-modal contrastive learning to generate modality-independent robust fusion feature representations. Combined with causal discovery algorithms and regularized loss functions, a high-precision tissue recognition model is constructed using a very small amount of coarsely labeled data.

Benefits of technology

It enables real-time tissue identification with high precision and high generalization ability without the need for massive amounts of precisely labeled data, thus improving the safety and accuracy of surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122005085A_ABST
    Figure CN122005085A_ABST
Patent Text Reader

Abstract

The invention provides a puncture tissue characteristic recognition system based on multi-modal fusion. The puncture tissue characteristic recognition system comprises a model training module. The module is used for acquiring label-free force, vibration and impedance signals and sample data with a rough tissue type label; an encoder network is trained through cross-modal contrast learning, and modal-independent robust fusion feature representation is generated by taking different modal signal features at the same moment as positive samples and taking different modal signal features at different moments as negative samples; determining causal and non-causal feature subsets by using a causal discovery algorithm; a causal knowledge-based regularization loss function is constructed according to the causal knowledge-based regularization loss function, wherein the causal knowledge-based regularization loss function comprises two items which enable a causal attention vector to approach 1 in a causal feature subset dimension weight and approach 0 in a non-causal feature subset dimension weight; and finally, supervising and finely adjusting the encoder network by using related data to generate an organization recognition model. According to the method, a high-precision and high-generalization-ability organization recognition model can be constructed without massive accurate annotation data, and the problem that model training application is difficult due to the fact that high-quality annotation data are difficult to obtain is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of ophthalmic characteristic recognition, and in particular relates to a multimodal fusion puncture tissue characteristic recognition system. Background Technology

[0002] In delicate ophthalmic puncture surgeries such as glaucoma drainage valve implantation, intravitreal injection, and intraocular tumor biopsy, existing technologies aim to identify the characteristics of different biological tissues (such as sclera, ciliary body, vitreous body, retina, etc.) along the puncture path by combining preoperative planning with intraoperative perception, so as to assist surgical navigation and improve surgical safety.

[0003] Current clinical practice relies primarily on two components: first, static path planning based on preoperative images such as optical coherence tomography (OCT) and ultrasound biomicroscopy (UBM); and second, the surgeon's subjective, experiential judgment based on "feel" during the procedure, assessing changes in tissue resistance encountered by the puncture needle. Preoperative images provide a "static map" of the tissue structure for planning the puncture path. During the procedure, the surgeon perceives and interprets changes in force signals through tactile feedback from holding the puncture needle, thereby inferring the type of tissue the needle tip may be contacting. Furthermore, existing research has attempted to quantify the surgeon's "feel" by integrating force sensors into surgical instruments (such as puncture needles) to convert tissue resistance into a one-dimensional force signal curve for analysis, aiming to provide objective evidence for tissue identification.

[0004] However, there are significant drawbacks in achieving organizational identification through the above methods:

[0005] First, preoperative images are static and cannot reflect the dynamic tissue deformation caused by instrument intervention and changes in intraocular pressure during surgery, resulting in a decrease in navigation accuracy.

[0006] Second, relying solely on a single sensing mode of one-dimensional force signals results in a severe lack of information dimension. Different tissues may exhibit similar mechanical properties (such as hardness), making it difficult for the system to make specific distinctions and provide accurate early warnings.

[0007] Third, while deep learning and other artificial intelligence models possess powerful potential for multimodal information fusion and pattern recognition, their supervised training heavily relies on massive, high-precision "input signal-output tissue type" paired labeled samples. However, in real-world ophthalmic puncture surgery scenarios, the "input signal" consists of multimodal temporal data such as force, vibration, and impedance generated in real-time during the procedure, while the corresponding "output label" (i.e., the precise tissue type contacted by the needle tip) is almost impossible to verify and label at the "gold standard" level without interfering with the surgery or increasing risks. Therefore, the extreme scarcity of high-quality training data has become a fundamental obstacle to applying artificial intelligence technology in this field to achieve accurate, objective, and real-time tissue identification. Summary of the Invention

[0008] The purpose of this application is to overcome the deficiencies in the prior art and provide a multimodal fusion puncture tissue characteristic identification system.

[0009] This application provides a multimodal fusion system for identifying puncture tissue characteristics, including: a model training module,

[0010] The model training module includes: acquiring unlabeled force signals, unlabeled vibration signals, unlabeled impedance signals, and sample data with coarse tissue type labels; training an encoder network through cross-modal contrastive learning based on the force signals, vibration signals, and impedance signals to generate modality-independent robust fusion feature representations, wherein cross-modal contrastive learning uses modal features corresponding to the force signals, vibration signals, and impedance signals acquired at the same time as positive samples and modal features corresponding to signals acquired at different times as negative samples; and determining causal and non-causal feature subsets based on the modality-independent robust fusion feature representations using a causal discovery algorithm, wherein the causal feature subset consists of feature dimensions causally related to tissue type, and the non-causal feature subset consists of feature dimensions causally related to tissue type. The causal feature subset is composed of feature dimensions that are not causally related. Based on the causal feature subset and the non-causal feature subset, a regularized loss function based on causal knowledge is constructed. This regularized loss function includes a first regularization term to make the weights of the learnable causal attention vector on the corresponding dimensions of the causal feature subset approach 1, and a second regularization term to make the weights of the learnable causal attention vector on the corresponding dimensions of the non-causal feature subset approach 0. Based on the sample data with coarse tissue type labels, the modality-independent robust fusion feature representation, the regularized loss function based on causal knowledge, and the learnable causal attention vector, the encoder network is supervised and fine-tuned to generate a tissue recognition model.

[0011] Optionally, acquiring unlabeled force signals, unlabeled vibration signals, and unlabeled impedance signals includes:

[0012] Acquire the unlabeled force signal collected by the fiber optic Fabry-Perot sensor, which is integrated into the sidewall of the puncture needle at a first distance from the needle tip;

[0013] The labelless vibration signal is acquired and collected by a piezoelectric ceramic element, which is integrated into the side wall of the puncture needle at a second distance from the needle tip, and the second distance is greater than the first distance;

[0014] Acquire the labelless impedance signal obtained by the measurement of the microelectrode pair, the microelectrode pair including a working electrode integrated into the tip of the puncture needle and a counter electrode located at a third distance behind the working electrode;

[0015] The piezoelectric ceramic element is located between the fiber optic Fabry-Perot sensor and the microelectrode pair.

[0016] Optionally, the encoder network is trained via cross-modal contrastive learning, including:

[0017] The force signal is resampled in time and scaled in amplitude to obtain an enhanced force signal;

[0018] The encoder network is trained by performing cross-modal contrastive learning based on the enhanced force signal, the vibration signal, and the impedance signal.

[0019] Optionally, the encoder network is trained via cross-modal contrastive learning, including:

[0020] A broadband noise with a preset signal-to-noise ratio is added to the vibration signal to obtain an enhanced vibration signal;

[0021] The encoder network is trained by performing cross-modal contrastive learning based on the force signal, the vibration signal, and the impedance signal.

[0022] Optionally, the encoder network includes:

[0023] The first one-dimensional convolutional layer, the second one-dimensional convolutional layer, the third one-dimensional convolutional layer, the global average pooling layer, and the multilayer perceptron projection head are connected in sequence.

[0024] The number of convolution kernels in the first, second, and third one-dimensional convolutional layers increases sequentially, while the kernel size decreases sequentially.

[0025] Optionally, a causal feature subset and a non-causal feature subset are determined using a causal discovery algorithm, including:

[0026] A permutation test based on the Hilbert-Schmidt independence criterion is used to test the conditional independence between different feature dimensions of the modality-independent robust fusion feature representation;

[0027] Based on the results of the conditional independence test, the causal feature subset and the non-causal feature subset are determined.

[0028] Optionally, a regularized loss function based on causal knowledge is constructed, including:

[0029] Construct the first regularization term, which is the norm of the difference between the weight of the learnable causal attention vector on the corresponding dimension of the causal feature subset and a first value, where the first value is 1;

[0030] Construct the second regularization term, which is the norm of the difference between the weight of the learnable causal attention vector on the corresponding dimension of the non-causal feature subset and the second value, where the second value is 0;

[0031] The causal knowledge-based regularization loss function is constructed based on the first regularization term and the second regularization term.

[0032] Optionally, it also includes:

[0033] Real-time identification based on the tissue identification model includes:

[0034] The force, vibration, and impedance signals acquired in real time are captured using a sliding window of fixed length.

[0035] The force signal, vibration signal, and impedance signal captured by the sliding window are input into the tissue recognition model;

[0036] The step size of the sliding window is much smaller than the length of the sliding window.

[0037] Optionally, the encoder network is trained through cross-modal contrastive learning based on the force signal, the vibration signal, and the impedance signal, including:

[0038] The vibration signal and the impedance signal are subjected to time-frequency transformation to generate a vibration time-frequency diagram representation and an impedance time-frequency diagram representation;

[0039] The vibration time-frequency diagram, the impedance time-frequency diagram, and the time series representation of the force signal are spliced ​​together in the channel dimension to generate a fused multi-channel signal.

[0040] The fused multi-channel signal is input into the encoder network to perform the cross-modal contrastive learning.

[0041] Optionally, real-time identification based on the tissue identification model further includes:

[0042] Obtain the recognition results and recognition confidence of the tissue recognition model for signals within multiple consecutive sliding windows;

[0043] When the identification result is a specific high-risk organization type and the identification confidence level exceeds a preset threshold, a security intervention instruction is generated.

[0044] The safety intervention commands include visual warning commands sent to the augmented reality device, and compliant control commands or soft locking commands sent to the surgical robot motion controller.

[0045] The beneficial effects of this application are:

[0046] This application provides a multimodal fusion-based puncture tissue characteristic recognition system, comprising: a model training module, the model training module comprising: acquiring unlabeled force signals, unlabeled vibration signals, unlabeled impedance signals, and sample data with coarse tissue type labels; training an encoder network through cross-modal contrastive learning based on the force signals, vibration signals, and impedance signals to generate modality-independent robust fusion feature representations, wherein the cross-modal contrastive learning uses modal features corresponding to the force signals, vibration signals, and impedance signals acquired at the same time as positive samples, and modal features corresponding to signals acquired at different times as negative samples; and determining causal feature subsets and non-causal feature subsets based on the modality-independent robust fusion feature representations through a causal discovery algorithm, wherein the causal feature subsets are composed of features related to tissue type. The encoder network is constructed using causally related feature dimensions, and the non-causal feature subset is composed of non-causally related feature dimensions. Based on the causal and non-causal feature subsets, a regularized loss function based on causal knowledge is constructed. This regularized loss function includes a first regularization term that makes the weights of the learnable causal attention vector on the corresponding dimensions of the causal feature subset approach 1, and a second regularization term that makes the weights of the learnable causal attention vector on the corresponding dimensions of the non-causal feature subset approach 0. The encoder network is then supervised and fine-tuned based on the sample data with coarse tissue type labels, the modality-independent robust fusion feature representation, the causal knowledge-based regularized loss function, and the learnable causal attention vector to generate a tissue recognition model. This application utilizes cross-modal contrastive learning to train an encoder to generate modality-independent robust fusion feature representations. Then, a causal discovery algorithm is used to separate a subset of features with stable causal associations with tissue type from this representation. The model is then supervised and fine-tuned using a very small amount of coarsely labeled data combined with a regularized loss function based on causal knowledge. This enables the construction of a high-precision, high-generalization real-time tissue identification model without relying on massive amounts of accurately labeled data, thus overcoming the core defect that artificial intelligence models cannot be effectively trained and applied due to the inability to obtain high-quality labeled data. Attached Figure Description

[0047] Figure 1 This is a diagram of the overall technical architecture of the system.

[0048] Figure 2 This is a schematic diagram of the internal structure of a neural network model. Detailed Implementation

[0049] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it is to be understood that various forms of implementation of the present disclosure are intended and should not be limited to the embodiments set forth herein. Rather, the embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0050] This application provides a multimodal fusion puncture tissue characteristic recognition system, applied in the fields of ophthalmic surgical robots and intelligent surgical navigation technology, to solve the problems of inaccurate real-time tissue recognition and low safety in ophthalmic puncture surgery due to reliance on preoperative static images and surgeon's tactile experience, as well as insufficient information from a single sensor modality and lack of training data for artificial intelligence models.

[0051] Please refer to Figure 1 and Figure 2 As shown, a multimodal fusion puncture tissue characteristic recognition system includes: a model training module, wherein the model training module includes:

[0052] S101. Acquire unlabeled force signals, unlabeled vibration signals, unlabeled impedance signals, and sample data with rough tissue type labels.

[0053] Acquiring unlabeled force signals refers to signals collected by a fiber optic Fabry-Perot interferometer force sensing unit, which measures the axial and radial micro-forces during puncture and senses the macroscopic stiffness and puncture resistance of the tissue.

[0054] Specifically, the label-free force signal is acquired by a fiber optic Fabry-Perot sensor integrated into the sidewall of the puncture needle at a first distance from the needle tip, for example, 2 mm from the needle tip. A femtosecond laser is used to etch a depth... A cylindrical microcavity with a diameter of approximately 50 μm and a diameter of 30 ± 5 μm is formed by inserting a 125 μm single-mode optical fiber with a flat end face into the microchannel inside the needle body, ensuring that its end face fits tightly against the bottom of the microcavity, thus forming the first reflecting surface of the Fabry interferometer. The outer wall of the needle, i.e., the top of the microcavity, forms a second reflective surface. The interior of the microcavity is filled with air.

[0055] When the needle tip is subjected to force At that time, the needle body undergoes micro-strain, resulting in the microcavity depth. occur Changes in reflected light intensity With cavity length The relationship is described by the following formula:

[0056]

[0057] in, It is the intensity of the incident light, and the unit can be watts (W) or any other unit of intensity. , These are the optical intensity reflectivities of the fiber end face and the microcavity top face, respectively, which are dimensionless and range from 0 to 1. It is the refractive index of the medium inside the cavity, relative to air. Dimensionless; the stated It is the center wavelength of the light source, and the unit is meters (m). It is the initial phase offset, in radians (rad). It is the instantaneous depth of the Fabry-Perot microcavity, measured in meters (m).

[0058] By demodulation The change can be calculated Furthermore, through a pre-defined force-displacement relationship:

[0059]

[0060] Gaining power ,in This is the local stiffness coefficient of the needle body, with units of Newtons per meter (N / m).

[0061] Acquiring unlabeled vibration signals refers to signals collected by a piezoelectric ceramic vibration excitation and sensing unit, which is used to excite and receive high-frequency micro-vibration signals.

[0062] Specifically, the labelless vibration signal excited and collected by the piezoelectric ceramic element is acquired. The piezoelectric ceramic element is integrated into the side wall of the puncture needle at a second distance from the needle tip, the second distance being greater than the first distance. For example, a ring-shaped PZT-5H ceramic with an outer diameter of 0.45 mm, an inner diameter of 0.35 mm, and a thickness of 0.2 mm is sleeved inside the needle body at a distance of 5 mm from the needle tip and sealed with biocompatible epoxy resin. The piezoelectric ceramic element is located between the fiber optic Fabry-Perot sensor and the microelectrode pair.

[0063] This unit serves as both an exciter and a sensor, providing the excitation signal. For frequency Sinusoidal voltage with linear frequency sweep from 10kHz to 50kHz:

[0064]

[0065] in, It is the voltage amplitude, and the unit is volts (V). , These are the start and end frequencies of the sweep, respectively, and the unit is Hertz (Hz). It is the frequency sweep period, and the unit is seconds (s). It's time, measured in seconds (s).

[0066] The PZT ring generates longitudinal mechanical vibrations at the same frequency, which are transmitted to the needle tip through the needle body. The mechanical impedance at the needle tip-tissue interface... The reflection of the vibration wave is determined, and the reflected wave is received by the same PZT ring to generate a voltage signal. .

[0067] right and Its spectrum is obtained by performing a fast Fourier transform. and Calculate the frequency response function:

[0068]

[0069] Equivalent mechanical resistance of tissue and It is related to the inherent parameters of the system, and the relationship curve can be obtained through pre-calibration.

[0070] Acquiring a label-free impedance signal refers to the signal obtained by measuring a pair of miniature bioelectrical impedance measurement electrodes, which are used to sense the electrophysiological properties of tissues.

[0071] Specifically, the label-free impedance signal obtained by measuring the microelectrode pair is acquired. The microelectrode pair includes a working electrode integrated into the tip of the puncture needle and a counter electrode located at a third distance behind the working electrode. For example, a platinum film with a thickness of about 100 nm is deposited by magnetron sputtering at the needle tip (working electrode, WE) and 0.5 mm behind the needle tip (counter electrode, CE), and then the film is processed into a width using focused ion beam etching technology. It is a ring electrode with a diameter of 10±2μm.

[0072] The four-terminal measurement method is used, and the frequency is generated by a precision impedance analysis chip. AC excitation current from 1 kHz to 1 MHz The flow of the tissue between the two electrodes was synchronized with the measurement of the voltage between them. .

[0073] Complex impedance of tissue for:

[0074]

[0075] in, The series resistance of an organization is measured in ohms (Ω). The series reactance of the structure is expressed in ohms (Ω).

[0076] , It is angular frequency. It is the equivalent series capacitance, and its unit is farad (F).

[0077] It is the imaginary unit; The spectrum reflects the dielectric properties of the tissue.

[0078] All sensor signals are sampled synchronously with high precision by a multi-channel synchronous data acquisition card, and then processed by a preprocessing module to form a time-aligned multimodal data stream.

[0079] Preprocessing includes power frequency notch filtering, bandpass filtering, noise reduction, and normalization. Normalization involves performing sample-by-sample z-score normalization on each signal segment of each mode.

[0080]

[0081] in, It is the mean of that signal segment. It is the standard deviation of the signal segment. This refers to the segment length, for example, 10,000 sampling points.

[0082] Obtaining sample data with rough tissue type labels refers to collecting a small dataset:

[0083]

[0084] in Very small, for example, 100. These are rough tissue category labels such as 1-sclera, 2-ciliary body / choroid, 3-vitreous body, 4-retina, 5-tumor / abnormal tissue. These labels can come from targeted puncture records of ex vivo pig eye experiments or from retrospective rough judgments of clinical surgical videos.

[0085] S102. Based on the force signal, the vibration signal, and the impedance signal, train the encoder network through cross-modal contrastive learning to generate a modality-independent robust fusion feature representation. In the cross-modal contrastive learning, the modal features corresponding to the force signal, the vibration signal, and the impedance signal collected at the same time are used as positive samples, and the modal features corresponding to the signals collected at different times are used as negative samples.

[0086] The encoder network was trained through cross-modal contrastive learning, including constructing an unlabeled dataset and collecting raw signals from a large number of clinical or in vitro experiments to obtain three synchronized discrete time series. , , , .

[0087] Generating training sample pairs involves sampling, randomly selecting a starting point on the time series. The cut length is Each sampling point corresponds to a physical duration ,For example , when Excerpt:

[0088]

[0089]

[0090]

[0091] These three segments constitute an original triplet sample. .

[0092] To prevent the model from learning meaningless low-level features, data augmentation applies random weak augmentation to each modality segment, generating two different views of the sample.

[0093] Define enhancement function ,in Represents a mode.

[0094] The unlabeled force signal is subjected to time resampling and amplitude scaling to obtain an enhanced force signal, specifically as follows: ,in It is a time resampling function. Simulate minute changes in puncture speed. It is the amplitude scaling factor.

[0095] A broadband noise with a preset signal-to-noise ratio is added to the unlabeled vibration signal to obtain an enhanced vibration signal, specifically as follows:

[0096] ,in

[0097] It is an additive white Gaussian noise vector, with a signal-to-noise ratio (SNR) controlled within the range of 20-30 dB, i.e.:

[0098]

[0099] in, This indicates the calculation of the root mean square.

[0100] For impedance signals The enhancement to ,in .

[0101] For an original sample Two independent enhancements were applied to obtain two enhanced samples:

[0102]

[0103]

[0104] The enhancement parameters are generated independently and randomly.

[0105] The cross-modal contrastive learning is performed based on the enhanced force signal, the unlabeled vibration signal, and the unlabeled impedance signal, or based on the unlabeled force signal, the enhanced vibration signal, and the unlabeled impedance signal.

[0106] like Figure 2 As shown, the encoder network includes a first one-dimensional convolutional layer, a second one-dimensional convolutional layer, a third one-dimensional convolutional layer, a global average pooling layer, and a multilayer perceptron projector connected in sequence. The number of convolutional kernels in the first one-dimensional convolutional layer, the second one-dimensional convolutional layer, and the third one-dimensional convolutional layer increases sequentially, and the size of the convolutional kernels in the first one-dimensional convolutional layer, the second one-dimensional convolutional layer, and the third one-dimensional convolutional layer decreases sequentially.

[0107] Specifically, the encoder network force signal Vibration signal Impedance signal They have the same structure, both being one-dimensional convolutional neural networks, but they do not share parameters. (Example: Force signal encoder) For example, the input layer receives a shape of... Force signal fragment , .

[0108] Number of convolutional kernels in convolutional layer 1 kernel size Step length ,filling Following batch normalization and the ReLU activation function, the output feature map size is [length]. , number of channels 64.

[0109] Number of convolutional kernels in a convolutional layer kernel size Step length ,filling Following batch normalization and ReLU, the output size length is calculated. Number of channels: 128.

[0110] Convolutional layer 3 Convolutional kernel number kernel size Step length ,filling Following batch normalization and ReLU, the output size length is calculated. Number of channels: 256.

[0111] Global average pooling layer for length dimension The average is then applied to output a 256-dimensional feature vector. The projection head is a two-layer multilayer perceptron used to map features to a contrastive learning space. The first layer is a linear transformation:

[0112]

[0113] in , ;

[0114] Second-level linear transformation:

[0115]

[0116] in , .

[0117] Final output And perform L2 normalization:

[0118]

[0119] Vibration and impedance encoders , The structures are exactly the same, and the final output is a normalized feature vector. .

[0120] For a batch A set of original samples, enhanced to obtain... Each augmented sample, after encoding and projection, yields the feature vector set of all samples:

[0121]

[0122] In the calculation of the cross-modal contrastive loss function, the goal of contrastive learning is that for any feature vector anchor point, the similarity with the positive feature vector sample from the same original sample and the same time but different modal should be as high as possible, and the similarity with all other negative feature vector samples in the batch should be as low as possible.

[0123] The feature vectors used in the similarity calculation have all been L2 normalized, and their cosine similarity is equal to their dot product:

[0124]

[0125] The loss function is defined using symmetric cross-modal InfoNCE loss, with anchor points as force modal features. For example, positive sample set Contains the same sample Same view Figure 1 Vibration and impedance characteristics under negative sample set That is, batch feature set The loss for a single anchor point is calculated as follows, excluding the anchor point itself and its positive samples, and all feature vectors.

[0126]

[0127] Among them, the It is a temperature hyperparameter, a scalar, dimensionless quantity, usually set to 0.1, used to adjust the degree of attention given to difficult negative samples.

[0128] Similarly, calculate with , , , , Loss at the anchor point:

[0129]

[0130] The total batch loss is averaged over the anchor point loss of all samples, all views, and all modalities within a batch:

[0131]

[0132] The model was trained using the Adam optimizer, with an initial learning rate set to... The encoder network is trained on an unlabeled dataset containing millions of triplet samples using cosine annealing learning rate scheduling, minimizing the contrastive loss. .

[0133] After training is complete, discard the projector head and retain the encoder. , , The convolutional backbone is used to extract general features. .

[0134] Furthermore, training the encoder network through cross-modal contrastive learning also includes performing time-frequency transformation on the unlabeled vibration signal and the unlabeled impedance signal to generate a vibration time-frequency map representation and an impedance time-frequency map representation. The vibration time-frequency map representation, the impedance time-frequency map representation, and the time series representation of the unlabeled force signal are then concatenated along the channel dimension to generate a fused multi-channel signal. The fused multi-channel signal is then input into the encoder network to perform the cross-modal contrastive learning.

[0135] S103. Based on the modality-independent robust fusion feature representation, a causal feature subset and a non-causal feature subset are determined by a causal discovery algorithm. The causal feature subset consists of feature dimensions that are causally related to the organization type, and the non-causal feature subset consists of feature dimensions that are not causally related.

[0136] Causal discovery algorithms are used to identify causal and non-causal feature subsets, including feature extraction and dataset construction, using a trained encoder. , , Freezing parameters, processing large amounts of unlabeled data, for a single sample Extracting general features , , These three 256-dimensional feature vectors are concatenated to obtain a 768-dimensional fused feature vector. ,use For example Such samples are used to construct causal analysis datasets ,in .

[0137] Constraint-based causal discovery algorithms utilize 768-dimensional features Each dimension , Treating it as an observed variable, we assume the existence of a common latent variable. Represents organization type, and possible confounding variables. Individual patient differences, sensor baselines, and other factors collectively generate observational features. The FastCausalInference algorithm is used for inference. The core of the causal structure between the dimensions is the conditional independence test.

[0138] Construct a complete undirected graph. Initialize a complete undirected graph with 768 nodes. Each node represents a feature dimension. Conditional independence test for each pair of adjacent nodes in the graph. and Test them given a certain set of conditions Whether the time intervals are independent is determined using a kernel-based Hilbert-Schmidt independence criterion permutation test, which can handle nonlinear relationships, and the null hypothesis is adopted. for and In a given Conditions are independent. The HSIC statistic is calculated, and the p-value is calculated by generating a null distribution using a permutation sample. If the p-value is greater than the significance level (e.g., 0.05), the result is accepted. In the diagram Remove connection and The edge.

[0139] After obtaining an undirected skeleton graph, the orienting edges are used in accordance with a series of orientation rules of the FCI algorithm, such as collision structure recognition. ,and and Not adjacent, and Not here and If the conditions are independent and centralized, then the orientation is... By determining the direction of some edges, a partially directed acyclic graph is obtained.

[0140] In the obtained PDAG, identify causal feature dimensions and search for structural patterns that satisfy the following conditions to infer latent variables. Using causal features, identify latent variable surrogates to find such node sets in the graph. They are a group of nodes that are highly connected to each other, and are a common feature of many other scattered nodes, namely, they have many directed edges connecting them. The nodes in the graph point to many other nodes. Verifying modal equalization involves checking whether the nodes being pointed to are uniformly distributed. This is achieved using three different modal encoders, i.e., indices. Roughly evenly distributed across three intervals: 1-256, 257-512, and 513-768. The true common cause should have a similar impact on all modalities. A causal feature set is defined to encompass all directly caused by... The node in the text points to the feature node The set is defined as a subset of causal features. The others in the picture are... Feature nodes without a direct causal relationship or primarily influenced by other unobserved variables are classified as non-causal feature subsets that may be affected by confounding factors. , and It is a set of indexes based on feature dimensions, satisfying:

[0141]

[0142] and In practice It may contain dozens to hundreds of indexes.

[0143] S104. Based on the causal feature subset and the non-causal feature subset, construct a regularized loss function based on causal knowledge, wherein the regularized loss function based on causal knowledge includes a first regularization term for making the weight of the learnable causal attention vector in the corresponding dimension of the causal feature subset approach 1, and a second regularization term for making the weight of the learnable causal attention vector in the corresponding dimension of the non-causal feature subset approach 0.

[0144] Constructing a regularized loss function based on causal knowledge includes constructing a first regularization term, wherein the first regularization term is the norm of the difference between the weights of the learnable causal attention vector on the corresponding dimension of the causal feature subset and a first value, wherein the first value is 1.

[0145] Specifically, a learnable causal attention vector is introduced. , where each element:

[0146]

[0147] in, These are learnable parameters that enable The learning of this vector will be constrained by the causal knowledge discovered in stage B.

[0148] Causal regularization loss function: This loss forces a causal attention vector. Consistent with the causal structure discovered in stage B, let the set of causal feature dimension indices identified in stage B be . The set of non-causal feature dimension indices is The second regularization term is constructed, which is the norm of the difference between the weights of the learnable causal attention vector on the corresponding dimension of the non-causal feature subset and a second value, where the second value is 0. Based on the first and second regularization terms, the causal knowledge-based regularization loss function is constructed, specifically as follows:

[0149]

[0150] in, and These represent the size and number of elements of the two sets, respectively. and It is a hyperparameter scalar, for example ;

[0151] The average attention weight for causal features is encouraged to approach 1. The average attention weight that penalizes non-causal features approaches 0.

[0152] S105. Based on the sample data with coarse tissue type labels, the modality-independent robust fusion feature representation, the regularized loss function based on causal knowledge, and the learnable causal attention vector, the encoder network is supervised and fine-tuned to generate a tissue recognition model.

[0153] Few-sample fine-tuning of objectives based on causal intervention utilizes a very small amount of coarsely labeled data to train a classifier whose decisions primarily rely on causal features. Prepare a small sample labeled dataset: Collect a small dataset.

[0154]

[0155] in, Very small, for example , These are rough tissue category labels such as 1-sclera, 2-ciliary body / choroid, 3-vitreous body, 4-retina, 5-tumor / abnormal tissue.

[0156] Network structure adjustment and causal intervention module feature extraction using frozen encoder , , Process each sample to obtain fusion features .

[0157] Causal feature masks introduce learnable causal attention vectors. Each element , These are learnable parameters that enable .

[0158] Feature weighting calculation of weighted features ,in This represents the Hadamard product, which is an element-wise product. Ideally, for... Causal Feature Index It should approach 1, for Non-causal feature index It should approach 0.

[0159] Classifier weighted features Mapped to category logical values ​​through a linear classification layer. ,in It is a learnable weight matrix, and then the class probability distribution is obtained through the softmax function:

[0160]

[0161] in, It is a sample Category The predicted probability.

[0162] Loss Function Design and Optimization: Total Loss Function The standard classification loss consists of two parts. and causal regularization loss :

[0163]

[0164] in, It is a scalar of the balancing hyperparameter, for example, set to 0.5.

[0165] The classification loss is the standard cross-entropy loss:

[0166]

[0167] in, It is an indicator function; its value is 1 when the condition inside the parentheses is true, and 0 otherwise.

[0168] Fine-tuning training in Optimize only the causal attention vector parameters on the dataset Right now and classifier weight matrix encoder , , The parameters are kept frozen, using a small learning rate such as The model is trained with the Adam optimizer for dozens of epochs until it converges. This process enables the model to learn to classify only those features that have a stable causal relationship with the organization type, thereby gaining strong generalization ability.

[0169] Furthermore, it includes real-time identification based on the trained tissue recognition model with high generalization ability, including capturing real-time force signals, vibration signals, and impedance signals using a sliding window of fixed length. The step size of the sliding window is much smaller than the length of the sliding window, for example, the window length... Step length Real-time processing of signal streams.

[0170] The force signal, vibration signal, and impedance signal captured by the sliding window are input into the tissue recognition model.

[0171] The specific online real-time inference process involves signal acquisition and preprocessing. This process acquires three raw signals within the current time window and performs real-time preprocessing to obtain standardized segments. .

[0172] Feature extraction inputs the preprocessed signal segments into the frozen encoder. , , Obtain the feature vector And spliced ​​together .

[0173] Causal feature selection and classification calculation of weighted features ,in It is the fixed attention vector obtained after fine-tuning. Input a trained linear classifier Obtain the logical value vector Then, the probability of each class is obtained by softmax. .

[0174] The decision and output take the category with the highest probability as the recognition result for the current window. Simultaneously calculate the confidence level .

[0175] The system acquires the identification results and confidence scores of the organization identification model for signals within multiple consecutive sliding windows. When the identification result indicates a specific high-risk organization type and the identification confidence scores all exceed a preset threshold, a security intervention command is generated. For example, when... For key tissues such as the retina and A safety warning is triggered when the value exceeds a threshold such as 0.9.

[0176] The safety intervention commands include visual warning commands sent to the augmented reality device, and compliant control commands or soft locking commands sent to the surgical robot's motion controller. Real-time display will... and The information transmitted to the AR display system is updated in real time, including overlaid information in the surgical field of view such as the color halo around the needle tip, tissue type text labels, and confidence progress bars.

[0177] The above description of the embodiments is provided to enable those skilled in the art to understand and apply this application. Those skilled in the art will readily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without inventive effort. Therefore, this application is not limited to the above embodiments, and any improvements and modifications made to this application based on the disclosure thereof should be within the scope of protection of this application.

Claims

1. A multimodal fusion system for identifying the characteristics of punctured tissue, characterized in that, include: Model training module, The model training module includes: acquiring unlabeled force signals, unlabeled vibration signals, unlabeled impedance signals, and sample data with rough tissue type labels; Based on the force signal, vibration signal, and impedance signal, an encoder network is trained through cross-modal contrastive learning to generate a modality-independent robust fusion feature representation. In this cross-modal contrastive learning, modal features corresponding to the force signal, vibration signal, and impedance signal acquired at the same time are considered positive samples, while modal features corresponding to signals acquired at different times are considered negative samples. Based on the modality-independent robust fusion feature representation, a causal feature subset and a non-causal feature subset are determined through a causal discovery algorithm. The causal feature subset consists of feature dimensions causally related to tissue type, and the non-causal feature subset consists of feature dimensions not causally related to tissue type. Based on the causal features… A causal knowledge-based regularized loss function is constructed using the subset and the non-causal feature subset. This causal knowledge-based regularized loss function includes a first regularization term that makes the weights of the learnable causal attention vector on the corresponding dimension of the causal feature subset approach 1, and a second regularization term that makes the weights of the learnable causal attention vector on the corresponding dimension of the non-causal feature subset approach 0. The encoder network is then supervised and fine-tuned based on the sample data with coarse tissue type labels, the modality-independent robust fusion feature representation, the causal knowledge-based regularized loss function, and the learnable causal attention vector to generate a tissue recognition model.

2. The system according to claim 1, characterized in that, Acquire unlabeled force signals, unlabeled vibration signals, and unlabeled impedance signals, including: Acquire the unlabeled force signal collected by the fiber optic Fabry-Perot sensor, which is integrated into the sidewall of the puncture needle at a first distance from the needle tip; The labelless vibration signal is acquired and collected by a piezoelectric ceramic element, which is integrated into the side wall of the puncture needle at a second distance from the needle tip, and the second distance is greater than the first distance; Acquire the labelless impedance signal obtained by the measurement of the microelectrode pair, the microelectrode pair including a working electrode integrated into the tip of the puncture needle and a counter electrode located at a third distance behind the working electrode; The piezoelectric ceramic element is located between the fiber optic Fabry-Perot sensor and the microelectrode pair.

3. The system according to claim 1, characterized in that, The encoder network is trained through cross-modal contrastive learning, including: The force signal is resampled in time and scaled in amplitude to obtain an enhanced force signal; The encoder network is trained by performing cross-modal contrastive learning based on the enhanced force signal, the vibration signal, and the impedance signal.

4. The system according to claim 1, characterized in that, The encoder network is trained through cross-modal contrastive learning, including: A broadband noise with a preset signal-to-noise ratio is added to the vibration signal to obtain an enhanced vibration signal; The encoder network is trained by performing cross-modal contrastive learning based on the force signal, the vibration signal, and the impedance signal.

5. The system according to claim 1, characterized in that, The encoder network includes: The first one-dimensional convolutional layer, the second one-dimensional convolutional layer, the third one-dimensional convolutional layer, the global average pooling layer, and the multilayer perceptron projection head are connected in sequence. The number of convolution kernels in the first, second, and third one-dimensional convolutional layers increases sequentially, while the kernel size decreases sequentially.

6. The system according to claim 1, characterized in that, Causal feature subsets and non-causal feature subsets are determined using causal discovery algorithms, including: A permutation test based on the Hilbert-Schmidt independence criterion is used to test the conditional independence between different feature dimensions of the modality-independent robust fusion feature representation; Based on the results of the conditional independence test, the causal feature subset and the non-causal feature subset are determined.

7. The system according to claim 1, characterized in that, Constructing a regularized loss function based on causal knowledge, including: Construct the first regularization term, which is the norm of the difference between the weight of the learnable causal attention vector on the corresponding dimension of the causal feature subset and a first value, where the first value is 1; Construct the second regularization term, which is the norm of the difference between the weight of the learnable causal attention vector on the corresponding dimension of the non-causal feature subset and the second value, where the second value is 0; The causal knowledge-based regularization loss function is constructed based on the first regularization term and the second regularization term.

8. The system according to claim 1, characterized in that, Also includes: Real-time identification based on the tissue identification model includes: The force, vibration, and impedance signals acquired in real time are captured using a sliding window of fixed length. The force signal, vibration signal, and impedance signal captured by the sliding window are input into the tissue recognition model; The step size of the sliding window is much smaller than the length of the sliding window.

9. The system according to claim 1, characterized in that, Based on the force signal, the vibration signal, and the impedance signal, an encoder network is trained through cross-modal contrastive learning, including: The vibration signal and the impedance signal are subjected to time-frequency transformation to generate a vibration time-frequency diagram representation and an impedance time-frequency diagram representation; The vibration time-frequency diagram, the impedance time-frequency diagram, and the time series representation of the force signal are spliced ​​together in the channel dimension to generate a fused multi-channel signal. The fused multi-channel signal is input into the encoder network to perform the cross-modal contrastive learning.

10. The system according to claim 8, characterized in that, Real-time identification based on the tissue identification model also includes: Obtain the recognition results and recognition confidence of the tissue recognition model for signals within multiple consecutive sliding windows; When the identification result is a specific high-risk organization type and the identification confidence level exceeds a preset threshold, a security intervention instruction is generated. The safety intervention commands include visual warning commands sent to the augmented reality device, and compliant control commands or soft locking commands sent to the surgical robot motion controller.