Surgical operation guiding method and system based on multi-modal fusion and storage medium
Through the multimodal fusion surgical operation guidance method, the LSTM model is used to predict the position of the surgical robot and the force information of the instrument, which solves the problem of artificial participation in surgical operation guidance in the prior art, and achieves the consistency and safety of surgical operation.
Patent Information
- Application Number
- CN202510586612.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, surgical operation guidelines require human participation, and the guidance data for subsequent operations cannot be directly obtained, resulting in an intraoperative suspension and increasing the risk factor.
Using a multimodal fusion-based surgical operation guidance method, by recording the standard position information of the surgical robot and the standard force information of the instrument, multimodal feature extraction and bidirectional timing prediction training of the LSTM model are carried out, and surgical guidance model is constructed to predict the next position posture of the surgical robot and the next moment force information of the instrument.
It realizes direct acquisition of subsequent operation guidance data, avoids intraoperative pause, ensures the consistency of surgical operations, and reduces the risk of surgery.
Smart Images

Figure CN120093435A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical operation guidance, and in particular to a surgical operation guidance method, system and storage medium based on multimodal fusion. Background Art
[0002] Surgical navigation is a technology that uses computers, precision measuring instruments, and image processing to guide surgeons in performing surgical operations. It can help surgeons accurately locate and track the position of surgical instruments in space to improve the safety and accuracy of surgery and help doctors perform surgery. It is currently widely used in neurosurgery, orthopedics, cardiac surgery, tumor interventional treatment, oral and maxillofacial surgery, ophthalmology, otolaryngology, etc. Surgical navigation technology is often likened to GPS navigation. Just like a driver can see his position on a map, surgical navigation can accurately guide doctors to the location of the lesion and achieve precise surgical operations.
[0003] Therefore, the current surgical operation guidelines still require human participation, and it is impossible to directly obtain the guidance data for subsequent operations, that is, the surgical robot posture and instrument force conditions for subsequent operations. Intraoperative analysis is required, which leads to intraoperative suspension and increases the risk factor. Summary of the invention
[0004] The purpose of the present invention is to provide a surgical operation guidance method based on multimodal fusion to solve the technical problems in the prior art that human participation is still required, guidance data for subsequent operations cannot be directly obtained, intraoperative analysis is required, intraoperative suspension is caused, and the risk factor is increased.
[0005] In order to solve the above technical problems, the present invention specifically provides the following technical solutions: A surgical operation guidance method based on multimodal fusion, comprising the following steps: During the entire surgical process, the standard posture information of the surgical robot at each process node and the standard force information of the instrument fed back by the force sensor will be recorded, and the standard posture information and the standard force information of the instrument will be combined into standard multi-source information for real-time surgical operations; Extracting multimodal features from the standard multi-source information at each process node to obtain multimodal features of the standard multi-source information at each process node; The multimodal features of the standard multi-source information at each process node are trained through a LSTM model for bidirectional time series prediction to obtain a surgical guidance model; The surgical guidance model is used to predict the next-moment position information of the surgical robot and the next-moment force information of the instrument fed back by the force sensor based on real-time multi-source information consisting of the real-time position information of the surgical robot and the real-time force information of the instrument fed back by the force sensor.
[0006] As a preferred solution of the present invention, the method for extracting multimodal features includes: The KL divergence is used to quantify the difference between the multidimensional information of the post-process node and the multidimensional information of the pre-process node in two adjacent process nodes in the whole process cycle of the surgery, and the KL divergence value between each type of information in the multidimensional information of the post-process node and each type of information in the multidimensional information of the pre-process node is obtained; The KL divergence value between each type of information in the multidimensional information of the post-process node and each type of information in the multidimensional information of the pre-process node is used as the time scale weight of each type of information in the standard multi-source information; The time scale weight is: ; ; ; In the formula, is the time scale weight of the i-th type of information in the standard multi-source information at the t-th process node, is the i-th type of information in the standard multi-source information at the t-th process node, is the i-th type of information in the standard multi-source information at the t-1th process node, for and The KL divergence between for and The KL divergence between The time scale weight of the i-th type of information in the standard multi-source information at the t-th process node Weighted to the i-th category of information in the standard multi-source information at the t-th process node , get the multi-source time combination information at the tth process node , n is the total number of information categories in the standard multi-source information; Combine the multi-source time information at the tth process node , through the CNN feature extraction network, the spatial features of the multi-source time combination information at the tth process node are obtained ; The spatial characteristics of the multi-source time combination information at the tth process node After being processed by the self-attention mechanism, the multimodal features at the tth process node are obtained; The multimodal features are: ; in, ; ; ; In the formula, , and They are the query Query, key Key and value Value in the attention mechanism. for The dimension of is the intermediate value operator, is a dimension extraction operator, is the spatial feature of the i-th type of information in the multi-source time combination information at the t-th process node, is the multimodal feature of the i-th category of information in the standard multi-source information at the t-th process node, and T is the transposition operator.
[0007] As a preferred embodiment of the present invention, the method for constructing the surgery guidance model includes: The standard multi-dimensional information at each process node in the whole process cycle of the operation is serialized by process nodes to obtain a standard multi-dimensional information sequence; The multimodal features of the standard multidimensional information are used as the input of the first LSTM network, and the standard multidimensional information located at the next process node of the multimodal features is used as the output of the first LSTM network; The standard multi-dimensional information is used as the input of the second LSTM network, and the standard multi-dimensional information located at the previous process node of the standard multi-dimensional information is used as the output of the second LSTM network; The first LSTM network and the second LSTM network are trained with a combined loss of a predictive loss between an output of the first LSTM network and a true value and a multi-dimensional information cycle consistency loss between an output of the second LSTM network and an input of the first LSTM network; The first LSTM network after training is used as the surgery guidance model; Wherein, the first LSTM network is: =LSTM1( ); In the formula, is the i-th type of information in the standard multi-source information at the t+1th process node, is the multimodal feature of the i-th type of information in the standard multi-source information at the t-th process node, LSTM1 is the LSTM network; The second LSTM network is: =LSTM2( ); In the formula, is the i-th type of information in the standard multi-source information at the t-th process node, is the i-th category of information in the standard multi-source information at the t+1th process node, LSTM2 is the LSTM network, and n is the total number of information categories in the standard multi-source information.
[0008] As a preferred solution of the present invention, the method for constructing the combined loss includes: Determine the predictive loss between the output of the first LSTM network and the second LSTM network and the true value as: ; In the formula, For predictive loss, Output of the first LSTM network , for The truth value of Output of the second LSTM network , for The truth value of is the L2 norm, is the L2 norm, m is the total number of process nodes in the whole surgical process cycle, and n is the total number of information categories in the standard multi-source information; Determine the multi-dimensional information cycle consistency loss between the output of the second LSTM network and the input of the first LSTM network as: ; In the formula, is the cycle consistency loss, is the input of the first LSTM network The corresponding , Output of the second LSTM network , is the L2 norm, m is the total number of process nodes in the whole surgical process cycle, and n is the total number of information categories in the standard multi-source information; Will and The combined loss is: ; In the formula, is the combined loss, For balance Hyperparameters of .
[0009] As a preferred solution of the present invention, the standard posture information at each process node is normalized, and the standard force information of the instrument at each process node is normalized.
[0010] As a preferred solution of the present invention, the first LSTM network and the second LSTM network have the same network structure.
[0011] As a preferred solution of the present invention, the hyperparameter for: ; In the formula, Output of the second LSTM network , for The truth value of .
[0012] As a preferred solution of the present invention, the real-time multi-source information and the standard multi-source information at each process node are normalized.
[0013] As a preferred solution of the present invention, the present invention provides a surgical operation guidance system based on multimodal fusion, which is applied to a surgical operation guidance method based on multimodal fusion. The system includes: A data acquisition unit is used to record the standard posture information of the surgical robot at each process node during the entire surgical process cycle, and the standard force information of the instrument fed back by the force sensor, and to combine the standard posture information and the standard force information of the instrument into standard multi-source information for real-time surgical operations, and to obtain real-time multi-source information composed of the real-time posture information of the surgical robot and the real-time force information of the instrument fed back by the force sensor; A feature processing unit, used for extracting multimodal features from the standard multi-source information at each process node to obtain multimodal features of the standard multi-source information at each process node; A model building unit, used to perform bidirectional time series prediction training on the multimodal features of the standard multi-source information at each process node through an LSTM model to obtain a surgery guidance model; The guidance prediction unit is used to use the surgical guidance model to predict the next moment posture information of the surgical robot and the next moment force information of the instrument fed back by the force sensor based on real-time multi-source information composed of the real-time posture information of the surgical robot and the real-time force information of the instrument fed back by the force sensor.
[0014] As a preferred embodiment of the present invention, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, a surgical operation guidance method based on multimodal fusion is implemented.
[0015] Compared with the prior art, the present invention has the following beneficial effects: The present invention constructs a surgical guidance model, and realizes the prediction of the surgical robot's position information at the next moment and the instrument's force information at the next moment fed back by the force sensor based on real-time multi-source information composed of the surgical robot's real-time position information and the instrument's real-time force information fed back by the force sensor, so as to directly obtain the guidance data for subsequent operations, that is, the surgical robot's position and instrument force conditions for subsequent operations, without the need to pause during the operation, thereby ensuring the continuity of the surgical operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the implementation methods or the description of the prior art. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.
[0017] Figure 1 A flow chart of a surgical operation guidance method based on multimodal fusion provided by an embodiment of the present invention; Figure 2 A block diagram of a surgical operation guidance system based on multimodal fusion provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] like Figure 1 As shown, the present invention provides a surgical operation guidance method based on multimodal fusion, comprising the following steps: During the entire surgical process, the standard posture information of the surgical robot at each process node and the standard force information of the instrument fed back by the force sensor will be recorded, and the standard posture information and the standard force information of the instrument will be combined into standard multi-source information for real-time surgical operations; Perform multimodal feature extraction in the standard multi-source information at each process node to obtain the multimodal features of the standard multi-source information at each process node; The multimodal features of the standard multi-source information at each process node are trained through the LSTM model for bidirectional time series prediction to obtain the surgical guidance model; The surgical guidance model is used to predict the next-moment position information of the surgical robot and the next-moment force information of the instrument fed back by the force sensor based on real-time multi-source information consisting of the real-time position information of the surgical robot and the real-time force information of the instrument fed back by the force sensor.
[0020] The present invention first aggregates all data related to the surgical operation, including the real-time posture information of the surgical robot, such as the robot's joint angles, position coordinates and other motion information, and the real-time force information of the instrument fed back by the force sensor, such as the depth, amplitude, speed and strength of the surgical instrument's tissue traction and cutting operations, and combines these surgery-related data into multi-source information to guide the surgical operation, thereby achieving comprehensive guidance of the surgical operation.
[0021] After combining to form multi-source information, the present invention mines the feature importance of the multi-source information on the time scale, thereby allocating attention to each type of information in the multi-source information on the time scale, and completing feature weighted fusion based on the attention allocated on the time scale to obtain the multimodal features of the multi-source information on the time scale. The present invention also highlights the important features on the time scale in the multimodal features, so that the subsequent surgical guidance model constructed based on the multimodal features can quickly capture the important features on the time scale and allocate high temporal attention to them, so as to obtain accurate surgical operation data prediction based on the important features on the time scale, that is, to obtain multimodal features for accurately predicting surgical operation data on the time scale.
[0022] Specifically, the present invention uses KL divergence to perform difference analysis on the time scale of multi-source information, so that at two adjacent process time nodes, if the difference of a certain category of information in the multi-source information is high, then this type of information changes, and more attention needs to be allocated on the time scale. The dynamic information contained therein increases, and the information provided for surgical status perception increases, so it is necessary to classify it with a high weight on the time scale to highlight it as an important feature on the time scale, and then the surgical guidance model can be used on the time scale to predict the surgical operation through the multimodal features on the time scale, so as to perform accurate surgical operation guidance.
[0023] The present invention obtains multi-source time combination information after time scale weighting of multi-modal features of multi-source information, and utilizes the self-attention mechanism to explore the feature importance of multi-source information on the spatial scale, thereby allocating attention to various types of information in the multi-source information on the spatial scale, and completing feature weighted fusion based on the attention allocated on the spatial scale to obtain the multi-modal features of the multi-source information on the spatial scale. It also highlights the important features on the spatial scale in the multi-modal features, so that the subsequent surgical guidance model constructed based on the multi-modal features can quickly capture the important features on the spatial scale and allocate high temporal attention to them, so as to achieve accurate surgical operation data prediction based on the important features on the spatial scale, that is, to achieve multi-modal features for accurately predicting surgical operation data on the spatial scale.
[0024] Among them, the Self-Attention Mechanism, also known as the Intra-Attention Mechanism, is a special attention mechanism that allows the model to focus on the relationship between elements within the sequence when processing sequence data, thereby capturing the complex dependencies within the sequence and obtaining feature relationships on a spatial scale.
[0025] In summary, the present invention performs feature extraction on multi-source information at both the temporal scale and the spatial scale, and can extract multimodal features from multi-source information that can ensure high-precision surgical prediction at both the temporal scale and the spatial scale, thereby providing a data basis for building an accurate surgical guidance model.
[0026] After completing the extraction of multimodal features from multi-source information, the present invention constructs a surgical guidance model with a time series prediction function, which can predict the multi-source information of the next surgical process node based on the real-time multi-source information of the current surgical process node, thereby grasping in advance the robot's joint angles, position coordinates and other motion information at the next surgical process node, as well as the depth, amplitude, speed and strength of the surgical instrument's tissue traction and cutting operations. At the next surgical process node, surgical operations can be performed based on the above data, thereby realizing surgical operation guidance.
[0027] The present invention adopts two model structures for joint training when constructing the surgical guidance model. The first LSTM network is used to establish a temporal mapping relationship between the multimodal features of real-time multidimensional information and the multidimensional information of the next process node, so that the multidimensional information of the next process node can be predicted according to the multimodal features of the real-time multidimensional information, which belongs to data prediction along the surgical process. The second LSTM network is used to establish a temporal mapping relationship between the multidimensional information of the next process node and the multidimensional information of the real-time process node, so that the multidimensional information at the real-time process node can be predicted according to the multidimensional information of the next process node, which belongs to data prediction performed against the surgical process.
[0028] The joint training uses the predictive loss of the first LSTM network and the second LSTM network, as well as the cycle consistency loss between the second LSTM network and the first LSTM network, as the loss function of the joint training, so as to ensure that the prediction of the first LSTM network is closest to the true value and the prediction accuracy is guaranteed. Similarly, the prediction of the second LSTM network is closest to the true value and the prediction accuracy is guaranteed.
[0029] On the basis of ensuring the prediction accuracy of the second LSTM network, the cycle consistency loss can ensure that the output of the first LSTM network is used as the input of the second LSTM network for reverse prediction, so that the input of the first LSTM network can be restored, that is, the real-time multi-source information prediction value obtained by restoring the multi-source information of the next process node predicted by the first LSTM network is consistent with the true value of the real-time multi-source information, that is, the multimodal characteristics of the real-time multi-source information can predict the multi-source information of the next process node, and the predicted multi-source information of the next process node can be reversely predicted and restored to the real-time multi-source information. In summary, the mapping relationship between the real-time multi-source information and the multi-source information of the next process node is performed in both the forward and reverse orders of the surgical process. This bidirectional mapping relationship makes the prediction have the mutual constraint characteristics of bidirectional data rules, avoids the generation of prediction randomness, and thus can further ensure the prediction accuracy of the first LSTM network.
[0030] After completing the joint training of the two model structures, the present invention uses the trained first LSTM network as a surgical guidance model to extract multimodal features of multi-source information in the surgical process, thereby performing accurate real-time surgical guidance.
[0031] The extraction methods of multimodal features include: The KL divergence is used to quantify the difference between the multidimensional information of the post-process node and the multidimensional information of the pre-process node in two adjacent process nodes in the whole process cycle of the surgery, and the KL divergence value between each type of information in the multidimensional information of the post-process node and each type of information in the multidimensional information of the pre-process node is obtained; The KL divergence value between each type of information in the multidimensional information of the post-process node and each type of information in the multidimensional information of the pre-process node is used as the time scale weight of each type of information in the standard multi-source information; The time scale weights are: ; ; ; In the formula, is the time scale weight of the i-th type of information in the standard multi-source information at the t-th process node, is the i-th type of information in the standard multi-source information at the t-th process node, is the i-th type of information in the standard multi-source information at the t-1th process node, for and The KL divergence between for and The KL divergence between The present invention uses KL divergence to perform difference analysis on the time scale of multi-source information, so that at two adjacent process time nodes, if the difference of a certain category of information in the multi-source information is high, then this type of information changes, and more attention needs to be allocated on the time scale. The dynamic information contained therein increases, and the information provided for surgical status perception increases, so it is necessary to classify it with a high weight on the time scale to highlight it as an important feature on the time scale, and then the surgical guidance model can be used on the time scale to predict the surgical operation through the multimodal features on the time scale, so as to perform accurate surgical operation guidance.
[0032] After combining to form multi-source information, the present invention mines the feature importance of the multi-source information on the time scale, thereby allocating attention to each type of information in the multi-source information on the time scale, and completing feature weighted fusion based on the attention allocated on the time scale to obtain the multimodal features of the multi-source information on the time scale. The present invention also highlights the important features on the time scale in the multimodal features, so that the subsequent surgical guidance model constructed based on the multimodal features can quickly capture the important features on the time scale and allocate high temporal attention to them, so as to obtain accurate surgical operation data prediction based on the important features on the time scale, that is, to obtain multimodal features for accurately predicting surgical operation data on the time scale.
[0033] The time scale weight of the i-th type of information in the standard multi-source information at the t-th process node Weighted to the i-th category of information in the standard multi-source information at the t-th process node , get the multi-source time combination information at the tth process node , n is the total number of information categories in the standard multi-source information; Combine the multi-source time information at the tth process node , through the CNN feature extraction network, the spatial features of the multi-source time combination information at the tth process node are obtained ; The spatial characteristics of the multi-source time combination information at the tth process node After being processed by the self-attention mechanism, the multimodal features at the tth process node are obtained; The multimodal features are: ; in, ; ; ; In the formula, , and They are the query Query, key Key and value Value in the attention mechanism. for The dimension of is the intermediate value operator, is a dimension extraction operator, is the spatial feature of the i-th type of information in the multi-source time combination information at the t-th process node, is the multimodal feature of the i-th category of information in the standard multi-source information at the t-th process node, and T is the transposition operator.
[0034] The present invention obtains multi-source time combination information after time scale weighting of multi-modal features of multi-source information, and utilizes the self-attention mechanism to explore the feature importance of multi-source information on the spatial scale, thereby allocating attention to various types of information in the multi-source information on the spatial scale, and completing feature weighted fusion based on the attention allocated on the spatial scale to obtain the multi-modal features of the multi-source information on the spatial scale. It also highlights the important features on the spatial scale in the multi-modal features, so that the subsequent surgical guidance model constructed based on the multi-modal features can quickly capture the important features on the spatial scale and allocate high temporal attention to them, so as to achieve accurate surgical operation data prediction based on the important features on the spatial scale, that is, to achieve multi-modal features for accurately predicting surgical operation data on the spatial scale.
[0035] The present invention performs feature extraction on multi-source information at both the temporal scale and the spatial scale, and is capable of extracting multimodal features from multi-source information that can ensure high-precision surgical prediction at both the temporal scale and the spatial scale, thereby providing a data basis for building an accurate surgical guidance model.
[0036] The construction method of the surgical guidance model includes: The standard multi-dimensional information at each process node in the whole process cycle of the operation is serialized by process nodes to obtain a standard multi-dimensional information sequence; The multimodal features of the standard multidimensional information are used as the input of the first LSTM network, and the standard multidimensional information located at the next process node of the multimodal features is used as the output of the first LSTM network; The standard multi-dimensional information is used as the input of the second LSTM network, and the standard multi-dimensional information located at the previous process node of the standard multi-dimensional information is used as the output of the second LSTM network; The first LSTM network and the second LSTM network are trained with a combined loss of a predictive loss between an output of the first LSTM network and a true value and a multi-dimensional information cycle consistency loss between an output of the second LSTM network and an input of the first LSTM network; The first LSTM network after training is used as the surgery guidance model; Among them, the first LSTM network is: =LSTM1( ); In the formula, is the i-th type of information in the standard multi-source information at the t+1th process node, is the multimodal feature of the i-th type of information in the standard multi-source information at the t-th process node, LSTM1 is the LSTM network; The second LSTM network is: =LSTM2( ); In the formula, is the i-th type of information in the standard multi-source information at the t-th process node, is the i-th category of information in the standard multi-source information at the t+1th process node, LSTM2 is the LSTM network, and n is the total number of information categories in the standard multi-source information.
[0037] The construction method of the combined loss includes: Determine the predictive loss between the output of the first LSTM network and the second LSTM network and the true value as: ; In the formula, For predictive loss, Output of the first LSTM network , for The truth value of Output of the second LSTM network , for The truth value of is the L2 norm, is the L2 norm, m is the total number of process nodes in the whole surgical process cycle, and n is the total number of information categories in the standard multi-source information; Determine the multi-dimensional information cycle consistency loss between the output of the second LSTM network and the input of the first LSTM network as: ; In the formula, is the cycle consistency loss, is the input of the first LSTM network The corresponding , Output of the second LSTM network , is the L2 norm, m is the total number of process nodes in the whole surgical process cycle, and n is the total number of information categories in the standard multi-source information; Will and The combined loss is: ; In the formula, is the combined loss, For balance Hyperparameters of .
[0038] The standard posture information at each process node is normalized, and the standard force information of the instrument at each process node is normalized.
[0039] The network structures of the first LSTM network and the second LSTM network are the same.
[0040] After completing the extraction of multimodal features from multi-source information, the present invention constructs a surgical guidance model with a time series prediction function, which can predict the multi-source information of the next surgical process node based on the real-time multi-source information of the current surgical process node, thereby grasping in advance the robot's joint angles, position coordinates and other motion information at the next surgical process node, as well as the depth, amplitude, speed and strength of the surgical instrument's tissue traction and cutting operations. At the next surgical process node, surgical operations can be performed based on the above data, thereby realizing surgical operation guidance.
[0041] The present invention adopts two model structures for joint training when constructing the surgical guidance model. The first LSTM network is used to establish a temporal mapping relationship between the multimodal features of real-time multidimensional information and the multidimensional information of the next process node, so that the multidimensional information of the next process node can be predicted according to the multimodal features of the real-time multidimensional information, which belongs to data prediction along the surgical process. The second LSTM network is used to establish a temporal mapping relationship between the multidimensional information of the next process node and the multidimensional information of the real-time process node, so that the multidimensional information at the real-time process node can be predicted according to the multidimensional information of the next process node, which belongs to data prediction performed against the surgical process.
[0042] The joint training uses the predictive loss of the first LSTM network and the second LSTM network, as well as the cycle consistency loss between the second LSTM network and the first LSTM network, as the loss function of the joint training, so as to ensure that the prediction of the first LSTM network is closest to the true value and the prediction accuracy is guaranteed. Similarly, the prediction of the second LSTM network is closest to the true value and the prediction accuracy is guaranteed.
[0043] On the basis of ensuring the prediction accuracy of the second LSTM network, the cycle consistency loss can ensure that the output of the first LSTM network is used as the input of the second LSTM network for reverse prediction, so that the input of the first LSTM network can be restored, that is, the real-time multi-source information prediction value obtained by restoring the multi-source information of the next process node predicted by the first LSTM network is consistent with the true value of the real-time multi-source information, that is, the multimodal characteristics of the real-time multi-source information can predict the multi-source information of the next process node, and the predicted multi-source information of the next process node can be reversely predicted and restored to the real-time multi-source information. In summary, the mapping relationship between the real-time multi-source information and the multi-source information of the next process node is performed in both the forward and reverse orders of the surgical process. This bidirectional mapping relationship makes the prediction have the mutual constraint characteristics of bidirectional data rules, avoids the generation of prediction randomness, and thus can further ensure the prediction accuracy of the first LSTM network.
[0044] Hyperparameters for: ; In the formula, Output of the second LSTM network , for The truth value of .
[0045] Among them, the hyperparameter is measured by the accuracy performance of the second LSTM network. Specifically, the worse the accuracy performance of the second LSTM network is, the higher the hyperparameter is, the higher the weight of the cycle consistency loss in the combined loss is, and the higher the training resources occupied by the cycle consistency loss are. Therefore, the second LSTM network is trained intensively until the accuracy performance of the second LSTM network is improved to the optimal level.
[0046] Normalize the real-time multi-source information with the standard multi-source information at each process node.
[0047] like Figure 2 As shown, the present invention provides a surgical operation guidance system based on multimodal fusion, which is applied to a surgical operation guidance method based on multimodal fusion. The system includes: A data acquisition unit is used to record the standard posture information of the surgical robot at each process node during the entire surgical process cycle, and the standard force information of the instrument fed back by the force sensor, and to combine the standard posture information and the standard force information of the instrument into standard multi-source information for real-time surgical operations, and to obtain real-time multi-source information composed of the real-time posture information of the surgical robot and the real-time force information of the instrument fed back by the force sensor; A feature processing unit, used for extracting multimodal features from the standard multi-source information at each process node to obtain multimodal features of the standard multi-source information at each process node; The model building unit is used to perform bidirectional time series prediction training on the multimodal features of the standard multi-source information at each process node through the LSTM model to obtain a surgical guidance model; The guidance prediction unit is used to use the surgical guidance model to predict the next moment posture information of the surgical robot and the next moment force information of the instrument fed back by the force sensor based on real-time multi-source information composed of the real-time posture information of the surgical robot and the real-time force information of the instrument fed back by the force sensor.
[0048] The present invention provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, a real-time surgical operation guidance method based on multimodal feature fusion is implemented.
[0049] The present invention constructs a surgical guidance model, and realizes the prediction of the surgical robot's position information at the next moment and the instrument's force information at the next moment fed back by the force sensor based on real-time multi-source information composed of the surgical robot's real-time position information and the instrument's real-time force information fed back by the force sensor, so as to directly obtain the guidance data for subsequent operations, that is, the surgical robot's position and instrument force conditions for subsequent operations, without the need to pause during the operation, thereby ensuring the continuity of the surgical operation.
[0050] The above embodiments are only exemplary embodiments of the present application and are not intended to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art may make various modifications or equivalent substitutions to the present application within the essence and protection scope of the present application, and such modifications or equivalent substitutions shall also be deemed to fall within the protection scope of the present application.
Claims
1. A surgical operation guidance method based on multimodal fusion, characterized in that: The following steps are involved: During the entire surgical process, the standard posture information of the surgical robot at each process node and the standard force information of the instrument fed back by the force sensor will be recorded, and the standard posture information and the standard force information of the instrument will be combined into standard multi-source information for real-time surgical operations; Extracting multimodal features from the standard multi-source information at each process node to obtain multimodal features of the standard multi-source information at each process node; The multimodal features of the standard multi-source information at each process node are trained through a LSTM model for bidirectional time series prediction to obtain a surgical guidance model; The surgical guidance model is used to predict the next-moment position information of the surgical robot and the next-moment force information of the instrument fed back by the force sensor based on real-time multi-source information consisting of the real-time position information of the surgical robot and the real-time force information of the instrument fed back by the force sensor.
2. The surgical operation guidance method based on multimodal fusion according to claim 1, characterized in that: The multimodal feature extraction method includes: The KL divergence is used to quantify the difference between the multidimensional information of the post-process node and the multidimensional information of the pre-process node in two adjacent process nodes in the whole process cycle of the surgery, and the KL divergence value between each type of information in the multidimensional information of the post-process node and each type of information in the multidimensional information of the pre-process node is obtained; The KL divergence value between each type of information in the multidimensional information of the post-process node and each type of information in the multidimensional information of the pre-process node is used as the time scale weight of each type of information in the standard multi-source information; The time scale weight is: ; ; ; In the formula, is the time scale weight of the i-th type of information in the standard multi-source information at the t-th process node, is the i-th type of information in the standard multi-source information at the t-th process node, is the i-th type of information in the standard multi-source information at the t-1th process node, for and The KL divergence between for and The KL divergence between The time scale weight of the i-th type of information in the standard multi-source information at the t-th process node Weighted to the i-th category of information in the standard multi-source information at the t-th process node , get the multi-source time combination information at the tth process node , n is the total number of information categories in the standard multi-source information; Combine the multi-source time information at the tth process node , through the CNN feature extraction network, the spatial features of the multi-source time combination information at the tth process node are obtained ; The spatial characteristics of the multi-source time combination information at the tth process node After being processed by the self-attention mechanism, the multimodal features at the tth process node are obtained; The multimodal features are: ; in, ; ; ; In the formula, , and They are the query Query, key Key and value Value in the attention mechanism. for The dimension of is the intermediate value operator, is a dimension extraction operator, is the spatial feature of the i-th type of information in the multi-source time combination information at the t-th process node, is the multimodal feature of the i-th category of information in the standard multi-source information at the t-th process node, and T is the transposition operator.
3. The surgical operation guidance method based on multimodal fusion according to claim 2, characterized in that: The method for constructing the surgical guidance model includes: The standard multi-dimensional information at each process node in the whole process cycle of the operation is serialized by process nodes to obtain a standard multi-dimensional information sequence; The multimodal features of the standard multidimensional information are used as the input of the first LSTM network, and the standard multidimensional information located at the next process node of the multimodal features is used as the output of the first LSTM network; The standard multi-dimensional information is used as the input of the second LSTM network, and the standard multi-dimensional information located at the previous process node of the standard multi-dimensional information is used as the output of the second LSTM network; The first LSTM network and the second LSTM network are trained with a combined loss of a predictive loss between an output of the first LSTM network and a true value and a multi-dimensional information cycle consistency loss between an output of the second LSTM network and an input of the first LSTM network; The first LSTM network after training is used as the surgery guidance model; Wherein, the first LSTM network is: =LSTM1( ); In the formula, is the i-th type of information in the standard multi-source information at the t+1th process node, is the multimodal feature of the i-th type of information in the standard multi-source information at the t-th process node, LSTM1 is the LSTM network; The second LSTM network is: =LSTM2( ); In the formula, is the i-th type of information in the standard multi-source information at the t-th process node, is the i-th category of information in the standard multi-source information at the t+1th process node, LSTM2 is the LSTM network, and n is the total number of information categories in the standard multi-source information.
4. The surgical operation guidance method based on multimodal fusion according to claim 3, characterized in that: The method for constructing the combined loss includes: Determine the predictive loss between the output of the first LSTM network and the second LSTM network and the true value as: ; In the formula, For predictive loss, Output of the first LSTM network , for The truth value of Output of the second LSTM network , for The truth value of is the L2 norm, is the L2 norm, m is the total number of process nodes in the whole surgical process cycle, and n is the total number of information categories in the standard multi-source information; Determine the multi-dimensional information cycle consistency loss between the output of the second LSTM network and the input of the first LSTM network as: ; In the formula, is the cycle consistency loss, is the input of the first LSTM network The corresponding , Output of the second LSTM network , is the L2 norm, m is the total number of process nodes in the whole surgical process cycle, and n is the total number of information categories in the standard multi-source information; Will and The combined loss is: ; In the formula, is the combined loss, For balance Hyperparameters of .
5. The surgical operation guidance method based on multimodal fusion according to claim 4, characterized in that: The standard posture information at each process node is normalized, and the standard force information of the instrument at each process node is normalized.
6. The surgical operation guidance method based on multimodal fusion according to claim 5, characterized in that: The first LSTM network and the second LSTM network have the same network structure.
7. The surgical operation guidance method based on multimodal fusion according to claim 6, characterized in that: The hyperparameters for: ; In the formula, Output of the second LSTM network , for The truth value of .
8. The surgical operation guidance method based on multimodal fusion according to claim 7, characterized in that: Normalize the real-time multi-source information with the standard multi-source information at each process node.
9. A surgical operation guidance system based on multimodal fusion, characterized in that: A surgical operation guidance method based on multimodal fusion applied to any one of claims 1 to 8, the system comprising: A data acquisition unit is used to record the standard posture information of the surgical robot at each process node during the entire surgical process cycle, and the standard force information of the instrument fed back by the force sensor, and to combine the standard posture information and the standard force information of the instrument into standard multi-source information for real-time surgical operations, and to obtain real-time multi-source information composed of the real-time posture information of the surgical robot and the real-time force information of the instrument fed back by the force sensor; A feature processing unit, used for extracting multimodal features from the standard multi-source information at each process node to obtain multimodal features of the standard multi-source information at each process node; A model building unit, used to perform bidirectional time series prediction training on the multimodal features of the standard multi-source information at each process node through an LSTM model to obtain a surgery guidance model; The guidance prediction unit is used to use the surgical guidance model to predict the next moment's posture information of the surgical robot and the next moment's force information of the instrument fed back by the force sensor based on real-time multi-source information composed of the real-time posture information of the surgical robot and the real-time force information of the instrument fed back by the force sensor.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the method according to any one of claims 1 to 8 is implemented.