Operation multi-dimensional sensing method and system based on multi-source information and storage medium

By combining multi-source information and self-attention mechanism, multi-dimensional spatio-temporal characteristics during the surgery are extracted and multi-dimensional surgical perception model is constructed, which solves the problem of insufficient surgical status perception efficiency and objectivity in the existing technology, and achieves efficient and accurate surgical status evaluation.

CN120089289AInactive Publication Date: 2025-06-03XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510579890.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art relies on artificial empirical surgical state perception methods, which have defects in efficiency and objectivity, making it difficult to achieve efficient objectivity state evaluation.

Method used

Using a multi-dimensional surgical perception method based on multi-source information, multi-dimensional spatio-temporal characteristics are extracted to construct a multi-dimensional surgical multi-dimensional perception model by recording the motion information of the surgical robot, the visual information feedback from the camera and the instrumental force information feedback from the force sensor.

Benefits of technology

It realizes efficient and objective assessment of surgical status, improves the efficiency and accuracy of multi-dimensional perception models for surgical status recognition, and reduces the risk of human subjective mistakes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089289A_ABST
    Figure CN120089289A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of operation process evaluation, in particular to an operation multi-dimensional sensing method and system based on multi-source information and a storage medium, and the method comprises the following steps: extracting the time difference of the multi-source information in the whole operation process period, and determining the time scale weight of the combined multi-source information; weighting the time scale weight to the multi-source information to obtain multi-source time combination information; extracting spatial characteristics of the multi-source time combination information through a self-attention mechanism to obtain multi-dimensional spatial-temporal characteristics which are combined and extracted from the multi-source information and are used for representing an operation state; and constructing a surgical multi-dimensional perceptual model based on the multi-dimensional spatial-temporal characteristics. According to the method, the multi-source information used for sensing the surgical state is combined, the multi-dimensional spatial-temporal characteristics representing the surgical state are extracted from the multi-source information, and the most effective spatial-temporal characteristics are reserved for establishing the multi-dimensional sensing model, so that the efficiency and accuracy of the multi-dimensional sensing model for recognizing the surgical state can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of surgical process evaluation, and specifically relates to a multi-dimensional perception method, system and storage medium for surgery based on multi-source information. Background Art

[0002] During the evaluation of the surgical situation in the surgical process, usually, by collecting data related to the surgery, such as medical image data, surgical record data, etc., experienced experts are organized to conduct a manual review and evaluation of the data related to the surgery to determine the intraoperative situation in the surgical process. In this way, relatively high professional requirements are imposed on the personnel evaluating the surgical state. When the amount of evaluation data is large, the human burden is too heavy, and it is easy to cause subjective human errors under pressure.

[0003] Therefore, the existing method of surgical state perception relying on human experience has defects in terms of efficiency and objectivity, and it is difficult to achieve high-efficiency objective state evaluation, which affects the effect of surgical state perception. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-dimensional perception method, system and storage medium for surgery based on multi-source information, so as to solve the technical problem that the existing method of surgical state perception relying on human experience has defects in terms of efficiency and objectivity, and it is difficult to achieve high-efficiency objective state evaluation.

[0005] To solve the above technical problem, the present invention specifically provides the following technical solutions: A multi-dimensional perception method for surgery based on multi-source information, comprising the following steps: During the entire surgical process cycle, record the motion information of the surgical robot, the visual information feedback by the camera, and the instrument force information feedback by the force sensor at each process node, and combine the motion information, visual information and instrument force information into multi-source information for perceiving the surgical state; Extract the time difference of the multi-source information during the entire surgical process cycle, and determine the time-scale weight of the combined multi-source information; Weight the time-scale weight to the multi-source information to obtain multi-source time combined information; perform spatial feature extraction on the multi-source time combined information through the self-attention mechanism to obtain multi-dimensional spatio-temporal features for combining and extracting to represent the surgical state from the multi-source information; Based on the multi-dimensional spatio-temporal features, construct a surgical multi-dimensional perception model for evaluating the surgical state of each process node during the entire surgical process cycle.

[0006] As a preferred solution of the present invention, the method for determining the time-scale weight includes: Quantify the difference between the multi-dimensional information of the subsequent process node and the multi-dimensional information of the previous process node in two adjacent process nodes during the entire surgical process using the KL divergence, and obtain the KL divergence values between various types of information in the multi-dimensional information of the subsequent process node and various types of information in the multi-dimensional information of the previous process node; Take the KL divergence value between various types of information in the multi-dimensional information of the subsequent process node and various types of information in the multi-dimensional information of the previous process node as the time-scale weight of various types of information in the multi-source information; The time-scale weight is: ; ; ; In the formula, is the time-scale weight of the i-th type of information in the multi-source information at the t-th process node, is the i-th type of information in the multi-source information at the t-th process node, is the i-th type of information in the multi-source information at the (t - 1)-th process node, is and the KL divergence between, is and the KL divergence between.

[0007] As a preferred solution of the present invention, the construction method of the multi-source time combination information includes: Weight the time-scale weight of the i-th type of information in the multi-source information at the t-th process node to the i-th type of information in the multi-source information at the t-th process node , to obtain the multi-source time combination information at the t-th process node, where n is the total number of information categories in the multi-source information.

[0008] As a preferred solution of the present invention, the extraction method of the multi-dimensional spatio-temporal features includes: Input the multi-source time combination information at the t-th process node into a CNN feature extraction network to obtain the spatial features of the multi-source time combination information at the t-th process node; Process the spatial features of the multi-source time combination information at the t-th process node through a self-attention mechanism to obtain the multi-dimensional spatio-temporal features at the t-th process node; The multi-dimensional spatio-temporal features are: ; where, ; ; ; In the formula, , and are the Query, Key, and Value in the attention mechanism respectively, is 's dimension, is the operator for taking the middle value, is the dimension extraction operator, is the spatial feature of the i-th type of information in the multi-source time combination information at the t-th process node, is the spatio-temporal feature of the i-th type of information in the multi-source information at the t-th process node, and T is the transpose operator.

[0009] As a preferred solution of the present invention, the construction method of the surgical multi-dimensional perception model: Take the multi-dimensional spatio-temporal features of the multi-source information as the input quantity of the classifier, and take the surgical state corresponding to the multi-source information as the output quantity of the classifier; Train the classifier to obtain the surgical multi-dimensional perception model as: ; In the formula, is the surgical state at the t-th process node, is the spatio-temporal feature of the i-th type of information in the multi-source information at the t-th process node, and SVM is the classifier.

[0010] As a preferred solution of the present invention, the present invention provides a surgical multi-dimensional perception system based on multi-source information, which is applied to a surgical multi-dimensional perception method based on multi-source information. The system includes: A data acquisition unit, which is used to record the motion information of the surgical robot, the visual information feedback by the camera, and the instrument force information feedback by the force sensor at each process node during the entire surgical process cycle, and combine the motion information, visual information, and instrument force information into multi-source information for perceiving the surgical state; A feature extraction unit, which is used to extract the time difference of the multi-source information during the entire surgical process cycle, determine the time-scale weight of the combined multi-source information; weight the time-scale weight to the multi-source information to obtain multi-source time combination information; extract the spatial characteristics of the multi-source time combination information through the self-attention mechanism to obtain multi-dimensional spatio-temporal features for combining and extracting the surgical state in the multi-source information; A model construction unit, which is used to construct a surgical multi-dimensional perception model for evaluating the surgical state of each process node during the entire surgical process cycle based on the multi-dimensional spatio-temporal features.

[0011] As a preferred embodiment of the present invention, the method for the feature processing unit to determine the time-scale weight includes: Quantify the difference between the multi-dimensional information of the subsequent process node and the multi-dimensional information of the previous process node in two adjacent process nodes during the entire surgical process cycle by using KL divergence, and obtain the KL divergence values between various types of information in the multi-dimensional information of the subsequent process node and various types of information in the multi-dimensional information of the previous process node; Take the KL divergence values between various types of information in the multi-dimensional information of the subsequent process node and various types of information in the multi-dimensional information of the previous process node as the time-scale weights of various types of information in the multi-source information; The time-scale weight is: ; ; ; In the formula, is the time-scale weight of the i-th type of information in the multi-source information at the t-th process node, is the i-th type of information in the multi-source information at the t-th process node, is the i-th type of information in the multi-source information at the (t - 1)-th process node, is and the KL divergence between, is and the KL divergence between.

[0012] As a preferred embodiment of the present invention, the method for the feature extraction unit to extract multi-dimensional spatio-temporal features includes: Input the multi-source time combination information at the t-th process node into the CNN feature extraction network to obtain the spatial features of the multi-source time combination information at the t-th process node; Process the spatial features of the multi-source time combination information at the t-th process node through the self-attention mechanism to obtain the multi-dimensional spatio-temporal features at the t-th process node; The multi-dimensional spatio-temporal features are: ; Among them, ; ; ; In the formula, , and are the query Query, key Key, and value Value in the attention mechanism respectively, is the dimension of, is the median operator, is the dimension extraction operator, is the spatial feature of the i-th type of information in the multi-source time combination information at the t-th process node, is the spatio-temporal feature of the i-th type of information in the multi-source information at the t-th process node, and T is the transpose operator.

[0013] As a preferred embodiment of the present invention, the method for constructing the surgical multi-dimensional perception model by the model construction unit is as follows: Take the multi-dimensional spatio-temporal features of the multi-source information as the input quantity of the classifier, and take the surgical state corresponding to the multi-source information as the output quantity of the classifier; Train the classifier to obtain the surgical multi-dimensional perception model as: ; In the formula, is the surgical state at the t-th process node, is the spatio-temporal feature of the i-th type of information in the multi-source information at the t-th process node, and SVM is the classifier.

[0014] As a preferred embodiment of the present invention, the present invention provides a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the surgical multi-dimensional perception method based on multi-source information is implemented.

[0015] The present invention has the following beneficial effects compared with the prior art: The present invention combines multi-source information for perceiving the surgical state, and extracts multi-dimensional spatio-temporal features representing the surgical state from the multi-source information, realizing double extraction of the features related to the surgical state in the multi-source information in terms of time and space, and retaining the most effective spatio-temporal features for establishing a multi-dimensional perception model, thereby being able to improve the efficiency and accuracy of the multi-dimensional perception model in identifying the surgical state. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained according to the provided drawings.

[0017] Figure 1 is the flowchart of the surgical multi-dimensional perception method based on multi-source information provided by the embodiments of the present invention; Figure 2The surgical multi-dimensional perception system based on multi-source information provided by the embodiments of the present invention. Detailed implementation manners

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] As Figure 1 shown, the present invention provides a surgical multi-dimensional perception method based on multi-source information, including the following steps: During the entire surgical process cycle, record the motion information of the surgical robot at each process node, the visual information feedback by the camera, and the instrument force information feedback by the force sensor, and combine the motion information, visual information, and instrument force information into multi-source information for perceiving the surgical state; Extract the time difference of the multi-source information during the entire surgical process cycle, and determine the time-scale weight of the combined multi-source information; Weight the time-scale weight to the multi-source information to obtain multi-source time-combined information; perform spatial feature extraction on the multi-source time-combined information through the self-attention mechanism to obtain multi-dimensional spatio-temporal features for combining and extracting to represent the surgical state from the multi-source information; Based on the multi-dimensional spatio-temporal features, construct a surgical multi-dimensional perception model for evaluating the surgical state of each process node during the entire surgical process cycle.

[0020] The present invention first aggregates all information categories related to evaluating the surgical state, including the motion information of the surgical robot, which can also be called the pose information of the surgical robot, the visual information feedback by the camera, which is the surgical image information captured by the camera, and the instrument force information feedback by the force sensor, which is the information such as the depth, amplitude, speed, and strength of the pulling and cutting operations of the surgical instrument on the tissue. Combine these information into multi-source information for surgical state perception, obtain information for perceiving the surgical state from multiple aspects, and the information comprehensiveness is higher, so as to ensure higher accuracy of surgical state perception.

[0021] After the present invention combines to form multi-source information, it mines the feature importance of the multi-source information on the time scale, so as to allocate attention to various information in the multi-source information on the time scale, highlighting the important features of the multi-source information on the time scale, that is, the subsequent constructed surgical multi-dimensional perception model can quickly capture the important features on the time scale, allocate high time attention to them, and achieve accurate surgical state perception based on the important features on the time scale, that is, achieve accurate surgical state perception on the time scale.

[0022] Specifically, the present invention uses KL divergence to perform difference analysis on the time scale of multi-source information, so that at two adjacent process time nodes, if the difference of a certain category of information in the multi-source information is high, then this type of information changes, and more attention needs to be allocated on the time scale. The dynamic information contained therein increases, and the information provided for surgical status perception increases, so it is necessary to classify it with a high weight on the time scale to highlight it as an important feature on the time scale, and then the surgical multidimensional perception model can be used on the time scale to perceive the surgical status of the important features on the time scale, so as to obtain accurate surgical status perception.

[0023] The present invention obtains multi-source time combination information after weighting the time scale of multi-source information, and utilizes the self-attention mechanism to explore the feature importance of multi-source information on the spatial scale, thereby allocating attention to various types of information in the multi-source information on the spatial scale, and highlighting the important features of the multi-source information on the spatial scale. That is, the subsequently constructed surgical multidimensional perception model can quickly capture the important features on the spatial scale and allocate high spatial attention to them, so as to obtain accurate surgical status perception on the spatial scale based on the important features on the spatial scale, that is, to obtain accurate surgical status perception on the spatial scale.

[0024] Among them, the Self-Attention Mechanism, also known as the Intra-Attention Mechanism, is a special attention mechanism that allows the model to focus on the relationship between elements within the sequence when processing sequence data, thereby capturing the complex dependencies within the sequence and obtaining feature relationships on a spatial scale.

[0025] In summary, the present invention performs feature extraction on multi-source information at both the temporal scale and the spatial scale, and is able to extract features from multi-source information that can ensure high-precision surgical perception at both the temporal scale and the spatial scale, thereby providing a data basis for building an accurate surgical perception model.

[0026] After obtaining multidimensional spatiotemporal features that combine the importance of time scale and space scale, the present invention constructs a multidimensional surgical perception model using the multidimensional spatiotemporal features as input, thereby achieving high-precision surgical status perception in multi-source information and realizing model-automated surgical status perception. Compared with manual evaluation, both efficiency and objectivity are effectively improved.

[0027] The present invention performs differential analysis on multi-source information on a time scale through KL divergence. If the difference in a certain type of information among multi-source information is high between two adjacent process time nodes, it indicates that this type of information has changed, and more attention needs to be allocated on the time scale. The dynamic information contained therein increases, and the information provided for surgical state perception increases. Therefore, a high weight needs to be assigned to it on the time scale to highlight its important features on the time scale. Furthermore, the surgical multi-dimensional perception model can be used on the time scale to perceive the surgical state based on the important features on the time scale, so as to obtain accurate surgical state perception, specifically as follows: The method for determining the time scale weight includes: Quantify the difference between the multi-dimensional information of the post-process node and the multi-dimensional information of the pre-process node in two adjacent process nodes during the entire surgical process cycle by using KL divergence, and obtain the KL divergence values between various types of information in the multi-dimensional information of the post-process node and various types of information in the multi-dimensional information of the pre-process node; Take the KL divergence values between various types of information in the multi-dimensional information of the post-process node and various types of information in the multi-dimensional information of the pre-process node as the time scale weights of various types of information in the multi-source information; The time scale weight is: ; ; ; In the formula, is the time scale weight of the i-th type of information in the multi-source information at the t-th process node, is the i-th type of information in the multi-source information at the t-th process node, is the i-th type of information in the multi-source information at the t - 1-th process node, is and the KL divergence between, is and the KL divergence between.

[0028] When the present invention quantifies the difference between the multi-dimensional information of two adjacent process nodes, KL divergence is used for measurement. The larger the KL divergence value, the greater the time scale weight. To eliminate the randomness of difference quantification, KL divergence is calculated once from the post-process node to the pre-process node direction and once from the pre-process node to the post-process node direction, and the results of the two KL divergences are balanced, so as to obtain an accurate difference quantification result and improve the extraction accuracy of important features on the time scale.

[0029] After the multi-source information is combined, the present invention mines the feature importance of the multi-source information on the time scale, thereby allocating attention to various types of information in the multi-source information on the time scale, highlighting the important features of the multi-source information on the time scale, that is, the subsequently constructed surgical multi-dimensional perception model can quickly capture the important features on the time scale, allocate high time attention to them, and achieve accurate surgical state perception based on the important features on the time scale, that is, achieve accurate surgical state perception on the time scale.

[0030] The construction method of the multi-source time combination information includes: Weight the time scale weight of the i-th type of information in the multi-source information at the t-th process node to the i-th type of information in the multi-source information at the t-th process node to obtain the multi-source time combination information at the t-th process node , where n is the total number of information categories in the multi-source information.

[0031] After the multi-source time combination information obtained by the time scale weighting of the multi-source information, the present invention uses the self-attention mechanism to mine the feature importance of the multi-source information on the spatial scale, thereby allocating attention to various types of information in the multi-source information on the spatial scale, highlighting the important features of the multi-source information on the spatial scale, that is, the subsequently constructed surgical multi-dimensional perception model can quickly capture the important features on the spatial scale, allocate high spatial attention to them, and achieve accurate surgical state perception based on the important features on the spatial scale, that is, achieve accurate surgical state perception on the spatial scale, specifically as follows: The extraction method of the multi-dimensional spatio-temporal features includes: The multi-source time combination information at the t-th process node is input into the CNN feature extraction network to obtain the spatial features of the multi-source time combination information at the t-th process node ; The spatial features of the multi-source time combination information at the t-th process node are processed by the self-attention mechanism to obtain the multi-dimensional spatio-temporal features at the t-th process node; The multi-dimensional spatio-temporal features are: ; where ; ; ; In the formula, , and are the query Query, key Key, and value Value in the attention mechanism respectively, is the dimension of, is the median operator, is the dimension extraction operator, is the spatial feature of the i-th type of information in the multi-source time combination information at the t-th process node, is the spatio-temporal feature of the i-th type of information in the multi-source information at the t-th process node, and T is the transpose operator.

[0032] After the present invention obtains multi-dimensional spatio-temporal features that take into account the importance of both time scale and space scale, a surgical multi-dimensional perception model is constructed with the multi-dimensional spatio-temporal features as the input, realizing high-precision surgical state perception in multi-source information and realizing automatic surgical state perception of the model. Compared with manual evaluation, both efficiency and objectivity are effectively improved, as follows: Method for constructing a surgical multi-dimensional perception model: Taking the multi-dimensional spatio-temporal features of multi-source information as the input quantity of the classifier, and taking the surgical state corresponding to the multi-source information as the output quantity of the classifier; Training the classifier to obtain the surgical multi-dimensional perception model as: ; In the formula, is the surgical state at the t-th process node, is the spatio-temporal feature of the i-th type of information in the multi-source information at the t-th process node, and SVM is the classifier.

[0033] As Figure 2 shown, the present invention provides a surgical multi-dimensional perception system based on multi-source information, which is applied to a surgical multi-dimensional perception method based on multi-source information. The system includes: A data acquisition unit, configured to record the motion information of the surgical robot, the visual information fed back by the camera, and the instrument force information fed back by the force sensor at each process node during the entire surgical process cycle, and combine the motion information, visual information, and instrument force information into multi-source information for perceiving the surgical state; A feature extraction unit, configured to extract the time difference of multi-source information during the entire surgical process cycle, determine the time scale weight of the combined multi-source information; weight the time scale weight to the multi-source information to obtain multi-source time combination information; extract the spatial characteristics of the multi-source time combination information through the self-attention mechanism to obtain multi-dimensional spatio-temporal features for combining and extracting the surgical state in multi-source information; A model construction unit, configured to construct a surgical multi-dimensional perception model for evaluating the surgical state of each process node during the entire surgical process cycle based on the multi-dimensional spatio-temporal features.

[0034] The method for the feature processing unit to determine the time scale weight includes: The KL divergence is used to quantify the difference between the multi-dimensional information of the subsequent process node and the multi-dimensional information of the previous process node in two adjacent process nodes during the entire surgical process, and the KL divergence value between various types of information in the multi-dimensional information of the subsequent process node and various types of information in the multi-dimensional information of the previous process node is obtained; The KL divergence value between various types of information in the multi-dimensional information of the subsequent process node and various types of information in the multi-dimensional information of the previous process node is used as the time-scale weight of various types of information in the multi-source information; The time-scale weight is: ; ; ; In the formula, is the time-scale weight of the i-th type of information in the multi-source information at the t-th process node, is the i-th type of information in the multi-source information at the t-th process node, is the i-th type of information in the multi-source information at the (t - 1)-th process node, is and the KL divergence between, is and the KL divergence between.

[0035] The extraction method of the feature extraction unit for the multi-dimensional spatio-temporal features includes: The multi-source time-combined information at the t-th process node is passed through the CNN feature extraction network to obtain the spatial features of the multi-source time-combined information at the t-th process node; The spatial features of the multi-source time-combined information at the t-th process node are processed through the self-attention mechanism to obtain the multi-dimensional spatio-temporal features at the t-th process node; The multi-dimensional spatio-temporal features are: ; Among them, ; ; ; In the formula, , and are the query Query, key Key, and value Value in the attention mechanism respectively, is the dimension of, is the operator for taking the intermediate value, is a dimensionality extraction operator, is the spatial feature of the i-th type of information in the multi-source time combination information at the t-th process node, is the spatio-temporal feature of the i-th type of information in the multi-source information at the t-th process node, and T is the transpose operator.

[0036] The method for constructing a surgical multi-dimensional perception model by the model construction unit: Use the multi-dimensional spatio-temporal features of multi-source information as the input quantity of the classifier, and use the surgical state corresponding to the multi-source information as the output quantity of the classifier; Train the classifier to obtain the surgical multi-dimensional perception model as: ; In the formula, is the surgical state at the t-th process node, is the spatio-temporal feature of the i-th type of information in the multi-source information at the t-th process node, and SVM is the classifier.

[0037] The present invention provides a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, a surgical multi-dimensional perception method based on multi-source information is implemented.

[0038] The present invention combines multi-source information for perceiving the surgical state, extracts multi-dimensional spatio-temporal features representing the surgical state from the multi-source information, realizes double extraction of the features related to the surgical state in the multi-source information in terms of time and space, and retains the most effective spatio-temporal features for establishing a multi-dimensional perception model, thereby improving the efficiency and accuracy of the multi-dimensional perception model for recognizing the surgical state.

[0039] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements within the essence and protection scope of the present application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present application.

Claims

1. A surgical multi-dimensional perception method based on multi-source information, characterized in that: The following steps are involved: During the whole surgical process, the motion information of the surgical robot at each process node, the visual information fed back by the camera, and the instrument force information fed back by the force sensor are recorded, and the motion information, visual information and instrument force information are combined into multi-source information for sensing the surgical status; Extract the time differences of multi-source information in the whole surgical process cycle and determine the time scale weight of combining multi-source information; The time scale weight is weighted to the multi-source information to obtain the multi-source time combination information; the multi-source time combination information is spatially characterized by using the self-attention mechanism to obtain a multi-dimensional spatiotemporal feature for combining and extracting the multi-source information to characterize the surgical status; Based on multidimensional spatiotemporal characteristics, a surgical multidimensional perception model is constructed to evaluate the surgical status of each process node in the entire surgical process cycle.

2. The method for multi-dimensional surgical perception based on multi-source information according to claim 1, characterized in that: The method for determining the time scale weight includes: The KL divergence is used to quantify the difference between the multidimensional information of the post-process node and the multidimensional information of the pre-process node in two adjacent process nodes in the whole process cycle of the surgery, and the KL divergence value between each type of information in the multidimensional information of the post-process node and each type of information in the multidimensional information of the pre-process node is obtained; The KL divergence value between each type of information in the multidimensional information of the post-process node and each type of information in the multidimensional information of the pre-process node is used as the time scale weight of each type of information in the multi-source information; The time scale weight is: ; ; ; In the formula, is the time scale weight of the i-th type of information in the multi-source information at the t-th process node, is the i-th type of information in the multi-source information at the t-th process node, is the i-th type of information in the multi-source information at the t-1th process node, for and The KL divergence between for and The KL divergence between .

3. The method for multi-dimensional surgical perception based on multi-source information according to claim 2, characterized in that: The method for constructing the multi-source time combination information includes: The time scale weight of the i-th type of information in the multi-source information at the t-th process node Weighted to the i-th category of information in the multi-source information at the t-th process node , get the multi-source time combination information at the tth process node , n is the total number of information categories in multi-source information.

4. The method for multi-dimensional surgical perception based on multi-source information according to claim 3, characterized in that: The method for extracting the multidimensional spatiotemporal features comprises: Combine the multi-source time information at the tth process node , through the CNN feature extraction network, the spatial features of the multi-source time combination information at the tth process node are obtained ; The spatial characteristics of the multi-source time combination information at the tth process node After being processed by the self-attention mechanism, the multi-dimensional spatiotemporal features at the t-th process node are obtained; The multidimensional spatiotemporal features are: ; in, ; ; ; In the formula, , and They are the query Query, key Key and value Value in the attention mechanism. for The dimension of is the intermediate value operator, is a dimension extraction operator, is the spatial feature of the i-th type of information in the multi-source time combination information at the t-th process node, is the spatiotemporal characteristics of the i-th type of information in the multi-source information at the t-th process node, and T is the transposition operator.

5. The method for multi-dimensional surgical perception based on multi-source information according to claim 4, characterized in that: Method for constructing the multidimensional surgical perception model: Using the multi-dimensional spatiotemporal features of the multi-source information as the input of the classifier, and using the surgical status corresponding to the multi-source information as the output of the classifier; Train the classifier and obtain the multi-dimensional perception model of surgery: ; In the formula, is the surgical status at the tth process node, is the spatiotemporal characteristics of the i-th type of information in the multi-source information at the t-th process node, and SVM is the classifier.

6. A surgical multi-dimensional perception system based on multi-source information, characterized in that: A multi-dimensional surgical perception method based on multi-source information as described in any one of claims 1 to 5, the system comprising: A data acquisition unit is used to record the motion information of the surgical robot at each process node, the visual information fed back by the camera, and the instrument force information fed back by the force sensor during the entire surgical process cycle, and to combine the motion information, visual information and instrument force information into multi-source information for sensing the surgical status; The feature extraction unit is used to extract the time difference of multi-source information in the whole process cycle of the operation, determine the time scale weight of the combined multi-source information; weight the time scale weight to the multi-source information to obtain the multi-source time combination information; extract the spatial characteristics of the multi-source time combination information through the self-attention mechanism, and obtain the multi-dimensional spatiotemporal features used to combine and extract from the multi-source information to characterize the operation status; The model building unit is used to build a surgical multidimensional perception model for evaluating the surgical status of each process node in the entire surgical process cycle based on multidimensional spatiotemporal characteristics.

7. The multi-dimensional surgical perception system based on multi-source information according to claim 6, characterized in that: The method for determining the time scale weight of the feature processing unit includes: The KL divergence is used to quantify the difference between the multidimensional information of the post-process node and the multidimensional information of the pre-process node in two adjacent process nodes in the whole process cycle of the surgery, and the KL divergence value between each type of information in the multidimensional information of the post-process node and each type of information in the multidimensional information of the pre-process node is obtained; The KL divergence value between each type of information in the multidimensional information of the post-process node and each type of information in the multidimensional information of the pre-process node is used as the time scale weight of each type of information in the multi-source information; The time scale weight is: ; ; ; In the formula, is the time scale weight of the i-th type of information in the multi-source information at the t-th process node, is the i-th type of information in the multi-source information at the t-th process node, is the i-th type of information in the multi-source information at the t-1th process node, for and The KL divergence between for and The KL divergence between .

8. The multi-dimensional surgical perception system based on multi-source information according to claim 7, characterized in that: The method for extracting multidimensional spatiotemporal features by the feature extraction unit includes: Combine the multi-source time information at the tth process node , through the CNN feature extraction network, the spatial features of the multi-source time combination information at the tth process node are obtained ; The spatial characteristics of the multi-source time combination information at the tth process node After being processed by the self-attention mechanism, the multi-dimensional spatiotemporal features at the t-th process node are obtained; The multidimensional spatiotemporal features are: ; in, ; ; ; In the formula, , and They are the query Query, key Key and value Value in the attention mechanism. for The dimension of is the intermediate value operator, is a dimension extraction operator, is the spatial feature of the i-th type of information in the multi-source time combination information at the t-th process node, is the spatiotemporal characteristics of the i-th type of information in the multi-source information at the t-th process node, and T is the transposition operator.

9. The multi-dimensional surgical perception system based on multi-source information according to claim 8, characterized in that: The model building unit constructs a multi-dimensional surgical perception model: Using the multi-dimensional spatiotemporal features of the multi-source information as the input of the classifier, and using the surgical status corresponding to the multi-source information as the output of the classifier; Train the classifier and obtain the multi-dimensional perception model of surgery: ; In the formula, is the surgical status at the tth process node, is the spatiotemporal characteristics of the i-th type of information in the multi-source information at the t-th process node, and SVM is the classifier.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the method according to any one of claims 1 to 5 is implemented.