Surgical robot availability data analysis method

By using a multi-source data synchronization and multi-modal fusion model, the accuracy and real-time performance of availability assessment for surgical robot systems have been improved. This solves the problems of multi-modal data synchronization and critical event identification in existing technologies and generates a quantitative availability index system.

CN121502684APending Publication Date: 2026-02-10NAT INST FOR FOOD & DRUG CONTROL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511783924.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing surgical robot systems lack multi-source objective data support in usability assessment, making it difficult to achieve time-series synchronization of multimodal data and accurate identification of key events. This results in highly subjective assessment results that fail to reflect quantitative analysis of operational behavior and system response, and the usability index system lacks a unified definition.

Method used

By collecting multi-source data and performing timestamp calibration and synchronization, and using machine learning and multimodal fusion models for data preprocessing and feature extraction, the system can automatically detect and label the surgical process, generate usability indicators, and visualize them.

Benefits of technology

It achieves precise time alignment of multimodal data and automatic identification of key events, generates a quantitative usability index system, improves the accuracy and real-time performance of usability assessment of surgical robot systems, and reduces subjective bias.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121502684A_ABST
    Figure CN121502684A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical robots, in particular to a surgical robot availability data analysis method, which is characterized by comprising the following steps: S1, acquiring multi-source data in the front stage, the middle stage and the rear stage of a surgery; s2, timestamp calibration and synchronization are carried out on the collected multi-source data, and time sequence alignment of the multi-modal data is achieved; s3, performing preprocessing and feature extraction on the synchronized data to obtain a structured time sequence feature vector; s4, key operations are detected and marked based on the operation logs, the video frames and the robot sensor events, and time sequence segments of the operation process are divided; s5, inputting the structured feature vector into a multi-modal fusion model for joint modeling, and outputting an availability index; and S6, generating usability evaluation reports before, during, after and overall operations according to the usability indexes, and visually displaying the usability evaluation reports. According to the multi-source data fusion and intelligent analysis, objective evaluation of the operation process of the surgical robot is realized, and the system performance and the operation safety are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical robots, in particular to a surgical robot usability data analysis method. BACKGROUND

[0002] With the development of intelligent medical treatment and robot-assisted surgery technology, surgical robots show significant advantages in complex minimally invasive operations, precise positioning and intraoperative stability. However, the existing surgical robot system still has problems such as insufficient usability evaluation means, data isolation and strong subjectivity in clinical use. On the one hand, traditional usability evaluation mainly relies on postoperative questionnaires or expert observation, lacks multi-source objective data support, and is difficult to fully reflect the operation load and system interaction efficiency. On the other hand, surgical process data is usually composed of multiple modalities such as video, sensor, log, etc. The time reference is not unified between different modalities, and the data volume is huge, which leads to high difficulty in data integration and analysis.

[0003] In addition, the existing method is difficult to accurately identify and label key events in the surgical process, which limits the quantitative analysis of the operator's operation behavior, cognitive state and system response. At the same time, the usability evaluation index system lacks unified definition, cannot establish an effective correlation model between subjective experience and objective performance, and is also difficult to realize intraoperative risk warning or postoperative visual review. Therefore, there is an urgent need for an analysis method that can integrate multi-modal data, realize time synchronization, automatically label key events and generate usability indicators to improve the usability evaluation accuracy and clinical applicability of surgical robot systems. In order to solve the above problems, we propose a surgical robot usability data analysis method. SUMMARY

[0004] In view of the deficiencies of the prior art, the present application provides a surgical robot usability data analysis method to solve the problems raised in the background art.

[0005] The above technical purpose of the present application is realized by the following technical scheme: The surgical robot usability data analysis method comprises the following steps: S1, collecting multi-source data before, during and after surgery, the multi-source data being: surgery time, bleeding volume, task completion and failure events, operation error labeling, robot joint pose, robot force and torque data, system response time, system failure log, surgery operation video, surgery control software interface interaction record, doctor eye movement trajectory, intraoperative voice conversation recording, postoperative subjective questionnaire; S2, time stamp calibration and synchronization of the multi-source data collected in step S1 to correct the time drift between the multi-source collected data and realize accurate time sequence alignment of the multi-modal data; S3, pre-processing and feature extraction are performed on the multi-source data synchronized in the step S2 to obtain structured time-series feature vectors available for analysis; S4, key operations in the surgery process are automatically detected and labeled based on the surgery log, video frames and robot sensor events, event boundaries are identified through threshold determination, pattern matching and machine learning classifiers for identifying event boundaries, and the surgery process is divided into several time-series segments with time continuity; S5, the structured time-series feature vectors obtained in the step S3 are input into a multi-modal fusion model for joint modeling, the multi-modal fusion model realizes correlated learning of different modal features through feature-level splicing and time-series convolution or attention mechanism, and outputs an availability index; S6, an availability evaluation report is generated according to the availability index of the step S5, and the availability index is visualized and displayed.

[0006] Preferably, the availability index includes task completion rate TCR and unit task error rate ER, and the index is calculated according to the following formula:

[0007]

[0008] Wherein is the number of successfully completed tasks, is the number of errors occurred, is the total number of tasks counted.

[0009] Preferably, the timestamp calibration and synchronization are realized by the following steps: A unified hardware trigger pulse is added to each data acquisition device and the pulse is recorded; Initial synchronization is performed on the local clock of each acquisition device based on the network time protocol; In the post-processing stage, the time-series drift of the collected data is compensated by cross-correlation analysis combined with dynamic time warping (DTW) method to realize sub-second time alignment; Wherein the average time alignment error TAE of time alignment is calculated according to the following formula:

[0010] Wherein is the number of reference events used for evaluation of alignment, and are the timestamps of the i-th reference event under two modalities, respectively.

[0011] ​Preferably, the matching of the doctor eye movement trajectory with the surgical operation video and the surgical control software interface comprises the following steps: mapping the calibrated doctor gaze point coordinates to a predefined interface area of interest (AOI); spatially corresponding the interface area of interest to the surgical robot end effector coordinate system through coordinate transformation; calculating the gaze operation coupling degree based on the mapping relationship, wherein the gaze operation coupling degree is calculated according to the following formula: .

[0012] Preferably, the method further comprises a step of calculating a cognitive load index, which is calculated according to a linear weighting formula after standardizing the pupil diameter change Δpupil, heart rate variability (HRV), and voice stress features respectively:

[0013] wherein represents a standardization function, and the weight , , is determined by a supervised learning method in a training set to minimize the error prediction error and satisfy .

[0014] Preferably, the method further comprises a key event labeling and timing segmentation step, and the key event labeling and timing segmentation adopts a hybrid process: first, automatically detecting event boundaries based on a preset threshold or pattern matching rule; then, applying a deep learning model based on sequence labeling to classify and correct the event type; the sequence labeling model expands the recognition ability for low-frequency events by using a small amount of expert labeling through semi-supervised training.

[0015] Preferably, the multi-modal fusion model is a model based on feature-level splicing and inputting a time series convolutional neural network, which splices the time series features of each modality through the model, extracts joint time series features through several time series convolutional layers and pooling layers, and outputs the task success probability, error prediction and comprehensive usability score through a fully connected layer, and the model training adopts cross-validation and hyperparameter search to obtain the optimal parameters.

[0016] Preferably, the multi-modal fusion model is a Transformer structure based on a cross-modal attention mechanism, wherein the input of the model includes an interface area of interest embedding, an event identifier vector, and a time series feature vector of each modality; In the training process, the subjective score and the objective label are simultaneously included in the loss function to optimize the consistency of subjective and objective, wherein the training objective function is defined according to the following formula:

[0017] wherein is a task prediction loss; is a regularization term, is a subjective-objective consistency loss, wherein the subjective-objective consistency loss is calculated as follows:

[0018] wherein is the number of samples, is the subjective score of the th sample, is the corresponding objective score output by the model, and is the weight hyperparameter determined by cross-validation. Preferably, the method further comprises a step of intraoperative real-time anomaly detection and warning:

[0019] calculating the availability index and the high-risk probability in real time during surgery; generating a warning when the high-risk probability exceeds a preset threshold and automatically recording a video clip, a gaze trajectory, sensor data and event annotation associated with the warning as postoperative review materials.

[0020] In summary, the present application mainly has the following beneficial effects: 1. The present application constructs a multi-modal data analysis framework for the availability evaluation of a surgical robot system by synchronously collecting and fusing preoperative, intraoperative and postoperative multi-source data, which realizes the accurate time alignment of video, log, sensor, eye movement and voice multi-modal data, and combines automatic key event recognition, cognitive load calculation and gaze operation coupling analysis to objectively reflect the operator's operation behavior, cognitive state and system response performance, thereby forming a quantitative and interpretable availability index system.

[0021] 2. The present application realizes subjective-objective consistency modeling by a multi-modal fusion model based on feature splicing or cross-modal attention mechanism, and realizes real-time calculation and warning of the availability index and the high-risk probability during surgery, supports postoperative visual review and risk analysis, significantly improves the precision and real-time performance of the availability evaluation of the surgical robot, reduces the subjective bias of manual evaluation, and has high clinical promotion and scientific research application value. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 is a method flow chart view of the present application. DETAILED DESCRIPTION

[0023] ​In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the embodiments of the present application will be described clearly and completely below with reference to the drawings of the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the described embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without any inventive effort fall within the protection scope of the present application.

[0024] The following examples are used to illustrate the present application, but cannot be used to limit the protection scope of the present application. The conditions in the examples can be further adjusted according to specific conditions, and simple improvements of the method of the present application under the concept of the present application all fall within the protection scope of the present application.

[0025] Embodiment 1

[0026] Background: In existing surgical robot systems, surgical usability evaluation mainly relies on single-dimensional data such as task completion time, bleeding volume and physician questionnaire, and such evaluation methods have the following shortcomings: 1. Different modal data acquisition systems (such as robot sensors, video recording systems, eye trackers) have different time bases, resulting in time drift and making it impossible to accurately align the data; 2. The modal data are not structured, and it is difficult to reflect the time sequence correlation and attention changes of the physician's operation behavior during the surgery; 3. The usability index calculation relies on manual statistics, is highly subjective and lacks modeling analysis.

[0027] Therefore, it is impossible to effectively reflect the operation performance and user experience consistency of the surgical robot system in the actual surgery process.

[0028] Implementation steps: S1, collect multi-source data before, during and after surgery, including surgery time, bleeding volume, task completion and failure events, operation error annotation, robot joint pose, force and torque signal, system response time, system log, operation video, physician eye movement trajectory, voice conversation recording and postoperative subjective questionnaire, etc. A unified hardware trigger interface is set for each collection device for time sequence synchronization.

[0029] S2, time stamp calibration and synchronization of the multi-source data, to solve the time drift problem between multi-modal data, a unified trigger pulse signal is embedded at the collection end and the pulse time is recorded, and the local clock of the collection terminal is initially synchronized through the network time protocol (NTP); in the post-processing stage, the inter-correlation analysis is combined with the dynamic time warping (DTW) algorithm to calculate the time sequence difference between the modes, and the precise synchronization is realized by minimizing the alignment error, and the average time alignment error (TAE) is calculated according to the following formula:

[0030] in The number of reference events used to evaluate alignment, and The first After synchronization processing, the system achieves sub-second timing alignment of the timestamps of events in modality A and modality B, ensuring the consistency of each modality feature on the time axis.

[0031] S3. After synchronization, the data is filtered, normalized and feature extracted to obtain a structured temporal feature vector describing the operation process. The extracted features include robot end effector speed, acceleration and trajectory smoothness, force control response time, force output fluctuation rate, video frame change rate and gaze point shift frequency, as well as doctor's voice energy change and speech rate fluctuation, etc. After being encapsulated with a unified time index, a multidimensional dataset that can be input into the model is formed.

[0032] S4. Based on surgical logs and robot sensor events, the system uses threshold judgment rules to detect event boundaries and then uses a machine learning classifier to identify event types and stages. For example, when the force feedback exceeds the set threshold and is accompanied by an increase in the rate of change of joint angle, the system automatically identifies it as a "suture tension adjustment" event and finally divides the surgical process into several time-continuous stages such as positioning, cutting, and suturing, forming a standardized temporal structure.

[0033] S5. Input the multi-source temporal features extracted from the feature set into the multimodal fusion model. The model uses feature-level concatenation combined with temporal convolution and attention mechanisms to achieve interactive learning of information from different modalities, and outputs usability metrics, including task completion rate (TCR) and error rate per unit task (ER). The calculation formulas are as follows:

[0034]

[0035] in The number of tasks successfully completed. The number of errors that occurred. This represents the total number of tasks to be counted.

[0036] S6. Generate availability assessment reports for the three stages of pre-operative, intra-operative, and post-operative periods, as well as the overall availability assessment, based on availability indicators. Output time series analysis charts, task completion rate curves, and error distribution charts to assist R&D personnel in diagnosing and improving system performance.

[0037] In summary, this embodiment achieves unified modeling of multimodal information such as surgical robot operation data, eye movement trajectory, voice interaction, and system response through multi-source data synchronization and fusion analysis. It solves the problems of asynchronous multimodal data time, strong subjectivity of manual evaluation, and difficulty in quantifying usability in traditional methods. By extracting temporal features and fusion modeling, it realizes the transformation from qualitative observation to quantitative analysis, significantly improving the accuracy and objectivity of surgical robot system usability assessment, and providing technical support for robot performance optimization and physician training.

[0038] Example 2

[0039] Background: In existing studies on surgical availability assessment, physicians' cognitive load is typically assessed using subjective questionnaires or single physiological signals (such as heart rate and pupillary changes). However, these methods have significant limitations: First, a single modal signal is insufficient to fully reflect the psychological load and stress state of doctors during complex surgical procedures; Second, physiological signals are greatly affected by ambient light, noise, and individual differences, resulting in a low signal-to-noise ratio. Third, subjective scoring results often deviate from objective indicators, resulting in a lack of consistency and stability in usability analysis results.

[0040] Therefore, a quantitative model that can integrate multi-source physiological signals and behavioral characteristics is needed to achieve objective calculation of cognitive load and reduce assessment bias through subjective-objective consistency optimization model.

[0041] Implementation steps: S1. During the surgery, the doctor's eye movement data, heart rate signal and voice signal are collected in real time. The eye movement data is obtained through a high-precision eye tracker with a sampling frequency of 250Hz; the heart rate signal is collected by a chest strap sensor; and the voice signal is recorded through a far-field microphone array.

[0042] S2. Preprocess the raw signal. Use median filtering to remove transient noise from the eye-tracking data and extract pupil diameter changes. Heart rate variability (HRV) is calculated from the heart rate signal using the RR interval; speech stress features such as fundamental frequency change rate, speech rate, and energy fluctuation rate are extracted from the speech signal.

[0043] S3. Standardize the three types of features respectively, and use the Z-score transformation to unify the scale, to obtain Z( Z(HRV), Z(voiceStress); then, the Comprehensive Cognitive Load Index (CLI) is calculated using a linear weighted model, and its calculation formula is as follows:

[0044] Where w1, w2, and w3 are model weight parameters, and satisfy:

[0045] The weights are determined by using a supervised learning method on the training set to minimize the cognitive load prediction error.

[0046] S4. To address the discrepancy between subjective ratings and objective indicators, subjective assessments such as postoperative NASA-TLX scores from doctors are integrated with objective features and incorporated into the training phase of the multimodal model. This model is based on a Transformer structure with a cross-modal attention mechanism, and the task prediction loss is optimized simultaneously during training. Loss of consistency between subjective and objective factors With regularization term The overall objective function is defined as:

[0047] Among them, the loss of consistency between subjective and objective factors Calculate using the following formula:

[0048] For the sample size, For the first Expert subjective ratings for a sample, The corresponding objective score output by the model, and and The weight hyperparameters are determined by cross-validation.

[0049] S5. During training, 5-fold cross-validation and the Adam optimizer are used, with the initial learning rate set to 1×10⁻. 4 After smoothing, the cognitive load prediction results output by the model had a correlation coefficient of r=0.91 with the subjective rating, which was significantly higher than that of the traditional single-modal HRV assessment method (r=0.63).

[0050] In summary, this embodiment achieves objective calculation and subjective-objective consistency optimization of multimodal cognitive load, overcoming the shortcomings of traditional surgical usability assessment that relies on a single physiological signal or subjective questionnaire evaluation. It significantly improves the accuracy and robustness of the assessment. By integrating eye movement, heart rate, and voice stress characteristics, this embodiment achieves dynamic quantitative monitoring of physician psychological load, enabling the usability assessment system to shift from static subjective evaluation to dynamic objective calculation. This is helpful for optimizing the human-computer interaction design and operational risk warning of surgical robot systems.

[0051] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that, unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains, and the terms such as “comprising” or “including” as used in this invention mean that the element or object preceding the word covers the element or object listed after the word and its equivalents.

[0052] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for analyzing the availability data of surgical robots, characterized in that, Includes the following steps: S1. Collect multi-source data before, during and after surgery. The multi-source data includes: operation time, blood loss, task completion and failure events, operation error annotations, robot joint pose, robot force and torque data, system response time, system fault log, surgical operation video, surgical control software interface interaction record, doctor's eye movement trajectory, recording of voice conversation during surgery, and postoperative subjective questionnaire. S2. Timestamp and synchronize the multi-source data collected in step S1 to correct the time drift between the multi-source data and achieve accurate time alignment of multimodal data. S3. Preprocess and extract features from the multi-source data synchronized in step S2 to obtain a structured time-series feature vector that can be used for analysis; S4. Based on surgical logs, video frames and robot sensor events, key operations during the surgical process are automatically detected and labeled. Event boundaries are identified through threshold judgment, pattern matching and machine learning classifiers for identifying event boundaries, and the surgical process is divided into several time-series segments with temporal continuity. S5. Input the structured temporal feature vector obtained in step S3 into the multimodal fusion model for joint modeling. The multimodal fusion model realizes the association learning of different modal features through feature-level concatenation and temporal convolution or attention mechanism, and outputs usability index. S6. Generate usability assessment reports for the preoperative, intraoperative, and postoperative stages and the overall situation based on the usability indicators in step S5, and visualize the usability indicators.

2. The surgical robot availability data analysis method according to claim 1, characterized in that, The availability metrics include Task Completion Rate (TCR) and Error Rate per Unit Task (ER), and these metrics are calculated using the following formulas: ; ; in The number of tasks successfully completed. The number of errors that occurred. This represents the total number of tasks to be counted.

3. The surgical robot availability data analysis method according to claim 1, characterized in that, The timestamp calibration and synchronization are achieved through the following steps: A unified hardware trigger pulse is added to each data acquisition device and the pulse is recorded. Perform initial synchronization of the local clocks of each acquisition device based on the network time protocol; In the post-processing stage, cross-correlation analysis combined with dynamic time warping (DTW) is used to compensate for the temporal drift of the acquired data in order to achieve sub-second temporal alignment. The average time alignment error (TAE) for time alignment is calculated using the following formula: ; in The number of reference events used to evaluate alignment, and The first The timestamps of a reference event in both modalities.

4. The surgical robot availability data analysis method according to claim 1, characterized in that, The matching of the doctor's eye movement trajectory with the surgical operation video and the surgical control software interface includes the following steps: Map the calibrated doctor gaze coordinates to a predefined interface region of interest (AOI); The region of interest on the interface is spatially mapped to the coordinate system of the surgical robot's end effector through coordinate transformation. The gaze operation coupling degree is calculated based on the mapping relationship, where the gaze operation coupling degree is calculated according to the following formula: 。 5. The surgical robot availability data analysis method according to claim 1, characterized in that, The method further includes a step of calculating a cognitive load index, which is obtained by standardizing pupil diameter change Δpupil, heart rate variability (HRV), and speech stress features, and then calculating them using a linear weighted formula. ; in Represents the standardization function, weights , , The method is determined in the training set by supervised learning to minimize the error in misprediction and satisfy the following conditions: .

6. The surgical robot availability data analysis method according to claim 1, characterized in that, The method further includes key event labeling and time series segmentation steps, wherein the key event labeling and time series segmentation adopt a hybrid process: First, event boundaries are automatically detected based on preset thresholds or pattern matching rules; Subsequently, a deep learning model based on sequence labeling was applied to classify and correct the event types; Sequence labeling models extend their ability to identify low-frequency events by using a small number of expert annotations through semi-supervised training.

7. The surgical robot availability data analysis method according to claim 1, characterized in that, The multimodal fusion model is based on feature-level concatenation and input into a temporal convolutional neural network. The model concatenates the temporal features of each modality and extracts joint temporal features through several temporal convolutional and pooling layers. It then outputs the task success probability, error prediction, and comprehensive usability score through a fully connected layer. The model training adopts cross-validation and hyperparameter search to obtain the optimal parameters.

8. The surgical robot availability data analysis method according to claim 1, characterized in that, The multimodal fusion model is a Transformer structure based on a cross-modal attention mechanism, wherein the model input includes the embedding of the interface region of interest, the event identifier vector, and the temporal feature vector of each modality; During training, both subjective ratings and objective labels are incorporated into the loss function to optimize the consistency between subjective and objective ratings. The training objective function is defined as follows: ; in Predict losses for the task; For regularization terms, The loss of consistency between subjective and objective elements is calculated using the following formula: ; in For the sample size, For the first Expert subjective ratings for a sample, The corresponding objective score output by the model, and and The weight hyperparameters are determined by cross-validation.

9. The surgical robot availability data analysis method according to claim 1, characterized in that, The method further includes a real-time anomaly detection and alarm step during surgery: Real-time calculation of availability metrics and high-risk probabilities during surgery; When the probability of high risk exceeds a preset threshold, an alarm is generated and the video clips, gaze trajectory, sensor data and event annotations associated with the alarm are automatically recorded as postoperative review materials.