Training methods for event monitoring models and anomaly detection methods for surgical videos
By building an event monitoring model, using the source domain and target domain sample video sets for feature enhancement and pseudo-label training, the problems of cross-domain universality and multi-center equipment differences in minimally invasive surgery are solved, and the accuracy and robustness of abnormal detection in complex surgical scenarios are improved.
Patent Information
- Application Number
- CN202510748813.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The prior art fails to fully cover potential abnormalities in minimally invasive surgery, lacks cross-domain universality, and it is difficult to effectively monitor abnormalities in complex surgical scenarios under multi-center equipment and environmental differences.
By building an event monitoring model, the feature extraction module, event memory unit and source domain classifier are trained using the source domain sample video set, and feature enhancement is performed in combination with the target domain sample video, and iteratively adjusts the event and scene feature prototypes to achieve cross-domain decoupling and pseudo-label fine training, and improves abnormal detection capabilities.
It improves the generalization ability and accuracy of abnormal detection of surgical videos, ensures consistent abnormal performance in different surgical fields, enhances the ability and robustness of subtle abnormalities, and improves the overall recognition ability and accuracy of handling various abnormal situations.
Smart Images

Figure CN120279469B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video processing technology, and more particularly to a training method for an event monitoring model and an anomaly detection method for surgical videos. Background Art
[0002] With the development of minimally invasive surgical techniques and surgical robotic systems, one of the main challenges is coping with unforeseen scenarios and unexpected events in unstructured, deformable surgical environments. The sensors used in minimally invasive laparoscopic surgical robots can acquire critical surgical information and infer the actual status of the surgical process, potential adverse events, and their inducing factors based on this information. This helps surgeons perform more complex surgeries, thereby improving patient safety.
[0003] However, current studies only monitor a single type of adverse event, failing to fully cover potential intraoperative abnormalities and lacking comprehensiveness. Furthermore, existing technologies fail to account for differences in equipment and environments across multiple centers, resulting in a lack of universal applicability. Summary of the Invention
[0004] In view of this, the present invention provides a training method for an event monitoring model and an anomaly detection method for surgical videos.
[0005] One aspect of the present invention provides a training method for an event monitoring model, comprising: using a source domain sample video set to train a feature extraction module, an event memory unit, and a source domain classifier to obtain a pre-trained model; inputting a target domain sample video into the feature extraction module of the pre-trained model for feature extraction, and outputting target domain sample features; inputting target domain sample features and source domain sample features into a scene memory unit and an event memory unit of the pre-trained model for feature enhancement, and outputting mixed domain event enhancement features and scene enhancement features respectively, wherein the source domain sample features are obtained by using the feature extraction module of the pre-trained model to extract the source domain sample video from the source domain sample video set. The method is obtained after feature extraction of a domain sample video; the mixed domain event enhancement feature, the target domain sample feature and the source domain sample feature are input into the target domain classifier, and the target domain prediction result for the target domain sample video is output; according to the difference between the target domain prediction result and the pseudo label for the target domain sample video, the event feature prototype in the event memory unit of the pre-trained model, the scene feature prototype in the scene memory unit and the network parameters of the target domain classifier are iteratively adjusted until the first iteration stopping condition is met, and a trained event monitoring model is obtained, wherein the pseudo label is the prediction result output after the source domain sample video is processed by the pre-trained model.
[0006] The second aspect of the present invention provides a method for detecting anomalies in surgical videos, comprising: acquiring a laparoscopic surgical video in real time, wherein the laparoscopic surgical video includes M video frames, M ≥ 1, and M is a positive integer; processing the M video frames using a feature extraction module, an event memory unit, and a target domain classifier in a trained event monitoring model to obtain a predicted value for each video frame, wherein the predicted value represents the probability that the video frame is an abnormal situation, and the event monitoring model is trained using the above-mentioned training method; and determining a video frame whose predicted value is greater than a preset value as a video frame for a target event.
[0007] According to the training method of the event monitoring model and the anomaly detection method for surgical videos provided by the present invention, by constructing a scene-decoupled event memory unit, the event memory unit stores event feature prototypes of normal events and different types of adverse events, and the scene memory unit stores scene feature prototypes of surgical scenes related to the source domain and the target domain, thereby achieving cross-domain decoupling. For complex surgical scenes such as cholecystectomy, prostatectomy and partial nephrectomy, it can effectively distinguish normal events from adverse events in the surgery, improve the generalization ability and accuracy of the event monitoring model in complex surgical scenes, ensure that the abnormal performance in different surgical fields remains consistent, achieve cross-domain invariance and inter-domain comparison, and improve comprehensiveness; in addition, the basic features of the source domain sample video set are learned through the pre-training model, and then based on the pre-training model, high-confidence abnormal areas are automatically selected as pseudo labels. The event monitoring model is further refined and trained based on the pseudo labels and the target domain prediction results, thereby improving the recognition ability of subtle anomalies. The two-stage training process not only improves the overall recognition ability of the event monitoring model for abnormal features, but also enhances the robustness and accuracy in handling various abnormal situations. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:
[0009] Figure 1 A structural block diagram of an intraoperative management and control system integrated into an endoscopic imaging system according to an embodiment of the present invention is shown.
[0010] Figure 2 A flowchart of a method for training an event monitoring model according to an embodiment of the present invention is shown.
[0011] Figure 3 An example diagram of a training event monitoring model according to an embodiment of the present invention is shown.
[0012] Figure 4 A flowchart of a method for detecting abnormalities in surgical videos according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0013] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.
[0014] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0015] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0016] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0017] Related research has found that early surgical anomaly detection solutions primarily focused on monitoring specific adverse events. For example, they employed hidden Markov models to detect blood flow, combined with pixel-cost-based shortest path search, to automatically guide suction tools to remove blood from the surgical field. However, these systems are limited by their focus on a single type of adverse event, making it difficult to comprehensively monitor and warn of potential intraoperative anomalies, resulting in poor detection effectiveness.
[0018] Existing methods have also made significant progress in the field of laparoscopic surgery video control systems. For example, by extracting video frame features and performing image recognition, combined with surgical operation time, blood loss, and patient physiological parameters, laparoscopic surgery can be monitored in real time and abnormality coefficients can be assessed, thereby improving monitoring accuracy. However, existing technologies have certain limitations, focusing primarily on specific surgical environments and equipment, and failing to fully consider the differences in equipment and environments across multiple centers, as well as the data heterogeneity caused by different patient anatomical locations.
[0019] In view of this, an embodiment of the present invention provides a method for training an event monitoring model and a method for detecting anomalies in surgical videos. The method includes: using a source domain sample video set to train a feature extraction module, an event memory unit, and a source domain classifier to obtain a pre-trained model; inputting a target domain sample video into a feature extraction module to output target domain sample features; inputting the target domain sample features and source domain sample features into a scene memory unit and an event memory unit to output mixed domain event enhancement features and scene enhancement features, respectively; inputting the mixed domain event enhancement features, target domain sample features, and source domain sample features into a target domain classifier to output a target domain prediction result; and iteratively adjusting the event feature prototype, scene feature prototype, and network parameters based on the target domain prediction result and pseudo-label until the first iteration stop condition is met, thereby obtaining a trained event monitoring model.
[0020] In the technical solution of the present invention, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0021] In scenarios where personal information is used for automated decision-making, the methods, devices, and systems provided by embodiments of the present invention provide users with corresponding operational portals, allowing them to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered. The term "automated decision-making" herein refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, or credit status through computer programs and making decisions. The term "expert decision-making" herein refers to the activity of decision-making by individuals who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.
[0022] Figure 1 A structural block diagram of an intraoperative management and control system integrated into an endoscopic imaging system according to an embodiment of the present invention is shown.
[0023] like Figure 1As shown in Figure 1, the intraoperative control system integrated into the endoscopic imaging system includes a data processing board, image signal conversion circuit, LED (Light Emitting Diode) light source driver circuit, endoscopic image sensor, LED lighting control module, SD (Solid Drive) card, HDMI (High-Definition Multimedia Interface) / VGA (Video Graphics Array) output, SDRAM (Synchronous Dynamic Random Access Memory) data buffer, and a display. The data processing board adopts a modular design, with a field-programmable gate array (FPGA) as the main control unit. The FPGA main control unit mainly includes an image data acquisition module, an image processing module, a real-time anomaly detection module, an anomaly localization module, an image display module, a cache control module, a display driver module, and the QSYS (Quartus System) system management module.
[0024] The endoscope's image sensor captures intraoperative video and converts it into a standard image signal format via an image signal conversion circuit for subsequent processing by the FPGA main control unit. Simultaneously, the LED illumination control module controls surgical field illumination, providing a stable light source through the LED light source driver circuit to ensure image quality. The converted image signal enters the image data acquisition module and then the image processing module for feature extraction and enhancement, providing effective feature input for the anomaly detection module. The event monitoring model analyzes these features to locate potential adverse events or abnormal conditions. The detected and annotated image is rendered by the image display module. The display driver module then outputs the processed results to the HDMI / VGA output module, ultimately transmitting them to an external display for real-time visual control. The intraoperative control system manages data caching through the buffer control module, coordinating the data rates between image processing and display. Buffered data is stored in the SDRAM data buffer to support high-bandwidth image processing and can also be written to an SD card via QSYS system management and scheduling, enabling intraoperative data recording and subsequent analysis. The coordinated operation of these modules within the system not only optimizes performance and resource utilization, but also ensures high efficiency and stability for medical monitoring and diagnostic applications.
[0025] It should be noted that the sequence numbers of the operations in the following method are only used to indicate the operation for the purpose of description, and should not be regarded as indicating the order in which the operations should be performed. Unless explicitly stated, the method does not need to be performed in the order shown.
[0026] Figure 2A flowchart of a method for training an event monitoring model according to an embodiment of the present invention is shown.
[0027] like Figure 2 As shown, the training method 200 of the event monitoring model includes operations S210 to S250.
[0028] In operation S210 , a feature extraction module, an event memory unit, and a source domain classifier are trained using a source domain sample video set to obtain a pre-trained model.
[0029] Optionally, the event monitoring model includes a feature extraction module, an event memory unit, a source domain classifier, a scene memory unit, and a target domain classifier. The feature extraction module can be constructed based on an attention mechanism, and both the source domain classifier and the target domain classifier can be binary classifiers.
[0030] Optionally, the source domain sample video set can be a sample surgical video of a cholecystectomy procedure, which can be saved to the platform library. The sample surgical videos in the source domain sample video set are labeled with a sample label, which is a video-level label. The sample label is either a normal category or an abnormal category. The normal category indicates that no adverse events occurred during the laparoscopic surgery, while the abnormal category indicates that adverse events occurred during the laparoscopic surgery.
[0031] Alternatively, an adverse event is an unintended injury or complication during surgery caused by medical management that may result in prolonged hospitalization, disability, or death of the subject and is unrelated to the subject's underlying medical condition, for example, unexpected bleeding during surgery.
[0032] Optionally, the source domain sample video set includes N video sets. The nth video set consists of a set of videos labeled as normal and a set of videos labeled as abnormal. Based on the concept of multi-instance learning, each video is considered as a packet, and every 16 frames in the video are grouped into a segment, with each segment being an instance in the packet.
[0033] Optional, video Includes multiple video clips, using the pre-trained feature extraction network to process the i-th video clip To generate the i-th segment feature .
[0034] In one embodiment, the video Sample labels of As shown in formula (1):
[0035] (1);
[0036] in, is the feature of the i-th segment The fragment-level label of the i-th fragment feature Corresponding to a label (Only exists in the testing phase). If there are abnormal clips in the video ( ),but ,otherwise .
[0037] Optionally, the event memory unit includes event feature prototypes, and the scene memory unit includes scene feature prototypes.
[0038] Optionally, the event memory unit is a network unit used to store and retrieve representative features of key events. The event memory unit stores event feature prototypes, which can be a collection, template, or example of event features. Event feature representation retains typical or key features related to adverse events or normal event states in laparoscopic surgery scenarios.
[0039] For example, the event feature prototype may be a collection of event features related to adverse bleeding events that occur during laparoscopic surgery, such as bleeding volume features, pathological features, etc.
[0040] Optionally, the scene memory unit is a network unit used to store and retrieve representative features of key surgical scenes. The scene memory unit stores scene feature prototypes, which can be a collection, template, or example of scene features. The scene feature representation retains style features related to the sample states of the source domain sample video set or the target domain sample video in the laparoscopic surgery scene.
[0041] For example, a scene feature prototype can be a collection of scene features related to multiple surgical scenarios such as cholecystectomy and liver resection, such as surgical tools, tool usage, operation efficiency, and resected tissue characteristics.
[0042] Optionally, the source domain sample video set is input into the feature extraction module, event memory unit and source domain classifier in sequence to obtain classification results. The classification results and video-level labels are used to train the feature extraction module, event memory unit and source domain classifier to obtain a pre-trained model.
[0043] In operation S220 , the target domain sample video is input into a feature extraction module of a pre-trained model for feature extraction, and target domain sample features are output.
[0044] Alternatively, the target domain sample video can be a video produced during a surgery in a domain different from the source domain. For example, if the source domain is a cholecystectomy, the target domain can be a nephrectomy. In this case, the target domain sample video can be an unlabeled sample surgical video obtained from a nephrectomy surgery scenario.
[0045] Optionally, a feature extraction module of a pre-trained model is used to extract features from a video clip in the target domain sample video to obtain target domain sample features.
[0046] Optionally, the target domain sample feature representation captures the temporal characteristics of the global correlation relationship of the normal state and the local dependency relationship of the abnormal state in the target domain sample video.
[0047] In operation S230, the target domain sample features and the source domain sample features are input into the scene memory unit and the event memory unit of the pre-trained model for feature enhancement, and the mixed domain event enhanced features and scene enhanced features are output respectively.
[0048] Optionally, the source domain sample features are obtained by extracting features from a plurality of video segments of the source domain sample videos in the source domain sample video set using a feature extraction module of a pre-trained model.
[0049] Optionally, the source domain sample feature representation captures temporal features of the global correlation relationship of normal states and the local dependency relationship of abnormal states in the source domain sample video.
[0050] Optionally, the target domain sample features and the source domain sample features are input together into the event memory unit of the pre-trained model for feature enhancement, and the mixed domain event enhanced features are output.
[0051] Optionally, the hybrid domain event enhancement feature is a retrieval feature of the event memory unit, which represents an enhanced representation of the matching between the input target domain sample features and source domain sample features and the event feature prototype.
[0052] Optionally, the target domain sample features and the source domain sample features are input into the scene memory unit for feature matching, and the scene enhancement features are output.
[0053] Optionally, the scene enhancement feature is a retrieval result of the scene memory unit, representing an enhanced representation of the matching between the input target domain sample features and source domain sample features and the scene style feature prototype.
[0054] In operation S240 , the hybrid domain event enhancement features, the target domain sample features, and the source domain sample features are input into a target domain classifier, and a target domain prediction result for the target domain sample video is output.
[0055] Optionally, the hybrid domain event enhancement features, the target domain sample features, and the source domain sample features are spliced, and the spliced features are input into a target domain classifier to output a target domain prediction result for the target domain sample video.
[0056] Optionally, the target domain prediction result includes whether the target domain sample video belongs to a normal category or an abnormal category.
[0057] In operation S250, based on the difference between the target domain prediction result and the pseudo label for the target domain sample video, the event feature prototype in the event memory unit of the pre-trained model, the scene feature prototype in the scene memory unit, and the network parameters of the target domain classifier are iteratively adjusted until the first iteration stopping condition is met, thereby obtaining a trained event monitoring model.
[0058] Optionally, pseudo-labels are the predictions output by the pre-trained model after processing the source domain sample video. The source domain sample video is fed into the pre-trained model for processing, resulting in a prediction for the source domain sample video. The target domain sample video has no video-level labels, and the pseudo-labels can be used as labels for the target domain sample video.
[0059] Optionally, a loss function is used to calculate the difference between the target domain prediction result and the pseudo label for the target domain sample video to obtain a loss value. When the loss value does not meet the first iteration stopping condition, the event feature prototype in the event memory unit, the scene feature prototype in the scene memory unit, and the network parameters of the target domain classifier are iteratively adjusted until the loss value meets the first iteration stopping condition. The iteration is stopped to obtain a trained event monitoring model.
[0060] Optionally, the first iteration stopping condition may be a preset loss threshold range.
[0061] Optionally, since a scene-decoupled event memory unit is constructed, orthogonality constraints are imposed on the scene memory unit and the event memory unit to ensure that the feature enhancement of the same feature input in the two memory units is updated along the most orthogonal direction. The event memory unit stores event feature prototypes of normal events and different types of adverse events, and the scene memory unit stores scene feature prototypes of surgical scenes related to the source domain and the target domain, thereby achieving cross-domain decoupling. For complex surgical scenes such as cholecystectomy, prostatectomy and partial nephrectomy, it can effectively distinguish normal events from adverse events during surgery, improve the generalization ability and accuracy of the event monitoring model in complex surgical scenes, ensure that the abnormal performance in different surgical fields remains consistent, achieve cross-domain invariance and inter-domain comparability, and improve comprehensiveness; in addition, the basic features of the source domain sample video set are learned through the pre-training model, and then based on the pre-training model, high-confidence abnormal areas are automatically selected as pseudo labels, and the event monitoring model is further refined and trained based on the pseudo labels and target domain prediction results, thereby improving the ability to recognize subtle abnormalities. The two-stage training process not only improves the overall anomaly recognition capability of the event monitoring model, but also enhances its robustness and accuracy in handling various abnormal situations.
[0062] Optionally, the scene memory unit includes a source domain scene memory sub-unit and a target domain scene memory sub-unit, the source domain scene memory sub-unit includes a source domain scene feature prototype, and the target domain scene memory sub-unit includes a target domain scene feature prototype; and further includes: performing similarity matching on the target domain sample features with the source domain scene feature prototype and the target domain scene feature prototype, and outputting a first scene matching result for the target domain sample features; performing similarity matching on the source domain sample features with the source domain scene feature prototype and the target domain scene feature prototype, and outputting a second scene matching result for the source domain sample features; determining a scene loss value for the scene memory unit according to the scene loss function, the first scene matching result and the second scene matching result; and using the scene loss value, iteratively adjusting the scene feature prototype in the scene memory unit until the scene loss value meets the second iteration stop condition.
[0063] Optionally, the scene memory unit includes a source-domain scene memory subunit and a target-domain scene memory subunit. The source-domain scene memory subunit is a network subunit that stores source-domain scene feature prototypes. For example, the source-domain scene memory subunit stores a collection of style state features for different source-domain sample video sets of surgical scenes with video-level labels.
[0064] For example, if the surgical scene of the source domain sample video set is a cholecystectomy surgical scene, then the source domain scene memory subunit stores the feature prototype of the cholecystectomy surgical scene.
[0065] Optionally, the target domain scene memory subunit is a network subunit that stores target domain scene feature prototypes. For example, the source domain scene memory subunit stores a collection of style state features of different target domain sample videos in an unlabeled video-level surgical scene.
[0066] For example, if the surgical scene of the target domain sample video is a liver resection surgery scene, the target domain scene memory subunit stores the liver resection surgery scene feature prototype.
[0067] Optionally, each memory subunit adopts a learnable structure to maintain scene features.
[0068] Optionally, the first scene matching result includes a result of similarity matching between the target domain sample feature and the source domain scene feature prototype and a result of similarity matching between the target domain sample feature and the target domain scene feature prototype. The result of the similarity matching may be a similarity score.
[0069] Optionally, the second scene matching result includes a result of similarity matching between the source domain sample feature and the source domain scene feature prototype and a result of similarity matching between the source domain sample feature and the target domain scene feature prototype.
[0070] Optionally, a higher similarity between the target domain sample feature and the source domain scene feature prototype is obtained, which indicates that the target domain sample feature and the source domain scene feature prototype are more closely matched.
[0071] For example, a sigmoid function is used to calculate the similarity between the target domain sample features and each source domain scene memory sub-unit to obtain multiple similarities. The similarities of the K source domain scene memory sub-units that are most relevant to the target domain sample features are selected from the multiple similarities, and the average of the K similarities is determined as the result of similarity matching between the target domain sample features and the source domain scene feature prototype.
[0072] In one embodiment, the scene loss function can be determined by four binary cross entropy (BCE) loss terms, as shown in formula (2):
[0073] (2);
[0074] in, Characterize the scene loss value, Characterize the binary cross entropy loss function, Characterizes the result of similarity matching between source domain sample features and source domain scene feature prototypes, Characterizes the result of similarity matching between source domain sample features and target domain scene feature prototypes, Characterizes the result of similarity matching between the target domain sample features and the source domain scene feature prototype, Characterizes the result of similarity matching between the target domain sample features and the target domain scene feature prototype, Represented as a vector of all 1s of length T, It is represented by a vector of all zeros of length T.
[0075] Optionally, the second iteration stop condition can be that the scene loss value is less than a preset threshold, and the optimization goal is: the result of the similarity matching between the source domain sample feature and the source domain scene feature prototype Close to 1, and the result of similarity matching with the target domain scene feature prototype Close to 0; the result of similarity matching between the target domain sample features and the target scene feature prototype Close to 1, and the result of similarity matching with the source domain scene feature prototype Close to 0.
[0076] Optionally, when the scene loss value does not satisfy the second iteration stopping condition, the scene feature prototype in the scene memory unit is iteratively adjusted until the scene loss value satisfies the second iteration stopping condition.
[0077] Optionally, the orthogonal loss is determined based on the orthogonal loss function, the mixed domain event enhancement features and the scene enhancement features; and the event feature prototypes in the event memory unit and the scene feature prototypes in the scene memory unit are iteratively adjusted using the orthogonal loss until the orthogonal loss meets the third iteration stop condition.
[0078] Optionally, the hybrid domain event enhancement feature includes a first enhancement feature and a second enhancement feature. The first enhancement feature is a retrieval feature of the event memory unit, representing the matching result between the input feature and the normal event state prototype; the second enhancement feature is a retrieval feature of the event memory unit, representing the matching result between the input feature and the abnormal event state feature prototype.
[0079] Optionally, the scene enhancement feature includes a third enhancement feature and a fourth enhancement feature. The third enhancement feature is a retrieval feature of the scene memory unit, representing the matching result between the input feature and the source domain scene style feature prototype; the fourth enhancement feature is a retrieval feature of the scene memory unit, representing the matching result between the input feature and the target domain scene style feature prototype.
[0080] In one embodiment, the fourth enhancement feature As shown in formula (3):
[0081] ;
[0082] (3);
[0083] in, represents the normalization function, Represents the transpose of the number of target domain scene feature prototypes in the target domain scene memory subunit, D represents the feature dimension, S represents the attention score matrix, , Characterizing features in input target domain scene memory subunits.
[0084] In one embodiment, the orthogonal loss function is shown in formula (4):
[0085] (4);
[0086] in, Characterizing the first enhancement feature, Characterizing the second enhancement feature, Characterizing the third enhancement feature, Characterizing the fourth enhancement feature, represents the splicing operation along the feature dimension, Characterize the orthogonal loss, Characterization feature transposition processing, Characterizes the Frobenius norm between the two splicing features obtained by two splicing operations.
[0087] Optionally, the optimization goal of the third iteration stopping condition is to minimize the norm of the two spliced features to constrain their orthogonality. Since abnormal manifestations in different surgical scene domains should be consistent, the state prototype should have cross-domain invariance. Therefore, a scene-decoupled event memory mechanism is constructed. By imposing orthogonality constraints on the scene enhancement features output by the scene memory unit and the mixed domain event enhancement features output by the event memory unit, it ensures that the similarity matching results of the same input in the two units are updated along the most orthogonal direction.
[0088] Optionally, when the orthogonal loss does not satisfy the third iteration stopping condition, the event feature prototype in the event memory unit and the scene feature prototype in the scene memory unit are iteratively adjusted until the orthogonal loss value satisfies the third iteration stopping condition.
[0089] Optionally, based on the difference between the target domain prediction result and the pseudo label for the target domain sample video, iteratively adjusting the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier include: determining a first loss value based on a first loss function, the target domain prediction result and the pseudo label; determining a total loss value based on the first loss value, the scene loss value and the orthogonal loss; and iteratively adjusting the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier based on the total loss value.
[0090] In one embodiment, the total loss value As shown in formula (5):
[0091] (5);
[0092] in, Characterize the first loss value, Characterize the scene loss value, Characterize the orthogonal loss, The weight representing the scene loss value, The weight representing the orthogonal loss.
[0093] For example, is 0.1, is 0.01.
[0094] Optionally, minimize the total loss value. When the total loss value is greater than a preset threshold, iteratively adjust the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier until the total loss value is less than the preset threshold.
[0095] Optionally, event anomaly detection is converted into a fully supervised learning training problem, and the pseudo-labels of the video clips are used to minimize the first loss values on the source domain and the target domain to iteratively adjust the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier.
[0096] Optionally, assuming that the manifestation of adverse events in different surgical environments is consistent, the source domain scene memory subunit and the target domain scene memory subunit are optimized by the scene loss value to maintain the intra-domain invariance and inter-domain contrast at the fragment level. On this basis, the orthogonal loss is combined to ensure that the scene features and event features are gradually decoupled during the training process, thereby optimizing the event memory unit and learning the domain-invariant representation of the state.
[0097] Optionally, a sharpening algorithm is used to adjust the pseudo-label to generate a low-entropy soft label; based on the difference between the target domain prediction result and the pseudo-label for the target domain sample video, the event feature prototype in the event memory unit of the pre-trained model, the scene feature prototype in the scene memory unit, and the network parameters of the target domain classifier are iteratively adjusted, including: based on the difference between the target domain prediction result and the low-entropy soft label for the target domain sample video, the event feature prototype in the event memory unit of the pre-trained model, the scene feature prototype in the scene memory unit, and the network parameters of the target domain classifier are iteratively adjusted.
[0098] Optionally, the pseudo labels are based on the prediction results obtained from the source domain sample video.
[0099] In one embodiment, low entropy soft labels As shown in formula (6):
[0100] (6);
[0101] in, Representing pseudo labels.
[0102] In one embodiment, the sharpening algorithm As shown in formula (7):
[0103] (7);
[0104] in, Representing pseudo labels, Characterizes the temperature parameter, which can be set to 0.5. The temperature parameter is used to control the sharpening intensity, thereby regulating the smoothness of the similarity distribution.
[0105] Optionally, the pseudo-label is the predicted distribution of the source domain classifier, as the temperature parameter τ is gradually reduced, The result will gradually approach the Dirac distribution.
[0106] Optionally, a first loss function may be used to calculate the difference between the target domain prediction result and the pseudo label to obtain a first loss value.
[0107] In one embodiment, the first loss function is shown in formula (8):
[0108] (8);
[0109] in, Represent the prediction results of the target domain, Characterize low entropy soft labels, Characterize the logarithmic function, Characterizes the first loss value.
[0110] Optionally, a total loss value is determined based on the first loss, and the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier are iteratively adjusted based on the total loss value.
[0111] Optionally, the goal of the training phase is to utilize a weakly labeled source domain sample video set and unlabeled target domain sample videos. Based on the pre-trained model, the pseudo-labels generated from the source domain sample video set are used to implement a self-learning mechanism, thereby optimizing weak supervision performance and achieving domain adaptation. Due to the distribution differences between the source domain sample video set and the target domain sample video set, pseudo-labels are often unreliable. Therefore, a sharpening algorithm is used to adjust the anomaly score of the source domain classifier, and low-entropy soft labels are refined through self-training by reducing the temperature parameter. This optimizes the target domain classifier based on the low-entropy soft labels and the target domain prediction results, thereby improving classification accuracy.
[0112] Optionally, the source domain sample video set includes multiple source domain sample videos and sample labels for each source domain sample video; the source domain sample video set is used to train the feature extraction module, the event memory unit and the source domain classifier to obtain a pre-trained model including: inputting multiple source domain sample videos into the feature extraction module for feature extraction, and outputting multiple source domain sample features; inputting multiple source domain sample features into the event memory unit for feature enhancement, and outputting multiple source domain event enhancement features; inputting multiple source domain sample features and multiple source domain event enhancement features into the source domain classifier, and outputting source domain prediction results for multiple source domain sample videos; according to the difference between the source domain prediction results and the sample labels of the source domain sample videos, iteratively adjusting the network parameters of the feature extraction module, the source domain classifier and the event feature prototype in the event memory unit until the third iteration stop condition is met to obtain the pre-trained model.
[0113] Optionally, the source domain sample video set includes multiple source domain sample videos, where the source domain sample video is a complete video collected during laparoscopic surgery. Each source domain sample video is marked with a sample label, which is a video-level label. The sample label is a normal category or an abnormal category. The normal category indicates that no adverse events occurred during the laparoscopic surgery, and the abnormal category indicates that adverse events occurred during the laparoscopic surgery.
[0114] Optionally, the feature extraction module can be a lightweight state-aware temporal encoder (Transformer), which inputs the source domain sample video into the feature extraction module for feature extraction and outputs source domain sample features. The source domain sample features are the temporal information of the global correlation relationship of normal states and the abnormal local dependency relationship in the captured video.
[0115] Optionally, the feature extraction module focuses on prior associations related to normal or abnormal states by leveraging contextual information to model temporal dependencies.
[0116] Optionally, the source domain sample features are input into the event memory unit for feature enhancement, and the source domain event enhanced features are output, which are implemented through the attention mechanism. The source domain event enhanced features are features that are enhanced based on the influence of the event feature prototype stored in the retrieval quantification event memory unit on the source domain sample features.
[0117] Optionally, the source domain sample features and the source domain event enhancement features are spliced, and the spliced features are input into a source domain classifier for classification prediction, and a source domain prediction result for the source domain sample video is output.
[0118] Optionally, the source domain prediction result is a prediction classification result of the source domain sample video. For example, a source domain prediction result of 1 represents an abnormal category.
[0119] Optionally, the third iteration stopping condition may be to maximize the separability between the source domain sample videos whose sample labels are normal categories and the source domain sample videos whose sample labels are abnormal categories.
[0120] Optionally, a loss function is used to process the difference between the source domain prediction results and the sample labels to obtain a loss value. When the loss value does not meet the third iteration stopping condition, the feature extraction module, the network parameters of the source domain classifier, and the event feature prototypes in the event memory unit are iteratively adjusted until the loss value meets the third iteration stopping condition to obtain a pre-trained model.
[0121] Alternatively, due to the scarcity of adverse events during surgery and the dominance of normal events, it is often difficult to establish a strong association between the abnormal state features caused by adverse events and the entire sequence. Instead, they rely primarily on the temporal continuity and similar feature distributions of adjacent time points. Based on the global-local attention mechanism, a state-aware temporal encoder is constructed for video temporal relationship modeling. By combining global associations and abnormal local dependencies, and introducing a learnable Gaussian kernel function temporal mask to adaptively adjust the focus, the global associations of normal states and the local dependencies of abnormal states can be effectively learned. This not only significantly improves the accuracy of anomaly detection, but also reduces false positives, thereby improving safety and efficiency during surgery.
[0122] Optionally, the source domain sample video includes at least one source domain video clip, and the feature extraction module includes a feature extraction submodule and a feature encoding submodule; the source domain sample video is input into the feature extraction module for feature extraction, and outputting the source domain sample feature includes: inputting at least one source domain video clip into the feature extraction submodule, extracting the clip feature corresponding to each source domain video clip, and outputting at least one source domain clip feature; inputting the source domain clip feature into the feature encoding submodule, utilizing the attention mechanism implemented by the Gaussian kernel function and the time mask to extract the source domain clip feature, and outputting the source domain sample feature.
[0123] Optionally, the source domain sample video is time-series data and can be divided into source domain video segments consisting of 16 frames of images. A source domain sample video includes one or more source domain video segments. Each source domain video segment has a segment-level label, and the sample label of the source domain sample video is determined based on the segment-level label of the source domain video segment.
[0124] Optionally, the feature extraction submodule can be constructed based on an inflated 3D convolutional network (Inflated 3D ConvNet).
[0125] Optionally, the source domain video clips are input into the feature extraction submodule, which extracts the segment features corresponding to each source domain video clip and outputs them to the source domain segment features. The source domain segment features are segment-level color channel features (Red, Green, Blue, RGB).
[0126] Optionally, the feature encoding submodule is constructed based on the self-attention mechanism implemented by Gaussian kernel function and temporal mask.
[0127] Optionally, a feature encoding submodule is used to extract the source domain segment features to obtain the timing dependency features of the abnormal state and the timing dependency features of the normal state, respectively.
[0128] Optional, learnable Gaussian kernel functions and temporal masks enable adaptive adjustment of focus. Based on the properties of normal distribution, the weight of the Gaussian kernel decreases smoothly with increasing time difference, adaptively adjusting the focus range of anomaly associations at different sequence positions and in different representation subspaces, more effectively modeling the temporal dependencies of anomaly state features.
[0129] In one embodiment, the timing-dependent characteristics of the abnormal state As shown in formula (9):
[0130] (9);
[0131] Among them, T represents the sequence length of the source domain segment feature, Characterizes the standard deviation parameter of the Gaussian kernel, , Representation activation function, represents the exponential function, H represents the number of attention heads, Characterizing the source domain segment features feature, Represents the u-th feature in the source domain segment features.
[0132] Optionally, the different associated weights of each source domain segment feature are normalized separately through an activation function (Softmax) to obtain query, key, normal value, and abnormal value.
[0133] In one embodiment, the query ,key , normal value , outliers As shown in formula (10):
[0134] ;
[0135] ;
[0136] ;
[0137] (10);
[0138] in, Represent different weight matrix parameters, Characterize source domain fragment features.
[0139] Optionally, the global dependencies in the source domain fragment features are mined through the self-attention mechanism in the feature encoding submodule, enabling the model to adaptively capture the most effective normal state dependencies.
[0140] In one embodiment, the timing dependence characteristics of the normal state As shown in formula (11):
[0141] (11);
[0142] in, Characterization query, represents the transpose of the key, The dimension that characterizes the source domain fragments.
[0143] Optionally, the temporal dependency features of the normal state and the abnormal state are concatenated through a concatenation operation and fed into a multi-head attention (MHA) mechanism to capture correlation features across different subspaces. These correlation features are then combined with the source domain segment features, and the dependencies between the state patterns are further strengthened through a multi-layer perceptron (MLP) to generate source domain sample features.
[0144] In one embodiment, the source domain sample features As shown in formula (12):
[0145] (12);
[0146] in, Characterize the multi-head attention mechanism, Characterize the multilayer perceptron, Characterize the time-dependent features of abnormal states, Characterize outliers, Characterize the timing dependence characteristics of the normal state, Characterizes normal values, Characterize the source domain fragment features, Characterize the splicing operation.
[0147] Optionally, by adopting a weakly supervised multi-instance learning framework, the detection of intraoperative adverse events is defined as a multi-instance learning task (splitting the source domain sample video into multiple source domain video segments for prediction), which optimizes data utilization and reduces dependence on large-scale labeled data. This not only enhances the system's ability to respond quickly to complex surgical environments, but also improves prediction accuracy, effectively improving surgical safety and efficiency.
[0148] Optionally, the event memory unit includes a normal event memory sub-unit and an abnormal event memory sub-unit; the normal event memory sub-unit includes a normal event feature prototype, and the abnormal event memory sub-unit includes an abnormal event feature prototype; and further includes: performing similarity matching on each source domain sample feature with the normal event feature prototype and the abnormal event feature prototype, and outputting the event matching result of each source domain sample feature; determining the event loss value for the event memory unit according to the event loss function and the event matching result; and using the event loss value, iteratively adjusting the event feature prototype in the event memory unit until the event loss value meets the fourth iteration stop condition.
[0149] Optionally, the event memory unit includes multiple normal event memory subunits and multiple abnormal event memory subunits. The normal event memory subunit is a memory bank that stores normal event feature prototypes. For example, the normal event memory subunit stores a set of different normal state features in normal events.
[0150] Optionally, the abnormal event memory subunit is a memory bank that stores abnormal event feature prototypes. For example, the abnormal event memory subunit stores a collection of different abnormal state features in adverse events.
[0151] Optionally, the similarity between the source domain sample feature and each normal event memory sub-unit is calculated to obtain multiple similarities, and the similarities of K source domain scene memory sub-units that are most relevant to the target domain sample feature are selected from the multiple similarities, and the average value of the K similarities is determined as the result of similarity matching between the target domain sample feature and the source domain scene feature prototype.
[0152] Optionally, each source domain sample feature is matched against both a normal event feature prototype and an abnormal event feature prototype, and an event matching result is output for each source domain sample feature. The event matching result includes the similarity matching result between the source domain sample feature and the normal event feature prototype, and the similarity matching result between the source domain sample feature and the abnormal event feature prototype.
[0153] Optionally, an event loss function is used to calculate the event matching result to obtain an event loss value for the event memory unit.
[0154] Optionally, the source domain sample features include normal features and abnormal features.
[0155] Optionally, the optimization goal of the fourth iteration stop condition is to maximize the results of similarity matching between normal features and normal event feature prototypes, and the results of similarity matching between abnormal features and abnormal event feature prototypes.
[0156] Optionally, when the event loss value does not meet the fourth iteration stopping condition, the event feature prototype in the event memory unit is iteratively adjusted until the event loss value meets the fourth iteration stopping condition, thereby obtaining the event memory unit.
[0157] Optionally, under the distribution shift of different surgical scenarios, due to the significant difference between the source domain sample features and the target domain sample features, the feature extraction of adverse events is not reliable enough, and intraoperative adverse events have cross-domain consistency. Therefore, the domain-specific scene information is decoupled from the event memory unit to construct a scene memory unit, and only the feature patterns related to abnormal or normal states are retained.
[0158] Optionally, the multiple source domain sample features include at least one source domain positive sample feature and at least one source domain negative sample feature, and the event matching result includes normal similarity and abnormal similarity, the normal similarity characterizes the similarity between the source domain sample feature and the normal event feature prototype, and the abnormal similarity characterizes the similarity between the source domain sample feature and the abnormal event feature prototype; according to the normal similarity, a preset number of source domain positive sample features are determined from at least one source domain positive sample feature to obtain an anchor sample feature; according to the normal similarity, a preset number of source domain negative sample features are determined from at least one source domain negative sample feature to obtain a first sample feature; according to the abnormal similarity, a preset number of source domain negative sample features are determined from at least one source domain negative sample feature to obtain a second sample feature; according to the preset loss function, a first distance between the first sample feature and the anchor sample feature, and a second distance between the second sample feature and the anchor sample feature, a triplet loss value is determined; and using the triplet loss value, the event feature prototype in the event memory unit is iteratively adjusted until the triplet loss value meets the fifth iteration stop condition.
[0159] Optionally, the source domain sample video set includes multiple video sets, each video set consisting of a source domain sample video marked with a normal category label and a source domain sample video marked with an abnormal category label.
[0160] Optionally, feature extraction is performed on the source domain sample video marked with normal category labels to obtain source domain positive sample features; feature extraction is performed on the source domain sample video marked with abnormal category labels to obtain source domain negative sample features.
[0161] Optionally, the normal similarity includes the similarity between the source domain positive sample feature and the normal event feature prototype and the similarity between the source domain negative sample feature and the normal event feature prototype.
[0162] Optionally, the anomaly similarity includes the similarity between the source domain positive sample feature and the abnormal event feature prototype and the similarity between the source domain negative sample feature and the abnormal event feature prototype.
[0163] For example, an event memory unit includes multiple normal event memory sub-units, and the similarity between the source domain positive sample feature and the normal event feature prototype in each normal event memory sub-unit is calculated to obtain multiple similarities. The multiple similarities are sorted in descending order, and K similarities from the 1st position to the Kth position are selected. The normal event feature prototype in the normal event memory sub-unit corresponding to each of the K similarities is most correlated with the source domain positive sample feature. The average value of the K similarities is determined as the normal similarity between the source domain positive sample feature and the normal event feature prototype.
[0164] In one embodiment, the normal similarity between the source domain positive sample feature and the normal event feature prototype As shown in formula (13):
[0165] (13);
[0166] in, Represents the k-th similarity, Characterize the top K similarities.
[0167] In one embodiment, the event loss function is shown in formula (14):
[0168] (14);
[0169] in, Characterize the event loss value, Represents the binary cross entropy loss function (Binary CrossEntropy), Characterizes the normal similarity between the source domain positive sample features and the normal event feature prototype, Characterizes the abnormal similarity between the source domain positive sample features and the abnormal event feature prototype, Characterizes the abnormal similarity between the source domain negative sample features and the abnormal event feature prototype, Characterizes the normal similarity between the source domain negative sample features and the normal event feature prototype, Represented as a vector of all 1s of length T, It is represented by a vector of all zeros of length T.
[0170] Optionally, the first two loss terms of formula (13) utilize the complete label information of the normal state sequence for fully supervised learning, and the last two loss terms focus on the key segments in the abnormal state sequence through the weak label characteristics in multi-instance learning to update the normal event memory subunit and the abnormal event memory subunit.
[0171] Optionally, since the source domain negative sample features are outliers in the normal distribution, some abnormal state features will activate the abnormal event memory subunit, and some normal state features will activate the normal event memory subunit. For example, the source domain negative sample features are [0, 0, 0, 0, 1, 1, 1, 0, 0], and the source domain positive sample features are [0, 0, 0, 0, 0, 0, 0, 0, 0].
[0172] Optionally, based on the normal similarity between each source domain positive sample feature and the normal event feature prototype, a preset number of most relevant normal similarities are selected from the multiple normal similarities, and the source domain positive sample features corresponding to each of the preset number of normal similarities are averaged to obtain the anchor sample feature. For example, the preset number may be 1.
[0173] Optionally, based on the normal similarity between each source domain negative sample feature and the normal event feature prototype, a preset number of most relevant normal similarities are selected from the multiple normal similarities, and the source domain negative sample features corresponding to each of the preset number of normal similarities are averaged to obtain the first sample feature. For example, the preset number may be an integer greater than 1.
[0174] Optionally, based on the anomaly similarity between each source domain negative sample feature and the abnormal event feature prototype, a preset number of most relevant anomaly similarities are selected from the multiple anomaly similarities, and the source domain negative sample features corresponding to each of the preset number of anomaly similarities are averaged to obtain the second sample feature. For example, the preset number may be 1.
[0175] Optionally, the anchor sample feature is the feature closest to the normal state extracted from the source domain sample video marked with the normal category label, the first sample feature is the feature closest to the normal state extracted from the source domain sample video marked with the abnormal category label, and the second sample feature is the feature closest to the abnormal state extracted from the source domain sample video marked with the abnormal category label.
[0176] In one embodiment, the preset loss function is shown in formula (15):
[0177] (15);
[0178] in, Represents the triplet loss value, Characterize the maximum function, Characterize the characteristics of anchor samples, Characterize the first sample characteristics, Characterize the second sample characteristics, represents the square of the Euclidean distance, is the margin parameter, which is used to control the distance between positive and negative samples.
[0179] For example, the margin parameter is 1.
[0180] Optionally, the optimization goal of the fifth iteration stopping condition is to minimize the distance between the first sample feature and the anchor sample feature in the feature space, while maximizing the distance between the second sample feature and the anchor sample feature.
[0181] Optionally, when the triple loss value does not satisfy the fifth iteration stopping condition, the event feature prototype in the event memory unit is iteratively adjusted until the triple loss value satisfies the fifth iteration stopping condition.
[0182] Optionally, both the source domain positive sample features and the source domain negative sample features are segment-level features, and the source domain prediction results corresponding to the segment-level features are processed using a loss function to obtain a segment-level loss value.
[0183] In one embodiment, the segment-level loss value As shown in formula (16):
[0184] (16);
[0185] in, Represents the source domain prediction result corresponding to the qth source domain positive sample feature, Represents the source domain prediction result corresponding to the qth source domain negative sample feature, Characterizes the preset number of selected prediction results, Characterize the margin parameters.
[0186] Optionally, a segment-level loss value is used to constrain the highest abnormal source domain prediction result in the source domain positive sample features to be higher than the source domain negative sample features, thereby maximizing the separability between positive and negative samples and maintaining the interval between positive and negative subsets to implicitly constrain the model to learn more discriminative features.
[0187] In one embodiment, the total loss value in the first stage is As shown in formula (17):
[0188] (17);
[0189] in, Characterize the segment-level loss value, Characterize the event loss value, Represents the triplet loss value, are all hyperparameters that balance the loss terms.
[0190] Optionally, the network parameters of the feature extraction module, the source domain classifier, and the event feature prototypes in the event memory unit are iteratively adjusted according to the total loss value of the first stage until the iteration stopping condition is met to obtain a pre-trained model.
[0191] Figure 3 An example diagram of a training method for an event monitoring model according to an embodiment of the present invention is shown.
[0192] like Figure 3 As shown, the implemented training method includes a first stage 310 and a second stage 320. The first stage performs pre-training based on multi-instance learning, specifically including: using the source domain sample video set to train the feature extraction submodule, the feature encoding submodule, the event memory unit and the source domain classifier. The source domain sample video set includes source domain sample videos marked with normal category labels and source domain sample videos marked with abnormal category labels. The source domain sample video set is input into the feature extraction submodule, and features are extracted from at least one source domain video segment in the source domain sample video, and source domain segment features are output; the source domain segment features are input into the feature encoding submodule, and the source domain segment features are extracted using the attention mechanism implemented by the Gaussian kernel function and the time mask, and the query ,key , normal value , outliers , the time series dependency features of normal state and abnormal state, by querying ,key , normal value , outliers , splicing the temporal dependency features of the normal state and the temporal dependency features of the abnormal state, and outputting the source domain sample features; inputting the source domain sample features into the event memory unit for feature enhancement, and outputting the source domain event enhancement features, wherein the event memory unit includes a normal event memory subunit and an abnormal event memory subunit; splicing the source domain sample features with the source domain event enhancement features to obtain a first splicing feature, inputting the first splicing feature into the source domain classifier, and outputting the source domain prediction result for the source domain sample video According to the difference between the source domain prediction results and the sample labels of the source domain sample video, the feature extraction submodule, the network parameters of the source domain classifier and the event feature prototype in the event memory unit are iteratively adjusted until the third iteration stop condition is met to obtain the pre-trained model. Adjusting source domain prediction results (i.e. pseudo labels), generating low entropy soft labels , completing the first phase of training.
[0193] The self-training for domain adaptation in the second stage specifically includes: using the target domain sample video to train the event memory unit, scene memory unit and target domain classifier. Specifically includes: inputting at least one target domain video clip in the unlabeled target domain sample video into the feature extraction submodule of the pre-training model for feature extraction, and outputting the target domain sample feature; after splicing the target domain sample feature and the source domain sample feature in the first stage, inputting the feature enhancement into the scene memory unit and the event memory unit of the pre-training model, and outputting the mixed domain event enhancement feature and scene enhancement feature respectively, wherein the scene memory unit includes the source domain scene memory subunit and the target domain scene memory subunit; splicing the mixed domain event enhancement feature, the target domain sample feature and the source domain sample feature to obtain the second splicing feature, inputting it into the target domain classifier, and outputting the target domain prediction result for the target domain sample video. ; Prediction results based on target domain and low entropy soft labels The difference between the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier are iteratively adjusted until the first iteration stopping condition is met, thereby obtaining a trained event monitoring model. It should be noted that the trained event monitoring module includes a feature extraction submodule, a feature encoding submodule, an event memory unit, and a target domain classifier.
[0194] Optionally, during the experimental phase of this invention, a standardized data preprocessing and model training process was designed and implemented to achieve efficient detection and robust modeling of abnormal events in surgical videos. During preprocessing, an inflated 3D convolutional neural network (I3D), pretrained on surgical stage recognition tasks, was used as the feature extraction module. This network, based on the deep convolutional neural network architecture (Inception-V1), is constructed by expanding the original 2D convolution kernel into a 3D form, capable of capturing information in both spatial and temporal dimensions. The feature extraction module consists of multiple consecutively stacked 3D convolutional layers (with kernel size of 3×3×3), a max pooling layer, 10 convolutional neural layers, and a global average pooling layer, providing powerful spatiotemporal feature encoding capabilities. During feature extraction, the surgical videos were segmented into 16-frame units, and high-dimensional spatiotemporal feature representations were extracted from each clip. This process uses exactly the same configuration and parameter settings in the source domain (such as cholecystectomy) and the target domain (such as prostate and partial nephrectomy) to ensure consistency in feature representation and improve the model's cross-domain generalization capabilities between different surgical scenario types.
[0195] Optionally, during the event detection model training phase, all input feature sequences were uniformly scaled to 200 frames in length and trained using the Adaptive Moment Estimation optimizer. The initial learning rate was set to 0.0001, the total number of training iterations was 3000, and the batch size was 64. A two-layer Transformer encoder was used, with each layer having a hidden state dimension of 512 and four attention heads. To enhance the temporal memory and state modeling capabilities of the event detection model, an event memory unit with M = 60 event feature prototypes was introduced. The K value in all similarity matching operations was set to ⌊M / 16⌋+1. In terms of loss function design, appropriate hyperparameters were set for different submodules to ensure training stability and model generalization. The experimental platform was implemented on an NVIDIA RTX 3090 GPU (24GB of video memory) to ensure high-performance modeling of temporal features in surgical videos and efficient detection of adverse events.
[0196] Optionally, during the training phase, the performance of the event detection model is comprehensively evaluated using three key metrics: frame-level AUC (area under the receiver operating characteristic curve) to measure overall detection capability, AP (average precision) to assess positive detection performance under class imbalance conditions, and FAR (false alarm rate) to measure the frequency of unnecessary alarms. A higher AUC / AP ratio and a lower FAR indicate better model performance. This multi-metric evaluation strategy combines the complementary strengths of AUC and AP, improving the robustness of detection for rare abnormal events. FAR also provides a practical reference with real clinical significance. On a sample source domain video dataset, the first-stage pre-training of the event detection model achieved an AUC of 92.85%, an AP of 70.28%, and a FAR of 1.40%. After introducing pseudo-label self-training in the second stage, the AUC reached 92.48%, the AP increased to 72.06%, and the FAR decreased to 0.80%. On target domain sample videos, the first phase achieved an AUC of 86.22%, an AP of 74.27%, and a FAR of 0.60%. This improvement was further achieved in the second phase, reaching an AUC of 92.65%, an AP of 85.38%, while maintaining a low FAR of 1.71%. These results demonstrate the excellent anomaly detection capabilities, cross-domain adaptability, and practical application value of this method.
[0197] Optionally, the event monitoring model is trained in two stages. In the first stage, the event monitoring model performs multi-instance learning on source domain sample videos with segment-level labels, distinguishes normal event feature prototypes from abnormal event feature prototypes, and trains the source domain classifier to generate segment-level pseudo labels. In the second stage, the knowledge of the pre-trained model is applied in parallel to the source domain sample video set and the target domain sample video set, and the segment-level pseudo labels are used for self-training to refine the low-entropy soft labels, thereby optimizing the target domain classifier based on the low-entropy soft labels and the target domain prediction results, and at the same time, the internal structure of the unlabeled target domain sample video set is mined to achieve domain adaptation and improve the accuracy of event prediction.
[0198] Figure 4 A flowchart of a method for detecting abnormalities in surgical videos according to an embodiment of the present invention is shown.
[0199] like Figure 4 As shown, the abnormality detection method 400 for surgical videos includes operations S410 to S430.
[0200] In operation S410 , a laparoscopic surgery video is acquired in real time.
[0201] In operation S420 , the M video frames are processed using the feature extraction module, the event memory unit, and the target domain classifier in the trained event monitoring model to obtain a prediction value for each video frame.
[0202] In operation S430 , it is determined that the video frame having the prediction value greater than a preset value is a video frame for the target event.
[0203] Optionally, the laparoscopic surgery video includes M video frames, where M≥1, and M is a positive integer.
[0204] Optionally, real-time video of the laparoscopic procedure during cholecystectomy can be acquired.
[0205] Optionally, the event monitoring model is trained using the above training method. M video frames are sequentially input into the feature extraction module, event memory unit, and target domain classifier in the trained event monitoring model to obtain a predicted value for each video frame.
[0206] Optionally, a trained event monitoring model is used to process the laparoscopic surgery video and output a prediction curve, where the prediction curve is constructed based on the time points and prediction values of the video frames.
[0207] Optionally, the prediction value represents the probability that the video frame is abnormal. For example, a prediction value of 0.6 represents a probability of 0.6 that the video frame is abnormal.
[0208] Optionally, the target event may be an intraoperative adverse event.
[0209] For example, the preset value may be 0.5, and video frames with prediction values greater than 0.5 are determined as video frames for the target event; and video frames with prediction values less than or equal to 0.5 are determined as video frames in which the target event does not occur.
[0210] Optionally, in complex surgical scenarios such as cholecystectomy, prostatectomy, and partial nephrectomy, the system can effectively distinguish between normal and abnormal states at each time point during the operation in real time based on the event monitoring model, significantly improving the accuracy of anomaly detection and reducing false alarms, thereby improving safety and efficiency during the operation.
[0211] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the boxes may occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, as well as the combination of boxes in the block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or may be implemented using a combination of dedicated hardware and computer instructions. It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or combined in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.
[0212] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.
Claims
1. A training method for an event monitoring model, characterized in that: The event monitoring model includes a feature extraction module, an event memory unit, a source domain classifier, a scene memory unit and a target domain classifier, wherein the event memory unit includes an event feature prototype and the scene memory unit includes a scene feature prototype; The method comprises: Using a source domain sample video set, the feature extraction module, the event memory unit, and the source domain classifier are trained to obtain a pre-trained model; Inputting the target domain sample video into the feature extraction module of the pre-trained model for feature extraction, and outputting the target domain sample features; Inputting the target domain sample features and the source domain sample features into the scene memory unit and the event memory unit of the pre-trained model for feature enhancement, and outputting a mixed domain event enhancement feature and a scene enhancement feature, respectively, wherein the source domain sample features are obtained by extracting features from the source domain sample videos in the source domain sample video set using the feature extraction module of the pre-trained model; Inputting the hybrid domain event enhancement feature, the target domain sample feature, and the source domain sample feature into the target domain classifier, and outputting a target domain prediction result for the target domain sample video; According to the difference between the target domain prediction result and the pseudo label for the target domain sample video, the event feature prototype in the event memory unit of the pre-trained model, the scene feature prototype in the scene memory unit and the network parameters of the target domain classifier are iteratively adjusted until the first iteration stopping condition is met, thereby obtaining a trained event monitoring model, wherein the pseudo label is the prediction result outputted after the source domain sample video is processed by the pre-trained model.
2. The method according to claim 1, characterized in that The scene memory unit includes a source domain scene memory subunit and a target domain scene memory subunit, the source domain scene memory subunit includes a source domain scene feature prototype, and the target domain scene memory subunit includes a target domain scene feature prototype; The method further comprises: Performing similarity matching on the target domain sample feature with the source domain scene feature prototype and the target domain scene feature prototype, and outputting a first scene matching result for the target domain sample feature; Performing similarity matching on the source domain sample feature with the source domain scene feature prototype and the target domain scene feature prototype, and outputting a second scene matching result for the source domain sample feature; Determining a scene loss value for the scene memory unit according to a scene loss function, the first scene matching result, and the second scene matching result; The scene feature prototype in the scene memory unit is iteratively adjusted using the scene loss value until the scene loss value satisfies a second iteration stop condition.
3. The method according to claim 2, characterized in that The method further comprises: Determining an orthogonal loss according to an orthogonal loss function, the mixed domain event enhancement feature, and the scene enhancement feature; The orthogonal loss is used to iteratively adjust the event feature prototypes in the event memory unit and the scene feature prototypes in the scene memory unit until the orthogonal loss meets a third iteration stopping condition.
4. The method according to claim 3, characterized in that Iteratively adjusting the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier according to the difference between the target domain prediction result and the pseudo label for the target domain sample video includes: Determine a first loss value according to a first loss function, the target domain prediction result, and the pseudo label; Determining a total loss value according to the first loss value, the scene loss value, and the orthogonal loss; According to the total loss value, the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier are iteratively adjusted.
5. The method according to claim 1, wherein The method further comprises: Adjusting the pseudo labels using a sharpening algorithm to generate low-entropy soft labels; The iteratively adjusting the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier according to the difference between the target domain prediction result and the pseudo label for the target domain sample video includes: According to the difference between the target domain prediction result and the low entropy soft label for the target domain sample video, the event feature prototype in the event memory unit of the pre-trained model, the scene feature prototype in the scene memory unit and the network parameters of the target domain classifier are iteratively adjusted.
6. The method according to claim 1, characterized in that The source domain sample video set includes multiple source domain sample videos and a sample label of each source domain sample video; The method of using the source domain sample video set to train the feature extraction module, the event memory unit, and the source domain classifier to obtain a pre-trained model includes: Inputting multiple source domain sample videos into the feature extraction module for feature extraction, and outputting multiple source domain sample features; Inputting the plurality of source domain sample features into the event memory unit for feature enhancement, and outputting a plurality of source domain event enhanced features; Inputting the plurality of source domain sample features and the plurality of source domain event enhancement features into the source domain classifier, and outputting source domain prediction results for the plurality of source domain sample videos; According to the difference between the source domain prediction result and the sample label of the source domain sample video, the feature extraction module, the network parameters of the source domain classifier and the event feature prototype in the event memory unit are iteratively adjusted until the third iteration stopping condition is met to obtain a pre-trained model.
7. The method according to claim 6, characterized in that The source domain sample video includes at least one source domain video clip, and the feature extraction module includes a feature extraction submodule and a feature encoding submodule; Inputting a plurality of source domain sample videos into the feature extraction module for feature extraction and outputting a plurality of source domain sample features comprises: Inputting the at least one source domain video segment into the feature extraction submodule, extracting segment features corresponding to each source domain video segment, and outputting at least one source domain segment feature; The source domain segment features are input into the feature encoding submodule, and the source domain segment features are extracted using an attention mechanism implemented by a Gaussian kernel function and a time mask, and the source domain sample features are output.
8. The method according to claim 7, characterized in that The event memory unit includes a normal event memory subunit and an abnormal event memory subunit; the normal event memory subunit includes a normal event feature prototype, and the abnormal event memory subunit includes an abnormal event feature prototype; The method further comprises: Performing similarity matching on each of the source domain sample features with the normal event feature prototype and the abnormal event feature prototype, and outputting the event matching result of each source domain sample feature; Determining an event loss value for the event memory unit according to an event loss function and the event matching result; The event feature prototype in the event memory unit is iteratively adjusted using the event loss value until the event loss value satisfies a fourth iterative stopping condition.
9. The method according to claim 8, characterized in that The multiple source domain sample features include at least one source domain positive sample feature and at least one source domain negative sample feature. The event matching result includes a normal similarity and an abnormal similarity. The normal similarity represents the similarity between the source domain sample feature and the normal event feature prototype, and the abnormal similarity represents the similarity between the source domain sample feature and the abnormal event feature prototype. The method further comprises: Determining a preset number of source domain positive sample features from the at least one source domain positive sample feature according to the normal similarity to obtain anchor sample features; Determining a preset number of source domain negative sample features from the at least one source domain negative sample feature according to the normal similarity to obtain a first sample feature; Determining a preset number of source domain negative sample features from the at least one source domain negative sample feature according to the abnormal similarity to obtain a second sample feature; Determine a triplet loss value according to a preset loss function, a first distance between the first sample feature and the anchor point sample feature, and a second distance between the second sample feature and the anchor point sample feature; The event feature prototype in the event memory unit is iteratively adjusted using the triplet loss value until the triplet loss value satisfies a fifth iteration stopping condition.
10. A method for detecting abnormalities in surgical videos, characterized in that: The method comprises: Acquire a laparoscopic surgery video in real time, wherein the laparoscopic surgery video includes M video frames, M ≥ 1, and M is a positive integer; Processing the M video frames using a feature extraction module, an event memory unit, and a target domain classifier in a trained event monitoring model to obtain a prediction value for each video frame, wherein the prediction value represents a probability that the video frame is an abnormal situation, the event monitoring model being trained using the training method of any one of claims 1 to 9; and The video frame whose predicted value is greater than a preset value is determined as a video frame for the target event.
Citation Information
Patent Citations
Robust field adaptive image learning method based on self-training noise label correction
CN114283287A
Abnormal event detection algorithm based on unsupervised domain self-adaption
CN116385935A