Event monitoring model training method and operation video anomaly detection method

By building an event monitoring model, using source domain and target domain sample video for feature extraction and enhancement, combined with scene memory units and classifiers, the abnormal monitoring problem caused by differences in multi-center equipment and environment in minimally invasive surgery is solved, cross-domain decoupling and pseudo-label self-training are achieved, and the accuracy and robustness of abnormal detection are improved.

CN120279469AActive Publication Date: 2025-07-08INST OF MEDICAL ROBOTICS & INTELLIGENT SYST TIANJIN UNIV

Patent Information

Application Number
CN202510748813.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-08
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to comprehensively monitor potential abnormalities caused by multi-center equipment and environmental differences in minimally invasive surgery, lacks versatility and comprehensiveness, and only monitors a single type of adverse events, and has poor detection results.

Method used

By building an event monitoring model, feature extraction and enhancement is performed using source domain and target domain sample video, combined with scene memory units and classifiers, iteratively adjust model parameters, realize cross-domain decoupling and pseudo-label self-training, and improve abnormal detection capabilities.

Benefits of technology

The generalization ability and accuracy of complex surgical scenarios are achieved, the ability to identify subtle abnormalities is improved, robustness and accuracy are enhanced, and abnormal performance in different surgical fields is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279469A_ABST
    Figure CN120279469A_ABST
Patent Text Reader

Abstract

The invention provides an event monitoring model training method and an operation video anomaly detection method, which can be applied to the technical field of video processing. The method comprises the following steps: training a feature extraction module, an event memory unit and a source domain classifier by using a source domain sample video set to obtain a pre-training model; inputting the target domain sample video into a feature extraction module and outputting target domain sample features; inputting the target domain sample features and the source domain sample features into a scene memory unit and an event memory unit, and respectively outputting mixed domain event enhancement features and scene enhancement features; inputting the mixed domain event enhancement feature, the target domain sample feature and the source domain sample feature into a target domain classifier, and outputting a target domain prediction result; and iteratively adjusting the event feature prototype, the scene feature prototype and the network parameters according to the target domain prediction result and the pseudo tag until a first iteration stop condition is met, and obtaining a trained event monitoring model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and more specifically, to a method for training an event monitoring model and a method for detecting anomalies in surgical videos. Background Art

[0002] With the development of minimally invasive surgery technology and surgical robot systems, one of the main challenges is to cope with unforeseen scenarios and unexpected events in unstructured and deformable surgical environments. The sensor devices involved in minimally invasive laparoscopic surgical robots can acquire key surgical information, and can infer the actual state of the surgical process, potential adverse events and their inducing factors based on the acquired information, helping surgeons perform more complex surgeries, thereby improving the surgical safety of patients.

[0003] However, current research only monitors a single type of adverse event, fails to comprehensively cover potential anomalies during surgery, and lacks comprehensiveness. At the same time, the prior art does not consider the differences in multi-center devices and environments, resulting in a lack of generality. Summary of the Invention

[0004] In view of this, the present invention provides a method for training an event monitoring model and a method for detecting anomalies in surgical videos.

[0005] One aspect of the present invention provides a method for training an event monitoring model, including: using a source domain sample video set to train a feature extraction module, an event memory unit, and a source domain classifier to obtain a pre-trained model; inputting a target domain sample video into the feature extraction module of the pre-trained model for feature extraction, and outputting target domain sample features; inputting the target domain sample features and the source domain sample features into a scene memory unit and the event memory unit of the pre-trained model for feature enhancement, and respectively outputting a hybrid domain event enhanced feature and a scene enhanced feature, wherein the source domain sample features are obtained by using the feature extraction module of the pre-trained model to extract features from the source domain sample videos in the source domain sample video set; inputting the hybrid domain event enhanced feature, the target domain sample features, and the source domain sample features into a target domain classifier, and outputting a target domain prediction result for the target domain sample video; according to the difference between the target domain prediction result and the pseudo-label for the target domain sample video, iteratively adjusting the event feature prototype in the event memory unit of the pre-trained model, the scene feature prototype in the scene memory unit, and the network parameters of the target domain classifier until the first iteration stop condition is satisfied, to obtain a trained event monitoring model, wherein the pseudo-label is the prediction result output after the source domain sample video is processed by the pre-trained model.

[0006] The second aspect of the present invention provides an abnormal detection method for surgical videos, including: obtaining a laparoscopic surgical video in real time, where the laparoscopic surgical video includes M video frames, M≥1 and M is a positive integer; using the feature extraction module, event memory unit, and target domain classifier in the trained event monitoring model to process the M video frames to obtain the prediction value of each video frame, where the prediction value represents the probability that the video frame is an abnormal situation, and the event monitoring model is trained using the above training method; and determining the video frames with prediction values greater than a preset value as the video frames for the target event.

[0007] According to the training method of the event monitoring model and the abnormal detection method for surgical videos provided by the present invention, by constructing a scene-decoupled event memory unit, the event memory unit stores the event feature prototypes of normal events and different types of adverse events, and the scene memory unit stores the scene feature prototypes of surgical scenes related to the source domain and the target domain respectively, so as to achieve cross-domain decoupling. For complex surgical scenes such as cholecystectomy, prostatectomy, and partial nephrectomy, it can effectively distinguish normal events and adverse events during surgery, improve the generalization ability and accuracy of the event monitoring model in complex surgical scenes, ensure that the abnormal manifestations in different surgical fields are consistent, achieve cross-domain invariance and inter-domain comparability, and improve comprehensiveness; in addition, by using the pre-trained model to learn the basic features of the source domain sample video set, and then based on the pre-trained model, automatically selecting the abnormal regions with high confidence as pseudo-labels, and further finely training the event monitoring model based on the pseudo-labels and the target domain prediction results, so as to improve the recognition ability of subtle abnormalities. The two-stage training process not only improves the overall recognition ability of the event monitoring model for abnormal features, but also enhances the robustness and accuracy in dealing with various abnormal situations. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0009] Figure 1 The structural block diagram of the intraoperative management and control system integrated in the endoscopic imaging system according to the embodiment of the present invention is shown.

[0010] Figure 2 The flowchart of the training method of the event monitoring model according to the embodiment of the present invention is shown.

[0011] Figure 3 The example diagram of training the event monitoring model according to the embodiment of the present invention is shown.

[0012] Figure 4 The flowchart of the abnormal detection method for surgical videos according to the embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, numerous specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0014] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0015] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0016] In cases where expressions similar to "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0017] In related research, it has been found that early surgical anomaly detection schemes mainly focused on the monitoring of specific adverse events. For example, a hidden Markov model was used to detect blood flow, and combined with the shortest path search based on pixel cost, to automatically guide the suction tool to clear the blood in the surgical field. However, the limitation of these systems is that they only focus on a single type of adverse event, making it difficult to comprehensively monitor and warn of possible anomalies during surgery, and the detection effect is poor.

[0018] Existing methods have also made significant progress in the field of laparoscopic surgery video management and control systems. For example, by extracting video frame characteristics and performing image recognition, combined with surgical operation time, blood loss, and patient physiological parameters, the laparoscopic surgery process is monitored in real time and the anomaly coefficient is evaluated, thereby improving the monitoring accuracy. However, there are certain limitations in the existing technologies, mainly concentrated on specific surgical environments and equipment, and the data heterogeneity caused by differences in multi-center equipment and environments, as well as different patient anatomical positions, has not been fully considered.

[0019] In view of this, embodiments of the present invention provide a method for training an event monitoring model and a method for anomaly detection of surgical videos. The method includes: using a source domain sample video set to train a feature extraction module, an event memory unit, and a source domain classifier to obtain a pre-trained model; inputting a target domain sample video into the feature extraction module to output target domain sample features; inputting the target domain sample features and the source domain sample features into a scene memory unit and an event memory unit to respectively output hybrid domain event enhanced features and scene enhanced features; inputting the hybrid domain event enhanced features, the target domain sample features, and the source domain sample features into a target domain classifier to output a target domain prediction result; and iteratively adjusting an event feature prototype, a scene feature prototype, and network parameters according to the target domain prediction result and the pseudo labels until a first iteration stop condition is satisfied to obtain a trained event monitoring model.

[0020] In the technical solution of the present invention, the involved user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application complies with relevant laws, regulations, and standards, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or refuse.

[0021] In the scenario of making automated decisions using personal information, the methods, devices, and systems provided by embodiments of the present invention all provide corresponding operation entrances for users to choose to agree or refuse the automated decision results; if the user chooses to refuse, an expert decision-making process will be entered. The expression "automated decision" here refers to an activity of automatically analyzing and evaluating a person's behavior habits, interests and hobbies, or economic, health, credit status, etc. through a computer program and making a decision. The expression "expert decision" here refers to an activity of making a decision by a person who specializes in a certain field of work, has specialized experience, knowledge, and skills, and has reached a certain professional level.

[0022] Figure 1 The structural block diagram of an intraoperative control system integrated in an endoscopic imaging system according to an embodiment of the present invention is shown.

[0023] As Figure 1As shown in the figure, the intraoperative control system integrated in the endoscopic imaging system includes a data processing board, an image signal conversion circuit, an LED (Light Emitting Diode) light source driving circuit, an endoscopic image sensor, an LED lighting control module, an SD (Solid Drive) card, an HDMI (High-Definition Multimedia Interface) / VGA (Video Graphics Array) output, an SDRAM (Synchronous Dynamic Random Access Memory) data buffer, and a display. The data processing board adopts a modular design and uses a Field-Programmable Gate Array (FPGA) as the main control unit. The FPGA main control unit mainly includes an image data acquisition module, an image processing module, a real-time anomaly detection module, an anomaly localization module, an image display module, a buffer control module, a display driving module, and a QSYS (Quartus System) system management.

[0024] The endoscopic image sensor captures intraoperative videos, and through the image signal conversion circuit, converts the intraoperative videos into image signals in a standard signal format for subsequent processing by the FPGA main control unit. At the same time, the LED lighting control module controls the illumination of the surgical field, providing a stable light source through the LED light source driving circuit to ensure image quality. The converted image signals enter the image data acquisition module and are then sent to the image processing module for feature extraction and enhancement, providing effective feature inputs for the anomaly detection module. The features are analyzed through an event monitoring model to locate possible adverse events or abnormal states. The detected and labeled images are rendered by the image display module. Subsequently, the display driving module outputs the processing results to the HDMI / VGA output module and finally transmits them to an external display for real-time visualization control. The intraoperative control system manages data caching through the buffer control module and coordinates the data rates of image processing and display. The buffered data is stored in the SDRAM data buffer to support high-bandwidth image processing and can also be written into the SD card through QSYS system management scheduling to achieve the recording and subsequent analysis of intraoperative data. The modules within the system work together, not only optimizing performance and resource utilization but also ensuring the high efficiency and stability of the system in medical monitoring and diagnostic applications.

[0025] It should be noted that the sequence numbers of each operation in the following methods are only used as representations of the operations for description and should not be regarded as indicating the execution order of each operation. Unless explicitly stated, the method does not need to be executed exactly in the order shown.

[0026] Figure 2The flowchart of the training method of the event monitoring model according to an embodiment of the present invention is shown.

[0027] As Figure 2 shown, the training method 200 of the event monitoring model includes operation S210 to operation S250.

[0028] In operation S210, using the source domain sample video set, train the feature extraction module, event memory unit and source domain classifier to obtain a pre-trained model.

[0029] Optionally, the event monitoring model includes a feature extraction module, an event memory unit, a source domain classifier, a scene memory unit and a target domain classifier. The feature extraction module can be constructed based on the attention mechanism, and both the source domain classifier and the target domain classifier can be binary classifiers.

[0030] Optionally, the source domain sample video set can be the sample surgical videos during the cholecystectomy surgery, and the sample surgical videos can be saved to the platform library. The sample surgical videos in the source domain sample video set are marked with sample labels, and the sample labels are video-level labels. The sample label is a normal category or an abnormal category. The normal category represents that no adverse events occur during the laparoscopic surgery, and the abnormal category is that adverse events occur during the laparoscopic surgery.

[0031] Optionally, the adverse event is an accidental injury or complication event caused by medical management factors during the surgery, which may lead to an extended hospital stay, disability or death of the target object, and is not related to the original disease of the target object. For example, an accidental intraoperative bleeding event.

[0032] Optionally, the source domain sample video set includes N groups of video sets. The nth group of video sets consists of a group of videos marked with the normal category and videos marked with the abnormal category. Based on the idea of multi-instance learning, each video is regarded as a bag, and every 16 frames in the video form a segment, and each segment is used as an instance in the bag.

[0033] Optionally, the video includes multiple segment videos, and use the pre-trained feature extraction network to process the ith segment video to generate the ith segment feature .

[0034] In one embodiment, the sample label of the video is as shown in formula (1): As shown in formula (1):

[0035] (1);

[0036] Wherein, is the segment-level label of the ith segment feature, and the ith segment feature of corresponds to a label (only existing in the test phase). If there are abnormal segments in the video ( ), then , otherwise .

[0037] Optionally, the event memory unit includes an event feature prototype, and the scene memory unit includes a scene feature prototype.

[0038] Optionally, the event memory unit is a network unit for storing and retrieving representative features of key events. The event feature prototype is stored in the event memory unit, and the event feature prototype can be a set, template, or example of event features, etc. The event features represent typical or key features related to the state of adverse events or normal events in the laparoscopic surgery scenario.

[0039] For example, the event feature prototype can be a set of event features related to the adverse event of bleeding that occurs during the laparoscopic surgery process. Such as, blood loss feature, pathological feature, etc.

[0040] Optionally, the scene memory unit is a network unit for storing and retrieving representative features of key surgery scenarios. The scene feature prototype is stored in the scene memory unit, and the scene feature prototype can be a set, template, or example of scene features, etc. The scene features represent style features related to the sample state of the source domain sample video set or the target domain sample video in the laparoscopic surgery scenario.

[0041] For example, the scene feature prototype can be a set of scene features related to multiple surgery scenarios such as cholecystectomy and hepatectomy. Such as, surgical tools, tool usage, operation efficiency, resection tissue features, etc.

[0042] Optionally, the source domain sample video set is sequentially input into the feature extraction module, the event memory unit, and the source domain classifier to obtain a classification result. Using the classification result and the video-level label, the feature extraction module, the event memory unit, and the source domain classifier are trained to obtain a pre-trained model.

[0043] In operation S220, the target domain sample video is input into the feature extraction module of the pre-trained model for feature extraction, and the target domain sample features are output.

[0044] Optionally, the target domain sample video can be a video generated during the surgery process in a domain different from the source domain. For example, if the source domain is a cholecystectomy surgery, then the target domain can be a nephrectomy surgery. At this time, the target domain sample video can be an unlabeled sample surgery video obtained based on the nephrectomy surgery scenario.

[0045] Optionally, the feature extraction module of the pre-trained model is used to extract features from the segment videos in the target domain sample video, and the target domain sample features are obtained.

[0046] Optionally, the target domain sample features characterize the temporal features that capture the global correlation relationship in the normal state and the local dependence relationship in the abnormal state in the target domain sample video.

[0047] In operation S230, the target domain sample features and the source domain sample features are input into the scene memory unit and the event memory unit of the pre-trained model for feature enhancement, and the hybrid domain event enhanced features and the scene enhanced features are output respectively.

[0048] Optionally, the source domain sample features are obtained by using the feature extraction module of the pre-trained model to extract features from multiple segment videos in the source domain sample videos in the source domain sample video set.

[0049] Optionally, the source domain sample features characterize the temporal features that capture the global correlation relationship in the normal state and the local dependence relationship in the abnormal state in the source domain sample video.

[0050] Optionally, the target domain sample features and the source domain sample features are input into the event memory unit of the pre-trained model together for feature enhancement, and the hybrid domain event enhanced features are output.

[0051] Optionally, the hybrid domain event enhanced features are the retrieval features of the event memory unit, representing the matching enhanced representation of the input target domain sample features and source domain sample features with the event feature prototype.

[0052] Optionally, the target domain sample features and the source domain sample features are input into the scene memory unit together for feature matching, and the scene enhanced features are output.

[0053] Optionally, the scene enhanced features are the retrieval results of the scene memory unit, representing the matching enhanced representation of the input target domain sample features and source domain sample features with the scene style feature prototype.

[0054] In operation S240, the hybrid domain event enhanced features, the target domain sample features and the source domain sample features are input into the target domain classifier, and the target domain prediction result for the target domain sample video is output.

[0055] Optionally, the hybrid domain event enhanced features, the target domain sample features and the source domain sample features are concatenated, and the concatenated features are input into the target domain classifier, and the target domain prediction result for the target domain sample video is output.

[0056] Optionally, the target domain prediction result includes that the target domain sample video belongs to the normal category or the abnormal category.

[0057] At operation S250, based on the difference between the target domain prediction result and the pseudo-label for the target domain sample video, iteratively adjust the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier until the first iteration stop condition is met, obtaining a trained event monitoring model.

[0058] Optionally, the pseudo-label is the prediction result output after processing the source domain sample video by the pre-trained model. Input the source domain sample video into the pre-trained model for processing to obtain the prediction result for the source domain sample video. The target domain sample video has no labeled video-level label, and the pseudo-label can be used as the label for the target domain sample video.

[0059] Optionally, use a loss function to calculate the difference between the target domain prediction result and the pseudo-label for the target domain sample video, obtaining a loss value. When the loss value does not meet the first iteration stop condition, iteratively adjust the event feature prototypes in the event memory unit, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier until the loss value meets the first iteration stop condition, stop the iteration, and obtain a trained event monitoring model.

[0060] Optionally, the first iteration stop condition can be a preset loss threshold range.

[0061] Optionally, since a scene-decoupled event memory unit is constructed, an orthogonality constraint is imposed on the scene memory unit and the event memory unit to ensure that the feature enhancement of the same feature input in the two memory units updates the event feature prototypes storing normal events and different types of adverse events in the event memory unit along the most orthogonal direction. The scene memory unit stores the scene feature prototypes related to the surgical scenes of the source domain and the target domain respectively, thereby achieving cross-domain decoupling. For complex surgical scenes such as cholecystectomy, prostatectomy, and partial nephrectomy, it can effectively distinguish normal events and adverse events during the operation, improve the generalization ability and accuracy of the event monitoring model in complex surgical scenes, ensure that the abnormal manifestations in different surgical fields are consistent, achieve cross-domain invariance and inter-domain comparability, and improve comprehensiveness. In addition, by learning the basic features of the source domain sample video set through the pre-trained model, and then based on the pre-trained model, automatically select the high-confidence abnormal regions as pseudo-labels, and further refine the training of the event monitoring model based on the pseudo-labels and the target domain prediction results, thereby enhancing the recognition ability for subtle abnormalities. The two-stage training process not only improves the overall recognition ability of the event monitoring model for heterogeneous features, but also enhances the robustness and precision in dealing with various abnormal situations.

[0062] Optionally, the scene memory unit includes a source domain scene memory subunit and a target domain scene memory subunit. The source domain scene memory subunit includes source domain scene feature prototypes, and the target domain scene memory subunit includes target domain scene feature prototypes. It further includes: performing similarity matching between the target domain sample features and the source domain scene feature prototypes and the target domain scene feature prototypes respectively, and outputting a first scene matching result for the target domain sample features; performing similarity matching between the source domain sample features and the source domain scene feature prototypes and the target domain scene feature prototypes respectively, and outputting a second scene matching result for the source domain sample features; determining a scene loss value for the scene memory unit according to the scene loss function, the first scene matching result, and the second scene matching result; and using the scene loss value to iteratively adjust the scene feature prototypes in the scene memory unit until the scene loss value satisfies a second iteration stop condition.

[0063] Optionally, the scene memory unit includes a source domain scene memory subunit and a target domain scene memory subunit. The source domain scene memory subunit is a network subunit that stores source domain scene feature prototypes. For example, the source domain scene memory subunit stores a set of style state features of different source domain sample video sets in a surgical scene with video-level labels.

[0064] For example, if the surgical scene of the source domain sample video set is a cholecystectomy surgical scene, the source domain scene memory subunit stores the cholecystectomy surgical scene feature prototype.

[0065] Optionally, the target domain scene memory subunit is a network subunit that stores target domain scene feature prototypes. For example, the source domain scene memory subunit stores a set of style state features of different target domain sample videos in a surgical scene without labeled video-level labels.

[0066] For example, if the surgical scene of the target domain sample video is a hepatectomy surgical scene, the target domain scene memory subunit stores the hepatectomy surgical scene feature prototype.

[0067] Optionally, each memory subunit maintains scene features using a learnable structure.

[0068] Optionally, the first scene matching result includes the result of similarity matching between the target domain sample features and the source domain scene feature prototypes and the result of similarity matching between the target domain sample features and the target domain scene feature prototypes. The result of similarity matching can be a similarity score.

[0069] Optionally, the second scene matching result includes the result of similarity matching between the source domain sample features and the source domain scene feature prototypes and the result of similarity matching between the source domain sample features and the target domain scene feature prototypes.

[0070] Optionally, the higher the similarity between the target domain sample features and the source domain scenario feature prototype, the more matching the target domain sample features are with the source domain scenario feature prototype.

[0071] For example, the similarity between the target domain sample features and each source domain scenario memory subunit is calculated using a logistic function (Sigmoid Function) to obtain multiple similarities. Among these multiple similarities, the similarities of the K source domain scenario memory subunits most relevant to the target domain sample features are selected, and the average value of the K similarities is determined as the result of the similarity matching between the target domain sample features and the source domain scenario feature prototype.

[0072] In one embodiment, the scenario loss function can be determined by four binary cross-entropy (BCE) loss terms, as shown in formula (2):

[0073] (2);

[0074] Where, represents the scenario loss value, represents the binary cross-entropy loss function, represents the result of the similarity matching between the source domain sample features and the source domain scenario feature prototype, represents the result of the similarity matching between the source domain sample features and the target domain scenario feature prototype, represents the result of the similarity matching between the target domain sample features and the source domain scenario feature prototype, represents the result of the similarity matching between the target domain sample features and the target domain scenario feature prototype, represents a vector of all 1s with length T, represents a vector of all 0s with length T.

[0075] Optionally, the second iteration stop condition can be that the scenario loss value is less than a preset threshold, and its optimization goal is: the result of the similarity matching between the source domain sample features and the source domain scenario feature prototype is close to 1, while the result of the similarity matching with the target domain scenario feature prototype is close to 0; the result of the similarity matching between the target domain sample features and the target scenario feature prototype is close to 1, while the result of the similarity matching with the source domain scenario feature prototype is close to 0.

[0076] Optionally, when the scenario loss value does not meet the second iteration stop condition, the scenario feature prototype in the scenario memory unit is iteratively adjusted until the scenario loss value meets the second iteration stop condition.

[0077] Optionally, an orthogonal loss is determined according to an orthogonal loss function, hybrid-domain event enhancement features, and scene enhancement features; the orthogonal loss is used to iteratively adjust the event feature prototype in the event memory unit and the scene feature prototype in the scene memory unit until the orthogonal loss meets the third iteration stop condition.

[0078] Optionally, the hybrid-domain event enhancement features include a first enhancement feature and a second enhancement feature. The first enhancement feature is the retrieval feature of the event memory unit, representing the matching result between the input feature and the normal event state prototype; the second enhancement feature is the retrieval feature of the event memory unit, representing the matching result between the input feature and the abnormal event state feature prototype.

[0079] Optionally, the scene enhancement features include a third enhancement feature and a fourth enhancement feature. The third enhancement feature is the retrieval feature of the scene memory unit, representing the matching result between the input feature and the source-domain scene style feature prototype; the fourth enhancement feature is the retrieval feature of the scene memory unit, representing the matching result between the input feature and the target-domain scene style feature prototype.

[0080] In one embodiment, the fourth enhancement feature is as shown in formula (3):

[0081] ;

[0082] (3);

[0083] where represents the normalization function, represents the transpose of the number of target-domain scene feature prototypes in the target-domain scene memory subunit, D represents the feature dimension, S represents the attention score matrix, , represents the feature input into the target-domain scene memory subunit.

[0084] In one embodiment, the orthogonal loss function is as shown in formula (4):

[0085] (4);

[0086] where represents the first enhancement feature, represents the second enhancement feature, represents the third enhancement feature, represents the fourth enhancement feature, represents the concatenation operation along the feature dimension, represents the orthogonal loss, represents the feature transpose process, represents the Frobenius norm between two concatenated features obtained by two concatenation operations.

[0087] Optionally, the optimization objective of the third iteration stop condition is to minimize the norm of two concatenated features to constrain their orthogonality. Since the abnormal manifestations in different surgical scenario domains should be consistent, the state prototype should have cross-domain invariance. Therefore, a scene-decoupled event memory mechanism is constructed, and by imposing an orthogonality constraint on the scene-enhanced feature output by the scene memory unit and the hybrid-domain event-enhanced feature output by the event memory unit, it is ensured that the similarity matching results of the same input in the two units are updated along the most orthogonal direction.

[0088] Optionally, when the orthogonality loss does not meet the third iteration stop condition, iteratively adjust the event feature prototype in the event memory unit and the scene feature prototype in the scene memory unit until the orthogonality loss value meets the third iteration stop condition.

[0089] Optionally, according to the difference between the target domain prediction result and the pseudo-label for the target domain sample video, iteratively adjust the event feature prototype in the event memory unit of the pre-trained model, the scene feature prototype in the scene memory unit, and the network parameters of the target domain classifier, including: determining a first loss value according to the first loss function, the target domain prediction result, and the pseudo-label; determining a total loss value according to the first loss value, the scene loss value, and the orthogonality loss; and iteratively adjusting the event feature prototype in the event memory unit of the pre-trained model, the scene feature prototype in the scene memory unit, and the network parameters of the target domain classifier according to the total loss value.

[0090] In one embodiment, the total loss value is as shown in formula (5):

[0091] (5);

[0092] wherein, represents the first loss value, represents the scene loss value, represents the orthogonality loss, represents the weight of the scene loss value, represents the weight of the orthogonality loss.

[0093] For example, is 0.1, is 0.01.

[0094] Optionally, minimize the total loss value. When the total loss value is greater than a preset threshold, iteratively adjust the event feature prototype in the event memory unit of the pre-trained model, the scene feature prototype in the scene memory unit, and the network parameters of the target domain classifier until the total loss value is less than the preset threshold.

[0095] Optionally, convert event anomaly detection into a fully supervised learning training problem, and use the pseudo-labels of the segment videos to minimize the first loss value on the source domain and the target domain, so as to iteratively adjust the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier.

[0096] Optionally, assume that the performance of adverse events is consistent in different surgical environments, and optimize the source domain scene memory subunit and the target domain scene memory subunit through the scene loss value to maintain intra-domain invariance and inter-domain contrast at the segment level. On this basis, combined with the orthogonal loss, ensure that the scene features and event features are gradually decoupled during training, thereby optimizing the event memory unit and learning domain-invariant representations of the states.

[0097] Optionally, use a sharpening algorithm to adjust the pseudo-labels to generate low-entropy soft labels; according to the difference between the target domain prediction result and the pseudo-labels for the target domain sample videos, iteratively adjust the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier, including: according to the difference between the target domain prediction result and the low-entropy soft labels for the target domain sample videos, iteratively adjust the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier.

[0098] Optionally, the pseudo-labels are based on the prediction results of the source domain sample videos.

[0099] In one embodiment, the low-entropy soft labels are as shown in formula (6):

[0100] (6);

[0101] where represents the pseudo-labels.

[0102] In one embodiment, the sharpening algorithm is as shown in formula (7):

[0103] (7);

[0104] where represents the pseudo-labels, represents the temperature parameter, which can be set to 0.5. The temperature parameter is a parameter used to control the sharpening intensity, and thus regulate the smoothness of the similarity distribution.

[0105] Optionally, the pseudo-labels are the prediction distributions of the source domain classifier. As the temperature parameter τ gradually decreases, The result will gradually approach the Dirac Distribution.

[0106] Optionally, the difference between the target domain prediction result and the pseudo-label can be calculated using the first loss function to obtain the first loss value.

[0107] In one embodiment, the first loss function is as shown in formula (8):

[0108] (8);

[0109] Where, represents the target domain prediction result, represents the low-entropy soft label, represents the logarithmic function, represents the first loss value.

[0110] Optionally, based on the first loss, the total loss value is determined, and the event feature prototype in the event memory unit of the pre-trained model, the scene feature prototype in the scene memory unit, and the network parameters of the target domain classifier are iteratively adjusted based on the total loss value.

[0111] Optionally, the goal of the training phase is to utilize the weakly labeled source domain sample video set and the unlabeled target domain sample videos to implement a self-learning mechanism through the pseudo-labels generated based on the source domain sample video set on the basis of the pre-trained model, thereby optimizing the weak supervision performance and achieving domain adaptation. Due to the distribution difference between the source domain sample video set and the target domain sample videos, the pseudo-labels are often not reliable. Therefore, the anomaly score of the source domain classifier is adjusted using the sharpening algorithm, and the low-entropy soft label is refined by self-training with a reduced temperature parameter, so as to optimize the target domain classifier based on the low-entropy soft label and the target domain prediction result, improving the classification accuracy.

[0112] Optionally, the source domain sample video set includes multiple source domain sample videos and the sample labels of each source domain sample video; using the source domain sample video set to train the feature extraction module, the event memory unit, and the source domain classifier, the pre-trained model obtained includes: inputting multiple source domain sample videos into the feature extraction module for feature extraction, and outputting multiple source domain sample features; inputting multiple source domain sample features into the event memory unit for feature enhancement, and outputting multiple source domain event enhanced features; inputting multiple source domain sample features and multiple source domain event enhanced features into the source domain classifier, and outputting the source domain prediction results for multiple source domain sample videos; according to the difference between the source domain prediction results and the sample labels of the source domain sample videos, iteratively adjusting the network parameters of the feature extraction module and the source domain classifier and the event feature prototype in the event memory unit until the third iteration stop condition is met, obtaining the pre-trained model.

[0113] Optionally, the source domain sample video set includes multiple source domain sample videos, which are complete videos collected during laparoscopic surgery. Each source domain sample video is labeled with a sample label, and the sample label is a video-level label. The sample label is either a normal category or an abnormal category. The normal category indicates that no adverse events occurred during the laparoscopic surgery, and the abnormal category indicates that adverse events occurred during the laparoscopic surgery.

[0114] Optionally, the feature extraction module can be a lightweight state-aware temporal encoder (Transformer). The source domain sample videos are input into the feature extraction module for feature extraction, and source domain sample features are output. The source domain sample features are temporal information that captures the global correlation relationship of the normal state and the local dependence relationship of the abnormality in the captured video.

[0115] Optionally, the feature extraction module models the temporal dependence relationship by utilizing context information and focuses on the prior correlations related to the normal or abnormal states.

[0116] Optionally, the source domain sample features are input into the event memory unit for feature enhancement, and source domain event-enhanced features are output. This is achieved through an attention mechanism. The source domain event-enhanced features are features whose influence on the source domain sample features is enhanced based on retrieving and quantifying the event feature prototypes stored in the event memory unit.

[0117] Optionally, the source domain sample features and the source domain event-enhanced features are concatenated, and the concatenated features are input into the source domain classifier for classification prediction, and the source domain prediction result for the source domain sample video is output.

[0118] Optionally, the source domain prediction result is the predicted classification result of the source domain sample video. For example, if the source domain prediction result is 1, it represents the abnormal category.

[0119] Optionally, the third iteration stopping condition can be to maximize the separability between the source domain sample videos with sample labels of the normal category and the source domain sample videos with sample labels of the abnormal category.

[0120] Optionally, the loss function is used to handle the difference between the source domain prediction result and the sample label to obtain a loss value. When the loss value does not meet the third iteration stopping condition, the network parameters of the feature extraction module, the source domain classifier, and the event feature prototypes in the event memory unit are iteratively adjusted until the loss value meets the third iteration stopping condition, and a pre-trained model is obtained.

[0121] Optionally, due to the scarcity of intraoperative adverse events and the dominance of normal events, it is usually difficult to establish a strong correlation between the abnormal state characteristics caused by adverse events and the entire sequence. It mainly relies on the temporal continuity of adjacent time points and similar feature distributions. Based on the global-local attention mechanism, a state-aware temporal encoder for video temporal relationship modeling is constructed. By combining global correlations and abnormal local dependencies, and introducing a learnable Gaussian kernel function time mask to adaptively adjust the focus of attention, the global correlations of normal states and the local dependencies of abnormal states can be effectively learned. This not only significantly improves the accuracy of anomaly detection but also reduces false alarms, thus enhancing the safety and efficiency during the surgical process.

[0122] Optionally, the source domain sample video includes at least one source domain video segment, and the feature extraction module includes a feature extraction sub-module and a feature encoding sub-module; the source domain sample video is input into the feature extraction module for feature extraction, and the output source domain sample features include: at least one source domain video segment is input into the feature extraction sub-module to extract the segment features corresponding to each source domain video segment, and at least one source domain segment feature is output; the source domain segment features are input into the feature encoding sub-module, and the attention mechanism implemented by the Gaussian kernel function and the time mask is used to extract features from the source domain segment features, and the source domain sample features are output.

[0123] Optionally, the source domain sample video is temporal data, which can be divided into source domain video segments composed of 16 frames of images. A source domain sample video includes one or more source domain video segments. Each source domain video segment has a segment-level label, and the sample label of the source domain sample video is determined according to the segment-level label of the source domain video segment.

[0124] Optionally, the feature extraction sub-module can be constructed based on the Inflated 3D ConvNet.

[0125] Optionally, the source domain video segment is input into the feature extraction sub-module to extract the segment features corresponding to each source domain video segment, and output to the source domain segment features. The source domain segment features are segment-level color channel features (Red Green Blue, RGB).

[0126] Optionally, the feature encoding sub-module is constructed based on the self-attention mechanism implemented by the Gaussian kernel function and the time mask.

[0127] Optionally, the feature encoding sub-module is used to extract features from the source domain segment features to obtain the temporal dependence features of the abnormal state and the temporal dependence features of the normal state respectively.

[0128] Optionally, a learnable Gaussian kernel function and a time mask are used to adaptively adjust the focus of attention. Based on the properties of the normal distribution, the weights of the Gaussian kernel smoothly decrease as the time difference increases, which can adaptively adjust the abnormal correlation focus range of different sequence positions and different representation subspaces, and more effectively model the temporal dependence of abnormal state features.

[0129] In one embodiment, the temporal dependence feature of the abnormal state As shown in formula (9):

[0130] (9);

[0131] Where, T represents the sequence length of the source domain segment features, represents the standard deviation parameter of the Gaussian kernel, , represents the activation function, represents the exponential function, H represents the number of attention heads, represents the feature in the source domain segment features, represents the u-th feature in the source domain segment features.

[0132] Optionally, the different association weights of each source domain segment feature are normalized respectively through the activation function (Softmax) to obtain the query, key, normal value, and abnormal value.

[0133] In one embodiment, the query , key , normal value , abnormal value As shown in formula (10):

[0134] ;

[0135] ;

[0136] ;

[0137] (10);

[0138] Where, respectively represent different weight matrix parameters, represents the source domain segment features.

[0139] Optionally, the global dependence in the source domain segment features is mined through the self-attention mechanism in the feature encoding sub-module, enabling the model to adaptively capture the most effective dependence relationship of the normal state.

[0140] In one embodiment, the temporal dependence feature of the normal state As shown in formula (11):

[0141] (11);

[0142] Wherein, represents a query, represents the transpose of a key, represents the dimension of the source domain segment feature.

[0143] Optionally, by concatenating the temporal dependence features in the normal state and the temporal dependence features in the abnormal state, the concatenated features are input into the multi-head attention mechanism (MHA) to capture the associated features in different subspaces. Then, the associated features are combined with the source domain segment features, and the dependence relationship of the state pattern is further strengthened through a multi-layer perceptron (MLP) to generate the source domain sample features.

[0144] In one embodiment, the source domain sample features are as shown in formula (12):

[0145] (12);

[0146] Wherein, represents the multi-head attention mechanism, represents the multi-layer perceptron, represents the temporal dependence features in the abnormal state, represents an outlier, represents the temporal dependence features in the normal state, represents a normal value, represents the source domain segment feature, represents the concatenation operation.

[0147] Optionally, by adopting a weakly supervised multi-instance learning framework, the detection of intraoperative adverse events is defined as a multi-instance learning task (splitting the source domain sample video into multiple source domain video segments for prediction), which optimizes the data utilization, reduces the dependence on large-scale labeled data, not only enhances the system's rapid response ability to complex surgical environments, but also improves the prediction accuracy, effectively improving the surgical safety and efficiency.

[0148] Optionally, the event memory unit includes a normal event memory subunit and an abnormal event memory subunit; the normal event memory subunit includes a normal event feature prototype, and the abnormal event memory subunit includes an abnormal event feature prototype; further included are: respectively performing similarity matching between each source domain sample feature and the normal event feature prototype and the abnormal event feature prototype, and outputting the event matching result of each source domain sample feature; determining an event loss value for the event memory unit according to the event loss function and the event matching result; using the event loss value to iteratively adjust the event feature prototype in the event memory unit until the event loss value meets the fourth iteration stop condition.

[0149] Optionally, the event memory unit includes multiple normal event memory subunits and multiple abnormal event memory subunits, and the normal event memory subunit is a memory bank for storing normal event feature prototypes. For example, the normal event memory subunit stores a set of different normal state features in normal events.

[0150] Optionally, the abnormal event memory subunit is a memory bank for storing abnormal event feature prototypes. For example, the abnormal event memory subunit stores a set of different abnormal state features in adverse events.

[0151] Optionally, calculate the similarity between the source domain sample feature and each normal event memory subunit to obtain multiple similarities, select the similarities of the K source domain scenario memory subunits most relevant to the target domain sample feature from the multiple similarities, and determine the average value of the K similarities as the result of the similarity matching between the target domain sample feature and the source domain scenario feature prototype.

[0152] Optionally, respectively perform similarity matching between each source domain sample feature and the normal event feature prototype and the abnormal event feature prototype, and output the event matching result of each source domain sample feature. The event matching result includes the result of the similarity matching between the source domain sample feature and the normal event feature prototype and the result of the similarity matching between the source domain sample feature and the abnormal event feature prototype.

[0153] Optionally, calculate the event matching result using the event loss function to obtain an event loss value for the event memory unit.

[0154] Optionally, the source domain sample feature includes normal features and abnormal features.

[0155] Optionally, the optimization objective of the fourth iteration stop condition is to maximize the result of the similarity matching between the normal feature and the normal event feature prototype, and the result of the similarity matching between the abnormal feature and the abnormal event feature prototype.

[0156] Optionally, in the case where the event loss value does not satisfy the fourth iteration stop condition, iteratively adjust the event feature prototype in the event memory unit until the event loss value satisfies the fourth iteration stop condition, and obtain the event memory unit.

[0157] Optionally, under the distribution shift of different surgical scenarios, due to the significant differences between the source domain sample features and the target domain sample features, the feature extraction of adverse events is not reliable enough. Intraoperative adverse events have cross-domain consistency. Therefore, decouple the domain-specific scenario information from the event memory unit to construct a scenario memory unit, and only retain the feature patterns related to abnormal or normal states.

[0158] Optionally, multiple source domain sample features include at least one source domain positive sample feature and at least one source domain negative sample feature. The event matching result includes a normal similarity and an abnormal similarity. The normal similarity represents the similarity between the source domain sample feature and the normal event feature prototype, and the abnormal similarity represents the similarity between the source domain sample feature and the abnormal event feature prototype; according to the normal similarity, determine a preset number of source domain positive sample features from at least one source domain positive sample feature to obtain the anchor sample feature; according to the normal similarity, determine a preset number of source domain negative sample features from at least one source domain negative sample feature to obtain the first sample feature; according to the abnormal similarity, determine a preset number of source domain negative sample features from at least one source domain negative sample feature to obtain the second sample feature; according to the preset loss function, the first distance between the first sample feature and the anchor sample feature, and the second distance between the second sample feature and the anchor sample feature, determine the triplet loss value; use the triplet loss value to iteratively adjust the event feature prototype in the event memory unit until the triplet loss value satisfies the fifth iteration stop condition.

[0159] Optionally, the source domain sample video set includes multiple groups of video sets, and each group of video sets consists of a source domain sample video marked with a normal category label and a source domain sample video marked with an abnormal category label.

[0160] Optionally, perform feature extraction on the source domain sample video marked with a normal category label to obtain the source domain positive sample feature; perform feature extraction on the source domain sample video marked with an abnormal category label to obtain the source domain negative sample feature.

[0161] Optionally, the normal similarity includes the similarity between the source domain positive sample feature and the normal event feature prototype and the similarity between the source domain negative sample feature and the normal event feature prototype.

[0162] Optionally, the abnormal similarity includes the similarity between the source domain positive sample feature and the abnormal event feature prototype and the similarity between the source domain negative sample feature and the abnormal event feature prototype.

[0163] For example, the event memory unit includes multiple normal event memory sub-units, calculates the similarity between the source domain positive sample features and the normal event feature prototypes in each normal event memory sub-unit to obtain multiple similarities, sorts the multiple similarities in descending order, selects the K similarities from the 1st position to the Kth position. The normal event feature prototypes in the K normal event memory sub-units corresponding to the K similarities are the most relevant to the source domain positive sample features, and the average value of the K similarities is determined as the normal similarity between the source domain positive sample features and the normal event feature prototypes.

[0164] In one embodiment, the normal similarity between the source domain positive sample features and the normal event feature prototypes is as shown in formula (13):

[0165] (13);

[0166] wherein, represents the kth similarity, represents the top K similarities.

[0167] In one embodiment, the event loss function is as shown in formula (14):

[0168] (14);

[0169] wherein, represents the event loss value, represents the Binary CrossEntropy loss function, represents the normal similarity between the source domain positive sample features and the normal event feature prototypes, represents the abnormal similarity between the source domain positive sample features and the abnormal event feature prototypes, represents the abnormal similarity between the source domain negative sample features and the abnormal event feature prototypes, represents the normal similarity between the source domain negative sample features and the normal event feature prototypes, represents a all-ones vector of length T, represents a all-zeros vector of length T.

[0170] Optionally, the first two loss terms in formula (13) use the complete label information of the normal state sequence for full supervision learning, and the last two loss terms utilize the weak label characteristics in multi-instance learning to focus on the key segments in the abnormal state sequence to update the normal event memory sub-units and the abnormal event memory sub-units.

[0171] Optionally, since the source domain negative sample features are outliers in the normal distribution, some abnormal state features activate the abnormal event memory subunit, and some normal state features activate the normal event memory subunit. For example, the source domain negative sample features are [0, 0, 0, 0, 1, 1, 1, 0, 0], and the source domain positive sample features are [0, 0, 0, 0, 0, 0, 0, 0, 0].

[0172] Optionally, according to the normal similarity between each source domain positive sample feature and the normal event feature prototype, the most relevant preset number of normal similarities are selected from multiple normal similarities, and the source domain positive sample features corresponding to the preset number of normal similarities are averaged to obtain the anchor sample features. For example, the preset number can be 1.

[0173] Optionally, according to the normal similarity between each source domain negative sample feature and the normal event feature prototype, the most relevant preset number of normal similarities are selected from multiple normal similarities, and the source domain negative sample features corresponding to the preset number of normal similarities are averaged to obtain the first sample features. For example, the preset number can be an integer greater than 1.

[0174] Optionally, according to the abnormal similarity between each source domain negative sample feature and the abnormal event feature prototype, the most relevant preset number of abnormal similarities are selected from multiple abnormal similarities, and the source domain negative sample features corresponding to the preset number of abnormal similarities are averaged to obtain the second sample features. For example, the preset number can be 1.

[0175] Optionally, the anchor sample features are the features closest to the normal state extracted from the source domain sample videos marked with normal class labels, the first sample features are the features closest to the normal state extracted from the source domain sample videos marked with abnormal class labels, and the second sample features are the features closest to the abnormal state extracted from the source domain sample videos marked with abnormal class labels.

[0176] In one embodiment, the preset loss function is shown in formula (15):

[0177] (15);

[0178] Wherein, represents the triplet loss value, represents the maximum value function, represents the anchor sample features, represents the first sample features, represents the second sample features, represents the square of the Euclidean distance, is the margin parameter, which is used to control the distance gap between positive and negative samples.

[0179] For example, the margin parameter is 1.

[0180] Optionally, the optimization objective of the fifth iteration stop condition is to minimize the distance between the first sample feature and the anchor sample feature in the feature space, while maximizing the distance between the second sample feature and the anchor sample feature.

[0181] Optionally, in the case where the triplet loss value does not satisfy the fifth iteration stop condition, iteratively adjust the event feature prototype in the event memory unit until the triplet loss value satisfies the fifth iteration stop condition.

[0182] Optionally, both the source domain positive sample feature and the source domain negative sample feature are segment-level features, and the loss function is used to process the source domain prediction results corresponding to the segment-level features to obtain the segment-level loss value.

[0183] In one embodiment, the segment-level loss value is as shown in formula (16):

[0184] (16);

[0185] where represents the source domain prediction result corresponding to the q-th source domain positive sample feature, represents the source domain prediction result corresponding to the q-th source domain negative sample feature, represents the preset number of selected prediction results, represents the margin parameter.

[0186] Optionally, by constraining the highest abnormal source domain prediction result in the source domain positive sample feature to be higher than the source domain negative sample feature through the segment-level loss value, the separability between positive and negative samples is maximized, and the interval between positive and negative subsets is maintained to implicitly constrain the model to learn more discriminative features.

[0187] In one embodiment, the total loss value of the first stage is as shown in formula (17):

[0188] (17);

[0189] where represents the segment-level loss value, represents the event loss value, represents the triplet loss value, are all hyperparameters for balancing each loss term.

[0190] Optionally, iteratively adjust the network parameters of the feature extraction module, the source domain classifier, and the event feature prototype in the event memory unit according to the total loss value of the first stage until the iteration stop condition is satisfied to obtain the pre-trained model.

[0191] Figure 3 The figure shows an example diagram of a method for training an event monitoring model according to an embodiment of the present invention.

[0192] As Figure 3 shown, the training method of this embodiment includes a first stage 310 and a second stage 320. The first stage performs pre-training based on multi-instance learning, specifically including: training a feature extraction sub-module, a feature encoding sub-module, an event memory unit, and a source domain classifier using a source domain sample video set. The source domain sample video set includes source domain sample videos labeled with normal class labels and source domain sample videos labeled with abnormal class labels. The source domain sample video set is input into the feature extraction sub-module to extract features from at least one source domain video segment in the source domain sample videos, and source domain segment features are output; the source domain segment features are input into the feature encoding sub-module, and an attention mechanism implemented by a Gaussian kernel function and a time mask is used to extract features from the source domain segment features to obtain queries , keys , normal values , abnormal values , temporal dependence features of the normal state, and temporal dependence features of the abnormal state. By concatenating the queries , keys , normal values , abnormal values , temporal dependence features of the normal state, and temporal dependence features of the abnormal state, source domain sample features are output; the source domain sample features are input into the event memory unit for feature enhancement, and source domain event enhanced features are output, where the event memory unit includes a normal event memory sub-unit and an abnormal event memory sub-unit; the source domain sample features and the source domain event enhanced features are concatenated to obtain a first concatenated feature, and the first concatenated feature is input into the source domain classifier to output a source domain prediction result for the source domain sample video ; according to the difference between the source domain prediction result and the sample label of the source domain sample video, the network parameters of the feature extraction sub-module, the source domain classifier, and the event feature prototypes in the event memory unit are iteratively adjusted until the third iteration stop condition is met to obtain a pre-trained model, and the sharpening algorithm is used to adjust the source domain prediction result (i.e., the pseudo-label) to generate a low-entropy soft label , and the training of the first stage is completed.

[0193] Self-training for domain adaptation in the second stage specifically includes: using the target domain sample videos to train the event memory unit, the scene memory unit, and the target domain classifier. Specifically, it includes: inputting at least one target domain video segment in the unlabeled target domain sample videos into the feature extraction sub-module of the pre-trained model for feature extraction, and outputting the target domain sample features; after splicing the target domain sample features and the source domain sample features in the first stage, inputting them into the scene memory unit and the event memory unit of the pre-trained model for feature enhancement, and respectively outputting the hybrid domain event enhanced features and the scene enhanced features, where the scene memory unit includes a source domain scene memory sub-unit and a target domain scene memory sub-unit; splicing the hybrid domain event enhanced features, the target domain sample features, and the source domain sample features to obtain the second splicing feature, inputting it into the target domain classifier, and outputting the target domain prediction result for the target domain sample videos ; According to the target domain prediction result and the low-entropy soft labels the differences between them are used to iteratively adjust the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier until the first iteration stop condition is met, and the trained event monitoring model is obtained. It should be noted that the trained event monitoring module includes a feature extraction sub-module, a feature encoding sub-module, an event memory unit, and a target domain classifier.

[0194] Optionally, in the experimental stage of the present invention, to achieve efficient detection and robust modeling of abnormal events in surgical videos, a standardized data preprocessing and model training process is designed and implemented. During the preprocessing process, an inflated three-dimensional convolutional neural network (Inflated 3D ConvNet, I3D) pre-trained on the surgical stage recognition task is used as the feature extraction module. This network is based on the deep convolutional neural network structure (Inception-V1) and is constructed by expanding the original two-dimensional convolutional kernel into a three-dimensional form, capable of capturing information in both spatial and temporal dimensions simultaneously. The feature extraction module includes multiple continuously stacked three-dimensional convolutional layers (convolution kernel size 3×3×3), max pooling layers, 10 convolutional neural layers, and a global average pooling layer, with powerful spatio-temporal feature encoding capabilities. During the feature extraction process, the surgical videos are segmented in units of 16 frames, and the corresponding high-dimensional spatio-temporal feature representations are extracted from each segment clip. This process uses exactly the same configuration and parameter settings in the source domain (such as cholecystectomy) and the target domain (such as prostate and kidney partial resection) to ensure the consistency of feature representations and improve the cross-domain generalization ability of the model between different surgical scene types.

[0195] Optionally, during the event monitoring model training phase, the length of all input feature sequences is unified to 200 frames, and the Adaptive Moment Estimation optimizer is used for training. The initial learning rate is set to 0.0001, the total number of training iterations is 3000 times, and the batch size is 64. Two layers of Transformer encoders are adopted, with the hidden state dimension set to 512 and the number of attention heads set to 4 for each layer. To enhance the temporal memory and state modeling capabilities of the event monitoring model, an event memory unit with the number of event feature prototypes M = 60 is introduced. Among them, the value of K in all similarity matching operations is set to ⌊M / 16⌋ + 1. In terms of the loss function design, reasonable hyperparameters are set for different sub-modules to ensure training stability and model generalization ability. The experimental platform is implemented in a computing environment configured with NVIDIA RTX 3090 GPUs (24GB of video memory) to ensure sufficient modeling of temporal features in surgical videos and efficient detection of adverse events under high-performance conditions.

[0196] Optionally, three key metrics are used to comprehensively evaluate the performance of the event monitoring model during the training phase, namely: frame-level AUC (area under the receiver operating characteristic curve) to measure the overall detection ability, AP (average precision) to evaluate the positive detection effect under class imbalance conditions, and FAR (false alarm rate) to measure the occurrence frequency of unnecessary alarms. The higher the AUC / AP and the lower the FAR, the better the model performance. The multi-metric evaluation strategy combines the complementary advantages of AUC and AP, improving the detection robustness for rare abnormal events. At the same time, FAR provides a practical reference with actual clinical significance. On the source domain sample video set, the event monitoring model achieved AUC 92.85%, AP 70.28%, and FAR 1.40% in the first-stage pre-training; after introducing pseudo-label self-training in the second stage, the AUC was 92.48%, the AP increased to 72.06%, and the FAR decreased to 0.80%. On the target domain sample videos, only AUC 86.22%, AP 74.27%, and FAR 0.60% were achieved in the first stage; in the second stage, it was further improved to AUC 92.65%, AP 85.38%, and still maintained a low FAR of 1.71%. The above results indicate that the present invention has excellent abnormal detection ability, cross-domain adaptability, and practical application value.

[0197] Optionally, the event monitoring model is trained in two stages. In the first stage, the event monitoring model performs multi-instance learning on the source domain sample videos with segment-level labels, differentiates the normal event feature prototypes from the abnormal event feature prototypes, and trains the source domain classifier to generate segment-level pseudo-labels. In the second stage, the knowledge of the pre-trained model is applied in parallel to the source domain sample video set and the target domain sample video set. Self-training is used to refine the low-entropy soft labels using the segment-level pseudo-labels, thereby optimizing the target domain classifier based on the low-entropy soft labels and the target domain prediction results. At the same time, the internal structure of the unlabeled target domain sample video set is mined to achieve domain adaptation and improve the event prediction accuracy.

[0198] Figure 4 The flowchart of the abnormal detection method for surgical videos according to an embodiment of the present invention is shown.

[0199] As Figure 4 shown, the abnormal detection method 400 for surgical videos includes operations S410 to S430.

[0200] In operation S410, the endoscopic surgical video is acquired in real time.

[0201] In operation S420, the M video frames are processed by using the feature extraction module, the event memory unit, and the target domain classifier in the trained event monitoring model to obtain the prediction value of each video frame.

[0202] In operation S430, the video frames with prediction values greater than a preset value are determined as the video frames for the target event.

[0203] Optionally, the endoscopic surgical video includes M video frames, M≥1, and M is a positive integer.

[0204] Optionally, the endoscopic surgical video during cholecystectomy is acquired in real time.

[0205] Optionally, the event monitoring model is trained by using the above training method. The M video frames are sequentially input into the feature extraction module, the event memory unit, and the target domain classifier in the trained event monitoring model to obtain the prediction value of each video frame.

[0206] Optionally, the endoscopic surgical video is processed by using the trained event monitoring model, and a prediction curve is output. The prediction curve is composed of the time points of the video frames and the prediction values.

[0207] Optionally, the prediction value represents the probability that the video frame is in an abnormal situation. For example, a prediction value of 0.6 represents that the probability that this video frame is in an abnormal situation is 0.6.

[0208] Optionally, the target event may be an intraoperative adverse event.

[0209] For example, the preset value can be 0.5. A video frame with a predicted value greater than 0.5 is determined as a video frame for the target event; a video frame with a predicted value less than or equal to 0.5 is determined as a video frame in which the target event has not occurred.

[0210] Optionally, in complex surgical scenarios such as cholecystectomy, prostatectomy, and partial nephrectomy, it is possible to effectively distinguish the normal and abnormal states at each time point during the surgery based on the event monitoring model, significantly improving the accuracy of anomaly detection and reducing false alarms, thereby enhancing the safety and efficiency during the surgical process.

[0211] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or by a combination of dedicated hardware and computer instructions. Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0212] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A training method for an event monitoring model, characterized in that The event monitoring model includes a feature extraction module, an event memory unit, a source domain classifier, a scenario memory unit, and a target domain classifier. The event memory unit includes an event feature prototype, and the scenario memory unit includes a scenario feature prototype; The method includes: Using a source domain sample video set to train the feature extraction module, the event memory unit, and the source domain classifier to obtain a pre-trained model; Inputting a target domain sample video into the feature extraction module of the pre-trained model for feature extraction, and outputting target domain sample features; Inputting the target domain sample features and the source domain sample features into the scenario memory unit and the event memory unit of the pre-trained model for feature enhancement, and respectively outputting hybrid domain event enhanced features and scenario enhanced features. Among them, the source domain sample features are obtained by using the feature extraction module of the pre-trained model to extract features from the source domain sample videos in the source domain sample video set; Inputting the hybrid domain event enhanced features, the target domain sample features, and the source domain sample features into the target domain classifier, and outputting a target domain prediction result for the target domain sample video; According to the difference between the target domain prediction result and the pseudo-label for the target domain sample video, iteratively adjust the event feature prototype in the event memory unit of the pre-trained model, the scenario feature prototype in the scenario memory unit, and the network parameters of the target domain classifier until the first iteration stop condition is met, and obtain a trained event monitoring model. Among them, the pseudo-label is the prediction result output after the source domain sample video is processed by the pre-trained model.

2. The method according to claim 1, wherein The scenario memory unit includes a source domain scenario memory subunit and a target domain scenario memory subunit. The source domain scenario memory subunit includes a source domain scenario feature prototype, and the target domain scenario memory subunit includes a target domain scenario feature prototype; The method further includes: Performing similarity matching on the target domain sample features with the source domain scenario feature prototype and the target domain scenario feature prototype respectively, and outputting a first scenario matching result for the target domain sample features; Performing similarity matching on the source domain sample features with the source domain scenario feature prototype and the target domain scenario feature prototype respectively, and outputting a second scenario matching result for the source domain sample features; Determining a scenario loss value for the scenario memory unit according to a scenario loss function, the first scenario matching result, and the second scenario matching result; Using the scenario loss value to iteratively adjust the scenario feature prototype in the scenario memory unit until the scenario loss value meets the second iteration stop condition.

3. The method according to claim 2, wherein The method further includes: Determining an orthogonal loss according to an orthogonal loss function, the hybrid domain event enhanced features, and the scenario enhanced features; Using the orthogonal loss to iteratively adjust the event feature prototype in the event memory unit and the scenario feature prototype in the scenario memory unit until the orthogonal loss meets the third iteration stop condition.

4. The method according to claim 3, characterized in that, Iteratively adjusting the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier according to the difference between the target domain prediction result and the pseudo-label for the target domain sample video includes: Determining a first loss value according to a first loss function, the target domain prediction result, and the pseudo-label; Determining a total loss value according to the first loss value, the scene loss value, and the orthogonality loss; Iteratively adjusting the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier according to the total loss value.

5. The method according to claim 1, wherein The method further includes: Adjusting the pseudo-label using a sharpening algorithm to generate a low-entropy soft label; The iteratively adjusting the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier according to the difference between the target domain prediction result and the pseudo-label for the target domain sample video includes: Iteratively adjusting the event feature prototypes in the event memory unit of the pre-trained model, the scene feature prototypes in the scene memory unit, and the network parameters of the target domain classifier according to the difference between the target domain prediction result and the low-entropy soft label for the target domain sample video.

6. The method according to claim 1, characterized in that, The source domain sample video set includes multiple source domain sample videos and the sample labels of each source domain sample video; The training the feature extraction module, the event memory unit, and the source domain classifier using the source domain sample video set to obtain a pre-trained model includes: Inputting multiple source domain sample videos into the feature extraction module for feature extraction to output multiple source domain sample features; Inputting the multiple source domain sample features into the event memory unit for feature enhancement to output multiple source domain event enhanced features; Inputting the multiple source domain sample features and the multiple source domain event enhanced features into the source domain classifier to output source domain prediction results for the multiple source domain sample videos; Iteratively adjusting the network parameters of the feature extraction module, the source domain classifier, and the event feature prototypes in the event memory unit according to the difference between the source domain prediction result and the sample labels of the source domain sample videos until a third iteration stop condition is met to obtain a pre-trained model.

7. The method according to claim 6, characterized in that, The source domain sample video includes at least one source domain video segment, and the feature extraction module includes a feature extraction sub-module and a feature encoding sub-module; The inputting multiple source domain sample videos into the feature extraction module for feature extraction to output multiple source domain sample features includes: Inputting the at least one source domain video segment into the feature extraction sub-module to extract segment features corresponding to each source domain video segment and output at least one source domain segment feature; Inputting the source domain segment features into the feature encoding sub-module and using an attention mechanism implemented by a Gaussian kernel function and a temporal mask to perform feature extraction on the source domain segment features and output the source domain sample features.

8. The method according to claim 7, wherein The event memory unit includes a normal event memory subunit and an abnormal event memory subunit; the normal event memory subunit includes a normal event feature prototype, and the abnormal event memory subunit includes an abnormal event feature prototype; The method further includes: Performing similarity matching on each of the source domain sample features with the normal event feature prototype and the abnormal event feature prototype respectively, and outputting the event matching result of each source domain sample feature; Determining an event loss value for the event memory unit according to an event loss function and the event matching result; Using the event loss value to iteratively adjust the event feature prototype in the event memory unit until the event loss value satisfies a fourth iteration stop condition.

9. The method according to claim 8, characterized in that, The multiple source domain sample features include at least one source domain positive sample feature and at least one source domain negative sample feature, the event matching result includes a normal similarity and an abnormal similarity, the normal similarity characterizes the similarity between the source domain sample feature and the normal event feature prototype, and the abnormal similarity characterizes the similarity between the source domain sample feature and the abnormal event feature prototype; The method further includes: Determining a preset number of source domain positive sample features from the at least one source domain positive sample feature according to the normal similarity to obtain anchor sample features; Determining a preset number of source domain negative sample features from the at least one source domain negative sample feature according to the normal similarity to obtain first sample features; Determining a preset number of source domain negative sample features from the at least one source domain negative sample feature according to the abnormal similarity to obtain second sample features; Determining a triplet loss value according to a preset loss function, a first distance between the first sample feature and the anchor sample feature, and a second distance between the second sample feature and the anchor sample feature; Using the triplet loss value to iteratively adjust the event feature prototype in the event memory unit until the triplet loss value satisfies a fifth iteration stop condition.

10. An abnormal detection method for surgical videos, characterized in that, The method includes: Obtaining a laparoscopic surgery video in real time, where the laparoscopic surgery video includes M video frames, M≥1 and M is a positive integer; Processing the M video frames by using a feature extraction module, an event memory unit, and a target domain classifier in a trained event monitoring model, and obtaining a prediction value for each video frame, where the prediction value characterizes the probability that the video frame is an abnormal situation, and the event monitoring model is trained by using the training method according to any one of claims 1 to 9; and Determining the video frames with the prediction value greater than a preset value as the video frames for the target event.

Citation Information

Patent Citations

  • Robust field adaptive image learning method based on self-training noise label correction

    CN114283287A

  • Abnormal event detection algorithm based on unsupervised domain self-adaption

    CN116385935A

  • Deep learning-based spine endoscopic surgery real-time auxiliary method and system

    CN116687561A

  • Self-adaptive bearing fault classification method and system in passive field

    CN118094367A

  • Devices, systems, methods, and media for domain adaptation using hybrid learning

    US20230082899A1

Cited By

  • Equipment energy consumption state monitoring method based on electric signals of multi-energy complementary energy supply system

    CN120994973A