Smil-based cholecystectomy cvs assessment system, method and apparatus
The SMIL-based CVS assessment system for cholecystectomy utilizes self-supervised distillation and multi-instance learning to adaptively extract keyframes and extract global context and local instance features, solving the problem of poor adaptability in existing technologies. This achieves accurate CVS assessment under different conditions, improving the safety and operational stability of cholecystectomy.
Patent Information
- Application Number
- CN202511434492.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-10-09
AI Technical Summary
Existing technologies rely on fixed inputs in cholecystectomy, resulting in poor method adaptability and insufficient generalization ability. They cannot effectively adapt to different hospitals, surgeon styles, and equipment data, and lack the ability to continuously model the dynamic evolution of CVS.
A CVS evaluation system based on SMIL is adopted. The system adaptively extracts keyframes through the image frame extraction module, and combines a label-free self-supervised distillation architecture and a student Transformer with a multi-instance learning architecture to extract global context and local instance features. The system then performs feature fusion and evaluation to achieve dynamic evaluation of the CVS standard.
This improves the system's adaptability and generalization ability, enabling accurate evaluation of CVS standards under different conditions, reducing reliance on high-quality annotations, and enhancing the safety and operational stability of cholecystectomy.
Smart Images

Figure CN120894734B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video processing technology, and more specifically, to a SMIL-based system, method, and device for evaluating CVS during cholecystectomy. Background Technology
[0002] Minimally invasive surgery typically involves making a small incision on the patient's skin and performing the procedure under the guidance of imaging equipment such as endoscopes. Compared to traditional surgery, minimally invasive surgery is characterized by smaller incisions, less trauma, faster recovery, and less pain, and has been widely used in clinical practice, becoming an important diagnostic and treatment method in modern medicine. Cholecystectomy (LC) is currently the gold standard surgical procedure for treating gallstones, cholecystitis, and other diseases. As one of the most widely performed minimally invasive surgeries globally, it has become an important application scenario in the field of general surgery.
[0003] Despite the numerous advantages of biliary tract injury (LC), the risk of intraoperative bile duct injury remains, seriously impacting patients' postoperative health. Studies show that approximately 97% of bile duct injuries originate from intraoperative visual cognitive errors, making the prevention of bile duct injury a significant challenge in current LC procedures. To reduce intraoperative cognitive errors, the Critical View of Safety (CVS) standard has become an internationally recognized standard for intraoperative anatomical verification. CVS requires three criteria to be met before gallbladder transection to clearly identify key anatomical structures.
[0004] However, existing technologies often rely on static images or continuous fixed frame sequences as fixed inputs, which are poorly adaptable to multi-source heterogeneous surgical data such as low frame rates and non-continuous images, and are prone to performance degradation, resulting in insufficient generalization ability of the model in different hospitals, different surgeon styles, and different equipment data. Summary of the Invention
[0005] The problem that this invention aims to solve is that existing technologies use fixed inputs, resulting in poor adaptability and insufficient generalization ability.
[0006] To address the aforementioned problems, in a first aspect, the present invention provides a SMIL-based CVS assessment system for cholecystectomy, comprising:
[0007] The image frame extraction module is used to segment the read cholecystectomy video, and adaptively extract keyframes according to the content of each video segment to obtain image frames.
[0008] The global feature extraction module is used to input image frames into the student Transformer in the unlabeled self-supervised distillation architecture within the SMIL network architecture to extract global contextual features. The SMIL network architecture includes an unlabeled self-supervised distillation architecture and a multi-instance learning architecture. The unlabeled self-supervised distillation architecture includes a student Transformer, and the multi-instance learning architecture includes a student Transformer shared with the unlabeled self-supervised distillation architecture and the MIL Attention module.
[0009] The instance feature extraction module is used to divide the image frame into multiple image blocks and input them into the student Transformer, extract the instance features of each image block, and form an instance feature matrix.
[0010] The local feature aggregation module is used to input the instance feature matrix into the MIL Attention module for weighted aggregation to obtain local instance features;
[0011] The feature fusion module is used to fuse features based on global context features, local instance features, and dynamic weights to obtain specific fused features corresponding to different CVS standards.
[0012] The evaluation and prediction module is used to input the specific fusion features corresponding to different CVS standards into the corresponding linear classifiers to obtain the CVS standard evaluation results for each image frame.
[0013] Secondly, the present invention also provides a SMIL-based method for assessing CVS after cholecystectomy, comprising:
[0014] The cholecystectomy video was segmented, and keyframes were adaptively extracted based on the content of each video segment to obtain image frames.
[0015] Image frames are input into the student Transformer in the unlabeled self-supervised distillation architecture within the SMIL network architecture to extract global contextual features. The SMIL network architecture includes an unlabeled self-supervised distillation architecture and a multi-instance learning architecture. The unlabeled self-supervised distillation architecture includes a student Transformer, and the multi-instance learning architecture includes a student Transformer shared with the unlabeled self-supervised distillation architecture and the MIL Attention module.
[0016] The image frame is divided into multiple image blocks and input into the student Transformer. The instance features of each image block are extracted to form an instance feature matrix.
[0017] The instance feature matrix is input into the MIL Attention module for weighted aggregation to obtain local instance features;
[0018] Feature fusion is performed based on global context features, local instance features, and dynamic weights to obtain specific fused features corresponding to different CVS standards;
[0019] The specific fusion features corresponding to different CVS standards are input into the corresponding linear classifiers to obtain the CVS standard evaluation results for each image frame.
[0020] Thirdly, the present invention provides an electronic device, including a memory and a processor;
[0021] The memory is used to store computer programs;
[0022] The processor is configured to, when executing the computer program, implement the SMIL-based CVS assessment method for cholecystectomy as described in the first aspect.
[0023] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the SMIL-based CVS assessment method for cholecystectomy as described in the first aspect.
[0024] This invention provides a SMIL-based system, method, and device for assessing CVS after cholecystectomy. Compared with existing technologies, it has the following advantages:
[0025] The image frame extraction module 10 segments the read cholecystectomy video and adaptively extracts keyframes based on the content of each video segment to obtain continuous or discontinuous image frames, allowing for flexible data extraction. The global feature extraction module 20 inputs the image frames into the student Transformer in the unlabeled self-supervised distillation architecture within the SMIL network architecture to extract global contextual features. The SMIL network architecture includes an unlabeled self-supervised distillation architecture and a multi-instance learning architecture. The unlabeled self-supervised distillation architecture includes a student Transformer, and the multi-instance learning architecture includes a student Transformer shared with the unlabeled self-supervised distillation architecture and the MIL Attention module. The unlabeled self-supervised distillation architecture can pre-train the student Transformer using minimally invasive video data without video labels to learn the spatiotemporal dynamic features in the cholecystectomy video. Using the pre-trained student Transformer model for multi-branch feature extraction in the subsequent SMIL network architecture can effectively utilize the spatiotemporal information of cholecystectomy videos to construct the CVS standard dynamic formation process. This allows the system to focus on key anatomical structures even without requiring fine segmentation and annotation, reducing the requirements for input data and improving the system's adaptability and generalization ability. The instance feature extraction module 30 divides the image frame into multiple image blocks and inputs them into the student Transformer to extract instance features from each image block, forming an instance feature matrix. Then, the local feature aggregation module 40 inputs the instance feature matrix into the MIL in the multi-instance learning architecture. The Attention module performs weighted aggregation to obtain local instance features of key anatomical structures after aggregation. The Feature Fusion module 50 performs feature fusion based on global context features, local instance features, and dynamic weights to obtain specific fusion features corresponding to different CVS standards, improving the model's ability to extract key region features of complex anatomical structures, thereby enhancing the system's dynamic understanding capability. The Evaluation and Prediction module inputs the specific fusion features corresponding to different CVS standards into the corresponding linear classifiers to obtain the CVS standard evaluation results for each image frame, which are used to evaluate the quality of key safety views. This is beneficial for improving the operation of cholecystectomy, facilitating the system's application in different hospitals, and adapting to surgeons with different styles or different equipment. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1A schematic diagram of a SMIL-based CVS assessment system for cholecystectomy is provided in an embodiment of the present invention.
[0028] Figure 2 This is a schematic diagram of a label-free self-supervised distillation architecture provided in an embodiment of the present invention;
[0029] Figure 3 This is a schematic diagram of a multi-instance learning architecture provided in an embodiment of the present invention;
[0030] Figure 4 This is a schematic diagram of a feature fusion stage provided in an embodiment of the present invention;
[0031] Figure 5 This is a schematic diagram of the CVS standard evaluation stage provided in an embodiment of the present invention;
[0032] Figure 6 A schematic flowchart of a SMIL-based cholecystectomy CVS assessment method provided in an embodiment of the present invention;
[0033] Figure 7 A schematic diagram showing the comparative experimental results of various evaluation methods provided in the embodiments of the present invention;
[0034] Figure 8 A schematic diagram of the ablation experiment results provided in an embodiment of the present invention;
[0035] Figure 9 A line graph comparing the indicators of ablation experiment results provided in an embodiment of the present invention;
[0036] Figure 10 This is a schematic diagram illustrating the experimental results of multiple indicators of SMIL provided in an embodiment of the present invention. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application are described clearly and completely. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0038] Besides relying on fixed inputs, which leads to poor adaptability, existing technologies also suffer from insufficient task focus and limited CVS evaluation results. Current technologies primarily evaluate anatomical structure recognition, surgical stage analysis, and the overall surgical process within surgical videos, using the CVS standard as a local key evaluation indicator rather than as an independent task for in-depth optimization modeling. Existing technologies often depend on static images or fixed frame rate sequences as input, lacking the ability to continuously model the dynamic evolution of CVS, resulting in insufficient stability and accuracy of CVS evaluation results in complex intraoperative scenarios. Furthermore, existing technologies heavily rely on frame-by-frame detailed annotation of large-scale surgical videos during model training, including information such as anatomical structures, CVS scores, and surgical stages. This annotation process consumes significant expert manpower and exhibits inconsistencies, limiting the quality and scalability of training data, further impacting the model's generalization ability and application dissemination.
[0039] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0040] like Figure 1 As shown in the embodiment of this application, a SMIL-based CVS assessment system for cholecystectomy includes:
[0041] The image frame extraction module 10 is used to segment the read cholecystectomy video, and adaptively extract key frames according to the content of each video segment to obtain image frames.
[0042] The global feature extraction module 20 is used to input image frames into the student Transformer in the unlabeled self-supervised distillation architecture within the SMIL network architecture to extract global contextual features. The SMIL network architecture includes an unlabeled self-supervised distillation architecture and a multi-instance learning architecture. The unlabeled self-supervised distillation architecture includes a student Transformer, and the multi-instance learning architecture includes a student Transformer shared with the unlabeled self-supervised distillation architecture and the MIL Attention module.
[0043] The instance feature extraction module 30 is used to divide the image frame into multiple image blocks and input them into the student Transformer to extract the instance features of each image block and form an instance feature matrix.
[0044] The local feature aggregation module 40 is used to input the instance feature matrix into the MIL Attention module for weighted aggregation to obtain local instance features.
[0045] The feature fusion module 50 is used to fuse features based on global context features, local instance features and dynamic weights to obtain specific fused features corresponding to different CVS standards.
[0046] The evaluation and prediction module 60 is used to input the specific fusion features corresponding to different CVS standards into the corresponding linear classifiers to obtain the CVS standard evaluation results for each image frame.
[0047] In this optional embodiment, the image frame extraction module 10 segments the read cholecystectomy video and adaptively extracts keyframes according to the content of each video segment to obtain continuous or discontinuous image frames, thus providing flexible data extraction. The global feature extraction module 20 inputs the image frames into the student Transformer in the unlabeled self-supervised distillation architecture within the SMIL network architecture to extract global contextual features. The SMIL network architecture includes an unlabeled self-supervised distillation architecture and a multi-instance learning architecture. The unlabeled self-supervised distillation architecture includes a student Transformer, and the multi-instance learning architecture includes a student Transformer shared with the unlabeled self-supervised distillation architecture and the MIL Attention module. The unlabeled self-supervised distillation architecture can pre-train the student Transformer using minimally invasive video data without video labels to learn the spatiotemporal dynamic features in the cholecystectomy video. Using the pre-trained student Transformer model for multi-branch feature extraction in the subsequent SMIL network architecture can effectively utilize the spatiotemporal information of the cholecystectomy video to construct the CVS standard dynamic formation process. This allows the system to focus on key anatomical structures even without fine segmentation and annotation, reducing the requirements for input data and improving the system's adaptability and generalization ability. The instance feature extraction module 30 divides the image frame into multiple image blocks and inputs them into the student Transformer in the multi-instance learning architecture to extract instance features from each image block, forming an instance feature matrix. Then, the local feature aggregation module 40 inputs the instance feature matrix into the MIL in the multi-instance learning architecture. The Attention module performs weighted aggregation to obtain local instance features of key anatomical structures after aggregation. The Feature Fusion module 50 performs feature fusion based on global context features, local instance features, and dynamic weights to obtain specific fusion features corresponding to different CVS standards, improving the model's ability to extract key region features of complex anatomical structures, thereby enhancing the system's dynamic understanding capability. The Evaluation and Prediction module inputs the specific fusion features corresponding to different CVS standards into the corresponding linear classifiers to obtain the CVS standard evaluation results for each image frame, which are used to evaluate the quality of key safety views. This is beneficial for improving the operation of cholecystectomy, facilitating the system's application in different hospitals, and adapting to surgeons with different styles or different equipment.
[0048] The following is a detailed description of each module of the system.
[0049] First, it should be noted that the Critical Safety View (CVS) criteria require clinicians to confirm three CVS criteria before performing any cross-sectional resection: C1: Clearing fat or connective tissue from the Calot's triangle to obtain an unobstructed view; C2: The lower part of the gallbladder is detached from the liver bed, exposing the lower 1 / 3 of the gallbladder plate; C3: Two and only two tubular structures connecting to the gallbladder are visible. The labels for the CVS criteria are 3D binary vectors: Let... The current image frame of a cholecystectomy video satisfies the i-th CVS criterion, where i = 1, 2, 3. This application proposes a CVS evaluation system for cholecystectomy based on self-supervised distillation and multi-instance learning (SMIL) to achieve automated CVS evaluation of videos during laparoscopic cholecystectomy.
[0050] The image frame extraction module 10 is used to segment the read cholecystectomy video, adaptively extract keyframes based on the content of each video segment, and obtain image frames. This module specifically includes the following:
[0051] Based on a specified time anchor point, the cholecystectomy video before the specified time anchor point is cropped and removed, and abnormal video frames are removed again according to preset indicators to obtain the processed cholecystectomy video.
[0052] Specifically, to focus on CVS-related stages and reduce redundancy and noise, the cholecystectomy video was preprocessed. First, surgical event anchoring and clipping were performed: for example, using "first use of hemostatic clip" as the designated time anchor. Remove the preparation phase prior to this time anchor point, and retain only... The segment is used for subsequent processing ( The video ends at the designated time point for the cholecystectomy procedure, thus focusing on the critical surgical phases requiring CVS evaluation. Then, abnormal frames or segments are removed: based on preset indicators such as the mean / variance of brightness, sharpness (Laplacian variance), and saturation of video frames, indistinguishable video frames such as black screens, severe distortion, heavy blur, or large-area obstruction are detected and deleted; consecutive abnormal short segments are directly removed to ensure the quality of the input video.
[0053] The content of the processed cholecystectomy video is processed, the cholecystectomy video is segmented, and the sampling frequency is adaptively changed according to the content of each video segment to extract key frames, thus obtaining the image data of the original video frames.
[0054] Specifically, content-aware keyframe extraction is performed, and adaptive frame extraction is conducted on the retained video segments: when adjacent frames have high structural similarity, low optical flow amplitude, and stable edge density (anatomical structures are basically static), sampling is performed at a low frequency (e.g., one frame per second); when traction, significant instrument movement, or rapid structural changes occur, the sampling frequency is temporarily increased to a higher frequency (e.g., one frame at 3 fps) to cover transitional segments. This strategy removes a large number of redundant frames while ensuring the integrity of key information.
[0055] For image data Normalization is performed to obtain image frame X.
[0056] Specifically, because cholecystectomy involves rich and continuously changing dynamic anatomical features, effective modeling and extraction of the spatiotemporal features throughout the entire procedure are necessary for accurate assessment of surgical quality. First, the laparoscopic cholecystectomy surgical video is read, segmented, and video image data is extracted. G=3 represents the RGB channels, T represents the number of frames in the segment, and J=K=224 represents the image size after preprocessing and adjustment. Image frames are obtained by normalizing the image data of the original video. The image dataset is then stitched together to restore the video data format of the dataset. This system can classify videos and perform CVS evaluation tasks on datasets of static images.
[0057] Image Frame ,
[0058] in, , Each frame of the image is divided into Non-overlapping patches, patch size .
[0059] like Figure 3 As shown, the global feature extraction module 20 is used to input image frames into the student Transformer in the unlabeled self-supervised distillation architecture within the pre-trained SMIL network architecture, and extract the main semantic identifier as global context features. Among them, such as Figure 2 and Figure 3 As shown, the SMIL network architecture includes an unlabeled self-supervised distillation architecture and a multi-instance learning architecture. The unlabeled self-supervised distillation architecture includes a student Transformer and a teacher Transformer. The multi-instance learning architecture includes a student Transformer shared with the unlabeled self-supervised distillation architecture and a MILAttention module connected to the student Transformer.
[0060] Specifically, the unlabeled self-supervised distillation DINO framework is adopted, and teacher and student Transformers are designed for video self-supervised pre-training. Two types of images are designed as inputs: a global image, covering the video segment to capture overall spatiotemporal information; and a local image, cropping local image regions to enhance the ability to learn local features.
[0061] The SMIL-based CVS assessment system for cholecystectomy also includes a pre-training module 70 for video self-supervised pre-training of the unlabeled self-supervised distillation architecture. The video self-supervised pre-training process is as follows.
[0062] 1. Input the global image into the teacher Transformer to obtain the spatiotemporal features of the video segments. After softmax processing, the target probability distribution is obtained. Divide the image of each frame of the cholecystectomy video into N non-overlapping patch tokens, and introduce an additional Primary Semantic Token (PST) to form the input sequence. Input the input sequence into the student Transformer to obtain the spatiotemporal features of the video segments. After softmax processing, the predicted probability distribution is obtained.
[0063] Specifically, the teacher's Transformer inputs a global view, while the student's Transformer inputs both a global and a local view, enabling consistency learning. Each frame is divided into patch tokens, plus a primary semantic identifier, forming the input sequence:
[0064]
[0065] in, Represents the input sequence. Represents the main semantic identifier data. This represents the first patch token in the t-th image. This represents the Nth patch token in the t-th image. This indicates the position code.
[0066] Multi-layer multi-head self-attention computation is performed using Transformer to extract spatiotemporal features of video segments.
[0067] The student and teacher Transformer outputs are processed using temperature-scaled softmax:
[0068]
[0069] in, This represents the output characteristics of the student Transformer; This represents the output features of the teacher's Transformer, used to provide target supervision; This is a temperature coefficient used to adjust the smoothness of the softmax output; This represents the predicted probability distribution of the student Transformer after temperature scaling. This represents the target probability distribution after the teacher Transformer has been scaled by temperature.
[0070] 2. Obtain the loss value based on the predicted probability distribution, the target probability distribution, and the cross-entropy loss function. If the loss value does not meet expectations, adjust the parameters of the student Transformer and update the teacher Transformer using the exponential moving average of the student Transformer. If the loss value meets expectations, stop the video self-supervised pre-training.
[0071] The cross-entropy loss function is used to measure the difference between the teacher's Transformer output and the student's Transformer output:
[0072]
[0073] Here, k represents the vector dimension index, and the loss measures the distributional consistency between the student Transformer output and the teacher Transformer output. In the self-supervised distillation framework, the training objective is to minimize this loss. After training, the student Transformer possesses the ability to represent spatiotemporal dynamic information in LC videos with high quality.
[0074] like Figure 2 As shown, the parameters of the teacher Transformer are updated using the exponential moving average (EMA) of the student Transformer:
[0075]
[0076] in, For the teacher's Transformer parameters, For the student Transformer parameter, m is the momentum factor, which is typically set to 0.996 or higher.
[0077] The pre-training module 70 is also used for weakly supervised training of the SMIL network architecture, including:
[0078] Weak training loss values are obtained using the image-level CVS standard labels, CVS standard evaluation results, and a three-term binary cross-entropy loss function. The image-level CVS standard labels are used as the basis for this calculation. , This indicates that the current image frame of the cholecystectomy video satisfies the i-th CVS criterion. =1, 2, 3.
[0079] Specifically, during training, only image-level CVS standard labels are used, and the loss function is a three-term binary cross-entropy loss function. This achieves the goal of eliminating the need for segmentation labeling, significantly reducing labeling costs, while maintaining good training efficiency and generalization ability. The three-term binary cross-entropy loss function is as follows.
[0080]
[0081] in, This represents the image-level label of the i-th CVS standard; This represents the evaluation result of the SMIL network architecture for the i-th CVS standard prediction.
[0082] In this optional embodiment, a multi-label weakly supervised training strategy is designed for the three CVS criteria (C1, C2, C3), requiring only image-level labels and eliminating the dependence on expensive pixel-level segmentation annotations. This strategy employs three independent binary loss functions for joint optimization, ensuring the classification independence and interpretability of the three CVS criteria in diverse scenarios.
[0083] After the above training, the SMIL network architecture can be directly used for feature extraction. For example... Figure 3 As shown, image frame X is taken as input, and the pre-trained student Transformer extracts the main semantic identifier as global contextual features:
[0084]
[0085] Where X is the image frame; This represents the function of the student Transformer to process the input image frames, and the output is a sequence of tokens containing spatiotemporal features; PST represents the main semantic identifier; Characterize the overall semantic and feature information of the surgical video, i.e., global contextual features.
[0086] In this embodiment, a teacher-student video Transformer architecture based on unlabeled self-supervised distillation is introduced to extract global features with spatiotemporal semantic consistency from surgical videos. Employing the DINO distillation method, without any manual annotation, a temporally consistent contextual representation is constructed through global-local view comparison learning, significantly improving the model's ability to model the dynamic anatomical structures of cholecystectomy.
[0087] The instance feature extraction module 30 is used to divide the image frame into multiple image blocks and input them into the student Transformer in the multi-instance learning architecture to extract the instance features of each image block and form an instance feature matrix.
[0088] like Figure 3 As shown, the input image frame is divided into N patches, and the instance features of each patch are extracted: .
[0089] Instance feature matrix: ,in, Let be the instance feature of the j-th image patch, N be the total number of image patches, and D be the dimension of the instance feature.
[0090] In this embodiment, a multi-instance learning (MIL) strategy is employed to mine local features in the anatomical structure discrimination region, and spatial saliency aggregation is achieved through an attention weighting mechanism. Unlike traditional CVS evaluation methods that rely on static or regular regions, this module enables the model to automatically focus on image regions with discriminative power, effectively improving the classification accuracy and robustness of the three CVS criteria.
[0091] The local feature aggregation module 40 is used to input the instance feature matrix into the MILAttention module in the multi-instance learning architecture for weighted aggregation, so as to obtain the aggregated local instance features.
[0092] like Figure 3 As shown, the instance feature matrix is then input into the MIL Attention module, as follows: Figure 4 As shown, the MIL Attention module performs weighted aggregation on the instance feature matrix, focusing on key regions related to CVS. Here, u represents the importance score of each patch feature, and the weights and MIL feature aggregation are obtained after softmax normalization, calculated as follows.
[0093]
[0094] Where V represents the linear transformation matrix, which performs feature mapping on the instance features of the image patch; b is the bias term; and tanh is the hyperbolic tangent activation function. represents the attention vector, used to map the intermediate representation to a scalar score; u represents the importance score of each patch feature; α represents the intermediate parameter; Represents local instance features.
[0095] The system further includes a dynamic weight allocation module 80, used to fuse global context features and local instance features to obtain initial fused features; and to obtain dynamic weights corresponding to different CVS standards based on the initial fused features and the CVS standard category, wherein the dynamic weights are... ,in, This represents the Sigmoid function. , and These are the parameters of the weight dynamic allocation module corresponding to the i-th CVS standard.
[0096] Specifically, such as Figure 3 and Figure 4 As shown, after obtaining the aggregated local instance features, they are concatenated and fused with the context features containing global semantic information according to the feature dimension to obtain the initial fused features:
[0097]
[0098] The dynamic weight allocation module 80 outputs dynamic weights corresponding to different CVS standards based on the category of CVS standards.
[0099] Specifically, the dynamic weight is ,in, This represents the Sigmoid function. , and The parameters are for the weight dynamic allocation module (which can be a perceptual gating) corresponding to the i-th CVS criterion. To adapt to the discrimination emphasis of different criteria, a weight dynamic allocation module is introduced, which can dynamically adjust the weights according to different criterion categories and different initial fusion features. This strategy allows different global-local ratios to be learned on C1, C2, and C3. For example, C3 relies more on the local geometric saliency of tubular structures, while C1 and C2 rely more on the global anatomical context.
[0100] Feature fusion module 50 is used to fuse features based on global context features, local instance features, and dynamic weights to obtain specific fused features corresponding to different CVS standards, specifically:
[0101] We perform weighted fusion based on the dynamic weights, global context features, and local instance features corresponding to different CVS standards to obtain specific fusion features corresponding to different CVS standards.
[0102] Specific fusion features ,in, This represents the specific fusion feature corresponding to the i-th CVS standard. This represents element-wise product.
[0103] In this embodiment, multimodal feature fusion involves concatenating and weighting context-level global features with instance-level local features to construct a specific fused feature for joint discrimination based on the three CVS standards. This fusion strategy preserves the overall semantic context of the surgery while emphasizing local structural features and the characteristics of different standards, thus overcoming the limitations of a single feature dimension.
[0104] The evaluation and prediction module 60 is used to input the specific fusion features corresponding to different CVS standards into the corresponding linear classifiers to obtain the CVS standard evaluation results for each image frame.
[0105] Specifically, such as Figure 5 As shown, the fused features are input into a linear classifier to obtain the decision results on the three CVS criteria. The CVS criteria decision results for each frame are output in real time during the surgical video inference stage. , The sigmoid function represents the linear classifier corresponding to the i-th CVS standard. When the confidence level is higher than a set threshold (e.g., 0.5), the current standard is considered to be achieved.
[0106] The final output is the binary classification result of the current frame based on the three CVS criteria (CVS standard evaluation result):
[0107] ,
[0108] in, γ represents the predicted probability of the SMIL network architecture for the i-th CVS standard; γ represents the decision threshold parameter of the CVS standard. This indicates the judgment result of each image frame on the three CVS standards. For example, [1,1,1] means that all three standards are met, [0,1,1] means that only two standards, C2 and C3, are met, and [1,0,0] means that only one standard, C1, is met.
[0109] In this optional embodiment, by introducing self-supervised video pre-training and the MIL framework, the adaptability of the SMIL network architecture under different video frame rates, surgical styles, and device data is improved, enhancing its generalization ability in practical clinical applications. For single-frame image input data, the CVS evaluation result of the single-frame image can be output. For video input data, video-level CVS standard evaluation results and keyframe evaluation results can be output. Combined with the reliability of the confidence evaluation results, this application can be used for intraoperative decision support and also for providing intelligent support for postoperative evaluation.
[0110] like Figure 6 As shown in the embodiment of this application, a SMIL-based method for assessing CVS after cholecystectomy includes:
[0111] S1: The cholecystectomy video is segmented, and keyframes are adaptively extracted based on the content of each video segment to obtain image frames.
[0112] S2: Input the image frame into the student Transformer in the unlabeled self-supervised distillation architecture within the SMIL network architecture to extract global contextual features. The SMIL network architecture includes an unlabeled self-supervised distillation architecture and a multi-instance learning architecture. The unlabeled self-supervised distillation architecture includes a student Transformer, and the multi-instance learning architecture includes a student Transformer shared with the unlabeled self-supervised distillation architecture and the MIL Attention module.
[0113] S3: Divide the image frame into multiple image blocks and input them into the student Transformer. Extract the instance features of each image block to form an instance feature matrix.
[0114] S4: Input the instance feature matrix into the MIL Attention module for weighted aggregation to obtain local instance features.
[0115] S5: Feature fusion is performed based on global context features, local instance features, and dynamic weights to obtain specific fused features corresponding to different CVS standards.
[0116] S6: Input the specific fusion features corresponding to different CVS standards into the corresponding linear classifiers to obtain the CVS standard evaluation results for each image frame.
[0117] In an optional embodiment of this application, the global context feature ,
[0118] Where X is the image frame; This represents the function of the student Transformer to process the input image frames, and the output is a sequence of tokens containing spatiotemporal features; PST represents the main semantic identifier;
[0119] The instance feature matrix ,
[0120] in, Let N be the instance feature of the j-th image patch, N be the total number of image patches, and D be the dimension of the instance feature;
[0121]
[0122] Where V represents the linear transformation matrix, which performs feature mapping on the instance features of the image patch; b is the bias term; and tanh is the hyperbolic tangent activation function. The attention vector is represented by u; the importance score of each patch feature is represented by α; and the intermediate parameter is represented by α. Represents local instance features;
[0123] The system also includes a weight dynamic allocation module, used to fuse global context features and local instance features to obtain initial fused features. Based on the initial fusion features and the category of CVS standard, dynamic weights corresponding to different CVS standards are obtained, wherein the dynamic weights are: ,in, This represents the Sigmoid function. , and The parameters for the weight dynamic allocation module corresponding to the i-th CVS standard;
[0124] The feature fusion based on global context features, local instance features, and dynamic weights yields specific fused features corresponding to different CVS standards, including:
[0125] We perform weighted fusion based on dynamic weights, global context features, and local instance features corresponding to different CVS standards to obtain specific fusion features corresponding to different CVS standards. ,in, This represents the specific fusion feature corresponding to the i-th CVS standard. This represents element-wise product.
[0126] In an optional embodiment of this application, the CVS standard evaluation result ,
[0127] in, This indicates that the SMIL network architecture is for the first The predicted probability of the CVS standard; γ represents the decision threshold parameter of the CVS standard; This indicates the judgment result of each image frame on the three CVS criteria.
[0128] In an optional embodiment of this application, the step of segmenting the read cholecystectomy video and adaptively extracting keyframes based on the content of each video segment to obtain image frames includes:
[0129] Based on a specified time anchor point, the cholecystectomy video before the specified time anchor point is cropped and removed, and abnormal video frames are removed again according to preset indicators to obtain the processed cholecystectomy video.
[0130] The content of the processed cholecystectomy video is processed, the cholecystectomy video is segmented, and the sampling frequency is adaptively changed according to the content of each video segment to extract key frames, thus obtaining the image data of the original video frames.
[0131] For image data Normalization is performed to obtain image frame X;
[0132] Image Frame ,
[0133] in, , .
[0134] In an optional embodiment of this application, the unlabeled self-supervised distillation architecture further includes a teacher Transformer; the SMIL-based cholecystectomy CVS evaluation method further includes: performing video self-supervised pre-training on the unlabeled self-supervised distillation architecture, wherein the video self-supervised pre-training specifically includes:
[0135] The global image is input into the teacher Transformer to obtain the spatiotemporal features of the video clip. After softmax processing, the target probability distribution is obtained.
[0136] Each frame of a cholecystectomy video is divided into N non-overlapping patch tokens, plus a main semantic identifier, to form the input sequence.
[0137] The input sequence is fed into the student Transformer to obtain the spatiotemporal features of the video clips. After softmax processing, the predicted probability distribution is obtained.
[0138] The loss value is obtained based on the predicted probability distribution, the target probability distribution, and the cross-entropy loss function.
[0139] If the loss value does not meet expectations, the parameters of the student Transformer are adjusted, and the teacher Transformer is updated using the exponential moving average of the student Transformer.
[0140] If the loss value reaches the expected value, stop the video self-supervised pre-training.
[0141] In an optional embodiment of this application, the cross-entropy loss function is:
[0142]
[0143]
[0144] Where k represents the vector dimension index, This represents the output characteristics of the student Transformer; This represents the output characteristics of the teacher's Transformer; This is a temperature coefficient used to adjust the smoothness of the softmax output; This represents the predicted probability distribution of the student Transformer after temperature scaling. This represents the target probability distribution after the teacher Transformer has been scaled by temperature.
[0145] In an optional embodiment of this application, the SMIL-based cholecystectomy CVS evaluation method further includes: weakly supervised training of the SMIL network architecture, wherein the weakly supervised training specifically includes:
[0146] Weak training loss values are obtained using the image-level CVS standard labels, CVS standard evaluation results, and a three-term binary cross-entropy loss function. The image-level CVS standard labels are used as the basis for this calculation. , This indicates that the current image frame of the cholecystectomy video satisfies the first... CVS standard, =1, 2, 3;
[0147] The three binary cross-entropy loss functions are:
[0148]
[0149] in, This represents the image-level label of the i-th CVS standard; This represents the evaluation result of the SMIL network architecture for the i-th CVS standard prediction.
[0150] In summary, compared with existing technologies, it has the following beneficial effects:
[0151] 1. This application utilizes a video understanding method based on self-supervised distillation and multi-instance learning to automatically identify whether the three CVS criteria are met during cholecystectomy. It can accurately determine key safety aspects of intraoperative anatomical procedures, effectively addressing the problems of existing methods such as heavy reliance on high-quality annotations, poor generalization of CVS evaluation, and insufficient utilization of spatiotemporal structural information. By introducing a self-supervised distillation pre-training mechanism that requires no manual annotation, and combining it with a multi-instance attention mechanism to model key anatomical regions, this application achieves efficient automatic evaluation of the C1, C2, and C3 criteria, and achieves superior recognition performance compared to existing technologies without using pixel-level segmentation labels.
[0152] 2. In intraoperative applications, this application provides surgeons with video-based intelligent auxiliary assessment capabilities. It can provide real-time or near-real-time feedback on whether the current operation meets safety assessment standards without interfering with the surgical procedure, thereby helping surgeons reduce the risk of bile duct injury due to visual misjudgment and improving surgical safety and operational stability. This application is applicable not only to continuous video input scenarios but also to static image or low frame rate video scenarios, demonstrating good adaptability and reliability under different device acquisition conditions.
[0153] 3. Regarding postoperative analysis, this application enables intelligent postoperative retrospective analysis of surgical video content, providing assistance for the assessment of surgical quality, skills, and risks. It can be used to automatically extract and evaluate key operational segments during the surgical process, achieving structured understanding and classification analysis of surgical video content, supporting quantitative assessment and refined management of surgical quality, physician skills, and risk events. This capability has broad application prospects in clinical review and surgical quality control, providing intelligent technical support for hospital postoperative quality management systems.
[0154] 4. In the teaching and training phase, this application can automatically analyze massive amounts of surgical video data, identify key frames and processes related to CVS assessment, and automatically mark whether safety standards are met, providing instructors with precise feedback criteria. It also assists trainees in understanding standard operating procedures and common error scenarios. Through structured and standardized video learning content, it effectively improves training efficiency and teaching quality, contributing to the intelligent development of surgical education.
[0155] Experiments were conducted using the SMIL-based CVS assessment method for cholecystectomy described above. To balance class detection capability and overall performance under class imbalance conditions, we selected Balanced Accuracy (Bacc) and Mean Precision (mAP) as the main evaluation metrics. Accuracy (Acc), Negative Predictive Value (NPV), and Positive Predictive Value (PPV) were also reported for comprehensive evaluation.
[0156] The SMIL method was compared with representative existing methods, including SwinCVS, DeepCVS, and LG-CVS, all based on the same public dataset and dataset splitting ratio, ensuring a fair and consistent evaluation. Figure 7 As shown in the table, a checkmark indicates pixel-level segmentation annotation, while an X indicates no segmentation annotation is needed. The comparison results show that, without segmentation annotation, SMIL (Ours represents SMIL in the figure) outperforms the best existing unlabeled segmentation method, SwinCVS (Frozen), on most metrics. It improves mAP by +5.55% and +7.71% on C1 and C2, respectively, and Bacc by +8.08% and +5.52%. Although there is a slight decrease in mAP on C3 (-3.65), it still achieves an overall improvement of +3.21 mAP and +7.25 Bacc, demonstrating better overall balance. Compared to segmentation-annotated supervised DeepCVS and LG-CVS, SMIL still achieves an average mAP of 70.66% and a Bacc of 79.30% on C1, even surpassing LG-CVS in some metrics, despite the lack of segmentation labels. This fully validates the advantages and robustness of the self-supervised spatiotemporal feature modeling and MIL local attention mechanism in the CVS evaluation task.
[0157] To evaluate the contributions of Multi-Instance Learning (MIL) and Self-Supervised Distillation Pre-training (SSL), we designed ablation experiments to analyze the performance under different component configurations. Figure 8 As shown in the table, the experimental results indicate that removing the MIL module significantly reduces the model's ability to focus on key anatomical regions, leading to a comprehensive decrease in mAP and Bacc. When only the MIL is retained and self-supervised distillation pre-training is removed, the model's spatiotemporal modeling and semantic representation capabilities are insufficient, and the Bacc remains fixed at 50.0% across all categories. The complete SSL+MIL structure significantly outperforms the single-component configuration in C1–C3 and overall average mAP and Bacc, verifying the complementarity between global spatiotemporal context and local instance attention. Figure 9 The visualized NPV, PPV, Sensitivity, and Specificity results further illustrate that SMIL outperforms single-module models in all metrics and maintains its advantage under class imbalance conditions.
[0158] A detailed analysis of SMIL's overall performance across all CVS standards and overall tasks was conducted (see...). Figure 10 The table shows multiple metrics including mAP, Acc, Bacc, F1-score, Specificity, Sensitivity, PPV, and NPV. In the overall task (CSV assessment), SMIL achieved 70.66% mAP, 77.50% Bacc, and 95.78% Specificity, with excellent performance in NPV (93.11%) and PPV (72.80%). Despite a Sensitivity of 59.21%, it achieved a high level of balance between accuracy and specificity, ensuring identification stability while reducing the risk of missed detections. All results collectively demonstrate that SMIL possesses stable and reliable classification capabilities under a multi-metric system, making it suitable for deployment in real-world clinical environments.
[0159] An electronic device provided in this application includes a memory and a processor; the memory is used to store a computer program; the processor is used to implement the SMIL-based CVS assessment method for cholecystectomy as described above when the computer program is executed.
[0160] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the SMIL-based CVS assessment method for cholecystectomy as described above.
[0161] In this embodiment, the beneficial effects of the electronic device and the computer-readable storage medium are similar to those of the SMIL-based CVS assessment method for cholecystectomy described above, and will not be repeated here.
[0162] The present invention describes electronic devices that can serve as servers or clients of this application, which are examples of hardware devices that can be applied to various aspects of this application. Electronic devices are intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the application described and / or claimed herein.
[0163] Electronic devices include a computing unit that can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) or loaded from a storage unit into random access memory (RAM). The RAM can also store various programs and data required for device operation. The computing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0164] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc. In this application, the separately described modules may or may not be physically separate. Some or all of the modules can be selected to achieve the purpose of the embodiments of this application according to actual needs. Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0165] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0166] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A SMIL-based cholecystectomy CVS evaluation system, characterized in that, The system comprises: an image frame extraction module configured to segment the read cholecystectomy video, and extract key frames from each segment of the video according to the content of the segment, to obtain image frames; a global feature extraction module configured to input the image frames into a student Transformer in a label-free self-supervised distillation architecture in an SMIL network architecture to extract global context features, wherein the SMIL network architecture comprises the label-free self-supervised distillation architecture and a multiple instance learning architecture, the label-free self-supervised distillation architecture comprises the student Transformer, and the multiple instance learning architecture comprises the student Transformer shared with the label-free self-supervised distillation architecture and a MIL Attention module; an instance feature extraction module configured to divide the image frames into a plurality of image blocks to input the image blocks into the student Transformer to extract instance features of each image block, to form an instance feature matrix; a local feature aggregation module configured to input the instance feature matrix into the MIL Attention module for weighted aggregation to obtain local instance features; a feature fusion module configured to perform feature fusion according to the global context features, the local instance features, and a dynamic weight to obtain specific fusion features corresponding to different CVS standards; an evaluation prediction module configured to input the specific fusion features corresponding to different CVS standards into corresponding linear classifiers respectively to obtain CVS standard evaluation results of each image frame; the label-free self-supervised distillation architecture further comprises a teacher Transformer; the system further comprises a pre-training module configured to perform video self-supervised pre-training on the label-free self-supervised distillation architecture, comprising: inputting a global image into the teacher Transformer to obtain spatiotemporal features of a video segment, and performing softmax processing to obtain a target probability distribution; dividing images of each frame of the cholecystectomy video into N non-overlapping patch tokens, plus a main semantic identifier, to form an input sequence; inputting the input sequence into the student Transformer to obtain spatiotemporal features of the video segment, and performing softmax processing to obtain a predicted probability distribution; obtaining a loss value according to the predicted probability distribution, the target probability distribution, and a cross-entropy loss function; if the loss value does not reach an expectation, adjusting parameters of the student Transformer, and simultaneously updating the teacher Transformer through an exponential moving average of the student Transformer; if the loss value reaches the expectation, stopping the video self-supervised pre-training.
2. The SMIL-based cholecystectomy CVS evaluation system of claim 1, wherein, the global context feature , Wherein, X is an image frame; A function representing that the student Transformer processes the input image frame, and the output is a token sequence containing spatiotemporal features; PST represents a principal semantic identifier; The example feature matrix , wherein, is the instance feature of the jth image patch, N is the total number of image patches, and D is the dimension of the instance feature. Wherein V represents a linear transformation matrix, the instance features of the image block are mapped; b is a bias term; tanh is a hyperbolic tangent activation function; represents an attention vector; u represents an importance score of each patch feature; a represents an intermediate parameter; represents a local instance feature; The system further comprises a weight dynamic allocation module configured to fuse the global context features and the local instance features to obtain initial fusion features ; and obtain dynamic weights corresponding to different CVS standards according to the initial fusion features and the categories of the CVS standards, wherein the dynamic weights are , wherein represents a Sigmoid function, , and is a parameter of the weight dynamic allocation module corresponding to the i-th CVS standard. the feature fusion according to the global context features, the local instance features, and the dynamic weight to obtain the specific fusion features corresponding to different CVS standards comprises: According to different CVS standards corresponding to dynamic weights, global context features and local instance features are weighted and fused to obtain specific fusion features corresponding to different CVS standards wherein, denotes the specific fusion feature corresponding to the i-th CVS standard, denotes an element-wise product.
3. The SMIL-based cholecystectomy CVS evaluation system of claim 1, wherein, the CVS standard evaluation result , wherein, represents the prediction probability of the SMIL network architecture on the first item CVS standard; γ represents the decision threshold parameter of the CVS standard; represents the decision result of each image frame on the three items of the CVS standard.
4. The SMIL-based cholecystectomy CVS evaluation system of claim 1, wherein, the segmentation of the read cholecystectomy video, and the adaptive extraction of key frames from each segment of the video according to the content of the segment to obtain image frames comprises: based on a specified time anchor, cutting and removing the cholecystectomy video before the specified time anchor, and according to a preset index, removing abnormal video frames again to obtain a processed cholecystectomy video; The content of the cholecystectomy video after perception processing is processed, the cholecystectomy video is segmented, and key frames are adaptively extracted according to the content of each video segment, so that image data of the original video frame is obtained; normalizing the image data to obtain an image frame X; Image frame , wherein , .
5. The SMIL-based cholecystectomy CVS evaluation system of claim 1, wherein, The cross-entropy loss function is: where k denotes the vector dimension index, denotes the output feature of the student Transformer; denotes the output feature of the teacher Transformer; is a temperature coefficient for adjusting the smoothness of the softmax output; denotes the prediction probability distribution of the student Transformer after temperature scaling, denotes the target probability distribution of the teacher Transformer after temperature scaling.
6. The SMIL-based cholecystectomy CVS evaluation system of claim 1, wherein, The pre-training module is further configured to perform weakly supervised training on the SMIL network architecture, including: The weak training loss value is obtained by using the image-level CVS standard label, the CVS standard evaluation result and three binary cross-entropy loss functions, wherein the image-level CVS standard label , represents that the current image frame of the cholecystectomy video satisfies the first CVS standard, =1, 2, 3. The three binary cross-entropy loss functions are: wherein, represents the image-level label of the ith CVS criterion; represents the evaluation result predicted by the SMIL network architecture for the ith CVS criterion.
7. A SMIL-based cholecystectomy CVS evaluation method, characterized in that, Including: The read cholecystectomy video is segmented, and key frames are adaptively extracted according to the content of each video segment, so that image frames are obtained; The image frames are input into a student Transformer in a label-free self-supervised distillation architecture in the SMIL network architecture to extract global context features, wherein the SMIL network architecture includes a label-free self-supervised distillation architecture and a multiple instance learning architecture, the label-free self-supervised distillation architecture includes a student Transformer, and the multiple instance learning architecture includes a student Transformer shared with the label-free self-supervised distillation architecture and a MIL Attention module; The image frames are divided into a plurality of image blocks and input into the student Transformer to extract instance features of each image block and form an instance feature matrix; The instance feature matrix is input into the MIL Attention module for weighted aggregation to obtain local instance features; According to the global context features, the local instance features and the dynamic weights, the specific fusion features corresponding to different CVS standards are obtained; The specific fusion features corresponding to different CVS standards are respectively input into corresponding linear classifiers to obtain CVS standard evaluation results of each image frame; The label-free self-supervised distillation architecture further includes a teacher Transformer; The SMIL-based cholecystectomy CVS evaluation method further includes video self-supervised pre-training of the label-free self-supervised distillation architecture, including: The global image is input into the teacher Transformer to obtain the spatio-temporal features of the video segment, and the target probability distribution is obtained after softmax processing; The image of each frame of cholecystectomy video is divided into N non-overlapping patch tokens, plus a subject semantic identifier, to form an input sequence; The input sequence is input into the student Transformer to obtain the spatio-temporal features of the video segment, and the prediction probability distribution is obtained after softmax processing; According to the prediction probability distribution, the target probability distribution and the cross-entropy loss function, a loss value is obtained; If the loss value does not reach the expectation, the parameters of the student Transformer are adjusted, and the teacher Transformer is updated through the exponential moving average of the student Transformer; If the loss value reaches the expectation, the video self-supervised pre-training is stopped.
8. An electronic device, comprising: The memory and the processor are included; The memory is configured to store a computer program; The processor is configured to implement the SMIL-based cholecystectomy CVS evaluation method of claim 7 when executing the computer program.
9. A computer-readable storage medium, characterized in that, The storage medium has stored thereon a computer program which, when executed by a processor, implements the SMIL-based cholecystectomy CVS evaluation method according to claim 7.
Citation Information
Patent Citations
Medical image segmentation method fusing SAM global modeling and U-Net local optimization
CN119832012A
Robot medical image segmentation and feature extraction method for precise operation
CN120236083A