Method, device, equipment, medium and program product for intraoperative risk report generation
Patent Information
- Application Number
- CN202610943674.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-06-29
AI Technical Summary
[0005]本申请提供一种术中风险报告生成的方法、装置、设备、介质及程序产品,用以解决现有手术视频分析中对出血、烟雾等术中风险事件难以连续识别、难以刻画其时序演化过程且缺少结构化表达和量化依据的问题,并提出一种基于手术视频对风险事件进行连续建模并生成风险报告的技术方案
[0039]本申请提供的术中风险报告生成方案,通过对手术视频中的风险相关信息进行持续提取、关联分析与结构化汇总,形成能够反映风险事件发生及演化特征的报告内容。该方案将原本停留于流程级的分析进一步延伸至风险事件级的连续建模,使风险信息在时间维度上具有可追踪性,在表达形式上具有结构化特征。由此,可减少依赖人工回放进行风险识别与整理的不足,提升术中风险感知的连续性。进一步地,该方案能够为术后复盘与风险评估提供相对统一的量化依据,从而改善现有方案难以直接识别术中风险事件、且不便进行连续分析的技术问题。
Smart Images

Figure CN122474247B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the medical field, and more particularly to a method, apparatus, device, medium, and procedure for generating intraoperative risk reports. Background Technology
[0002] With the popularization of minimally invasive surgical techniques, surgical videos have become a core data carrier for recording surgical procedures. In clinical practice, surgical videos are not only used for postoperative review and teaching, but also serve as an important basis for evaluating surgical quality and optimizing operational procedures.
[0003] However, traditional surgical video analysis relies heavily on manual review, resulting in inefficiency, subjectivity, and the potential for missed risks. For example, surgeons need to examine video frame by frame to identify risk events such as bleeding and smoke obstruction, a time-consuming process prone to errors due to fatigue. Furthermore, the dynamic evolution of risk events during surgery (such as the spread of bleeding points and the recovery from smoke obstruction) is difficult to capture fully through manual recording, leading to a lack of quantitative basis for risk assessment. While existing technologies have incorporated artificial intelligence for surgical instrument identification and procedure inference, these methods primarily focus on process recording and cannot directly perceive intraoperative safety risks, let alone model the temporal evolution of risk events. For instance, in laparoscopic cholecystectomy, the surgeon may make mistakes due to blurred vision caused by smoke obstruction, but existing systems cannot detect this risk in real time and issue warnings.
[0004] Therefore, there is an urgent need for a technical solution that can automatically identify risk events from surgical videos and generate structured risk reports to improve surgical safety, reduce the workload of doctors, and provide data support for postoperative quality assessment. Summary of the Invention
[0005] This application provides a method, apparatus, equipment, medium, and program product for generating intraoperative risk reports, which addresses the problems in existing surgical video analysis where it is difficult to continuously identify intraoperative risk events such as bleeding and smoke, difficult to characterize their temporal evolution process, and lack of structured expression and quantitative basis. It also proposes a technical solution for continuously modeling risk events based on surgical videos and generating risk reports.
[0006] In a first aspect, embodiments of this application provide a method for generating an intraoperative risk report, including:
[0007] Acquire video data during the surgical procedure;
[0008] Based on the video data, identify the risk areas in each frame of the image;
[0009] Based on the location and time of occurrence of the risk area, generate risk event instances;
[0010] During the operation, the newly detected risk area is matched with the existing risk event instance. If the match is successful, the newly detected risk area is associated with the existing risk event instance, and the risk event instance is updated.
[0011] Based on historical time-series data of risk event instances, predict the risk level at future moments;
[0012] The output includes a visual surgical report showing the risk level.
[0013] In some embodiments, the current stage of surgery is determined based on the surgical instruments in the video data;
[0014] Determine the risk detection threshold based on the stage of surgery;
[0015] Based on the risk detection threshold, the parameters in the deep learning detection network are adjusted, and the risk areas in the video data are identified through the deep learning detection network.
[0016] In some embodiments, if the distance between two adjacent risk areas is less than a preset distance, an associated risk area is determined;
[0017] By using deep learning networks to identify associated risk areas, the probability of multi-region linkage risks within a single frame can be determined.
[0018] Output the probability of multi-regional linkage risks.
[0019] In some embodiments, the similarity of appearance features between the newly detected risk region and existing risk event instances and the intersection-union ratio of their bounding boxes are calculated;
[0020] If the appearance feature similarity is greater than the similarity threshold, and / or the crossover ratio is greater than the crossover ratio threshold, then frame interpolation is performed based on the historical time series data of the risk event instances and the newly detected risk areas to obtain continuous risk event instances.
[0021] In some embodiments, the future area at a future time is determined based on area changes in historical time-series data of risk event instances;
[0022] A comprehensive risk score is calculated based on the future area, the current area, the number of currently concurrent risk areas, and the duration of risk event instances from their start to a future time.
[0023] The risk level is determined based on the comprehensive risk score.
[0024] In some embodiments, if the total area of multiple sub-risk areas in a risk event instance is greater than a first preset multiple of the initial total area, then the multiple sub-risk areas are treated as a new risk event instance; wherein, the first preset multiple is greater than 1;
[0025] If the combined area of multiple sub-risk areas in a risk event instance is less than a second preset multiple of the sum of the areas of the individual sub-risk areas, then the sub-risk area with the largest area will be taken as the area of the risk event instance; wherein, the second preset multiple is less than 1.
[0026] If the Mahalanobis distance between the location of the newly emerging risk area and the predicted location of a historical risk event instance is less than a preset distance value, and the appearance similarity is greater than a preset similarity threshold, then the two are identified as the same risk event instance, and the interrupted trajectory frame is interpolated to complete the frame.
[0027] Secondly, embodiments of this application provide an apparatus for generating an intraoperative risk report, the apparatus comprising:
[0028] The acquisition module is used to acquire video data during the surgical procedure;
[0029] The risk identification module is used to determine the risk areas in each frame of the video based on the video data.
[0030] The risk event association module is used to generate risk event instances based on the location, area, and time of occurrence of the risk area;
[0031] The matching and updating module is used to perform feature matching between newly detected risk areas and existing risk event instances during surgery. If the match is successful, the newly detected risk areas are associated with the existing risk event instances, and the risk event instances are updated.
[0032] The prediction module is used to predict the risk level at future moments based on historical time-series data of risk event instances;
[0033] The output module is used to output a visual surgical report that includes the risk level.
[0034] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0035] The memory stores the instructions that the computer executes;
[0036] The processor executes computer execution instructions stored in memory, causing the processor to perform the methods described above.
[0037] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method provided above.
[0038] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method described above.
[0039] The intraoperative risk report generation solution provided in this application continuously extracts, correlates, and structurally summarizes risk-related information from surgical videos to generate reports that reflect the occurrence and evolution of risk events. This solution extends analysis from the process level to continuous modeling at the risk event level, making risk information traceable over time and structured in its presentation. This reduces reliance on manual playback for risk identification and processing, improving the continuity of intraoperative risk perception. Furthermore, this solution provides a relatively unified quantitative basis for postoperative review and risk assessment, thereby addressing the technical challenges of existing solutions in directly identifying intraoperative risk events and conducting continuous analysis. Attached Figure Description
[0040] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0041] Figure 1 Schematic diagram of the method for generating intraoperative risk reports provided in this application Figure 1 ;
[0042] Figure 2 A flowchart illustrating a method for generating an intraoperative risk report provided in this application. Figure 2 ;
[0043] Figure 3 A flowchart illustrating a method for generating an intraoperative risk report provided in this application. Figure 3 ;
[0044] Figure 4 A flowchart illustrating a method for generating an intraoperative risk report provided in this application. Figure 4 ;
[0045] Figure 5 A schematic diagram of a device for generating intraoperative risk reports provided in this application;
[0046] Figure 6 A schematic diagram of the structure of the electronic device provided in this application.
[0047] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0048] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0049] Minimally invasive surgery, especially laparoscopic surgery, single-port surgery, endoscopic surgery, and robot-assisted surgery, has become an important and widely used surgical procedure in clinical practice.
[0050] Currently, surgical videos not only serve as intraoperative records but are also increasingly becoming an important source of information for postoperative review, teaching and training, quality control, and evidence in medical disputes. In practice, the video acquisition system in the operating room typically needs to work in conjunction with equipment such as laparoscopes, endoscopes, and surgical robots to continuously receive high-definition image streams. After the surgery, doctors or quality control personnel review and analyze the videos. For minimally invasive procedures such as cholecystectomy, gastrointestinal surgery, gynecology, and urology, the videos often contain risks such as bleeding, smoke obstruction, abnormal tissue traction, instrument residue, or interrupted vision. These risks directly affect surgical safety and postoperative recovery.
[0051] Therefore, establishing an analytical workflow around surgical videos that can perceive intraoperative safety risks, support event retrospection, and structured management has become an important development direction for intelligent surgical assistance systems and a fundamental capability that urgently needs to be improved in current clinical digital management.
[0052] In existing technologies, the analysis of surgical videos is mostly focused on information extraction at the process level. Some solutions combine video analysis results with physiological parameters, medical record information or template text to automatically generate surgical reports.
[0053] The existing technology has the following drawbacks:
[0054] Deficiency 1: Lack of direct identification of intraoperative risk events; existing methods infer surgical steps based on surgical instrument identification, detecting the instruments rather than the risks themselves. They cannot directly detect and model intraoperative risk events such as bleeding, surgical smoke obstruction, and foreign bodies, thus failing to reflect the safety risk status during the surgical process.
[0055] Defect 2: Lack of modeling capability from "frame-level detection" to "event-level representation"; existing technologies remain at the level of single-frame target detection or simple statistics, lacking the ability to integrate discrete detection results into risk event instances with temporal continuity. It cannot accurately depict the complete lifecycle of risk events (occurrence → development → end), and cannot obtain key information such as risk duration and development trends.
[0056] Deficiency 3: Lack of a time-series evolution analysis mechanism for risk events; existing technologies lack effective modeling methods for the complex evolutionary processes of risk events, such as splitting, merging, occlusion, and recovery over time. This makes it impossible to reflect the correlations and dynamic characteristics between risk events, and thus difficult to support causal analysis and trend prediction of risks.
[0057] Defect 4: Lack of risk-oriented structured report generation capabilities; existing surgical reports are mostly process descriptions or text summaries, lacking structured data analysis and timeline representation based on risk events. They cannot provide quantitative analysis tools such as risk statistics, time-series distribution, and spatial heatmaps, and do not support surgical quality assessment and risk review.
[0058] This application proposes a method for generating intraoperative risk reports: First, video data during the surgical procedure is acquired. Then, risk regions in each frame of the video data are determined. Further, risk event instances are generated based on the location and occurrence time of the risk regions. During the surgery, newly detected risk regions are continuously matched with existing risk event instances, and when a match is successful, an association is established and the risk event instances are updated. Finally, the risk level for future moments is predicted based on the historical time-series data of the risk event instances, and a visualized surgical report containing the risk level is output. This technical approach shifts surgical video analysis from simple target identification and process recording to continuous modeling and time-series prediction of risk events, thereby providing more direct and structured support for intraoperative safety management, postoperative review, and quality assessment.
[0059] Example 1
[0060] Figure 1 Schematic diagram of the method for generating intraoperative risk reports provided in this application Figure 1 This includes the following steps:
[0061] S101: Acquire video data during the surgical procedure.
[0062] Video data refers to a continuous sequence of images captured by cameras from laparoscopy, endoscopy, or surgical robots during surgery. Its function is to serve as the raw input for risk identification and event modeling, and to detect risk areas such as bleeding and smoke frame-by-frame, further supporting subsequent time-series tracking, prediction, and report generation. In this embodiment, the executing entity can be an image processing server deployed on a local workstation in the operating room, an edge computing unit communicating with surgical equipment, or a dedicated analysis module in a hospital information platform. Specifically, the surgical field images can be received in real time through the video output interface of the laparoscopic host, the endoscope image acquisition card, or the video stream forwarding interface of the surgical robot control system. The received video streams are written into a buffer queue in chronological order to form continuous video data. In some embodiments, the video capture interface can be HDMI (High-Definition Multimedia Interface), SDI (Serial Digital Interface), DVI (Digital Visual Interface), network video streaming protocols, or proprietary interfaces of the device manufacturer, thereby adapting to the output formats of different surgical devices; in other embodiments, the main lens video, auxiliary lens video, and external monitoring camera video can be simultaneously accessed, and multiple video inputs can be formed by timestamp alignment, so that a certain environmental context can still be preserved when the main video is momentarily obstructed.
[0063] In some embodiments, to ensure the stability of subsequent recognition, the acquired video data may undergo a preprocessing process. This preprocessing process includes resolution unification, color space transformation, brightness normalization, gamma correction, and noise suppression of image frames. For example, 1920×1080, 1280×720, or higher resolution images output from different devices are uniformly adjusted to a fixed resolution required by the model input, and the RGB format is converted to a standard tensor format suitable for inference by the detection network. Simultaneously, to avoid color drift caused by different light sources, lens white balance settings, or intra-abdominal reflections, color normalization processing can be performed on the images to keep the color differences between blood areas, smoke areas, and the tissue background within a identifiable range.
[0064] In some embodiments, adaptive frame rate adjustment can also be performed based on the surgical device status signal. For example, when the electrosurgical unit, ultrasonic scalpel, or energy platform is activated, the sampling frame rate can be increased to enhance the capture of smoke bursts and bleeding diffusion processes, while the frame rate can be appropriately reduced during the exploration phase or when the image changes little to reduce the computational load.
[0065] Furthermore, the video data acquisition stage can also simultaneously record video frame numbers, system timestamps, surgical stage identifiers, instrument activation status, and basic anonymity information of patients, thereby providing a time reference and auxiliary semantic information for the subsequent generation of risk event instances and contextual analysis.
[0066] In this embodiment, video data is not only image input but also the time carrier in the entire risk modeling chain. By establishing a stable timestamp mapping relationship during the acquisition stage, each subsequent risk area can be accurately mapped to a specific time, thereby giving discrete detection results traceable temporal attributes. Based on the above analysis, it is clear that standardizing video acquisition, synchronization, and preprocessing can reduce missed detections and false detections caused by unstable input quality, providing a consistent data foundation for subsequent risk area identification and continuous event modeling, thus alleviating the problems of fragmented single-frame results and cross-frame breaks in existing technologies.
[0067] S102: Based on the video data, determine the risk areas in each frame of the image.
[0068] A risk region refers to an image area in a video frame that is identified by the model as having potential intraoperative safety hazards. Its function is to carry the risk detection results and serve as the basic object for event modeling. In this embodiment, the risk region can be a bleeding point, a smoke-covered area, a foreign object area, an abnormal anatomical structure area, or a region with abnormal visual interruption. Specifically, image frames can be read sequentially from the video cache queue formed in step S101, and each image frame can be input into a pre-trained risk detection model. This model can be an object detection network, an instance segmentation network, or a joint detection and segmentation network. For example, a YOLO network can be used to quickly output candidate risk boxes, and a Mask R-CNN network can be used to obtain a more refined risk region mask. The two can also be used in series, i.e., the detection network first performs coarse localization, and then the segmentation network refines the contour. For different risk types in the surgical scenario, the model output can include risk category labels (i.e., whether it belongs to a risk region), bounding box coordinates, pixel-level masks, confidence scores, and feature vector representations.
[0069] In the specific detection process, color, texture, edge, and deep semantic features can be extracted for each frame to enhance the recognition ability of complex surgical field environments. For example, for bleeding areas, red component enhancement, color histogram statistics, and saturation change analysis can be combined to improve the model's sensitivity to fresh bleeding and oozing areas; for smoke-covered areas, low-contrast features, texture blurring, brightness diffusion features, and semantic texture features extracted by deep networks can be combined to improve the accuracy of identifying smoke clumps during electrosurgical operation; for areas with foreign objects or instrument residues, edge geometry and metallic reflective features can be combined to assist in discrimination.
[0070] To reduce false alarms caused by single-frame fluctuations, post-processing can be performed after the detection output, including non-maximum suppression, confidence threshold filtering, region area threshold filtering, and morphological smoothing. Small, isolated noise regions that are clearly inconsistent with clinical scenarios can be directly removed; for adjacent or contacting risk regions of the same type, connected component merging can be used to form a more realistic risk region expression.
[0071] In some embodiments, the system can also dynamically adjust the detection threshold based on information from the surgical stage. For example, during the incision or separation stage, the model can maintain high sensitivity to small-scale bleeding to capture risk spread trends as early as possible; during the suturing or irrigation stage, the confidence threshold for certain risk categories can be appropriately increased to avoid misclassifying normal fluid flow as bleeding. For cases where multiple risk regions exist within a single frame, the system can further calculate the center distance, boundary adjacency, and co-occurrence relationships between each risk region to establish a spatial association graph within the current frame. When several risk regions are spatially close to each other in the image, and the distance between them is lower than a threshold calculated based on lens calibration or empirical scaling, it can be determined that there is a potential for linkage risk, and this spatial association result is written into the subsequent instance modeling data structure.
[0072] The key to this step lies in transforming the raw image stream into a computable and traceable set of risk objects. Compared to solutions that only identify instruments or classify surgical procedures, this embodiment directly targets the detection of safety-affecting risk phenomena and provides a unified input format for subsequent continuous association by outputting structured information such as location, area, category, and confidence level. Based on the above analysis, it can be seen that by stably identifying risk areas in each frame of the image, safety hazards in the surgical field can be separated from background information, thus laying the foundation for the conversion from single-frame labels to continuous event representation. It should be understood that the above example is only for demonstration and not a limitation.
[0073] S103: Generate risk event instances based on the location of the risk area and the time of occurrence.
[0074] A risk event instance refers to a continuous event unit formed by associating discrete risk regions belonging to the same risk object across multiple frames. Its function is to elevate the detection result of a single frame into a risk representation with temporal continuity. In this embodiment, the location of the risk region can be determined based on the bounding box coordinates, centroid coordinates, or the spatial location of the segmentation mask, and the occurrence time can be obtained based on the mapping relationship between the video frame number and the timestamp. Specifically, after step S102 outputs one or more risk regions in a frame, the system first reads the timestamp corresponding to that frame and extracts the center point coordinates, the upper left and lower right corner coordinates of the bounding box, the region's pixel area, category label, and appearance feature vector from each risk region. If the risk region has not yet been associated with any existing event instance, a new risk event instance is created using that risk region as the initial observation. For example, if a risk region is detected in the first frame of the video, that risk region is considered a risk event instance. If a risk region is identified in the next frame, and the risk regions have the same location and shape, they are associated as a single risk event instance until the risk region disappears. In some scenarios, bleeding and blood mist occur at the same location within a single frame, indicating a correlation between the two. These two risk areas can be considered as one instance of a risk event.
[0075] When creating a new instance, a globally unique event identifier can be assigned to it, and initial attribute fields can be written, including first occurrence time, most recent occurrence time, current duration, initial location, initial area, risk category, initial confidence level, and keyframe index. For easier backtracking, the instance can also be synchronously bound to a screenshot of the frame, the index address of video segments several seconds before and after, and the surgical stage, instrument status, and equipment signal information at that time. Regarding location determination, if a bounding box representation is used, the center point of the bounding box can be used as the trajectory point, and the width and height can be recorded to describe morphological changes; if a segmentation mask representation is used, the mask centroid, principal axis direction, perimeter, and area change rate can be further calculated to maintain continuous description of the same event even when the shape of the risk object changes significantly. The occurrence time can be calculated using the formula "Actual time = Start surgery time + Frame number / Sampling frame rate," or the timestamp from the video stream can be directly read.
[0076] In some embodiments, risk event instances can be stored as structured data records, such as relational database tables, key-value stores, or lists of in-memory objects. Each instance contains at least an instance ID, risk type, status flag, set of time-series observation points, set of spatial trajectories, area sequence, and set of contextual information. To support complex evolutionary relationships, instances can also be configured with fields such as parent event ID, child event ID, and merge source ID to represent event splitting, merging, and recovery. If a risk region has a clear linkage with neighboring regions at creation, a list of concurrent events and a spatial adjacency matrix can be stored at the instance level. In this way, the subsequent system can not only know whether a risk exists at a certain moment, but also when the risk started, where it developed, whether it co-occurred with other risks, and the evolutionary trajectory of the risk throughout the surgical procedure.
[0077] This step elevates the semantics from "detection result" to "event object," addressing the critical issue that existing technologies cannot accurately express the duration, spread, and recovery of risks. Based on the above analysis, organizing discrete risk regions into risk event instances according to location and occurrence time transforms previously independent single-frame annotations in the video into continuous risk units with lifecycles and evolutionary trajectories. This enables the system to possess the fundamental capabilities for tracking, statistically analyzing, and predicting risk events. It should be understood that the above example is merely illustrative and not limiting.
[0078] S104: During the operation, the newly detected risk area is matched with the existing risk event instance. If the match is successful, the newly detected risk area is associated with the existing risk event instance, and the risk event instance is updated.
[0079] Feature matching refers to the matching process used to determine whether a newly detected risk region belongs to the same event as an existing risk event instance. Its function is to ensure the continuity of risk tracking. In this embodiment, feature matching integrates indicators such as appearance feature similarity, bounding box intersection-union ratio, and Mahalanobis distance to associate risk regions that reappear after occlusion, have slight positional drift, or have morphological changes. Specifically, after step S102 outputs the newly detected risk region of the current frame, the system reads existing instances that are still active or temporarily disappeared in the recent risk event instance pool, and constructs a candidate matching set with risk category consistency as the first screening condition. Then, for each candidate instance and the current newly detected region, appearance feature similarity, spatial overlap, and motion prediction deviation are calculated respectively.
[0080] In one specific implementation, appearance feature similarity can be represented by the cosine similarity of the feature vectors output by the deep neural network, or by a weighted sum of color histogram similarity and texture feature similarity. The intersection-over-union (IoU) ratio of bounding boxes measures spatial overlap; it is calculated by dividing the area of the intersection of two regions by the area of their union. A higher IoU value indicates greater spatial continuity between the newly detected region and historical instances. Motion prediction bias can be assessed by first using a Kalman filter to predict the next moment's position based on the historical trajectories of existing instances, and then using Mahalanobis distance to measure the deviation between the center point of the newly detected region and the predicted position. Mahalanobis distance can consider the covariance distribution in trajectory estimation, making it more suitable than Euclidean distance for evaluating dynamic bias in scenarios with slight lens shake or local deformation caused by tissue stretching. This embodiment can employ a three-level cascaded matching strategy: first, candidates with obvious dissimilarities are filtered out based on appearance features; then, spatially continuous candidates are selected based on IoU; and finally, the remaining candidates are finely judged through motion prediction. Alternatively, a weighted scoring method can be used, converting appearance similarity, IoU, and motion bias into a unified matching score. A successful match is considered achieved when the comprehensive score exceeds a preset threshold.
[0081] In one specific implementation, matching can be performed using only the appearance feature similarity and the intersection-union ratio of spatial overlap; or matching can be performed using the Mahalanobis distance between appearance feature similarity and motion prediction deviation.
[0082] Upon successful matching, the system adds the current risk area as the latest observation point for the corresponding risk event instance and updates the instance's most recent occurrence time, trajectory point sequence, area change sequence, appearance template, and duration. To reduce the impact of instantaneous noise, instance features can be updated using an exponential moving average method, which proportionally merges current and historical values, enabling the model to respond to morphological changes without identity jumps caused by single-frame errors.
[0083] For existing instances that fail to match new observations in several consecutive frames, they can be marked as temporarily missing rather than terminated immediately, and the system continues to retain their predicted trajectory. If a new risk area that meets the dual conditions is detected again in subsequent frames, that is, both appearance similarity and Mahalanobis distance meet the threshold, the association with the original instance is restored to solve the problem of cross-frame breakage caused by smoke obstruction, equipment briefly blocking the lens, or rapid lens movement.
[0084] In cases where risk areas split, child instances can be created while retaining the original main instance, and a parent-child relationship field can be established. In cases where multiple adjacent instances gradually merge into a larger area, the main instance with a longer duration or larger area can be retained, while other instances are marked as the source of the merge. For trajectory points missing during occlusion, gaps in the time series can be filled through trajectory interpolation, thereby maintaining the smooth continuity of historical data.
[0085] By introducing joint matching of appearance, spatial, and motion information, and combining it with event management mechanisms such as temporary disappearance, occlusion recovery, and split merging, the probability of missed and false associations can be effectively reduced, ensuring the consistency of risk event instances across frames, thereby providing a guarantee for the reliable construction of subsequent historical time-series data. It should be understood that the above example is merely illustrative and not limiting.
[0086] S105: Based on historical time-series data of risk event instances, predict the risk level at future moments.
[0087] Historical time-series data refers to the continuous recording of risk event instances over time. In this embodiment, historical time-series data includes changes in the area of the risk region, duration, location evolution, expansion rate, number of times obstruction is restored, and contextual information related to the surgical stage and instrument status. Risk level refers to the classification result of the severity of intraoperative risk. Its function is to convert the detection and prediction results into quantitative indicators that are easy for clinical understanding and intervention, and to reflect the severity of risks such as blood spread and increased smoke obstruction.
[0088] In practice, the system extracts the observation sequence within the most recent time window from the updated risk event instances, such as the area sequence, position drift sequence, confidence sequence, and concurrent event count sequence of the last 5, 10, or 30 seconds, and generates corresponding temporal feature vectors based on the risk type. For hemorrhage events, the area growth rate, boundary diffusion direction, and duration can be used as the main features; for smoke events, the occlusion ratio, the decrease in image clarity, and the persistence density can be used as the main features.
[0089] For predictive models, Long Short-Term Memory (LSTM) networks, lightweight gated recurrent networks, exponential smoothing models, Holt linear trend models, or hybrid prediction frameworks based on rule-based and model fusion can be used. If an LSTM model is used, the input is a time-ordered sequence of risk features, and the output is a risk score for a future preset time point or within a future preset time window. If exponential smoothing or Holt models are used, short-term risks can be quickly estimated based on the current area value and the trend of area changes, making them more suitable for real-time execution on edge devices.
[0090] Risk scores can be mapped to risk levels through normalization, for example, dividing the comprehensive score into four levels: low risk, medium risk, high risk, and severe risk. The comprehensive score can be expressed as "Comprehensive Score = α × Area Change Rate + β × Duration Factor + γ × Complication Quantity Factor + δ × Context Enhancement Factor," where α, β, γ, and δ are weighted parameters set through training or experience. This formula transforms temporal characteristics of different dimensions into a unified expression of risk intensity and highlights clinically sensitive factors through weight adjustment. For example, when the bleeding area rapidly expands, increasing the weight corresponding to the area change rate can make the score more timely in reflecting risk escalation; when persistent but small-area smoke obstruction affects vision, increasing the duration factor helps prevent the system from underestimating its harm.
[0091] In some embodiments, risk level prediction can also be context-enhanced by incorporating surgical stage, instrument operation signals, and other multimodal data. For example, if the system detects an increase in electrosurgical power and a continued expansion of the smoke event, the risk level can be increased by one level based on the original prediction; if the current procedure is in the irrigation or hemostasis stage and the bleeding area is rapidly decreasing, the probability of future escalation can be reduced. For different surgical procedures, such as cholecystectomy, gastrointestinal anastomosis, gynecological dissection, or urological reconstruction, model parameters can also be dynamically calibrated using historical data from similar surgeries, thereby reducing prediction bias caused by differences in surgical procedures. When the system predicts that the risk level will rise from medium risk to high risk within a preset time window, it can immediately generate a real-time warning signal and write it to the report cache to alert the surgeon.
[0092] This step advances risk analysis from post-operative identification to trend prediction, directly addressing clinical needs for intraoperative safety management. Based on the above analysis, by utilizing continuous historical time-series data of risk event instances, the system can not only identify when a risk has occurred but also estimate whether it will expand, stabilize, or diminish in the future. This allows the report to move beyond static recording and become forward-looking and quantitative. It should be understood that the above example is for illustrative purposes only and is not limiting.
[0093] S106: Output a visual surgical report including risk level.
[0094] A visualized surgical report is a structured report that presents intraoperative risk information in the form of charts, timelines, keyframes, etc. Its function is to provide an intuitive medium for postoperative review, quality assessment, and clinical communication. In this embodiment, the report integrates risk levels, event statistics, heatmaps, Gantt charts, and key video evidence to form a browsable and traceable summary of surgical risks. Specifically, the system periodically calls the report generation module after or during surgery to read event-level, stage-level, and surgical-level information from the risk event instance database and generate a structured report object based on a unified template. This report object can contain at least an objective data layer, a text summary layer, and an evidence layer. The objective data layer includes the overall risk level distribution of this surgery, the number of various risk events, the total duration, the longest duration event, the peak of concurrent risks, a Gantt chart of risks arranged by time, and a spatial heat map of the surgical field; the text summary layer automatically generates a risk overview based on structured fields, such as describing the occurrence of smoke obstruction for a certain period of time and its duration for several seconds, or the rapid escalation of a certain risk event at a certain stage and its decline after treatment; the evidence layer links keyframe images, fragment links, thumbnails, and corresponding timestamps, allowing doctors to click to replay and verify.
[0095] In terms of visualization, the system can plot risk event instances on a timeline, using different colors to represent different risk categories, line segment lengths to represent duration, and color intensity or labels to indicate predicted or actual risk levels. For spatial distribution information, risk areas can be cumulatively overlaid in a standardized surgical field coordinate system to generate a heatmap reflecting high-risk locations. For risk evolution trends, a risk score line graph can be plotted, showing the rise, stabilization, and decline of the score over time. If the system supports an interactive front-end interface, users can also click on any risk event on the timeline to directly jump to the corresponding segment of the original video and view the event's first occurrence frame, peak frame, and fading frame. There are no restrictions on the report output format.
[0096] In some embodiments, the system can also generate differentiated report templates based on the type of surgery and the user's role. For example, reports for surgeons highlight key risk events and changes before and after treatment, reports for hospital quality control departments emphasize statistical indicators, trend comparisons, and standardized scores, and reports for teaching and training emphasize key video evidence and examples of event evolution. For multiple analyses of the same surgery, the system can also retain version information for comparing results before and after model upgrades or for preserving traces of manual revisions.
[0097] This step aggregates the results of the aforementioned video capture, risk detection, event modeling, correlation updates, and risk level prediction into a readable, reviewable, and searchable clinical document, realizing the practical application of the algorithm output in clinical scenarios. Based on the above analysis, by outputting a visualized surgical report containing risk levels, risk information scattered throughout the surgical procedure can be organized into a structured, evidence-based, and time-sensitive result, thereby supporting postoperative review, quality assessment, medical safety management, and clinical communication. It should be understood that the above example is merely illustrative and not limiting.
[0098] Based on the above analysis, this disclosure provides a method for generating intraoperative risk reports, including: acquiring video data during the surgical procedure; determining risk regions in each frame of the video data; generating risk event instances based on the location and occurrence time of the risk regions; during the surgical procedure, performing feature matching between newly detected risk regions and existing risk event instances; if a match is successful, associating the newly detected risk regions with existing risk event instances and updating the risk event instances; predicting the risk level at future moments based on historical time-series data of the risk event instances; and outputting a visualized surgical report including the risk level. In this embodiment, by further elevating the discrete risk detection results in the surgical video to a risk event modeling process with temporal continuity, spatial continuity, and dynamic evolutionary relationships, the system can continuously correlate, stably track, and predict trends of bleeding, smoke, foreign body residue, and other intraoperative safety hazards, and output visualized reports through a unified structured report. This overcomes the problems of existing technologies that only focus on single-frame detection and cannot accurately describe the duration, diffusion process, concurrency relationship, and future development trend of risks, improving the continuity of intraoperative risk identification, the quantification of risk assessment, and the usability of postoperative review and quality control. It should be understood that the above examples are for illustrative purposes only and are not intended to be limiting.
[0099] Example 2
[0100] Based on the above embodiment 1, the criteria for identifying risk areas are different for different surgical stages. The following is a detailed description of how to identify risk areas in step S102 using an embodiment.
[0101] Figure 2 A flowchart illustrating a method for generating an intraoperative risk report provided in this application. Figure 2 ,like Figure 2 As shown, it includes the following steps:
[0102] S1021: Determine the current stage of surgery based on the surgical instruments in the video data;
[0103] Surgical instruments refer to electrosurgical units, suture devices, ultrasonic scalpels, or end effectors of surgical arms. Their location, activation status, and operational status in the surgical field reflect the current surgical intent. Surgical stages are surgical contexts defined based on instrument type and its dynamic behavior, used to characterize different operational states such as cutting, suturing, or exploration. Risk detection thresholds are boundary values used to determine the presence of risk in the image, such as the risk of bleeding or smoke obstruction. These thresholds can vary with different surgical stages to accommodate the differences between normal tissue bleeding and abnormal bleeding within each stage. Deep learning detection networks are used for target recognition or region segmentation of video images to output location and confidence information for bleeding, smoke, or other risk areas.
[0104] In its implementation, after acquiring the surgical video, the system first identifies instruments in consecutive frames of images. Combining the type of instrument, the duration of its presence in the field of view, and whether it is active, the system infers the current surgical stage. When the system detects continuous operation of the electrosurgical unit accompanied by tissue separation, the current stage is identified as the cutting stage. When the system detects the suture device entering and closing and releasing, the current stage is identified as the suturing stage. When the screen mainly shows instrument exploration, observation, and cleaning, the current stage is identified as the exploration stage.
[0105] S1022: Determine the risk detection threshold based on the stage of surgery;
[0106] In one specific implementation, a risk detection threshold corresponding to a surgical stage is determined based on a pre-established stage mapping table. This risk detection threshold may include at least one of a bleeding detection threshold, a smoke detection threshold, and a foreign body detection threshold.
[0107] S1023: Adjust the parameters in the deep learning detection network according to the risk detection threshold, and identify risk areas in the video data through the deep learning detection network.
[0108] The determined risk detection thresholds are incorporated into the inference parameters of the detection network, allowing the network's detection confidence threshold, class determination threshold, non-maximum suppression threshold, and feature fusion weights to adjust synchronously with the stage. For the cutting stage, the threshold for detecting minor bleeding can be lowered to promptly capture early bleeding; for the exploration stage, the sensitivity to abnormal red areas and diffused fluid areas is increased to reduce missed detections. After the network parameters are updated, the system inputs the current frame into the deep learning detection network, outputting candidate risk boxes, risk masks, and corresponding class labels, and uses this information to determine the risk region.
[0109] This approach links instrument identification results to surgical stages and further drives adaptive changes in risk detection thresholds and network parameters, enabling risk detection to adapt to the visual characteristics and clinical tolerance of different surgical stages. Because the detection network uses different judgment parameters at different stages, the system can reduce false alarms caused by normal operations and improve the sensitivity of identifying unexpected bleeding, thereby enhancing the accuracy and stability of risk area identification and strengthening the reliability of subsequent risk event modeling and reporting.
[0110] Example 3
[0111] Based on the above embodiment 1, the risk level within a single frame can be determined, and then a risk warning can be output to remind the surgeon to pay attention during the operation.
[0112] Figure 3 A flowchart illustrating a method for generating an intraoperative risk report provided in this application. Figure 3 ,like Figure 3 As shown, it includes the following steps:
[0113] S301: Determine whether the distance between two adjacent risk areas is less than the preset distance.
[0114] Among them, the risk area refers to a local area detected as bleeding or suspected bleeding in a single frame image, or a blood fog area, etc.; the associated risk area refers to multiple risk areas that are spatially close to each other after distance determination and may jointly reflect the same bleeding situation. The preset distance is usually obtained based on the image calibration results and the surgical field scale, and is used to map the pixel distance to the actual spatial distance threshold.
[0115] S302: If the distance between two adjacent risk areas is less than a preset distance, then the two risk areas are determined to be related risk areas;
[0116] In practical implementation, the system first extracts the bounding box, centroid coordinates, and region mask for the risk regions within a single frame, and calculates the center distance or minimum circumscribed distance between any two risk regions. When the distance between two risk regions is less than a preset distance, the system combines them into an associated risk region.
[0117] S303: Use deep learning networks to identify risks in associated risk areas and determine the probability of multi-region linkage risks within a single frame;
[0118] The combined region is fed into a deep learning network. This network, employing convolutional neural networks, attention networks, or graph neural networks, jointly analyzes the color distribution, texture features, area features, and relative positional relationships of the associated risk regions to output the probability of multi-regional linkage risks. The network jointly models the continuity of blood color, edge diffusion patterns, and spatial coupling relationships between multiple bleeding patches within the combined region, outputting a probability value representing the likelihood of multi-regional linkage. This probability value can be written into a single-frame risk result cache and further used for risk labeling, alarm triggering, or report generation.
[0119] S304: Output the probability of multi-regional linkage risks.
[0120] In this embodiment, the deep learning network can be deployed as a trained model parameter file, which is called by the image processing unit at runtime. Through the above processing, the system can identify whether multiple adjacent bleeding areas are linked at the single-frame level and quantify the linkage in probabilistic form. This approach avoids making local judgments only for isolated areas, enabling a unified expression of complex risks such as multi-point bleeding, diffuse oozing, and adjacent bleeding clusters, thereby improving the accuracy of bleeding risk identification and the timeliness of clinical prompts.
[0121] Example 4
[0122] Based on the above embodiment 1, the process of feature matching between newly detected risk areas and existing risk event instances will be described below.
[0123] In one specific implementation, the similarity of appearance features between the newly detected risk region and the existing risk event instances and the intersection-union ratio of the bounding boxes are calculated; if the appearance feature similarity is greater than the similarity threshold, and / or the intersection-union ratio is greater than the intersection-union ratio threshold, then frame interpolation is performed based on the historical time-series data of the risk event instances and the newly detected risk region to obtain continuous risk event instances.
[0124] Among them, appearance feature similarity refers to the degree of visual consistency between a newly detected risk region and an existing risk event instance. It can be calculated from semantic feature vectors, color histogram features, or texture feature vectors extracted by a deep neural network, and is used to determine whether the two belong to the same risk object. The intersection-union ratio (IUR) of bounding boxes is the ratio of the intersection area to the union area of two bounding boxes, used to measure the degree of spatial overlap, to help confirm whether the newly detected region and historical instances are in the same or similar positions. The similarity threshold and IUR threshold are used to limit the minimum matching standards for appearance consistency and spatial consistency, respectively, thereby reducing the probability of false association.
[0125] In practical processing, the system reads historical time-series data from existing risk event instances and extracts newly detected risk areas as current observation data. Historical time-series data includes the first occurrence time, last occurrence time, spatial trajectory, area change trend, and interval frame information. After comparing the features of the current observation with historical instances, if the similarity of their appearance features and / or the intersection-over-union (IoU) ratio simultaneously meet the threshold conditions, the newly detected area is considered a continuation observation of the same risk event at an adjacent time. The system then performs frame interpolation to complete the missing time periods based on the historical time-series data. Preferably, only when the similarity of appearance features and the IoU ratio simultaneously meet the threshold conditions are they considered the same risk event. Frame interpolation can employ linear interpolation, spline interpolation, or trajectory completion based on motion models to restore continuous trajectories in time gaps caused by occlusion, short-term missed detections, or detection jitter. After frame interpolation, the system writes the completed observation points into the corresponding risk event instance and updates its time axis, spatial trajectory, and area sequence to maintain temporal continuity for the instance.
[0126] This implementation method, by jointly utilizing appearance feature similarity and bounding box intersection-union ratio for matching, can reliably identify the same risk object even when the surgical field is obscured by smoke, instruments are briefly obscured, or the viewpoint is jittery. Furthermore, by combining historical time-series data for frame interpolation, risk event instances can maintain a continuous lifecycle representation, avoiding erroneous splitting into multiple discrete instances, thereby improving the completeness and continuity of risk tracking.
[0127] By adopting this implementation method, the system can reduce cross-frame breaks, missing correlations, and erroneous splits, accurately record the duration, diffusion process, and recovery process of risk events such as bleeding and smoke, and provide a continuous and reliable time-series basis for subsequent risk level prediction and visualization report generation.
[0128] Example 5
[0129] Based on the above Example 1, this paper introduces how to predict the risk level at future moments.
[0130] Figure 4 A flowchart illustrating a method for generating an intraoperative risk report provided in this application. Figure 4 ,like Figure 4 As shown, it includes the following steps:
[0131] S1051: Determine the future area at future moments based on area changes in historical time-series data of risk event instances;
[0132] Among them, area change refers to the difference or rate of change in the size of the region between adjacent time points of risk event instances, which is used to characterize the trend of risk expansion or contraction; future area is the predicted size of the risk region at the predicted time point calculated based on the historical area sequence, which is used to reflect the evolution of risk in a short period of time.
[0133] In practice, the system reads the area sequence arranged by time from the established risk event instances and smooths the area changes of adjacent frames to eliminate fluctuations caused by single-frame noise. When the area sequence shows an upward trend, the system can use exponential smoothing, linear trend extrapolation, or a lightweight temporal network to estimate the size of the region at the next prediction time, thereby obtaining the future area.
[0134] S1052: Calculate the comprehensive risk score based on the future area, the current area, the number of currently concurrent risk areas, and the duration of risk event instances from the start to a future time.
[0135] The current area is the actual size of the region corresponding to the prediction baseline time, used to characterize the immediate risk range; the number of concurrent risk regions is used to characterize the number of risk events existing at the same time, reflecting the degree of risk superposition; the duration refers to the length of time from the first occurrence of a risk event instance to the predicted future time, used to reflect the cumulative effect of risk exposure; the comprehensive risk score is an intermediate evaluation value after uniformly quantifying the area, number of concurrent events, and duration, used to output the final risk level.
[0136] S1053: Determine the risk level based on the comprehensive risk score.
[0137] The future area and the current area together constitute the area evolution characteristics, which are input into the scoring model along with the number and duration of currently concurrent risk areas. The scoring model can be a weighted summation model or a nonlinear mapping model fitted from training samples. The comprehensive risk score output by the scoring model can be mapped to multiple risk level threshold intervals. When the score is below the low-risk threshold, it is classified as low-level; when the score is in the middle interval, it is classified as medium-level; and when the score exceeds the high-risk threshold, it is classified as high-level or critical-level. By incorporating the continuously expanding area, the number of concurrent risk areas, and the long duration into the score, the risk level can be made more consistent with the actual intraoperative risk status.
[0138] This method enables forward estimation of the risk range at future moments based on continuously updated risk event instances, and unifies multiple temporal characteristics into a comparable comprehensive risk score, thereby achieving early judgment of risk escalation trends. Because it simultaneously considers future area, current area, number of concurrent risks, and duration, the resulting risk level reflects not only the current state but also the speed and accumulation of risk development. Therefore, it improves the continuity and accuracy of intraoperative risk assessment and provides consistent quantitative evidence for real-time alerts and postoperative review.
[0139] Example 6
[0140] Based on Example 1 above, dynamic and complex events may occur during the surgery, such as the splitting, merging, and restoration of occlusion in risk areas. In such cases, it is necessary to update the risk event instances, which will be described below with an example.
[0141] In one specific implementation, if the total area of multiple sub-risk regions in a risk event instance is greater than a first preset multiple of the initial total area, then the multiple sub-risk regions are regarded as a new risk event instance; wherein, the first preset multiple is greater than 1; if the area of the merged region of multiple sub-risk regions in a risk event instance is less than a second preset multiple of the sum of the areas of each sub-risk region, then the sub-risk region with the largest area is regarded as the region of the risk event instance; wherein, the second preset multiple is less than 1; if the Mahalanobis distance between the location of the newly emerging risk region and the predicted location of the historical risk event instance is less than a preset distance value, and the appearance similarity is greater than a preset similarity threshold, then the two are determined to be the same risk event instance, and the interrupted trajectory frame is interpolated and completed.
[0142] Multiple sub-risk areas refer to local risk portions formed by the same risk event instance, used to characterize the multi-regional association state caused by bleeding spread, smoke dispersion, or field-of-view obstruction. The merged area refers to the total area obtained after masking, boundary joining, or connected component merging of multiple sub-risk areas, used to determine whether multiple local areas have formed an overall risk. The initial total area refers to the area baseline recorded when the risk event instance is first established, used as a comparison object for subsequent area change judgments. The predicted location refers to the future spatial location of historical risk event instances calculated based on motion models, used for continuity matching with newly emerging areas after short-term disappearance. Mahalanobis distance is a distance index used to measure the statistical deviation between the actual observed location and the predicted location, which can combine trajectory covariance to suppress misjudgments caused by lens shake and local deformation. Interrupted trajectory frames refer to trajectory frames that were not continuously recorded due to obstruction, missed detection, or short-term disappearance, used to represent missing segments in the event trajectory.
[0143] In its implementation, after obtaining a risk event instance, the system extracts the mask, bounding box, and centroid coordinates of multiple sub-risk regions within the same instance, and calculates their directly summed total area, as well as the area of the merged region. If the total area of the sub-risk regions is significantly larger than the initial total area and exceeds the range defined by a first preset multiple (greater than 1), it indicates that these local regions have visually presented an overall interconnected state. This is represented by risk point splitting, which the system reconstructs into a new risk event instance to avoid misrecording the aggregated risk as multiple independent events. When the area of the merged region is less than a second preset multiple (less than 1) of the sum of the areas of the sub-risk regions, the risk points need to be merged, retaining only the sub-risk region with the largest area as the region of the current risk event instance, while the remaining regions are recorded as subordinate regions or temporarily ignored to maintain the stability of the event representation.
[0144] When a new risk area reappears, the system first outputs the predicted location based on the historical trajectory model, then calculates the Mahalanobis distance between that location and the centroid of the new area, and simultaneously extracts color histograms, texture features, or depth features to calculate appearance similarity. If the Mahalanobis distance is less than a preset distance value and the appearance similarity is greater than a preset similarity threshold, then the two are determined to belong to the same risk event instance, and the new area is merged into the original instance. Simultaneously, missing trajectory frames during occlusion are filled in using linear interpolation, spline interpolation, or kinematic model-based point completion methods, thus forming a continuous trajectory record. This processing method enables risk event instances to maintain a unified identifier in splitting, merging, and short-term interruption scenarios, and continuously outputs stable spatial location and temporal series information.
[0145] This implementation method, by setting criteria for merging enlarged areas, retaining reduced areas, and jointly matching locations and appearances, enables the system to automatically distinguish three complex scenarios: risk aggregation, dominant region retention, and occlusion recovery. This improves the continuity and consistency of risk event instances. After trajectory interpolation completion, the timeline of risk events is more complete, and area changes and location drifts are easier to statistically analyze, which is beneficial for subsequent risk level prediction and visualization report generation. It also reduces redundant modeling and erroneous associations caused by fragmented detection.
[0146] By analyzing surgical videos, identifying events, generating structured data, and outputting reports, the system automatically detects and records potential risks during surgery, forming a complete technological closed loop from video data to risk report generation. The entire process is described below.
[0147] Step 1: Acquisition of surgical video and preprocessing of video frames;
[0148] 1.1 Acquisition of surgical videos;
[0149] The surgical procedure is recorded in real time using video acquisition equipment (laparoscopy, endoscopy, etc.) in the operating room to obtain complete surgical video data.
[0150] 1.2 Frame Extraction;
[0151] (1) Extract video frames according to the preset frame rate;
[0152] (2) Adaptive frame rate strategy:
[0153] When a surgical device activation signal is received (electrosurgical, ultrasonic scalpel, etc.), switch to high frame rate mode (15-30fps).
[0154] When there is no device signal, the base frame rate (1-5fps) is used.
[0155] 1.3 Image normalization processing;
[0156] The image size is normalized to meet the model input requirements (640×640 or 1024×1024), and the colors are standardized.
[0157] Step 2: Risk detection and intra-frame relationship modeling;
[0158] This step expands isolated risk events into a network of interconnected events, and by analyzing the spatial relationships between these events, it upgrades the detection process from individual risk events to a global risk situation awareness.
[0159] 2.1 Single-frame risk detection and node construction;
[0160] (1) Input the current frame image into a deep learning detection network (YOLO or Mask R-CNN) to identify risk areas;
[0161] (2) Treat each risk area as a risk node and record it:
[0162] Visual features: High-dimensional semantic features of risk regions are extracted using a deep feature extraction network;
[0163] Geometric features: center coordinates, area;
[0164] 2.2 Single-frame risk relationship construction;
[0165] (1) Treat each risk area (bleeding point, smoke, etc.) detected in the current frame as a risk node and record its visual features (color, texture) and spatial location.
[0166] (2) Establish spatial association edges between risk nodes: If the distance between two risk areas on the image is less than the threshold (corresponding to an actual distance of 5cm), then establish an association edge to indicate that they may affect each other (such as smoke generated by bleeding points).
[0167] 2.3 Risk Situation Inference within a Single Frame: Based on a risk association network, the following information is inferred:
[0168] Risk level: Isolated bleeding vs. multi-regional bleeding (the latter is more serious);
[0169] Spatial trend: The direction of expansion of the risk area (e.g., bleeding penetrating into surrounding tissues);
[0170] 2.4 Adaptive calibration during the surgical phase;
[0171] (1) Obtain the current surgical stage identifier:
[0172] The activation signal of the electrosurgical unit corresponds to the cutting stage;
[0173] The suture device signal corresponds to the suturing stage;
[0174] No equipment signal corresponds to the exploration phase;
[0175] (2) Dynamically adjust the detection threshold according to the surgical stage:
[0176] Cutting phase: Lower the bleeding detection threshold (expect bleeding to be normal);
[0177] Exploration phase: Increase the bleeding detection threshold (unexpected bleeding is considered a true risk);
[0178] Step 3: Time-series tracking and evolution analysis of risk events;
[0179] This step associates the discrete risk regions detected in step two with continuous risk event instances, extracts complete time dimension information, and solves the problem that single-frame detection cannot obtain the duration and development trend of risks.
[0180] 3.1 Risk event instance initialization and matching;
[0181] 3.1.1 New instance initialization;
[0182] For risk areas in the current frame that do not match existing instances, generate new event instance IDs and record the initial timestamp T_start and initial features.
[0183] 3.1.2 Three-level cascading matching strategy;
[0184]
[0185] 3.1.3 Instance status management;
[0186] Match successful: Update instance features (exponential moving average), last occurrence time T_end, and spatial trajectory;
[0187] If there is no match for 3 consecutive frames: mark it as "temporarily disappeared", retain the instance but do not update the features;
[0188] If there are 30 consecutive frames (approximately 1 second) without a match or the video ends: mark it as "End" state and calculate the total duration ΔT;
[0189] 3.2 Handling complex event patterns;
[0190] To address the dynamic changes in risk events during surgery, the following special cases should be handled:
[0191]
[0192] 3.3 Temporal parameter extraction and structuring;
[0193] For each risk event instance, extract five-dimensional feature parameters:
[0194]
[0195] 3.4 Risk evolution prediction and real-time early warning;
[0196] Predicting future evolution based on historical time-series data:
[0197]
[0198] Application of prediction results: If the predicted risk level will escalate within 5 seconds (e.g., from level 2 to level 4), a real-time alarm will be triggered to alert the doctor.
[0199] 3.5 Calculation of comprehensive risk score;
[0200] A comprehensive risk score is calculated based on factors such as the duration, geographical scope, and concurrency of risk events.
[0201] Risk Score = f(duration, area, number of concurrent events);
[0202] There are no restrictions on the specific function.
[0203] Step 4: Generation of structured data for risk events;
[0204] 4.1 Structured risk event data;
[0205] Each risk event instance is converted into a predefined structured record, building a hierarchical data model:
[0206]
[0207] 4.2 Surgical-grade summary generation;
[0208] Total number of risk events, distribution of each type, and total duration of risk events;
[0209] Risk density index (number of risk events per unit time).
[0210] Maximum concurrent risk number and risk level change curve;
[0211] 4.3 Data persistence and standardized output;
[0212] Structured records are stored in a time-series database, and can be exported in standard formats such as JSON / XML.
[0213] Step 5: Adaptive risk report generation;
[0214] 5.1 The report template is adaptively selectable;
[0215] A generation strategy is selected from the report template library based on surgical metadata (surgery type, risk level history, physician preferences).
[0216] 5.2 Report content is generated in layers;
[0217]
[0218] Details of text summarization layer generation:
[0219] Input: Structured event logs + surgical metadata;
[0220] Solution: Inject clinical report writing guidelines through Few-shot Prompting to ensure that the generated content only contains objective descriptions (avoiding diagnostic conclusions);
[0221] Output: Chapter-by-chapter natural language text;
[0222] 5.3 Multimodal report assembly;
[0223] Combine text, charts, keyframe images, and video clip links into a unified report document;
[0224] Supported output formats: PDF, HTML, JSON;
[0225] Step Six: Visual Interactive Interface;
[0226] 6.1 Timeline view;
[0227] A horizontal timeline displays the entire surgical process, with risk events marked with color blocks;
[0228] Supports zooming (multi-level zoom from 1 second to 10 minutes to 1 hour) and dragging;
[0229] Clicking on a color block will jump to the corresponding frame in the video, automatically playing 10-second clips before and after it.
[0230] 6.2 Event Card View;
[0231] The grid layout displays all risk events, with cards including: thumbnail, type label, duration, and risk level;
[0232] Supports filtering (by type / level / time period) and sorting (by time / severity);
[0233] Drag and drop cards to adjust time boundaries (corrects T_start / T_end);
[0234] 6.3 Report editing view;
[0235] Edit the text summary layer content;
[0236] Annotation function: Doctors can add text annotations or voice annotations (automatically converted to text);
[0237] Version Comparison: Highlights the differences between the AI-generated version and the doctor-modified version.
[0238] The following is a complete example:
[0239] Example 1: Laparoscopic cholecystectomy
[0240] Application scenario: Laparoscopic cholecystectomy in the general surgery department of a tertiary hospital
[0241] Implementation steps:
[0242] 1. Video acquisition and preprocessing;
[0243] Obtain the entire surgical procedure via laparoscopic equipment;
[0244] Frames are sampled at 1fps (probing phase) / 30fps (cutting phase).
[0245] The uniform image size is 1024×1024;
[0246] 2. Risk detection and spatial modeling;
[0247] Input a pre-trained model (such as YOLOv8) to identify bleeding points and surgical smoke;
[0248] Construct a risk association network to identify multi-regional linked bleeding;
[0249] Based on the electrosurgical unit signal, the cutting stage is identified, and the bleeding detection threshold is reduced;
[0250] 3. Time series tracking and evolutionary analysis;
[0251] A three-level cascaded matching method is used to track risk areas;
[0252] Managing hematoma rupture and spread (split event) and multiple bleeding pools (combined event);
[0253] Extract five-dimensional feature parameters and establish a structured record;
[0254] 4. Risk prediction and early warning;
[0255] Run an LSTM prediction model on persistent bleeding events;
[0256] If the probability of risk escalation within 10 seconds is predicted to be greater than 70%, a real-time alarm will be triggered.
[0257] 5. Report generation;
[0258] Choose a template specifically designed for laparoscopic cholecystectomy;
[0259] Generate an intraoperative risk report containing the following:
[0260] a) Summary of the procedure: The operation lasted 65 minutes, the total risk time was 8 minutes, and there were 5 risk events;
[0261] b) Event statistics table: 3 bleeding events (2 level 2, 1 level 3), 2 smoke events;
[0262] c) Timeline Gantt Chart: X-axis represents time, Y-axis represents risk events, and the length of the colored block indicates the duration;
[0263] d) Risk heatmap: A semi-transparent red mask overlaid on surgical keyframes;
[0264] Output a PDF report;
[0265] 6. Visual interaction;
[0266] The doctor located the risk event at the 23rd minute using the timeline view;
[0267] View the 10-second video clips before and after autoplay;
[0268] Fine-tune and correct the event duration, and add text annotations;
[0269] In summary, the method of this application has the following advantages:
[0270] (1) Automatic identification mechanism for intraoperative risk events based on surgical videos.
[0271] This invention analyzes surgical videos frame by frame and uses deep learning object detection models (such as YOLO and Mask R-CNN) to automatically identify potential risk events during surgery, including but not limited to bleeding, surgical smoke, and foreign objects. This enables automatic detection of potential risks during surgery. Compared to existing technologies that infer surgical steps solely by identifying surgical instruments, this invention directly identifies risk events during surgery, improving the ability to perceive surgical safety.
[0272] (2) Extraction of risk event time information and construction of time axis.
[0273] This invention automatically extracts the occurrence time, duration, and frequency of detected risk events through continuous frame analysis, and constructs a risk event sequence on the surgical video timeline. By structurally recording the temporal information of risk events, intraoperative risk information can be intuitively displayed in the form of a timeline, facilitating postoperative analysis and surgical quality assessment.
[0274] (3) Structured representation of risk events and data management mechanism.
[0275] This invention converts detected risk events into a structured data format for storage. For example, it records information such as event type, start time, end time, and duration through a unified data structure, thereby forming a standardized intraoperative risk event dataset. This structured representation facilitates subsequent data analysis, statistics, and system expansion.
[0276] (4) An automated risk report generation method based on risk event information.
[0277] After obtaining structured risk event data, this invention further utilizes an automated report generation module to comprehensively analyze and summarize the risk events occurring during the surgery, and automatically generate an intraoperative risk report. This report reflects information such as the type, frequency, and duration of risk events during the surgery, providing doctors with crucial information for surgical risk assessment and postoperative review.
[0278] (5) Enable surgical video analysis and medical document generation.
[0279] This invention combines intelligent surgical video analysis technology with automatic report generation technology to realize a complete technical process from surgical video data to risk event identification and risk report generation. This improves surgical risk analysis capabilities while reducing the workload of doctors in manually recording and organizing information, thereby improving the efficiency of medical information processing.
[0280] Figure 5 A schematic diagram of a device for generating intraoperative risk reports provided in this application is shown below. Figure 5 As shown, the device 40 for generating intraoperative risk reports provided in this embodiment includes:
[0281] Acquisition module 401 is used to acquire video data during the surgical procedure;
[0282] The risk identification module 402 is used to determine the risk areas in each frame of the video based on the video data;
[0283] The risk event association module 403 is used to generate risk event instances based on the location, area, and time of occurrence of the risk area;
[0284] The matching update module 404 is used to perform feature matching between newly detected risk areas and existing risk event instances during the operation. If the match is successful, the newly detected risk areas are associated with existing risk event instances, and the risk event instances are updated.
[0285] Prediction module 405 is used to predict the risk level at future moments based on historical time-series data of risk event instances;
[0286] Output module 406 is used to output a visual surgical report including risk level.
[0287] Optionally, the risk identification module 402 is specifically used for:
[0288] Determine the current stage of surgery based on the surgical instruments shown in the video data;
[0289] Determine the risk detection threshold based on the stage of surgery;
[0290] Based on the risk detection threshold, the parameters in the deep learning detection network are adjusted, and the risk areas in the video data are identified through the deep learning detection network.
[0291] Optionally, the device also includes a single-frame risk identification module for:
[0292] If the distance between two adjacent risk areas is less than a preset distance, then the two risk areas are determined to be related risk areas.
[0293] By using deep learning networks to identify associated risk areas, the probability of multi-region linkage risks within a single frame can be determined.
[0294] Output the probability of multi-regional linkage risks.
[0295] Optionally, the 404 matching update module is specifically used for:
[0296] Calculate the similarity of appearance features between the newly detected risk regions and existing risk event instances, and the intersection-union ratio of their bounding boxes;
[0297] If the appearance feature similarity is greater than the similarity threshold, and / or the crossover ratio is greater than the crossover ratio threshold, then frame interpolation is performed based on the historical time series data of the risk event instances and the newly detected risk areas to obtain continuous risk event instances.
[0298] Optionally, the prediction module 405 is specifically used for:
[0299] Determine the future area at future moments based on area changes in historical time-series data of risk event instances;
[0300] A comprehensive risk score is calculated based on the future area, the current area, the number of currently concurrent risk areas, and the duration of risk event instances from their start to a future time.
[0301] The risk level is determined based on the comprehensive risk score.
[0302] Optionally, the device also includes a dynamic processing module for:
[0303] If the total area of multiple sub-risk areas in a risk event instance is greater than a first preset multiple of the initial total area, then the multiple sub-risk areas are treated as a new risk event instance; wherein, the first preset multiple is greater than 1.
[0304] If the combined area of multiple sub-risk areas in a risk event instance is greater than a second preset multiple of the sum of the areas of the individual sub-risk areas, then the sub-risk area with the largest area will be taken as the area of the risk event instance; wherein the second preset multiple is less than 1.
[0305] If the Mahalanobis distance between the location of the newly emerging risk area and the predicted location of a historical risk event instance is less than a preset distance value, and the appearance similarity is greater than a preset similarity threshold, then the two are identified as the same risk event instance, and the interrupted trajectory frame is interpolated to complete the frame.
[0306] The apparatus provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0307] Figure 6 A schematic diagram of the structure of the electronic device provided in this application. Figure 6 As shown, the electronic device 50 provided in this embodiment includes at least one processor 501 and a memory 502. Optionally, the device 50 further includes a communication component 503. The processor 501, memory 502, and communication component 503 are connected via a bus 504.
[0308] In a specific implementation, at least one processor 501 executes computer execution instructions stored in memory 502, causing at least one processor 501 to perform the above-described method.
[0309] The specific implementation process of processor 501 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0310] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0311] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0312] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0313] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0314] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0315] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0316] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0317] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0318] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0319] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0320] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0321] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0322] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A method for generating an intraoperative risk report, characterized in that, The method includes: Acquire video data during the surgical procedure; Based on the video data, the risk areas in each frame of the image are determined; Based on the location and occurrence time of the risk area, a risk event instance is generated, which includes an instance ID, a set of time series observation points, a set of spatial trajectories, and an area sequence. The risk event instance represents a continuous event unit formed by associating discrete risk areas belonging to the same risk object in multiple frames. During the procedure, newly detected risk regions are matched with existing risk event instances. If a match is successful, the newly detected risk regions are associated with existing risk event instances, and the risk event instances are updated. The feature matching includes: calculating the appearance feature similarity and the intersection-union ratio (IUR) of the bounding boxes between the newly detected risk regions and existing risk event instances. If the appearance feature similarity is greater than a similarity threshold and the IUR is greater than the IUR threshold, then the match is successful. Based on the area changes in historical time-series data of risk event instances, the future area at a future time is determined; based on the future area, the current area, the number of currently concurrent risk regions, and the duration of the risk event instance from its start to the future time, a comprehensive risk score is calculated; based on the comprehensive risk score, the risk level at the future time is determined. The output includes a visualized surgical report showing the risk level.
2. The method according to claim 1, characterized in that, The step of determining the risk area in each frame of the image based on the video data includes: Based on the surgical instruments in the video data, determine the current stage of the surgery; Based on the surgical stage, determine the risk detection threshold; Based on the risk detection threshold, the parameters in the deep learning detection network are adjusted, and the risk areas in the video data are identified through the deep learning detection network.
3. The method according to claim 2, characterized in that, The method further includes: If the distance between two adjacent risk areas is less than a preset distance, then the two risk areas are determined to be related risk areas. The associated risk areas are identified using a deep learning network to determine the probability of multi-region linkage risks within a single frame. Output the probability of the multi-regional linkage risk.
4. The method according to any one of claims 1-3, characterized in that, If a match is successful, the newly detected risk area is associated with an existing risk event instance, and the risk event instance is updated, including: After a successful match, frame interpolation is performed based on the historical time-series data of the risk event instances and the newly detected risk areas to obtain continuous risk event instances.
5. The method according to any one of claims 1-3, characterized in that, The method further includes: If the total area of multiple sub-risk areas in a risk event instance is greater than a first preset multiple of the initial total area, then the multiple sub-risk areas are treated as a new risk event instance; wherein, the first preset multiple is greater than 1. If the combined area of multiple sub-risk areas in a risk event instance is less than a second preset multiple of the sum of the areas of the sub-risk areas, then the sub-risk area with the largest area is taken as the area of the risk event instance; wherein, the second preset multiple is less than 1. If the Mahalanobis distance between the location of the newly emerging risk area and the predicted location of a historical risk event instance is less than a preset distance value, and the appearance similarity is greater than a preset similarity threshold, then the two are identified as the same risk event instance, and the interrupted trajectory frame is interpolated to complete the frame.
6. A device for generating intraoperative risk reports, characterized in that, The device includes: The acquisition module is used to acquire video data during the surgical procedure; The risk identification module is used to determine the risk areas in each frame of the video based on the video data; The risk event association module is used to generate risk event instances containing an instance ID, a set of time series observation points, a set of spatial trajectories, and an area sequence based on the location, area, and occurrence time of the risk area. The risk event instance represents a continuous event unit formed by associating discrete risk areas belonging to the same risk object in multiple frames. The matching and updating module is used during surgery to perform feature matching between newly detected risk regions and existing risk event instances. If the match is successful, the newly detected risk regions are associated with the existing risk event instances, and the risk event instances are updated. The feature matching includes: calculating the appearance feature similarity and the intersection-union ratio (IUR) of the bounding boxes between the newly detected risk regions and existing risk event instances; if the appearance feature similarity is greater than a similarity threshold and the IUR is greater than the IUR threshold, then the match is successful. The prediction module is used to determine the future area at a future time based on the area changes in the historical time-series data of risk event instances; calculate a comprehensive risk score based on the future area, the current area, the number of currently concurrent risk areas, and the duration of the risk event instance from its start to the future time; and determine the risk level at the future time based on the comprehensive risk score. The output module is used to output a visual surgical report that includes the risk level.
7. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-5.
9. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-5.
Citation Information
Patent Citations
Hemorrhagic spot detection method and device and computer equipment
CN117670814A
Urinary operation risk AI intelligent assessment method and system
CN121075635A