A medical video content semantic segmentation labeling method fusing a knowledge graph
Patent Information
- Application Number
- CN202610840027.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-06-11
AI Technical Summary
现有基于滑动窗口分类的医学视频分段标注方法在窗口级合并阶段通常更关注阶段候选置信结果或相邻窗口标签相似度,未充分利用医学知识图谱对相邻窗口的流程顺序、实体连续性和操作关联进行约束,在腔镜手术质控视频中,血液遮挡、水雾遮挡、运动模糊、反光、器械短时进入以及边界不稳定等因素可能使某一视频窗口产生局部高置信错误阶段标签;若该窗口被直接作为阶段切换依据,则容易在持续同一医学流程阶段中插入短时跳转片段,造成片段边界不准确、阶段标签不一致以及全局流程顺序异常;
本发明通过在滑动窗口阶段分类之后引入跳转风险结果、目标医学知识子图、窗口候选流程状态、状态可靠结果、转移代价结果和局部重解码结果,使医学视频分段标注不再仅依据单个窗口的最高阶段候选置信结果或相邻窗口标签相似度进行合并,而是将窗口视觉扰动、候选阶段置信差距、医学实体匹配关系、相邻阶段顺序关系和实体连续关系共同纳入流程状态判断。由此,当腔镜手术质控视频中出现血液遮挡、水雾遮挡、边界不稳定或器械短时进入等情况时,识别该类窗口可能引发的局部高置信异常,并通过医学流程状态图和路径代价递推抑制其直接形成独立片段;同时,对于已经生成的短时跳转片段、流程冲突片段或低可信片段,能够在异常窗口集合内执行局部重解码,对片段边界和片段标签进行修正,使输出的医学视频语义分段标注结果与目标医学知识子图中的阶段先后关系、解剖实体关系、器械实体关系及操作实体关系保持一致,提高腔镜手术质控视频结构化标注、流程复核和片段检索的稳定性与可解释性。
Smart Images

Figure CN122416455B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent annotation technology, and more specifically, to a method for semantic segmentation annotation of medical video content that integrates knowledge graphs. Background Technology
[0002] Existing medical video analysis technologies typically begin by extracting frames, dividing into sliding windows, or segmenting segments from laparoscopic surgery videos, endoscopic examination videos, or medical teaching videos. Then, image classification models, video classification models, temporal classification models, semantic segmentation models, or multi-scale visual networks are used to identify anatomical regions, instruments, target boundaries, noise levels, and stage labels within the video windows. Based on the stage classification results, adjacent video windows are merged through voting or similarity merging, thereby forming stage segments and label results for the medical video. This allows for the organization of medical knowledge through stage entities, anatomical entities, instrument entities, operational entities, as well as stage sequence relationships, stage-corresponding anatomical relationships, stage-corresponding instrument relationships, and stage-corresponding operational relationships. This provides a structured knowledge foundation for medical entity normalization, process relationship expression, and semantic verification.
[0003] The existing technology has the following shortcomings: Existing medical video segmentation and annotation methods based on sliding window classification typically focus more on candidate confidence results or the similarity of adjacent window labels during the window-level merging stage. They do not fully utilize medical knowledge graphs to constrain the process sequence, entity continuity, and operational associations of adjacent windows. In laparoscopic surgery quality control videos, factors such as blood occlusion, water mist occlusion, motion blur, reflection, short-term instrument entry, and unstable boundaries may cause a video window to generate a locally high-confidence erroneous stage label. If this window is directly used as the basis for stage switching, it is easy to insert short-term jump segments in the same continuous medical process stage, resulting in inaccurate segment boundaries, inconsistent stage labels, and abnormal global process sequence. There is an urgent need for a medical video semantic segmentation and annotation method that can combine target medical knowledge subgraphs, jump risk judgment, global process path decoding, and local re-decoding after the sliding window stage classification, so as to reduce the impact of local high-confidence abnormal windows on the segmentation and labeling of medical process segments.
[0004] To address the above problems, this invention proposes a solution. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a semantic segmentation and annotation method for medical video content that integrates knowledge graphs, in order to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: A semantic segmentation and annotation method for medical video content that integrates knowledge graphs, comprising the following steps; Step S1: Acquire the target medical video and divide it into video windows according to the preset window length and window step size. Perform visual analysis and stage classification on the video windows to obtain the window visual analysis results and window stage candidate results. Generate jump risk results based on the visual perturbation information in the window visual analysis results and the candidate confidence gap in the window stage candidate results. Step S2: Based on the candidate results of the window stage and the visual analysis results of the window, the target medical knowledge subgraph is selected from the medical knowledge graph. The candidate stages of each video window are mapped to the window candidate process state, which includes stage entities, anatomical entities, instrument entities and operation entities. Reliable state results are generated based on the knowledge path relationship between the candidate stages and the corresponding entities. Step S3: Construct a medical process state graph based on the candidate process states and reliable state results of the window, establish state transition relationships between the candidate process states of adjacent video windows, determine the transition cost results based on the process sequence relationship, entity continuity relationship and jump risk results in the target medical knowledge subgraph, and perform global process path decoding based on the reliable state results and transition cost results to obtain the global process path results and generate the initial semantic fragment annotation results. Step S4: Perform abnormal segment detection on the initial semantic segment annotation results. For segments that meet the abnormal conditions, perform local re-decoding within the corresponding abnormal window set. Correct the segment boundaries or segment stage labels based on the local re-decoding results, and output the corrected medical video semantic segmentation annotation results.
[0007] In a preferred embodiment, step S1 includes the following: Obtain multiple video frames arranged in chronological order and the corresponding time position of each video frame, and identify the multiple video frames as the target medical video; The window length is determined based on the shortest effective duration of the stage to be identified, the window step size is determined based on the window length, and the target medical video is divided into sliding segments according to the window length and window step size to obtain video windows arranged in chronological order. For each video window, extract the window representative frame. The window representative frame is determined by the frame clarity result and the inter-frame variation result of the video frames within the video window. Perform surgical scene visual analysis on the window representative frame to obtain the window visual analysis result; The window visual analysis results include anatomical region results, instrument object results, target boundary results, noise status results, and pixel label confidence results; The noise occlusion result is determined based on the noise state result, the boundary instability result is determined based on the target boundary result, and the visual perturbation result is determined based on the noise occlusion result and the boundary instability result. The video window is classified into stages to obtain candidate window stages, including candidate stage labels and candidate stage confidence results; The confidence gap result is determined based on the difference between the highest-level candidate confidence result and the second-highest-level candidate confidence result, and the jump risk result is determined based on the confidence gap result and the visual perturbation result.
[0008] In a preferred embodiment, step S2 includes the following: Obtain a medical knowledge graph that includes a set of medical entities and a set of medical relationships; The medical entity set includes stage entities, anatomical entities, instrument entities, and operational entities; The medical relationship set includes stage sequence relationships, stage-corresponding anatomical relationships, stage-corresponding instrument relationships, stage-corresponding operational relationships, entity synonym relationships, and entity hierarchical relationships; Based on the surgical type, window stage candidate results, anatomical region results, and instrument object results corresponding to the target medical video, entities and relationships related to the target medical video are filtered from the medical knowledge graph to obtain the target medical knowledge subgraph; Map the candidate stages in the window stage candidate results to stage entities in the target medical knowledge subgraph; Select the anatomical entities that have a stage-corresponding anatomical relationship with the stage entities based on the anatomical region results in the video window. Select the device entities that have a corresponding device relationship with the stage entity based on the device object results in the video window; Based on the spatial contact relationship between the instrument and the anatomical region, the target boundary results, and the positional changes of the instrument in adjacent windows, select the operation entity that has a stage-corresponding operation relationship with the stage entity; The window candidate process status consists of stage entities, anatomical entities, instrument entities, and operation entities; For each candidate stage, calculate the shortest path from the stage entity to the corresponding anatomical entity, instrument entity, and operation entity to obtain the knowledge path distance result. The state reliability result is determined based on the candidate confidence result of the stage and the knowledge path distance result, such that the greater the knowledge path distance between the candidate stage and the corresponding entity, the lower the state reliability result.
[0009] In a preferred embodiment, step S3 includes the following: A medical process state diagram is constructed based on the candidate process states of windows and the reliable results of the states. The medical process state diagram includes window layers arranged in chronological order, and each window layer includes the candidate process states of the corresponding video window. Establish state transition relationships between adjacent window layers; The knowledge order distance result is determined based on the stage sequence relationship of adjacent candidate process states corresponding to the stage entities in the target medical knowledge subgraph. The continuity of entities is determined based on the degree of continuity between the anatomical entities and instrument entities corresponding to adjacent candidate process states, and the transfer cost is determined in combination with the jump risk result. Based on the reliable state results, the path cost results of the previous video window, and the transition cost results between adjacent candidate process states, the path cost results of each candidate process state in each video window are recursively obtained. Record the state of the previous candidate process that minimizes the path cost, and form the path backtracking result; The global process path result is obtained based on the path backtracking result, and the video windows are merged based on the global process path result to generate the initial semantic segment annotation result.
[0010] In a preferred embodiment, step S4 includes the following: Anomaly detection is performed on the initial semantic segment annotation results to obtain short-term jump segments, process conflict segments, or low-confidence segments; Construct an abnormal window set based on the video windows inside the abnormal segment, the video window at the end of the semantic segment preceding the abnormal segment, and the video window at the beginning of the semantic segment following the abnormal segment; The segment risk result is determined based on the jump risk result of each video window in the abnormal window set, and the segment window number result is determined based on the number of video windows contained in the abnormal segment; When the segment risk result meets the risk condition and the segment window quantity result meets the short segment condition or the abnormal segment has a process reversal relationship; Re-execute global path decoding within the abnormal window set, and correct segment boundaries or segment stage labels based on the local re-decoding results, outputting the corrected medical video semantic segmentation annotation results.
[0011] The technical effects and advantages of the semantic segmentation and annotation method for medical video content integrating knowledge graphs in this invention are as follows: This invention introduces jump risk results, target medical knowledge subgraphs, window candidate process states, state reliability results, transition cost results, and local re-decoding results after the sliding window stage classification. This makes medical video segmentation annotation no longer based solely on the highest stage candidate confidence results of a single window or the similarity of adjacent window labels. Instead, it incorporates window visual perturbation, candidate stage confidence gap, medical entity matching relationship, adjacent stage sequence relationship, and entity continuity relationship into the process state judgment. Therefore, when blood occlusion, water mist occlusion, unstable boundaries, or short-term instrument entry occur in the quality control video of laparoscopic surgery, the system identifies local high-confidence anomalies that may be caused by such windows and suppresses their direct formation into independent segments through medical process state graphs and path cost recursion. At the same time, for already generated short-term jump segments, process conflict segments, or low-confidence segments, local re-decoding can be performed within the set of anomaly windows to correct segment boundaries and segment labels. This ensures that the output medical video semantic segmentation annotation results are consistent with the stage sequence relationships, anatomical entity relationships, instrument entity relationships, and operation entity relationships in the target medical knowledge subgraph, thereby improving the stability and interpretability of structured annotation, process review, and segment retrieval of laparoscopic surgery quality control videos. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of the process of a semantic segmentation and annotation method for medical video content that integrates knowledge graphs according to the present invention. Detailed Implementation The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] Example Please see Figure 1 As shown, this invention discloses a semantic segmentation and annotation method for medical video content that integrates knowledge graphs, including the following steps: Step S1: Acquire the target medical video and divide it into video windows according to the preset window length and window step size. Perform visual analysis and stage classification on the video windows to obtain the window visual analysis results and window stage candidate results. Generate jump risk results based on the visual perturbation information in the window visual analysis results and the candidate confidence gap in the window stage candidate results. Step S2: Based on the candidate results of the window stage and the visual analysis results of the window, the target medical knowledge subgraph is selected from the medical knowledge graph. The candidate stages of each video window are mapped to the window candidate process state, which includes stage entities, anatomical entities, instrument entities and operation entities. Reliable state results are generated based on the knowledge path relationship between the candidate stages and the corresponding entities. Step S3: Construct a medical process state graph based on the candidate process states and reliable state results of the window, establish state transition relationships between the candidate process states of adjacent video windows, determine the transition cost results based on the process sequence relationship, entity continuity relationship and jump risk results in the target medical knowledge subgraph, and perform global process path decoding based on the reliable state results and transition cost results to obtain the global process path results and generate the initial semantic fragment annotation results. Step S4: Perform abnormal segment detection on the initial semantic segment annotation results. For segments that meet the abnormal conditions, perform local re-decoding within the corresponding abnormal window set. Correct the segment boundaries or segment stage labels based on the local re-decoding results, and output the corrected medical video semantic segmentation annotation results.
[0014] In step S1, the target medical video is acquired and divided into video windows according to a preset window length and window step size. Visual analysis and stage classification are performed on the video windows to obtain the window visual analysis results and candidate stage results. Based on the visual perturbation information in the window visual analysis results and the candidate confidence gaps in the candidate stage results, a jump risk result is generated, specifically including: The quality control video of the laparoscopic surgery to be processed is obtained as the target medical video, which includes multiple video frames arranged in chronological order and the time position corresponding to each video frame; The target medical video is divided into sliding segments according to the preset window length and window step size to obtain the window sequence result. The window sequence result includes several video windows arranged in chronological order. Each video window is a continuous time segment in the target medical video. The window length is determined according to the shortest effective duration of the stage to be identified, and the window step size is less than or equal to the window length. Preferably, the window step size is smaller than the window length, allowing overlapping areas between adjacent video windows. This enables continuous visual evidence from the same stage to be repeatedly acquired and processed across multiple adjacent video windows, thereby reducing the impact of a single anomalous window on the overall segmentation results. For each video window, extract the window representative frame result. The window representative frame result is used to represent several image frames within the video window that can reflect the main visual state. Calculate the frame sharpness and inter-frame variation for each video frame within the video window; The frame sharpness result is used to indicate the degree to which the video frame is affected by motion blur, water mist, blood obscuring, or reflection. The inter-frame variation result is used to indicate the intensity of visual change of the video frame relative to the preceding and following video frames. At the beginning, middle, and end of the video window, select video frames that meet the clarity criteria and do not significantly deviate from the average visual state of the window as the representative frames of the window. If no video frame meets the clarity requirement within a certain time segment, the video frame with the lowest noise occlusion in that time segment is selected as the alternative representative frame. Perform surgical scene visual analysis on the window representative frame results to obtain window visual analysis results, which include anatomical region results, instrument object results, target boundary results, noise state results, and pixel label confidence results; Anatomical region results are used to represent tissue regions, target anatomical regions, or surgical operation areas appearing in the video window; The instrument object result is used to represent the surgical instruments appearing in the video window; The target boundary results are used to represent the boundary distribution between anatomical regions, between instruments and anatomical regions, or between instruments and the background; The noise status result is used to indicate whether there is motion blur, water mist occlusion, blood occlusion, reflection occlusion, or instrument occlusion in the video window; Pixel label confidence results are used to represent the reliability of the visual analysis model in recognizing anatomical region results and instrument object results; Based on the window visual analysis results, the window visual evidence results are constructed, which include noise occlusion results and boundary instability results. The noise occlusion results are used to represent the degree of interference of blood occlusion, water mist occlusion, motion blur and reflection on the visual content of the video window. Specifically, the noise occlusion result is determined based on the percentage of pixels marked as blood occlusion, water mist occlusion, motion blur, or reflective areas in the noise state result, as well as the proportion of such areas appearing continuously in the window representative frame result. The larger the noise masking area and the longer the duration, the higher the noise masking result. The boundary instability result is used to represent the degree of discontinuity of the instrument boundary, tissue boundary, or instrument-tissue contact boundary in the window representative frame result. Specifically, the boundary instability result is determined based on the positional offset of the target boundary result between consecutive representative frames, the boundary breakage ratio, and the change magnitude of the boundary confidence. The more drastic the change in boundary position, the more obvious the fracture, and the greater the fluctuation in boundary confidence, the higher the boundary instability result. Based on the noise masking results and boundary instability results, the first... The visual perturbation results of each video window are denoted as This is used to indicate whether the video window contains visual anomalies sufficient to interfere with stage classification; It should be noted that the noise occlusion result is determined based on the spatial proportion and duration of the occluded area within the video window. The areas in the representative frames of the window that are identified as blood occlusion, water mist occlusion, motion blur, reflection occlusion, or instrument occlusion are identified as noise areas. The area ratio of the noise area in the representative frames of the window is calculated, and the noise occlusion result is obtained by combining the frequency or duration ratio of the noise area in multiple representative frames. The noise occlusion result is high when the noise region is large and appears consecutively in multiple representative frames; the noise occlusion result is low when the noise region appears only briefly in a single representative frame and does not cover the main anatomical area or instrument. The boundary instability results are determined based on the positional changes, continuity changes, and confidence changes of the target boundary between adjacent representative frames. The anatomical region boundary, the instrument object boundary, and the contact boundary between the instrument and the anatomical region are extracted from the target boundary results. Then, the positional offset, breakage ratio, and boundary confidence decrease of the same target boundary in adjacent representative frames are compared. When the target boundary shifts significantly between adjacent representative frames, the proportion of boundary breaks is high, or the boundary confidence continues to decline, the boundary instability result is high. When the target boundary is stable, the boundary is continuous, and the confidence level changes little, the boundary instability result is low. The visual perturbation result is jointly determined by the noise occlusion result and the boundary instability result, and is used to indicate whether the visual evidence in the video window is sufficient to affect the credibility of the stage classification result; Preferably, Take the larger value between the noise masking result and the boundary instability result; For each video window, input the window stage classifier to obtain the window stage candidate results. Use a trained image classification model, video classification model, temporal classification model or visual feature-based classification model. The window stage candidate results include multiple candidate stage labels and their corresponding stage candidate confidence results. Candidate stage labels include one of the following: regional location, tissue exposure, tissue separation, local treatment, hemostasis confirmation, suturing, and result confirmation; For the Each video window is placed in the candidate stage. The stage candidate confidence result is denoted as The value ranges from zero to one; a higher value indicates that the classifier is more likely to consider the first value to be the correct value during the window phase. Each video window belongs to the candidate stage. ; The first The difference between the highest-level candidate confidence result and the second-highest-level candidate confidence result in each video window is determined as the confidence gap, denoted as . , used to indicate the degree of bias of the window stage classifier towards the highest candidate stage; like A larger value indicates that the classifier has a significant bias towards a certain candidate stage; like A smaller value indicates uncertainty in the classifier across multiple candidate stages; Existing window voting merging algorithms often directly use the highest-confidence candidate result as the window stage label. However, in laparoscopic surgery quality control videos, high confidence does not necessarily represent a true stage switch; it could also be an erroneous bias caused by visual perturbations. Therefore, this step further... and The generated redirection risk result is recorded as follows: ; Based on confidence gap results and visual perturbation results Calculate the first Risk of video window redirection results The calculation method is as follows: ;in, For the first Risks associated with redirecting individual video windows; For the first The confidence gap result for each video window is obtained by the difference between the highest-stage candidate confidence result and the second-highest-stage candidate confidence result for that video window. For the first The visual perturbation result of a video window is determined by the noise occlusion result and the boundary instability result of that video window. The risk of skipping only increases significantly when the classifier is highly confident about a certain stage and there is a significant perturbation in the visual evidence of that window at the same time. If a video window has a large confidence gap in its classification but stable visual evidence, it is not directly identified as an anomaly. If a video window has strong visual perturbation but the classifier does not show obvious stage bias, it will not be used as a window that is likely to trigger process jumps. This step ultimately yields window sequence results, window visual analysis results, window visual evidence results, window stage candidate results, and jump risk results. Window stage candidate results are used to retain multiple possible stages for each video window, and jump risk results are used to identify windows that may experience erroneous stage switching due to visual disturbances. By retaining multiple candidate stages and calculating jump risks, subsequent steps can reselect a more reasonable window state under the constraints of the medical process, avoiding the direct formation of independent semantic fragments by a single local high-confidence abnormal window. This step employs a data processing approach combining sliding window partitioning, representative frame filtering, surgical scene visual analysis, and candidate confidence gap calculation. It transforms the target medical video from continuous video frames into a temporally ordered sequence of video windows. Within each video window, it generates window visual analysis results, window visual evidence results, and candidate window stage results. Visual perturbation results are constructed using noise occlusion and boundary instability results. These visual perturbation results are then multiplied with the confidence gap results to obtain the jump risk result. This process identifies locally high-confidence anomalous windows where the classifier is highly confident about a particular stage but simultaneously exhibits visual perturbation. This step retains multiple candidate stages and simultaneously outputs jump risk results, enabling subsequent steps to reassess the window state under medical workflow constraints. This avoids single anomalous windows caused by blood occlusion, water mist occlusion, boundary fluctuations, or short-term instrument entry directly triggering independent semantic fragments.
[0015] In step S2, based on the candidate window stage results and the window visual analysis results, target medical knowledge subgraphs are selected from the medical knowledge graph. The candidate stages of each video window are mapped to window candidate flow states including stage entities, anatomical entities, instrument entities, and operational entities. Reliable state results are generated based on the knowledge path relationships between candidate stages and corresponding entities. Specific content includes: Acquiring a medical knowledge graph includes a set of medical entities and a set of medical relationships. The set of medical entities includes stage entities, anatomical entities, instrument entities, and operational entities. Stage entities are used to represent stages in the surgical quality control process, such as area localization, tissue exposure, tissue separation, local treatment, hemostasis confirmation, suturing, and result confirmation. Anatomical entities are used to represent tissue regions or anatomical structures involved in the target medical video; Instrument entities are used to represent surgical instruments that may appear in the video; Operational entities are used to represent operations such as traction, separation, clamping, cutting, hemostasis, suturing, and rinsing. The medical relation set includes stage sequence relations, stage-corresponding anatomical relations, stage-corresponding instrument relations, stage-corresponding operational relations, entity synonym relations, and entity hierarchical relations. Each medical relation in the medical knowledge graph is represented by a triple, which includes a head entity, relation type, and tail entity. Based on the surgical type corresponding to the target medical video, the candidate results of the window stage obtained in step S1, the anatomical region results, and the instrument object results; The entities and relationships related to the current target medical video are filtered out from the medical knowledge graph to obtain the target medical knowledge subgraph. The target medical knowledge subgraph includes the stage entities, anatomical entities, instrument entities, and operation entities involved in the current surgical quality control task, as well as the stage sequence relationship, stage-corresponding anatomical relationship, stage-corresponding instrument relationship, and stage-corresponding operation relationship between them. The selection scope of the target medical knowledge subgraph is determined based on the type of surgery or examination in the target medical video. Taking the entity of the type of surgery or examination corresponding to the target medical video as the starting entity, the stage entities that have a direct process relationship with the starting entity are retained first, and the stage entities corresponding to the candidate stages are retained based on the candidate stage results of the window stage. For the stage entities that are retained, further retain the anatomical entities that have anatomical relationships with the stage they exist in, the instrument entities that have instrument relationships with the stage they exist in, and the operational entities that have operational relationships with the stage they exist in. For the identified anatomical region results and instrument object results, if there are synonymous entities or hierarchical entities in the medical knowledge graph, they are mapped to the standard entities in the target medical knowledge subgraph through entity synonym relationships or entity hierarchical relationships. The target medical knowledge subgraph retains only entities and relationships related to the current medical process identification task, and does not retain remote entities and relationships unrelated to the current surgery type or examination type, thereby reducing the interference of irrelevant medical knowledge on candidate process state construction and path decoding; Specifically, the candidate stage labels in the window stage candidate results are first mapped to stage entities in the medical knowledge graph, the anatomical region results are mapped to anatomical entities, and the instrument object results are mapped to instrument entities. Using these entities as starting entities, retrieve first-order and second-order adjacent entities that match the current surgical type in the medical knowledge graph; Then, retain the medical relationships related to the sequence of stages, the objects of anatomical treatment, the use of instruments and the operation actions to form a target medical knowledge subgraph. For synonyms, abbreviations or hierarchical concepts, normalize them through entity synonyms and entity hierarchical relationships so that the same medical object is mapped to a consistent standard entity in different video windows. For the Each video window constructs a candidate process state based on its candidate stage set. The candidate process state is composed of stage entities, anatomical entities, instrument entities, and operation entities. For the candidate stage First, the candidate stage Mapped to stage entities in the target medical knowledge subgraph, and anatomical entities that have a stage-corresponding anatomical relationship with the stage entity are selected based on the anatomical region results of the video window. Based on the instrument object results in the video window, select the instrument entity that has a corresponding instrument relationship with the entity in this stage; Based on the spatial contact relationship between the instrument and the anatomical area, the target boundary results, and the positional changes of the instrument in adjacent windows, select the operation entity that has a stage-corresponding operation relationship with the entity in this stage. Candidate Phase A corresponding candidate process state is generated, if the candidate stage If no matching relationship can be found in the target medical knowledge subgraph for the current anatomical entity, instrument entity, and operation entity, the candidate stage will still be retained as a low-reliability candidate state, but the reliability of its subsequent states will be reduced. To determine whether a candidate process state conforms to the target medical knowledge subgraph, the knowledge path distance result is calculated and denoted as... The knowledge path distance result is used to represent the candidate stage. The connection distance between the anatomical entity, instrument entity, and manipulation entity in the video window and the target medical knowledge subgraph; It should be noted that the knowledge path distance result is used to characterize the degree of medical relationship between the candidate stage entity and the anatomical entity, instrument entity and operation entity in the current video window; When there is a direct stage correspondence between a candidate stage entity and its corresponding anatomical entity, instrument entity, or operation entity, the knowledge path distance between them is considered to be small. When two entities need to be connected through synonymous entities, hierarchical entities, or intermediate operational entities, it is believed that their knowledge path distance increases with the number of intermediate entities. When there is no interpretable path between a candidate stage entity and the current anatomical entity, instrument entity, or operation entity, this type of relationship is marked as unreachable and assigned a preset large distance. If the same candidate stage involves anatomical entities, instrumental entities, and operational entities, the maximum value among the path distances of each type of entity can be taken as the knowledge path distance result of the candidate process state, so that the candidate stage can obtain high state reliability only when the anatomical object, instrumental object, and operational action all have reasonable medical relationships. Specifically, for the candidate stage For each stage entity, calculate the shortest path from its current window's anatomical entity, instrument entity, and operation entity. If there are direct relationships between the stage entity and the anatomical entity, the instrument entity, and the operational entity, then Take the smaller value; If a connection is required through an intermediate entity, then It increases with increasing path length; If no interpretable relation exists, then... Set to the default larger value; The maximum value among the three types of relationship path distances is taken, so that the candidate stage has a small knowledge path distance only when the three types of relationships in the dissection stage, instrument stage and operation stage are relatively matched. Based on stage candidate confidence results Distance between knowledge path and results Calculate the first Select candidate process status in each video window. The reliable state result at that time is denoted as The calculation method is as follows: ;in, For the first Select candidate process status in each video window. Reliable results at that time; For the first Each video window belongs to the candidate stage. The stage candidate confidence results; For the first Candidate phase in a video window The knowledge path distance results between the anatomical entity, instrument entity, and operation entity in this window; If a candidate stage has high visual confidence, but lacks medical relationship support between it and the current anatomical entity, instrument entity, and operation entity, its reliable state result will be reduced by the knowledge path distance. This step ultimately yields the target medical knowledge subgraph, the candidate flow state of the window, and the reliable state result. The target medical knowledge subgraph provides the basis for stage sequence relationships and entity matching. The candidate flow state of the window expands each video window from an isolated stage label to a composite medical state consisting of stage entities, anatomical entities, instrument entities, and operational entities. The reliable state result indicates whether the candidate flow state simultaneously satisfies visual classification evidence and medical knowledge relationships. Through this step, subsequent global path decoding is no longer based on a single stage label, but on the candidate flow state constrained by medical entity relationships. This step employs a graph modeling approach that combines medical knowledge graph screening, entity normalization, candidate stage entity mapping, and knowledge path distance calculation. The window stage candidate results obtained in step S1 are transformed from single visual labels into window candidate process states composed of stage entities, anatomical entities, instrument entities, and operational entities. A target medical knowledge subgraph is constructed based on the surgical type corresponding to the target medical video, the window stage candidate results, the anatomical region results, and the instrument object results. The knowledge path distances between the candidate stages and the anatomical entities, instrument entities, and operational entities are calculated separately. These distances are then combined with the stage candidate confidence results to generate reliable state results. This step is also constrained by the entity relationships within the target medical knowledge subgraph. For candidate stages with high visual confidence but lacking medical relationship support with the current anatomical object, instrument object, or operational action, the knowledge path distance can reduce their reliable state results, providing a more stable candidate state foundation for the subsequent construction of the medical process state graph.
[0016] In step S3, a medical process state graph is constructed based on the candidate process states and reliable state results of the windows. State transition relationships are established between the candidate process states of adjacent video windows. The transition cost result is determined based on the process sequence relationship, entity continuity relationship, and jump risk result in the target medical knowledge subgraph. Global process path decoding is then performed based on the reliable state result and the transition cost result to obtain the global process path result. Initial semantic fragment annotation results are generated, including: Based on the obtained candidate window states and reliable state results, a medical workflow state diagram is constructed. The medical workflow state diagram includes multiple window layers arranged in chronological order. Each window layer corresponds to a video window, and each window layer includes multiple candidate window states for that video window. State transition edges are established between adjacent window layers to represent the states of the first and second windows. Candidate flow status of each video window Can it be transferred to the first Candidate flow status of each video window Each candidate process state has a reliable state result calculated in step S2. Each state transition edge has a transition cost result. ; For adjacent candidate process states and The result of calculating the knowledge order distance is denoted as The knowledge order distance result is used to represent the distance between two stage entities in the medical process sequence within the target medical knowledge subgraph; If the state and state For entities in the same stage, then Take the smaller value; If the state For state A reasonable subsequent stage, then It is also relatively small; If the state It requires traversing multiple intermediate stages to reach a state. Upon arrival, As the number of stages increases; If the state Relative to state If it's a reverse jump, or if there's no interpretable flow relationship between the two, then... Take the larger preset value; Meanwhile, the continuous results of the computation entity are denoted as Entity continuous results are used to represent the first Candidate flow status of each video window With the Candidate flow status of each video window The degree of continuity in anatomical and instrumental entities; Specifically, if the state and state If the anatomical entities correspond to the same anatomical entity or have a hierarchical relationship, then the anatomical continuity is relatively high. If the state and state If the corresponding medical device entities are the same, or if the change from one medical device entity to the next conforms to the stage-to-device relationship in the target medical knowledge subgraph, then the continuity of the devices is relatively high. If the anatomical entity changes abruptly and lacks support for the sequence of stages, or if the instrument entity does not match the candidate stage, the entity continuity result will be low. It can be determined based on the smaller value between anatomical continuity and instrumental continuity, so that adjacent states only have a high degree of physical continuity when both the anatomical object and the instrumental object have reasonable continuity; Based on knowledge order distance results Entity continuous results and redirection risk results The result of calculating the transfer cost is denoted as The calculation method is as follows: ;in, For the first Candidate flow status of each video window Transfer to the Candidate flow status of each video window The result of the transfer cost; Candidate process status With candidate process status The knowledge order distance result in the target medical knowledge subgraph; For the first Video window status With the Video window status Entity continuity results between them; To prevent stable constants with a denominator of zero; For the first Risks associated with redirecting individual video windows; This is the result of the phase switching indication; When the candidate process status With candidate process status If the corresponding stage entities are inconsistent, take one; otherwise, take zero. Switching between stages with unreasonable knowledge order, low entity continuity, and high risk of jumping to the current window has a high transfer cost, thereby suppressing medical process jumps caused by local high-confidence abnormal windows. Next, global process path decoding is performed, and the first... Select candidate process status in each video window. The path cost result at time is denoted as The path cost result is jointly determined by the unreliability of the current state, the cost of the optimal path in the previous window, and the cost of the current transition. The state reliability result... The higher the value, the lower the unreliability of the current state; The path cost result is derived recursively in the following way: ;in, For the first Select candidate process status in each video window. The path cost result at that time This is a reliable result regarding the status of the candidate process. For the first A set of candidate process states for each video window; For the first The status of any candidate process in a video window; For the first Select candidate process status in each video window. The path cost result at that time; To start from the candidate process status Transition to candidate process status The result of the transfer cost; For each candidate process state, record the previous candidate process state that minimizes the path cost result to form a path backtracking result. After recursively pushing to the last video window, select the termination state with the minimum path cost result, and obtain the global process path result in reverse based on the path backtracking result. Based on the global process path results, the video windows are semantically segmented. If the final stage entities of adjacent video windows are the same and the anatomical and instrumental entities remain continuous, the adjacent video windows are merged into the same semantic segment. If the final stage entities of adjacent video windows are different, but the knowledge order distance is small and the later stage is supported by multiple consecutive video windows, the position is determined as the semantic segment boundary. If the final stage entities of adjacent video windows are different, but the change in this stage is mainly triggered by a single video window or a small number of video windows with a high jump risk, the position is marked as the boundary result to be reviewed and handed over to step S4 for processing. This step outputs the global workflow path results and the initial medical video semantic segment annotation results. The initial medical video semantic segment annotation results include multiple semantic segments and the corresponding segment start time, segment end time, segment stage label, segment anatomy label, segment instrument label, and segment operation label for each semantic segment. Through this step, the existing technology's processing method of selecting the highest confidence label window by window and voting to merge is replaced by a global path decoding method based on medical knowledge subgraph and transfer cost, so that the output results are more in line with the medical workflow sequence. This step employs a global decoding algorithm that combines medical process state graph modeling, knowledge sequence distance calculation, entity continuity result calculation, transfer cost calculation, and dynamic path recursion. It organizes the candidate process states of each video window into a time-ordered state search space. The knowledge sequence distance result is determined by the stage sequence relationship in the target medical knowledge subgraph, and the entity continuity result is determined by the continuity of anatomical and instrument entities between adjacent windows. The transfer cost result is calculated in conjunction with the jump risk result obtained in step S1. Subsequently, the global process path result is obtained through path cost recursion and path backtracking. Based on the final process state, video windows are merged to generate initial medical video semantic segment annotation results. This step replaces the existing window-by-window highest confidence label selection and simple voting merging method with global process path decoding. This allows window stage determination to simultaneously consider the reliability of the current state, the rationality of adjacent state transitions, and the risk of local high-confidence anomalies. This suppresses medical process jumps caused by single or a few abnormal windows, making the initial segment boundaries and stage labels more consistent with the process sequence relationship in the target medical knowledge subgraph.
[0017] In step S4, abnormal segment detection is performed on the initial semantic segment annotation results. For segments that meet the abnormal conditions, local re-decoding is performed within the corresponding abnormal window set. Based on the local re-decoding results, the segment boundaries or segment stage labels are corrected, and the corrected medical video semantic segmentation annotation results are output. The specific content includes: Anomaly detection is performed on the initial medical video semantic segment annotation results to obtain anomaly segment results. The anomaly segment results include short-term jump segments, process conflict segments, and low-confidence segments. Short-term jump segments refer to segments with a small number of continuous windows and whose preceding and following segments correspond to the same stage entities. Conflicting process segments refer to segments in which there is a knowledge sequence reversal or a lack of interpretable process relationships between entities in the corresponding stage and entities in the preceding and following stages. A low-confidence segment refers to a segment in which multiple video windows have a high risk of jumping to other segments, and the reliable results of the state are insufficient to support the existence of an independent stage. For each abnormal segment result, an abnormal window set is determined. The abnormal window set includes all video windows inside the abnormal segment, several video windows at the end of the semantic segment preceding the abnormal segment, and several video windows at the beginning of the semantic segment following the abnormal segment. The abnormal window set is used to limit the processing range of local re-decoding so that local correction will not disrupt the already stable and determined preceding and following medical processes. The number of neighborhood windows is determined based on the window length result and the window step size result. It should be noted that the abnormal window set is used to limit the processing scope of local re-decoding. For each abnormal segment, the abnormal window set includes all video windows within the abnormal segment, several video windows at the end of the preceding semantic segment, and several video windows at the beginning of the following semantic segment; the number of windows at the end of the preceding semantic segment and the number of windows at the beginning of the following semantic segment can be determined based on the window length, window step size, and the shortest duration of the real stage. If the abnormal segment is located at the beginning of the target medical video, the abnormal window set does not include the end window of the previous semantic segment; If the abnormal segment is located at the end of the target medical video, the abnormal window set does not include the starting window of the next semantic segment; By adding stable neighborhood windows before and after abnormal segments, local re-decoding can constrain the boundaries and labels of abnormal segments by utilizing the stable process states on both sides of the abnormal segments without reprocessing the entire video, thus avoiding local corrections from disrupting the already stable and determined overall medical process structure. The result of calculating the fragment risk of anomaly fragments is denoted as: The segment risk result is determined by the jump risk result of each video window within the abnormal segment. Sure; Preferably, the maximum value of the jump risk result within the abnormal segment is taken to capture the window most likely to trigger the error stage jump, and the result of calculating the number of segment windows of the abnormal segment is denoted as... The segment window count result indicates the number of video windows contained in the abnormal segment, and a risk threshold is set. and minimum number of windows Risk threshold The minimum number of windows is determined based on the risk distribution of jumps to abnormal segments confirmed in historically annotated videos. Determined based on the shortest duration of the actual phase and the window step size; When fragment risk results Not less than the risk threshold And the result of the number of fragment windows Less than the minimum number of windows When this occurs, it indicates that the segment is more likely to be triggered by a local high-confidence anomaly window; If the abnormal segment still has a reverse flow relationship in the target medical knowledge subgraph, triggering local re-decoding, the global path decoding in step S3 is re-executed within the abnormal window set, but the starting state is limited to the final flow state of the previous stable segment in the abnormal window set, and the ending state is limited to the final flow state of the next stable segment in the abnormal window set. Local re-decoding only re-selects candidate flow states within the abnormal window set, without changing the stable flow structure that has been determined in the target medical video. Local re-decoding only reselects candidate process states within the abnormal window set, without changing the already determined global process path results outside the abnormal window set; During local re-decoding, the end candidate process state of the previous stable semantic segment before the abnormal window set is used as the starting point constraint of local decoding, and the starting candidate process state of the next stable semantic segment after the abnormal window set is used as the ending point constraint of local decoding. If the abnormal segment is located at the beginning or end of the target medical video, only the stable state on one side that exists will be used as a constraint. The candidate process state, state reliability result, knowledge order distance result, entity continuity result, and jump risk result used in local re-decoding are consistent with those used in global process path decoding, but the processing scope is limited to the set of anomaly windows. If the local re-decoding result allows the abnormal segment to be merged into the previous or next semantic segment, and the local path cost is lower than the original local path cost, then the independent segment boundary of the abnormal segment is deleted, and the segment start time, segment end time and segment stage label of the merged segment are updated. If the local re-decoding result supports the retention of abnormal segments, but the stage entity corresponding to its best candidate process state is different from the original segment stage label, then the segment stage label of the abnormal segment is corrected to the candidate stage entity with the lowest local path cost. If the local re-decoding result still supports the independent existence of the abnormal segment, and the number of segment windows is not less than the minimum number of windows, then the abnormal segment is retained, and a low confidence mark or a mark to be reviewed is added to it. The final output is the corrected medical video semantic segmentation and annotation results. The corrected medical video semantic segmentation and annotation results include multiple semantic segments. Each semantic segment includes segment start time, segment end time, segment stage label, segment anatomy label, segment instrument label, segment operation label, and annotation traceability results. The annotation traceability results include the main video window on which the segment is based, the main candidate process status, the stage sequence relationship in the target medical knowledge subgraph, whether it has undergone local re-decoding, and the label changes before and after local re-decoding. Through the above four steps, step S1 identifies local high-confidence anomaly windows and generates jump risk results. Step S2 maps candidate stages to candidate process states in a medical knowledge subgraph and generates reliable state results. Step S3 is based on the transfer cost result and path cost results After obtaining initial segmentation results consistent with the process, step S4 reuses the segment risk results. Results of fragment window count Local re-decoding correction of abnormal segments, with parameters linked before and after each step and consistent terminology, can address the problem of medical process jumps caused by local high-confidence abnormal windows after the sliding window stage classification of laparoscopic surgery quality control videos, forming a complete, implementable and targeted medical video semantic segmentation annotation method. This step employs a correction algorithm combining abnormal segment detection, abnormal window set construction, segment risk result calculation, segment window quantity judgment, and local re-decoding to locally correct the initial medical video semantic segment annotation results obtained in step S3. Abnormal segment results are determined based on the criteria for short-term jump segments, process conflict segments, and low-confidence segments, and an abnormal window set is constructed within the abnormal segment and its preceding and following neighborhoods. Furthermore, the jump risk results of each video window within the abnormal segment are used to determine the segment risk result, and the number of video windows contained in the abnormal segment is used to determine the segment window quantity result. When a high segment risk is met and the number of windows is insufficient, or a process reversal relationship exists, local path decoding is re-executed within the abnormal window set to correct segment boundaries and segment stage labels. This step does not disrupt the already stable and determined preceding and following medical process structure. For short-term erroneous segments formed by locally high-confidence abnormal windows, they can be merged into adjacent semantic segments or corrected into candidate stages that better match the path cost, thereby improving the process consistency and traceability of the final medical video semantic segmentation annotation results.
[0018] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0019] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.
[0020] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and inventive constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0021] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0022] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0023] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for semantic segmentation and annotation of medical video content integrating knowledge graphs, characterized in that, Includes the following steps: Step S1: Acquire the target medical video and divide it into video windows according to the preset window length and window step size. Perform visual analysis and stage classification on the video windows to obtain the window visual analysis results and window stage candidate results. The window visual analysis results include anatomical region results, instrument object results, target boundary results, noise status results, and pixel label confidence results; The noise occlusion result is determined based on the noise state result, the boundary instability result is determined based on the target boundary result, and the visual perturbation result is determined based on the noise occlusion result and the boundary instability result. The video window is classified into stages to obtain candidate window stages, including candidate stage labels and candidate stage confidence results; The confidence gap result is determined based on the difference between the highest-level candidate confidence result and the second-highest-level candidate confidence result, and the jump risk result is determined based on the confidence gap result and the visual perturbation result. Step S2: Based on the candidate results of the window stage and the visual analysis results of the window, select the target medical knowledge subgraph from the medical knowledge graph, and map the candidate stages of each video window into the candidate flow state of the window, which includes stage entities, anatomical entities, instrument entities and operation entities. For each candidate stage, calculate the shortest relationship path from the stage entity to the corresponding anatomical entity, instrument entity and operation entity to obtain the knowledge path distance result. The state reliability result is determined based on the candidate confidence result of the stage and the knowledge path distance result, so that the greater the knowledge path distance between the candidate stage and the corresponding entity, the lower the state reliability result. Step S3: Construct a medical process state graph based on the candidate process states and reliable state results of the window, establish state transition relationships between the candidate process states of adjacent video windows, determine the transition cost results based on the process sequence relationship, entity continuity relationship and jump risk results in the target medical knowledge subgraph, and perform global process path decoding based on the reliable state results and transition cost results to obtain the global process path results and generate the initial semantic fragment annotation results. Step S4: Perform abnormal segment detection on the initial semantic segment annotation results to obtain short-term jump segments, process conflict segments, or low-confidence segments. Construct an abnormal window set based on the video windows inside the abnormal segment, the video window at the end of the semantic segment preceding the abnormal segment, and the video window at the beginning of the semantic segment following the abnormal segment; The segment risk result is determined based on the jump risk result of each video window in the abnormal window set, and the segment window number result is determined based on the number of video windows contained in the abnormal segment; When the segment risk result meets the risk condition and the segment window quantity result meets the short segment condition or the abnormal segment has a process reversal relationship; Re-execute global path decoding within the abnormal window set, and correct segment boundaries or segment stage labels based on the local re-decoding results, outputting the corrected medical video semantic segmentation annotation results.
2. The method for semantic segmentation and annotation of medical video content based on knowledge graphs according to claim 1, characterized in that, Obtain multiple video frames arranged in chronological order and the corresponding time position of each video frame, and identify the multiple video frames as the target medical video; The window length is determined based on the shortest effective duration of the stage to be identified, and the window step size is determined based on the window length. The target medical video is then divided into sliding sections according to the window length and window step size to obtain video windows arranged in chronological order.
3. The method for semantic segmentation and annotation of medical video content based on knowledge graphs according to claim 2, characterized in that, For each video window, extract the window representative frame. The window representative frame is determined by the frame clarity result and the inter-frame variation result of the video frames within the video window. Perform surgical scene visual analysis on the window representative frame to obtain the window visual analysis result.
4. The method for semantic segmentation and annotation of medical video content based on knowledge graphs according to claim 1, characterized in that, Obtain a medical knowledge graph that includes a set of medical entities and a set of medical relationships; The medical entity set includes stage entities, anatomical entities, instrument entities, and operational entities; The medical relationship set includes stage sequence relationships, stage-corresponding anatomical relationships, stage-corresponding instrument relationships, stage-corresponding operational relationships, entity synonym relationships, and entity hierarchical relationships; Based on the surgical type, window stage candidate results, anatomical region results, and instrument object results corresponding to the target medical video, entities and relationships related to the target medical video are filtered from the medical knowledge graph to obtain the target medical knowledge subgraph.
5. The method for semantic segmentation and annotation of medical video content based on knowledge graphs according to claim 4, characterized in that, Map the candidate stages in the window stage candidate results to stage entities in the target medical knowledge subgraph; Select the anatomical entities that have a stage-corresponding anatomical relationship with the stage entities based on the anatomical region results in the video window. Select the device entities that have a corresponding device relationship with the stage entity based on the device object results in the video window; Based on the spatial contact relationship between the instrument and the anatomical region, the target boundary results, and the positional changes of the instrument in adjacent windows, select the operation entity that has a stage-corresponding operation relationship with the stage entity; The window candidate process status consists of stage entities, anatomical entities, instrument entities, and operation entities.
6. The method for semantic segmentation and annotation of medical video content based on knowledge graphs according to claim 1, characterized in that, A medical process state diagram is constructed based on the candidate process states of windows and the reliable results of the states. The medical process state diagram includes window layers arranged in chronological order, and each window layer includes the candidate process states of the corresponding video window. Establish state transition relationships between adjacent window layers; The knowledge order distance result is determined based on the stage sequence relationship of adjacent candidate process states corresponding to the stage entities in the target medical knowledge subgraph. The continuity of entities is determined based on the degree of continuity between the anatomical entities and instrument entities corresponding to adjacent candidate process states, and the transfer cost is determined in combination with the jump risk result.
7. The method for semantic segmentation and annotation of medical video content based on knowledge graphs according to claim 6, characterized in that, Based on the reliable state results, the path cost results of the previous video window, and the transition cost results between adjacent candidate process states, the path cost results of each candidate process state in each video window are recursively obtained. Record the state of the previous candidate process that minimizes the path cost, and form the path backtracking result; The global process path result is obtained based on the path backtracking result, and the video windows are merged based on the global process path result to generate the initial semantic segment annotation result.
Citation Information
Patent Citations
Operation process identification method, device and system and computer readable storage medium
CN112818959A
Auxiliary decision-making method and device for minimally invasive surgery
CN114724682A
Endoscopic video timeline interest level prediction
CN120391960A