A clinical guideline structured diagnosis and treatment skill automatic construction method and system
Patent Information
- Application Number
- CN202611097078.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-07-23
AI Technical Summary
[0005]有鉴于此,本发明提供一种临床指南结构化诊疗技能自动构建方法及系统,以解决现有技术中临床指南文档多模态内容处理能力不足、结构化产物与指南原文之间溯源关联缺失、以及自动化构建质量缺乏系统性验证机制的问题
(1)本发明提出了一种基于多模态大模型的临床指南结构化诊疗技能自动构建方法及系统,构建了以L1原文锚定层、L2决策图、L3可执行技能层为核心的三层知识架构,通过多模态统一语义化解析、有向无环决策图构建、语义感知子图切分与封装、反渲染比对质量验证、专家校验闭环优化、路径证据等级透明化输出及缺省入参降级调用等模块的协同作用,实现了从多模态临床指南文档到可供外部系统直接调用的可执行诊疗技能单元的自动化、高质量、全程可追溯构建,在显著降低指南知识工程化成本的同时,有效保障了结构化产物的语义准确性与原文一致性。
Smart Images

Figure CN122599099B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical artificial intelligence technology, and in particular to a method and system for automatically constructing structured clinical guideline-based diagnostic and treatment skills. Background Technology
[0002] Clinical guidelines are authoritative knowledge carriers that standardize diagnostic and treatment behaviors in clinical medical practice. They compile evidence-based treatment processes, graded recommendation schemes, and evidence level evaluations for specific diseases, and are important knowledge sources for digital medical applications such as clinical decision support systems, intelligent consultation systems, and intelligent agent systems. However, clinical guidelines are usually published in document format for human readers, with highly diverse content formats. They include extensive continuous textual descriptions, multi-dimensional comparative information presented in tables, branching decision logic presented in flowcharts, and image examples presented in medical images and their captions. These various types of content complement each other and jointly express the complete diagnostic and treatment decision logic. To transform the above content into structured diagnostic and treatment knowledge that can be automatically invoked by computer systems, accurate semantic parsing of multimodal content is required. However, this transformation process in practical applications usually relies on a large amount of manual annotation and expert collation work, resulting in low automation and severely restricting the large-scale application of clinical guideline knowledge.
[0003] With the rapid development of large language model technology, researchers have begun to explore the use of natural language processing and artificial intelligence methods to automate knowledge extraction and structured transformation of clinical guidelines, aiming to reduce manual compilation costs and improve the computer usability of guideline knowledge. Existing methods mainly focus on extracting entities, relationships, or constructing knowledge graphs from the textual content of clinical guidelines. Some works attempt to transform the diagnostic and treatment rules in the guidelines into decision trees or decision graphs, thereby providing more operational knowledge representations for downstream clinical decision support systems. These explorations have effectively promoted the progress of digital applications of clinical guidelines, but overall, they still face problems such as insufficient multimodal content processing capabilities, difficulties in tracing the structured products from the original guideline text, and a lack of systematic verification mechanisms for the quality of automated construction. A complete automated transformation path from multimodal guideline documents to executable diagnostic and treatment skill units has not yet been formed.
[0004] Patent publication number CN119227784A discloses a decision tree generation method based on a large language model. This method collects clinical guideline document information as training samples, fine-tunes the large language model to extract node and decision path information from the target medical document, and adjusts the node correlation based on historical feedback data from the target hospital, ultimately generating decision tree information, thus achieving a certain degree of automated conversion from clinical guidelines to decision trees. However, the guideline parsing process of this method focuses on textual content and fails to perform targeted structured parsing of non-textual modalities such as tables and flowcharts commonly found in guidelines, resulting in significant limitations in multimodal content processing capabilities. Furthermore, the generated decision tree nodes lack clear locational relationships with the original guideline text, making it difficult to trace the source of the structured product. In addition, this method does not provide a systematic verification method for the consistency between the automated construction results and the original guideline content, relying on manual verification for construction quality and lacking an automated assurance mechanism. Summary of the Invention
[0005] In view of this, the present invention provides a method and system for automatically constructing structured diagnostic and treatment skills in clinical guidelines, in order to solve the problems of insufficient multimodal content processing capabilities of clinical guideline documents, lack of traceability and correlation between structured products and original guideline texts, and lack of systematic verification mechanism for the quality of automated construction in the prior art.
[0006] The technical solution of this invention is implemented as follows: On the one hand, this invention provides a method for automatically constructing structured clinical practice skills, comprising the following steps: S1. The layout parsing engine is used to process the multimodal clinical guideline document page by page, dividing the document content into content blocks with different modal types, and generating original text anchor points carrying page number, position and modal type information for each content block; S2. Call the multimodal large model to perform semantic parsing on each content block obtained in S1. Extract clinical concepts, structured representations and inter-block relationships according to the modal features of each content block. Bind the parsing results with the corresponding original text anchor point identifiers to form a set of evidence units carrying semantic information and original text positioning information, which constitutes the L1 original text anchoring layer. S3. Based on the set of evidence units in the L1 original text anchoring layer, a combination of large model and rules is used to extract and organize the clinical decision conditions, decision results and their dependencies in the evidence units into a directed acyclic graph composed of nodes and directed edges. Each node maintains a reverse pointer to its source evidence unit, forming an L2 decision graph representing the clinical decision process. S4. The L2 decision graph is divided into subgraphs according to clinical application scenarios. Each subgraph is encapsulated into an executable diagnostic and treatment skill unit consisting of input parameter mode, output parameter mode and internal decision structure. All executable diagnostic and treatment skill units constitute the L3 executable skill layer. S5. Based on the L1 original text anchoring layer and the L2 decision graph, verify the consistency of each executable diagnostic and treatment skill unit in the L3 executable skill layer with the original clinical guidelines, correct the executable diagnostic and treatment skill units that fail the verification, and output a set of verified executable diagnostic and treatment skill units, which can be called by external systems.
[0007] Based on the above technical solutions, preferably, in step S1, the content blocks are divided into text blocks, table blocks, flowchart blocks, and image example blocks according to modal type; the original text anchor point identifier is composed of document identifier field, version field, page number field, block identifier field, block boundary box field, block type field, and reading order field. The block boundary box field describes the position of the content block with the coordinates of the upper left corner and the lower right corner of the page, and the block type field records the modal type to which the content block belongs.
[0008] Based on the above technical solutions, preferably, step S2, which calls the multimodal large model to perform semantic parsing on each content block obtained in S1, specifically includes: S21. For the text block, call the large model to extract the set of clinical concepts contained therein, and perform standard medical terminology mapping on each clinical concept to obtain the set of clinical concepts and standard terminology mapping results of the text block. S22. For the table block, call the multimodal large model to perform structural restoration of the table image to obtain a structured table representation containing the header hierarchy, cell content and merged cell information. S23. For the flowchart block, call the multimodal large model to identify nodes and edges in the flowchart image, and obtain a structured graph representation containing the node set, edge set and edge discrimination conditions; S24. For the image example block, call the multimodal large model to jointly understand the image and its captions, extract the description of clinical signs and the illustrated clinical situation, and establish cross-modal associations through the evidence units corresponding to the semantically related text blocks and flowchart blocks identified by the original anchor points. The above-obtained analysis results are bound to the original text anchor points of the corresponding content blocks to form a set of evidence units, constituting the L1 original text anchoring layer.
[0009] Based on the above technical solutions, the preferred method involves calling a multimodal large model to identify nodes and edges in the flowchart blocks, including the following sub-steps: S231. Call the multimodal large model to perform element detection on the flowchart image block and identify the set of nodes and the set of edges; S232. For each node, call the multimodal large model to classify the node as a candidate category, including the discriminant node, action node, start node, and end node. Combine the original text anchor point identifier to which the node belongs to perform context-aware text recognition on the text within the node, and extract the set of clinical concepts associated with the node from the recognized text. S233. For each edge, call the multimodal large model to extract the discrimination conditions on the edge from the label image of that edge; S234. Perform cycle checks on the obtained node set and edge set. For decision paths containing cycles, prioritize the explicit termination condition strategy for processing. That is, identify the exit judgment condition in the cycle and replace the back edges pointing to visited nodes in the cycle with edges pointing to the loop exit node caused by the exit judgment condition. For cycle segments that still cannot be resolved after explicit termination condition processing, adopt the cycle node encapsulation strategy to encapsulate the entire cycle segment into a cycle node. The cycle node retains the iterative decision logic inside and is connected to the structured graph representation as a single node. Record its iterative exit condition and maximum iteration number constraint. The output is a structured graph representation of the flowchart blocks formed by the node set, edge set, and edge discrimination conditions obtained from S231 to S234.
[0010] Based on the above technical solutions, preferably, the construction of the L2 decision graph in S3 includes the following sub-steps: S31. Traverse all evidence units in the L1 original text anchoring layer, and use different subgraph transformation strategies to convert each evidence unit into a decision subgraph according to the modality type of each evidence unit. Establish a reverse pointer to the source L1 evidence unit for each node in each subgraph, and merge all subgraphs into the initial L2 decision graph. S32. Perform deduplication and merging on nodes that are determined to be semantically equivalent after normalization in the initial L2 decision graph. Take the union of the reverse pointers of each node to be merged and assign it to the merged node. Reconnect the inbound and outbound edges of the merged nodes to the merged node to obtain the normalized L2 decision graph. S33. Perform loop elimination processing on the loop-containing structures from the flowchart blocks in the normalized L2 decision graph. Prioritize the explicit termination condition strategy to eliminate loops. For loop segments that cannot be directly eliminated, use the loop node encapsulation strategy to encapsulate them as loop nodes to maintain the overall directed acyclicity of the final L2 decision graph; thus obtaining the final L2 decision graph.
[0011] Based on the above technical solutions, preferably, the normalization process in step S32 specifically includes: Clinical concepts in each node of the initial L2 decision graph are mapped to standard terms based on a standard medical terminology database. For the discriminant nodes and branch nodes in the initial L2 decision graph that express clinical discrimination conditions, their discrimination conditions are structured and standardized, and uniformly represented as standard triplet forms consisting of comparison objects, comparison operators, and comparison values. When two nodes are of the same type and the standard terminology set after the above normalization is consistent with the standard triplet discrimination conditions, they are determined to be semantically equivalent.
[0012] Based on the above technical solutions, preferably, step S4 includes the following sub-steps: S41. Traverse the L2 decision graph, using the recommended action node as the segmentation anchor point. For each segmentation anchor point, collect the minimum set of predecessor nodes that make it complete and decidable along the incoming edge direction. The subgraph induced by the minimum set of predecessor nodes and the anchor point is taken as the candidate skill subgraph. Merge the candidate skill subgraphs with strong semantic association across boundaries, and verify the segmentation results with the clinical scene ontology as a constraint to obtain the set of candidate skill subgraphs after segmentation. S42. For each subgraph in the candidate skill subgraph set, construct its internal decision structure and perform static dependency analysis to establish a field-node dependency mapping table. Then, encapsulate the subgraph into an executable diagnostic and treatment skill unit. The executable diagnostic and treatment skill unit consists of an input parameter mode, an output parameter mode, an internal decision structure, and an evidence pointer set. The evidence pointer set records the L1 evidence unit identifier corresponding to each node. All executable diagnostic and treatment skill units constitute the L3 executable skill layer.
[0013] Based on the above technical solutions, preferably, the static dependency analysis in S42 includes: Traverse all the decision nodes in the internal decision structure, extract the input parameter fields on which the decision conditions of each decision node depend, establish the mapping relationship between each required input parameter field and the set of decision nodes it affects, and pre-calculate the default influence degree of each required field. The default influence degree comprehensively considers the proportion of decision paths affected after the field is omitted and the position sensitivity of the associated decision node in the internal decision structure. The field-node dependency mapping table and the default influence of each field are stored together with the internal decision structure in the executable diagnostic and treatment skill unit. This allows the skill unit to perform local reachability analysis based on the field-node dependency mapping table when there are default input parameters during runtime. It can output deterministic recommendations for reachable paths and conditional recommendations with default dependency fields marked for unreachable paths. Even when the input parameters are incomplete, it can still output local recommendation results.
[0014] Based on the above technical solutions, preferably, step S5 specifically includes: S51. For each executable diagnostic and treatment skill unit in the L3 executable skill layer, based on the node type of the subgraph corresponding to the skill unit in the L2 decision graph, restore the internal decision structure of the skill unit to a natural language fragment according to the predefined template, and splice them in topological order to obtain the inverse rendering text of the skill unit. S52. The anti-rendered text is concatenated with the original content of the L1 evidence unit referenced by the evidence pointer set of the executable diagnostic and treatment skill unit to form a reference text. A comprehensive consistency score is obtained by multi-dimensional similarity calculation and weighted fusion. S53. Perform tiered processing according to the comprehensive consistency score: if the score is higher than the high confidence threshold, it is determined to be consistent and automatically passes; if the score is in the middle range, a large model local correction is triggered and the score is re-scored. If it still fails, it is transferred to expert verification; if the score is lower than the low confidence threshold, it is directly transferred to expert verification; the correction records of the executable diagnostic and treatment skill units that have been corrected by expert verification are written back to the sample library in a structured form, and the set of verified executable diagnostic and treatment skill units is output.
[0015] Based on the above technical solution, preferably, S53 further includes: after outputting the verified set of executable diagnostic and treatment skill units, for each executable diagnostic and treatment skill unit in the set, traversing each complete decision path p in its internal decision structure, and calculating the comprehensive strength score of path evidence level: ; in, The overall strength score for path evidence level, Assign integrity weights to path evidence levels; It is the arithmetic mean of the internal scale values of the evidence level of each node on the path; The normalized information entropy for the path evidence level distribution; The entropy penalty coefficient; The comprehensive strength score of the path evidence level is extended and written into the output parameter mode of the corresponding executable diagnostic and treatment skill unit, which is used to provide clinical users with a graded and ranked recommendation result.
[0016] Based on the above technical solutions, preferably, the construction method further includes version management of the executable diagnostic and treatment skill unit set output by S5, specifically including: A three-level semantic versioning specification—major version number, minor version number, and revision number—is used to version-mark each executable clinical skill unit. When the underlying clinical guidelines are updated, a difference analysis is performed on the old and new versions of the guidelines to identify newly added, modified, and obsolete guideline content fragments. Based on the difference analysis results, incremental version changes are performed on the affected executable clinical skill units: the major version number increments when the content of the L1 evidence unit referenced by the skill unit undergoes substantial changes; the minor version number increments when the internal decision-making structure of the skill unit is adjusted but the core recommendations remain unchanged; and the revision number increments when only metadata or evidence pointers are adjusted. Each version change generates a difference record, which includes the change type, change fields, content before the change, content after the change, and related guideline difference anchors to support incremental synchronization and version tracking of executable clinical skill units in downstream systems.
[0017] In addition, the present invention also provides an automated system for constructing structured clinical guidelines and diagnostic and treatment skills to implement the above-mentioned method, comprising: The multimodal parsing module receives multimodal clinical guideline documents, performs layout parsing, identifies content units such as text blocks, table blocks, flowchart blocks, and image example blocks, and generates original text anchor point identifiers carrying page numbers, positions, and modal types; and calls the multimodal large model to perform semantic parsing on each content unit to obtain a multimodal semantic intermediate representation indexed by the original text anchor points. The L1 / L2 / L3 conversion module is used to sequentially construct the L1 original text anchoring layer, L2 decision graph, and L3 executable skill layer based on the multimodal semantic intermediate representation. Specifically, the L1 evidence unit construction unit encapsulates the original content, modal information, and semantic parsing results of each content unit into evidence units and stores them in the evidence unit library; the L2 decision graph construction unit extracts clinical decision conditions, decision results, and their dependencies from the evidence units, organizes them into a directed acyclic graph composed of nodes and directed edges, and maintains a reverse pointer to the source evidence unit in each node; the L3 skill segmentation and encapsulation unit segments the L2 decision graph into subgraphs according to clinical application scenarios and encapsulates the subgraphs into executable diagnostic and treatment skill units composed of input parameter modes, output parameter modes, and internal decision structures. The consistency verification module is used to perform anti-rendering comparison and multi-dimensional similarity calculation on the executable diagnostic and treatment skill unit, output a comprehensive consistency score, and trigger automatic correction or push the skill unit to be corrected to the expert verification workbench based on the score result. The expert verification workbench is used to present decision image segments and diagnostic and treatment skill units to be verified at preset access points, receive expert verification opinions, and write back structured correction records to the sample library and diagnostic and treatment skill unit library. The version management module is used to perform semantic version marking, change difference recording, and incremental release management of the executable diagnostic and treatment skill units to support incremental and traceable maintenance in clinical guideline update scenarios.
[0018] The present invention has the following advantages over the prior art: (1) This invention proposes an automatic construction method and system for structured diagnosis and treatment skills based on a multimodal large model of clinical guidelines. It constructs a three-layer knowledge architecture with L1 original text anchoring layer, L2 decision graph and L3 executable skill layer as the core. Through the synergistic effect of modules such as multimodal unified semantic parsing, directed acyclic decision graph construction, semantically aware subgraph segmentation and encapsulation, anti-rendering comparison quality verification, expert verification closed-loop optimization, transparent output of path evidence level and default input parameter downgraded call, it realizes the automated, high-quality and fully traceable construction from multimodal clinical guideline documents to executable diagnosis and treatment skill units that can be directly called by external systems. While significantly reducing the engineering cost of guideline knowledge, it effectively ensures the semantic accuracy of the structured product and the consistency with the original text.
[0019] (2) This invention employs targeted semantic parsing strategies for text blocks, table blocks, flowchart blocks, and image example blocks, unifying the parsing results of each modality into a semantic intermediate representation with the original text anchor point identifier as the key. Furthermore, in the L1 original text anchoring layer, each evidence unit is bound to the original text page number and content block location information, ensuring that each structured product can be traced back to the specific content of the original guideline. Compared to methods that only process text content, this invention effectively avoids the loss of flowchart branch conditions, multi-dimensional table recommendation information, and clinical evidence from image examples during the structuring process, improving the complete capture rate of guideline content and the verifiability of the evidence chain.
[0020] (3) This invention addresses the common issue of loop-containing decision paths in clinical guideline flowcharts by designing a hierarchical resolution mechanism that includes an explicit termination condition strategy and a loop node encapsulation strategy. The former replaces back edges with directed edges of loop exit nodes by identifying exit criteria, thus resolving loops while preserving iterative clinical semantics. For complex loop segments that cannot be resolved by the former, the latter encapsulates them as loop nodes and retains exit criteria and maximum iteration constraints to maintain the overall directed acyclicity of the L2 decision graph. This mechanism solves the problem that existing structured methods generally lack effective means of processing guideline flowcharts containing loops, enabling iterative diagnostic and treatment processes to be fully preserved and enter the subsequent executable encapsulation stage.
[0021] (4) This invention uses an anti-rendering comparison mechanism to restore the internal decision-making structure of the executable diagnostic and treatment skill unit into a natural language fragment according to a predefined template. This fragment is then compared with the original content of the corresponding evidence unit in the L1 original text anchoring layer using a three-way weighted fusion comparison of semantic embedding similarity, rule matching similarity, and large model scoring similarity. A quantifiable comprehensive consistency score is output, and automatic correction or push to the expert verification workbench is triggered based on the score classification. This mechanism precisely defines the scope of expert verification as skill units with genuine uncertainty. While ensuring the systematic verification of the construction quality, it effectively controls the workload of expert intervention. Furthermore, the expert verification results are written back to the sample database as structured correction records, supporting continuous incremental optimization of the multimodal large model and forming a positive closed loop of construction-verification-optimization. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a schematic diagram of the three-layer knowledge architecture of the present invention; Figure 3 This is a comparison diagram of the block loop elimination strategies of the present invention. Figure 4 This is a flowchart of the consistency verification and hierarchical processing of the present invention; Figure 5 This is a system module architecture diagram of the present invention. Detailed Implementation
[0024] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0025] like Figure 1 and Figure 2 As shown, this invention provides a method for automatically constructing structured clinical treatment skills based on clinical guidelines, comprising the following steps: S1. The layout parsing engine is used to process the multimodal clinical guideline document page by page, dividing the document content into content blocks with different modal types, and generating original text anchor points carrying page number, position and modal type information for each content block; S2. Call the multimodal large model to perform semantic parsing on each content block obtained in S1. Extract clinical concepts, structured representations and inter-block relationships according to the modal features of each content block. Bind the parsing results with the corresponding original text anchor point identifiers to form a set of evidence units carrying semantic information and original text positioning information, which constitutes the L1 original text anchoring layer. S3. Based on the set of evidence units in the L1 original text anchoring layer, a combination of large model and rules is used to extract and organize the clinical decision conditions, decision results and their dependencies in the evidence units into a directed acyclic graph composed of nodes and directed edges. Each node maintains a reverse pointer to its source evidence unit, forming an L2 decision graph representing the clinical decision process. S4. The L2 decision graph is divided into subgraphs according to clinical application scenarios. Each subgraph is encapsulated into an executable diagnostic and treatment skill unit consisting of input parameter mode, output parameter mode and internal decision structure. All executable diagnostic and treatment skill units constitute the L3 executable skill layer. S5. Based on the L1 original text anchoring layer and the L2 decision graph, verify the consistency of each executable diagnostic and treatment skill unit in the L3 executable skill layer with the original clinical guidelines, correct the executable diagnostic and treatment skill units that fail the verification, and output a set of verified executable diagnostic and treatment skill units, which can be called by external systems.
[0026] The "diagnostic and treatment skills" described in this invention refer to discrete decision-making units extracted and encapsulated from clinical guidelines, possessing independent functional boundaries, and capable of being activated and invoked by external systems through predefined interfaces. Each diagnostic and treatment skill consists of an input schema, an output schema, an internal decision structure, an evidence pointer, metadata, and a version number, and can independently complete a specific type of decision-making task in a clinical scenario, such as "initial staging determination of non-small cell lung cancer" or "recommendation of first-line treatment regimens for HER2-positive breast cancer."
[0027] In one embodiment of the present invention, in step S1, the content blocks are divided into text blocks, table blocks, flowchart blocks and image example blocks according to modal type; the original text anchor point identifier is composed of a document identifier field, a version field, a page number field, a block identifier field, a block bounding box field, a block type field and a reading order field, wherein the block bounding box field describes the position of the content block with the coordinates of the upper left corner and the lower right corner of the page, and the block type field records the modal type to which the content block belongs.
[0028] Specifically, this step receives multimodal clinical guideline documents (e.g., PDF files of oncology treatment guidelines) and processes them page by page using a layout parsing engine. The layout parsing engine performs block-level detection on each page based on a layout analysis model, identifying four types of content units: text blocks, table blocks, flowchart blocks, and image example blocks. It outputs the category label, bounding box coordinates, and reading order for each content block. A globally unique original text anchor identifier is generated for each content unit. This original text anchor identifier is composed of a document identifier field, a version field, a page number field, a block identifier field, a block bounding box field, a block type field, and a reading order field. The block bounding box field uses the coordinates of the top-left corner... and the coordinates of the bottom right corner This describes the location of the content block on the page. The block type field can be one of "text", "table", "flowchart" or "image_example", and the reading order field records the reading order number of the content block on the page. The original text anchor point identifier serves as the basic index for traceable association between subsequent processing layers, spanning the L1, L2, and L3 three-layer architecture.
[0029] In one embodiment of the present invention, step S2 employs a multimodal large model with image understanding and long context processing capabilities to perform the following semantic parsing on different types of content blocks obtained in S1, specifically including: S21. For the text block, call the large model to extract the set of clinical concepts contained therein, and perform standard medical terminology mapping on each clinical concept to obtain the set of clinical concepts and standard terminology mapping results of the text block.
[0030] Specifically, the large model is invoked to extract the clinical concepts contained therein, including disease entities, anatomical locations, examination items, treatment plans, drug names, dosages, evidence level labels, etc., and each extracted clinical concept is uniquely identified by UMLS or mapped to ICD-10 and SNOMED-CT standard terms to obtain the set of clinical concepts and standard term mapping results for the text block.
[0031] S22. For the table block, call the multimodal large model to perform structural restoration of the table image, and obtain a structured table representation containing the header hierarchy, cell content and merged cell information.
[0032] Specifically, a multimodal large model is invoked to restore the structure of the table image, and a structured table representation containing the header hierarchy, cell content and merged cell information is output to preserve the hierarchical correspondence between rows and columns in the table and multi-dimensional comparison information.
[0033] S23. For the flowchart block, call the multimodal large model to identify the nodes and edges of the flowchart image, and obtain a structured graph representation containing the node set, edge set and edge discrimination conditions.
[0034] The process of identifying nodes and edges by calling a multimodal large model on flowchart blocks includes the following sub-steps: S231. Call the multimodal large model to perform element detection on the flowchart image block and identify the set of nodes and the set of edges; S232. For each node, call the multimodal large model to classify the node as a candidate category, including the discriminant node, action node, start node, and end node. Combine the original text anchor point identifier to which the node belongs to perform context-aware text recognition on the text within the node, and extract the set of clinical concepts associated with the node from the recognized text. S233. For each edge, call the multimodal large model to extract the discrimination conditions on the edge from the label image of that edge; S234. Perform cycle checks on the obtained node set and edge set. For decision paths containing cycles, prioritize the explicit termination condition strategy for processing. That is, identify the exit judgment condition in the cycle and replace the back edges pointing to visited nodes in the cycle with edges pointing to the loop exit node caused by the exit judgment condition. For cycle segments that still cannot be resolved after explicit termination condition processing, adopt the cycle node encapsulation strategy to encapsulate the entire cycle segment into a cycle node. The cycle node retains the iterative decision logic inside and is connected to the structured graph representation as a single node. Record its iterative exit condition and maximum iteration number constraint. The output is a structured graph representation of the flowchart blocks formed by the node set, edge set, and edge discrimination conditions obtained from S231 to S234.
[0035] Specifically, such as Figure 3 As shown, Figure 3 In the middle (a), the original decision path contains a loop, where there are back edges pointing to visited nodes; Figure 3 (b) shows the result of the termination condition explicit strategy. After identifying the exit condition in the loop, the back edge is replaced with a directed edge pointing to the loop exit node, thus eliminating the loop while preserving the iterative clinical semantics. Figure 3(c) shows the result of the loop node encapsulation strategy. For complex loop segments that cannot be directly resolved, they are encapsulated as a single loop node, retaining the complete iterative decision logic internally, and externally connected to the structured graph representation as a single node. In the loop resolution process of S234, the determination condition for a loop segment that still cannot be resolved after explicit termination condition processing is any of the following: First, there are multiple interdependent exit judgment conditions in the loop, making it impossible to fully express the exit semantics through a single loop exit node; Second, the loop contains nested sub-loops, making the layer depth of the loop structure exceed 1; Third, the number of back edges in the loop exceeds the preset threshold 2, indicating that the loop segment corresponds to a composite iterative process with multi-path convergence. The above determination is obtained by the multimodal large model performing a structural check on the loop structure after identifying the exit judgment conditions, and the output is a binary judgment of "can be directly resolved" or "needs encapsulation", and the specific situation triggering encapsulation is written into the metadata of the loop node to ensure that the encapsulation decision is interpretable to the operator. Regardless of the resolution strategy employed, the nodes obtained through resolution or encapsulation all maintain a reverse pointer to the corresponding L1 evidence unit. This ensures that the results of loop processing are also traceable to the original guidelines.
[0036] By performing the above node classification, discrimination condition extraction, and loop resolution processing on the flowchart blocks, this step transforms the branch decision logic presented in graphical form in clinical guidelines into a structured graph representation that can be directly used for subsequent L2 construction. This solves the limitation that pure text parsing methods cannot fully capture the semantics of flowchart branches. At the same time, the loop resolution strategy preserves the semantics of guideline iterative diagnosis and treatment while maintaining the directed acyclic structure.
[0037] S24. For the image example block, call the multimodal large model to jointly understand the image and its captions, extract the description of clinical signs and the illustrated clinical situation, and establish cross-modal associations through the original text anchor mark and the evidence units corresponding to the semantically related text blocks and flowchart blocks.
[0038] Specifically, a multimodal large model is invoked to jointly understand the image and its captions, extracting descriptions of clinical signs and the clinical situations illustrated in the image. The image example block is then linked across modalities with the corresponding evidence units in the text blocks and flowcharts that are semantically related to the illustrated clinical situation, using its original text anchor markers. The output is a structured image example representation containing image content descriptions, caption text, and associated evidence unit markers. It should be noted that the image example block exists as a reference illustration in this method, providing intuitive image evidence and interpretability support for corresponding diagnostic and treatment decisions. Its parsing results are only linked to the corresponding L1 evidence units via original text anchor markers, and do not directly generate discrimination nodes, branch nodes, or recommended action nodes in the L2 decision graph, thus ensuring that the decision logic strictly originates from the text and flowchart content explicitly stated in the guidelines. In special cases where clinical guidelines use image feature descriptions as a criterion, when the criterion appears simultaneously as caption text in the associated text block of the image example block, the corresponding criterion node is generated from the parsing result of the text block, while the parsing result of the image example block is retained as supporting evidence to ensure the textual traceability of the decision-making logic and the integrity of the image evidence.
[0039] The above-obtained analysis results are bound to the original text anchor points of the corresponding content blocks to form a set of evidence units, constituting the L1 original text anchoring layer.
[0040] All parsing results are uniformly organized into a semantic intermediate representation, denoted as SemanticIR. Its structure is an index set with original text anchor identifiers as keys and semantic parsing results as values. The value of each record in SemanticIR is a structured object containing the following fields: modal type field (modal_type), with a value of one of "text", "table", "flowchart" or "image_example"; semantic parsing result field (parsed_content), for text blocks it is the extracted set of clinical concepts and their standard terminology mapping results, for table blocks it is a structured table object containing table headers and cell content, for flowchart blocks it is a local graph structure object composed of a set of nodes and a set of directed edges, and for image example blocks it is the discrimination condition description corresponding to the caption text; cross-modal association field (cross_modal_refs), recording other anchor identifiers semantically related to this record; evidence level field (evidence_grade), recording the recommendation level explicitly marked in this content block, set to null if no explicit marking is found.
[0041] It should be noted that, in this embodiment of the invention, the multimodal large model refers to a large-scale pre-trained model capable of jointly understanding and semantically parsing multiple modalities such as images, tables, and text. Its input may include a mixture of images and text, and its output is a structured semantic parsing result. This invention does not limit the specific type of multimodal large model used. Those skilled in the art can select existing multimodal large models with the aforementioned capabilities according to actual needs, such as GPT-4o (OpenAI), Claude 3.5 (Anthropic), Gemini 1.5 Pro (Google), or Deepseek. These models all possess the basic ability to semantically understand medical document images, tables, and flowcharts, and can be directly called or called after domain-specific adjustments in the steps described in this invention.
[0042] In one embodiment of the present invention, step S3 includes: S31. Traverse all evidence units in the L1 original text anchoring layer, and use different subgraph transformation strategies to convert each evidence unit into a decision subgraph according to the modality type of each evidence unit. Establish a reverse pointer to the source L1 evidence unit for each node in each subgraph, and merge all subgraphs into the initial L2 decision graph.
[0043] The L2 decision graph uses a directed acyclic graph structure, and its node types include the following four: **Decision Node:** Represents a clinical decision condition, such as "whether the patient's TNM stage is IIIA." The decision condition is standardized as a standard triple of "comparison object—comparison operator—comparison value." **Branch Node:** Represents a branch selection based on the decision result. **Action Node:** Represents a specific diagnostic action, such as "perform radical surgical resection," and records the associated evidence level attribute field `node_evidence_grade`. **Loop Node:** Encapsulates structures in the guideline flowchart containing iterative decision segments. Internally, it carries the sub-decision logic of the iterative segment. Externally, it connects to the L2 decision graph as a single node, recording its iteration exit condition and maximum iteration count constraint. Each node maintains a reverse pointer to an L1 evidence unit, denoted as `evidence_refs`.
[0044] Specifically, an empty directed acyclic graph (DAG) is first initialized as the initial L2 decision graph. Then, each evidence unit in the SemanticIR set obtained from S2 is traversed, and a corresponding subgraph transformation strategy is adopted based on its modality type: for evidence units corresponding to flowchart blocks, the structured graph representation obtained from S2, containing node sets, edge sets, and edge discrimination conditions, is directly converted into a decision subgraph, with the type and content of each node directly using the node classification and text recognition results from S2; for evidence units corresponding to table blocks, they are converted into decision subgraphs according to the semantic alignment results of the table header and table items. The decision subgraph, specifically, uses the dimensions expressed in the table header (such as treatment plan category, applicable population conditions) as discriminant nodes and the specific content corresponding to each table item (such as recommended treatment plan, evidence level) as recommended action nodes, and establishes directed edges according to the row and column correspondence of the table header and table item. For the evidence units corresponding to text blocks and image example blocks, the multimodal large model is invoked to extract the decision logic from their semantic parsing results to generate a decision subgraph. The subgraph transformation of the image example block only generates auxiliary evidence association nodes and does not directly generate discriminant nodes or recommended action nodes. For each node in each subgraph, by performing standard terminology alignment and matching between the node content and the L1 original text anchoring layer, a reverse pointer evidence_refs pointing to the corresponding L1 evidence unit is established. Then, the subgraph is merged into the initial L2 decision graph.
[0045] S32. Perform deduplication and merging on nodes that are determined to be semantically equivalent after normalization in the initial L2 decision graph. Take the union of the reverse pointers of the merged nodes and assign it to the merged node. Reconnect the incoming and outgoing edges of the merged nodes to the merged node to obtain the normalized L2 decision graph.
[0046] The normalization process specifically includes: mapping the clinical concepts in the content of each node in the initial L2 decision graph to standard terms based on the standard medical terminology database; for the discriminant nodes and branch nodes in the initial L2 decision graph that express clinical discriminant conditions, their discriminant conditions are structured and standardized, and uniformly represented as standard triplet forms consisting of comparison objects, comparison operators, and comparison values; when two nodes are of the same type and the standard terminology set and standard triplet discriminant conditions after the above normalization are consistent, they are determined to be semantically equivalent.
[0047] Specifically, the normalization process comprises two levels: First, for each node, the clinical concepts within its content are mapped to standard terms based on UMLS concept unique identifiers or SNOMED-CT concept codes to eliminate differences in synonyms, abbreviations, and order of expression. Second, the discrimination conditions for discriminant nodes and branch nodes are structurally standardized, normalizing different expressions such as "TNM stage is IIIA," "stage = IIIA," and "belongs to stage IIIA" into the same standard triplet form. During deduplication and merging, semantic embedding discrimination is also introduced as a supplement: when the cosine similarity of the concept embedding vectors of two nodes is higher than a preset deduplication threshold, they are also determined to be semantically equivalent. For nodes from different modalities that are determined to be semantically equivalent, a merging strategy is adopted to merge them into a single node. The union of the evidence_refs of the merged nodes is assigned to the merged node, so that the node retains traceable links pointing to the L1 evidence units corresponding to all its source modalities. The incoming and outgoing edges of the merged nodes are reconnected to the merged node. If the reconnection results in duplicate edges, the duplicate edges are deduplicated.
[0048] S33. Perform loop elimination processing on the loop-containing structures from the flowchart blocks in the normalized L2 decision graph. Prioritize the explicit termination condition strategy to eliminate loops. For loop segments that cannot be directly eliminated, use the loop node encapsulation strategy to encapsulate them as loop nodes to maintain the overall directed acyclicity of the final L2 decision graph; thus obtaining the final L2 decision graph.
[0049] Specifically, a full depth-first traversal is first performed on the normalized L2 decision graph to detect all path segments containing back edges (i.e., edges pointing to visited ancestor nodes), and the node set and back edge set of each loop segment are recorded. For each loop segment, a termination condition explicitness strategy is preferentially adopted: the multimodal large model is invoked to identify the exit criteria in the loop. The exit criteria refer to the clinical criteria that determine whether to end the cycle, such as "whether complete remission has been achieved", "whether intolerable adverse reactions have occurred", or "whether the preset number of treatment cycles has been reached". After identifying the exit criteria, the back edges pointing to visited nodes in the loop are replaced with directed edges pointing to the loop exit node derived from the exit criteria. Thus, while eliminating the loop, the clinical semantics of the original iterative path are preserved through the loop exit node and its labeled exit criteria.
[0050] For loop segments that cannot be directly resolved after explicit termination condition processing, a loop node encapsulation strategy is adopted. The determination condition is that any of the following conditions are met: First, the loop contains multiple interdependent exit conditions, i.e., the fulfillment of exit condition A depends on the historical value of another exit condition B, making it impossible to fully express the exit semantics through a single loop exit node; Second, the loop contains nested sub-loops, causing the loop structure's hierarchical depth to exceed 1; Third, the number of back edges in the loop exceeds a preset threshold (set to 2), indicating that the loop segment corresponds to a composite iterative process with multi-path convergence, and the exit condition is not unique. The above determination is obtained by the multimodal large model performing a structural check on the loop structure after identifying the exit conditions, outputting a binary judgment of "directly resolvable" or "requires encapsulation," along with a description of the specific circumstances triggering encapsulation. This description is also written into the loop node's metadata to ensure that the encapsulation decision itself is interpretable to the operator. For loop segments determined to "need encapsulation," the entire segment is encapsulated into a loop node. This loop node fully preserves the iterative decision logic (including all branch structures and temporal relationships within the loop) and connects to the L2 decision graph as a single node. The loop node also records its iteration exit condition and maximum iteration count constraint. Regardless of the resolution strategy used, the resolved or encapsulated nodes maintain a reverse pointer `evidence_refs` pointing to the corresponding L1 evidence unit to ensure that the loop processing results are also traceable back to the original guideline.
[0051] After the above processing, the final L2 decision graph is output, whose overall structure strictly maintains directed acyclicity. By sequentially executing three processing stages—subgraph transformation of multimodal evidence units, semantic normalization deduplication, and loop resolution—this step integrates clinical decision information scattered across various modalities of guideline text, tables, and flowcharts into a directed acyclic decision graph with a complete original text traceability chain. At the same time, it fully preserves the iterative diagnostic and treatment semantics in the guideline flowchart while ensuring that the overall structure can be used for subsequent subgraph segmentation.
[0052] In one embodiment of the present invention, step S4 includes: S41. Traverse the L2 decision graph, using the recommended action node as the splitting anchor point. For each splitting anchor point, collect the minimum set of predecessor nodes that make it complete and decidable along the incoming edge direction. The subgraph induced by the minimum set of predecessor nodes and the anchor point is taken as the candidate skill subgraph. Merge the candidate skill subgraphs with strong semantic association across boundaries, and verify the splitting results with the clinical scene ontology as a constraint to obtain the set of candidate skill subgraphs after splitting.
[0053] Specifically, S41 is implemented through a semantically aware subgraph segmentation algorithm for clinical decision graphs, and its execution process is as follows: The first step, target anchor point identification: Traverse the L2 decision graph and mark all recommended action nodes as splitting anchor points, denoted as the set of anchor points. .
[0054] The second step is backtracking using the minimal complete predecessor set: for each anchor point... Traverse the decision graph backwards along the incoming edges, collecting all decision nodes and branch nodes that are reachable from the anchor point, and removing those that are not. The discrimination logic does not generate redundant nodes in the constraints, thus achieving... The complete and determineable minimum set of predecessor nodes ;Will and The jointly induced subgraph is denoted as the candidate skill subgraph. .
[0055] The third step is boundary semantic verification: for any two candidate skill subgraphs... and Calculate the semantic similarity between cross-boundary adjacent nodes u and v. Two similarity measures are used simultaneously: one is the Jaccard similarity based on the clinical concept sets C(u) and C(v) associated with nodes. ; in, Let u be the Jaccard similarity between node u and node v based on the set of clinical concepts; u and v are cross-boundary adjacent nodes of any two candidate skill subgraphs. The set of clinical concepts associated with node u; This is the set of clinical concepts associated with node v.
[0056] The second is the cosine similarity based on the clinical concept embedding vectors e(u) and e(v): ; in Let cosine similarity be the clinical concept embedding vector between node u and node v. Let u be the clinical concept embedding vector. Clinical concept embedding vector for node v Take the weighted fusion value of the two: ; in, This represents the comprehensive semantic similarity between node u and node v (a weighted fusion value). The fusion weights for the similarity between the two paths; when Greater than the preset strong correlation threshold When u and v are in the candidate subgraph, they are merged to avoid semantically strongly related nodes being split into different skill units.
[0057] The fourth step involves hard constraints on the clinical scenario ontology: During the merging and segmentation process, a clinical scenario ontology is introduced as a hard constraint. Each node's associated disease entity is mapped to an ICD-10 disease classification tree, and the clinical scenario is mapped to the SNOMED-CT clinical scenario level. The constraint stipulates that all nodes within the same diagnostic skill unit must fall within the subtree of the same clinical scenario ontology node. When a decision path crosses the subtrees of different ontology nodes, it is forcibly cut off at the crossing point. Semantic similarity merging is only performed when two nodes belong to the same ontology node's subtree. Regarding constraints on the granularity of invocation, the number of decision structure nodes within a single skill is limited to no more than 50. When a clinical scenario is semantically cohesive but has more than 50 nodes, the semantic integrity of the clinical scenario is prioritized, and it is further subdivided according to the boundaries of decision subtasks. Simultaneously, the parent scenario identifier is recorded in the metadata of each subdivided unit to support on-demand combination and invocation by external systems.
[0058] S42. For each subgraph in the candidate skill subgraph set, construct its internal decision structure and perform static dependency analysis to establish a field-node dependency mapping table. Then, encapsulate the subgraph into an executable diagnostic and treatment skill unit. The executable diagnostic and treatment skill unit consists of an input parameter mode, an output parameter mode, an internal decision structure, and an evidence pointer set. The evidence pointer set records the L1 evidence unit identifier corresponding to each node. All executable diagnostic and treatment skill units constitute the L3 executable skill layer.
[0059] Specifically, each subgraph is encapsulated as an executable diagnostic and treatment skill unit. Its data structure includes a skill identifier field, a skill name field, a skill version field, a guideline source field, an input parameter mode field, an output parameter mode field, an internal decision structure field, an evidence pointer field, an applicable population constraint field, and a metadata field. The input parameter mode field defines the set of input parameters required when the skill is invoked, with each parameter distinguished between required (required is true) and optional (required is false). The output parameter mode field defines the set of output result fields after the skill is executed. The evidence pointer field records the L1 evidence unit identifier corresponding to each node, supporting complete traceability of the structured product to the original guideline text. The applicable population constraint field includes dimensions such as disease category, age range, organ function status, physical fitness score, pregnancy and lactation status, and previous treatment history. Each dimension is divided into hard constraints and soft constraints according to its constraint strength: for hard constraint dimensions, the diagnostic and treatment skill unit is not activated when the case to be decided does not meet the constraint; for soft constraint dimensions, it can still be invoked, but an applicability prompt is added to the result.
[0060] The static dependency analysis includes: traversing all discriminant nodes in the internal decision structure, extracting the input parameter fields on which the discriminant conditions of each discriminant node depend, establishing a mapping relationship between each required input parameter field and the set of discriminant nodes it affects, and pre-calculating the default influence degree of each required field. The default influence degree comprehensively considers the proportion of decision paths affected by the field's absence and the positional sensitivity of the associated discriminant node in the internal decision structure. The field-node dependency mapping table and the default influence degree of each field are stored together with the internal decision structure in the executable diagnostic skill unit. This allows the skill unit to perform local reachability analysis based on the field-node dependency mapping table when the input parameters are missing during runtime. It outputs deterministic recommendations for reachable paths and conditional recommendations with marked default dependency fields for unreachable paths. Even when the input parameters are incomplete, it can still output local recommendation results.
[0061] Specifically, the field-node dependency mapping is denoted as FieldNodeMap, and its construction method is as follows: traverse all discriminant nodes in the decision structure, extract the comparison object from the discriminant condition standard triplet of each discriminant node, align the comparison object with the field names of the input parameter pattern using standard terminology, and establish the field... f To the set of discriminant nodes N that depend on this field ( f Mapping relationship: ; in, For fields f The mapping relationship from the field to the set of discriminant nodes that depend on it, i.e., the field-node dependency mapping table; f This is a required input parameter field in the input parameter pattern of the diagnostic and treatment skills unit. It serves as a decision-making node within the internal decision-making structure.
[0062] For each required field f Pre-calculate its default impact during the encapsulation stage. Let the set of all complete paths from the root node to the recommended action node within the decision-making structure of the diagnostic skill unit be denoted as . The maximum depth of the decision structure is D, and the root node depth is defined as 0. Fields f The default influence degree is defined as: ; in, For fields f The default impact level; For path coverage loss rate, ; For depth-weighted position sensitivity, ; The larger the value, the more likely the field is to be larger. fThe default setting severely impacts skill unit availability. The formula employs a complementary multiplicative merging structure: For fields f The percentage of paths that are unaffected by the default settings. For fields f The degree to which associated nodes are not in critical positions; multiplying the two values represents the degree of union where both favorable conditions are simultaneously met. Taking the complement yields the default influence. Deterioration in either dimension will lead to... Significantly increased. FieldNodeMap and each field It is stored along with the skill package and is only available for runtime engine access; it is not exposed to the outside world.
[0063] When multiple required fields are omitted simultaneously, the overall impact of the omissions in this call is the maximum value of the impact of each omission. ; in The maximum value of the influence of each default field is taken as the overall default influence for this call. This is the set of all required fields that are omitted in this call. Based on the weakest link principle, the availability of the entire call is determined by the most significantly affected omitted field. The runtime engine then... The magnitude determines the integrity level of this call: ; in The integrity level for this call is determined by the runtime engine. The magnitude is determined; The lower limit of the preset grading threshold; This is the upper limit of the preset grading threshold.
[0064] When a runtime fallback call is triggered, the runtime engine first collects all the default required fields in this call to form a default field set. Calculate the set of undecidable nodes affected by default fields based on FieldNodeMap. , Each node in the process is marked as undecidable in this execution, and all its outgoing edges leading to its successor sub-paths are also marked as conditional paths. Subsequently, local reachability analysis is performed on the decision structure: the internal decision structure is traversed in a depth-first manner to identify all paths that are not traversed under the current known input parameters. The set of recommended action nodes that can be reached from any node in the network is denoted as the deterministic path set. All must go through The set of paths corresponding to recommended action nodes that can be reached only from at least one node is denoted as the conditional path set. .right The path is executed normally and a deterministic recommendation is generated. Simultaneously, based on the comprehensive strength of the path evidence level described in step S5, the strength of each deterministic path is calculated. And output it along with the result; The path generation conditional recommendation in the model annotates the set of default fields and the corresponding set of undecidable nodes on which it depends, and outputs partial evidence level information that can be calculated under the assumption that the default fields are known, for clinicians to refer to after supplementing the information.
[0065] The aforementioned downgraded invocation mechanism changes the handling of diagnostic skill units when required input parameters are missing from "full rejection" to "tiered output after local reachability analysis." This ensures effective recommendations are still provided even in scenarios with incomplete information, while also... The pre-calculation and integrity level classification provide downstream systems and clinicians with quantifiable default risk assessment criteria, improving the practical usability of diagnostic and treatment skill units in real clinical scenarios.
[0066] In one embodiment of the present invention, such as Figure 4 As shown, step S5 specifically includes: S51. For each executable diagnostic and treatment skill unit in the L3 executable skill layer, based on the node type of the subgraph corresponding to the skill unit in the L2 decision graph, restore the internal decision structure of the skill unit to natural language fragments according to the predefined template, and splice them in topological order to obtain the inverse rendering text of the skill unit.
[0067] Specifically, the de-rendering employs a node-based decomposition and generation strategy. For each node in the decision structure within a skill unit, the following operations are performed based on the node type: For discriminative nodes, their discriminative condition triplet (comparison object—comparison operator—comparison value) is restored to a natural language conditional sentence using a predefined clinical language template, for example, "TNM staging—equal to—Stage IIIA" is restored to "When the patient's TNM stage is Stage IIIA"; for recommended action nodes, their action description and evidence level attributes are restored to a recommendation statement using a predefined template, for example, "Recommend radical surgical resection (NCCN Class 1 recommendation)"; for loop nodes, their exit condition and maximum iteration count constraint are restored to an iterative treatment description statement. The de-rendered text of each node is concatenated according to the topological order of the decision structure and connected with conjunctions such as "if...then...otherwise..." to obtain the complete de-rendered text R(s). The reference text O(s) takes the original content field of the L1 evidence unit pointed to by all evidence_refs of the skill unit, sorts it according to the page number field and reading order field of its original text anchor, and then concatenates them in order so that R(s) and O(s) describe different forms of expression of the same clinical decision-making scenario, thus providing a clear semantic alignment basis for subsequent three-way similarity comparison.
[0068] S52. The anti-rendered text is concatenated with the original content of the L1 evidence unit referenced by the evidence pointer set of the executable diagnostic and treatment skill unit to form a reference text. A comprehensive consistency score is obtained by multi-dimensional similarity calculation and weighted fusion.
[0069] Specifically, the multi-dimensional similarity calculation comprises three paths. The first path is semantic embedding similarity (sim_emb), which uses a text embedding model adapted for the medical field to vectorize R(s) and O(s) before calculating cosine similarity. The second path is rule-matching similarity. sim_rule The matching rate, calculated based on structured elements such as key clinical concept sets, evidence level labels, and numerical ranges, is defined as follows: ; in sim_rule The similarity score is calculated based on the matching rate of structured elements such as key clinical concept sets, evidence level labels, and numerical ranges. Weights for similarity in the clinical entity dimension; The weights for similarity in the evidence level dimension; The weights of the three sub-dimensions are numerically constrained to determine the similarity of the dimensions. ; For clinical entity similarity, the ratio of the number of clinical entities that match the rendered segment and the original text (including exact matches and fuzzy matches scored by discounting the hierarchical relationship of SNOMED-CT) to the total number of entities is taken; For the similarity of evidence level, the value is 1 when the evidence level of the de-rendered fragment is completely consistent with that of the original text after mapping with the internal level scale. If the difference is one level, the score is calculated with a preset discount. If the difference is two levels or more, the value is 0. To constrain dimensional similarity, numerical constraints such as dosage, age, and stage are uniformly represented as intervals, with precise values degenerated into single-point intervals. The intervals of the inverse rendering fragments are then calculated. Interval with the original text overlap: ; in, This represents the interval representation of numerical constraints in the de-rendered fragment (precise values are degenerated into single-point intervals). This represents the interval representation corresponding to the numerical constraints in the original text.
[0070] The third approach is to score the similarity of the large model. The large model is invoked to perform semantic consistency scoring on R(s) and O(s), outputting the score within the interval [0,1]. The overall consistency score is defined as: ; in, The overall consistency score is a weighted fusion result of the three similarity scores. The fusion weights for semantic embedding similarity The fusion weights for rule-matching similarity. The fusion weights for scoring similarity in large models satisfy the following conditions: R(s) is the de-rendered text of skill unit s, which is the text obtained by restoring the internal decision structure to natural language fragments according to a predefined template and then splicing them in topological order; O(s) is the reference text of skill unit s, which is the text obtained by splicing the original content of the L1 evidence unit pointed to by all evidence_refs of this skill unit according to page number and reading order.
[0071] S53. Perform tiered processing according to the comprehensive consistency score: if the score is higher than the high confidence threshold, it is determined to be consistent and automatically passes; if the score is in the middle range, a large model local correction is triggered and the score is re-scored. If it still fails, it is transferred to expert verification; if the score is lower than the low confidence threshold, it is directly transferred to expert verification; the correction records of the executable diagnostic and treatment skill units that have been corrected by expert verification are written back to the sample library in a structured form, and the set of verified executable diagnostic and treatment skill units is output.
[0072] Specifically, the hierarchical processing rules are as follows: When Furthermore, if the difference in scores across the three pathways is less than 0.15, it is considered consistent, and the skill unit status is set to automatic pass; when If the range of the three-way scores is greater than or equal to 0.15, an automatic correction attempt is triggered. The large model is re-called to perform local corrections on the internal decision structure based on the differences. After correction, the overall consistency score is recalculated. If it still fails, it is transferred to the expert verification workbench. At that time, the system will directly switch to the expert verification workbench, and the skill unit status will be set to pending expert review.
[0073] The expert verification workbench presents the content to be verified at three access points: the first is the L2 to L3 conversion access point, which can be triggered after the subgraph segmentation and encapsulation results are generated, allowing experts to confirm the rationality of the segmentation boundaries; the second is the anti-rendering conflict access point, which is the skill unit in S53 that has not passed automatic correction; and the third is the coverage blind spot access point. Coverage blind spots are determined using a two-stage cross-analysis method: The first stage involves set difference calculation based on evidence pointers. The complete set is constructed by taking the evidence identifiers of all evidence units in the L1 original text anchoring layer, and the covered set is constructed by taking the union of the evidence pointer sets of all diagnostic and treatment skill units. The difference between these two sets is calculated to obtain evidence units not referenced by any skill unit's evidence pointer, which serves as the explicit coverage blind spot. The second stage involves coverage retrieval based on semantic embedding. For each evidence unit determined to be referenced in the first stage, the similarity between its original text semantic embedding vector and the semantic representation corresponding to the internal decision structure of the diagnostic and treatment skill unit that references it is further calculated. When the similarity is lower than a preset coverage threshold, the evidence unit is determined to be referenced, but its essential decision content is not covered, serving as the implicit coverage blind spot. Both explicit and implicit coverage blind spots are presented at the coverage blind spot access point for expert verification. The expert verification results are written back in the form of structured correction records. On the one hand, this is used to correct the diagnostic and treatment skill unit ontology; on the other hand, it enters the sample library as training samples for incremental fine-tuning of the multimodal large model, forming a positive closed loop of construction-verification-optimization.
[0074] Furthermore, S53 also includes: after outputting the verified set of executable diagnostic and treatment skill units, for each executable diagnostic and treatment skill unit in the set, traversing each complete decision path p in its internal decision structure, and calculating the comprehensive strength score of the path evidence level: ; in, The overall strength score for path evidence level, p It represents a complete decision-making path from the root node to the recommended action node within the internal decision-making structure of the diagnostic and treatment skill unit. The integrity weight of the path evidence level is labeled as follows: k is the number of nodes on path p whose node_evidence_grade field is not empty, n(p) is the total number of nodes on the path, and node_evidence_grade is the evidence level attribute field of the recommendation action node, recording the recommendation level label associated with that node. The sparser the annotations, the better. The smaller; The arithmetic mean of the internal scale values of the evidence levels of each node on the path is defined as follows: , Let be the internal scale value of the evidence level for the i-th node with an evidence level label on path p. The smaller the value, the higher the overall evidence level of the path; To be The mapping result linearly normalized to the [0,1] interval; H(p) is the normalized information entropy of the path evidence level distribution, defined as... ,in Rank scale value The frequency of occurrence in G(p), Rank scale value l In the path p G(p) represents the number of times each node appears in the path p, and is the multiset of evidence level scale values for all nodes with evidence level annotations on the path p. Higher entropy indicates that the quality of evidence at each stage of the path is more dispersed. This is the entropy penalty coefficient, used to control the penalty applied to the overall strength due to uneven distribution of evidence levels. A larger value indicates a higher overall recommendation strength for the path. When all nodes on path p have no explicit evidence level labeling, G(p) is an empty set. Set to null. When multiple complete paths exist within a single skill unit, each path is calculated independently. They are not merged, and the differences in the overall recommendation strength between paths are preserved.
[0075] The comprehensive strength score of the path evidence level is extended and written into the output parameter mode of the corresponding executable diagnostic and treatment skill unit. The new fields include: path_evidence_grade (string type, outputting the original grade label corresponding to the node that has the greatest constraint on the comprehensive recommendation strength of the path, after inverse mapping of the internal grade scale) and path_evidence_score (numeric type, with a value of The fields are: `path_evidence_grade_detail` (an array containing the node identifier, node content summary, evidence level label, and source L1 evidence unit identifier for each valid evidence level node on the path, arranged in order of node order on the path), and `path_evidence_grade_source` (a string indicating the node with the highest `node_evidence_grade`, representing the weakest evidence link in the path; when multiple nodes are listed, they are listed sequentially). These fields provide clinical users with a basis for ranking and classifying recommendation results and support review of the complete chain of evidence step by step.
[0076] In this step, through anti-rendering comparison and three-way similarity fusion verification mechanism, the quality of automated build is quantifiable and systematically verified without relying on full manual review. The hierarchical conflict handling strategy focuses expert verification work on skill units with real uncertainties, effectively reducing the cost of expert intervention.
[0077] In one embodiment of the present invention, the construction method further includes version management of the set of executable diagnostic and treatment skill units output by S5, specifically including: using a three-level semantic version specification of major version number, minor version number, and revision number to mark each executable diagnostic and treatment skill unit; when the clinical guidelines on which it is based are updated, performing a difference analysis on the old and new versions of the guidelines to identify newly added, modified, and obsolete guideline content fragments, and performing incremental version changes on the affected executable diagnostic and treatment skill units based on the difference analysis results: when the content of the L1 evidence unit referenced by the skill unit undergoes substantial changes, the major version number is incremented; when the internal decision-making structure of the skill unit is adjusted but the core recommendations remain unchanged, the minor version number is incremented; when only metadata or evidence pointers are adjusted, the revision number is incremented; each version change generates a difference record, which includes the change type, change fields, content before the change, content after the change, and related guideline difference anchors to support the incremental synchronization and version tracking of executable diagnostic and treatment skill units by the downstream system.
[0078] Specifically, differential analysis uses the original anchor points as the smallest granularity unit. By performing anchor-level set difference operations and semantic similarity comparisons on the SemanticIR of the old and new versions of the guidelines, it identifies the set of anchor points where the content has undergone substantial changes. For each affected anchor point, the impact of the change is propagated layer by layer along the reverse pointer link from L1 evidence unit → L2 decision graph node → L3 executable diagnostic and treatment skill unit. This accurately locates the set of skill units that need to be modified, avoiding unnecessary reconstruction of skill units that are not affected by the guideline update, and supporting incremental and traceable knowledge maintenance in guideline update scenarios.
[0079] like Figure 5 As shown, the present invention also provides an automated system for constructing structured clinical treatment skills based on clinical guidelines, including a multimodal parsing module, an L1 / L2 / L3 conversion module, a consistency verification module, an expert verification workbench, and a version management module.
[0080] The multimodal parsing module receives multimodal clinical guideline documents (e.g., PDF format oncology treatment guidelines). Internally, it integrates a layout parsing engine and a multimodal large-scale model calling interface, executing the process in two phases. The first phase is layout parsing: the layout parsing engine performs block-level detection page by page, identifying and segmenting the content of each page into four types of content units: text blocks, table blocks, flowchart blocks, and image example blocks. For each content unit, it generates a source anchor tag carrying document identifiers, page numbers, block bounding box coordinates, modality type, and reading order, serving as a basic index throughout the three-layer architecture. The second phase is semantic parsing: for different modal content units, the multimodal large-scale model is invoked to perform semantic parsing. Clinical concepts are extracted from text blocks and mapped to standard terminology; table blocks are restored to a structured table representation including header hierarchy and merged cell information; flowchart blocks undergo node classification, discrimination condition extraction, and loop resolution; and image example blocks are jointly parsed with captions and image content, establishing cross-modal associations with related text blocks and flowchart blocks. The parsing results are organized into a multimodal semantic intermediate representation (SemanticIR) using the original text anchor identifier as the key, and then passed to the L1 / L2 / L3 transformation module.
[0081] The L1 / L2 / L3 transformation module contains three processing units that are executed sequentially. The L1 evidence unit construction unit traverses SemanticIR and encapsulates the original content field, modality type field, semantic parsing result field, cross-modal association field, and evidence level field of each content unit into an evidence unit, which is then stored in the evidence unit library. All evidence units together constitute the L1 original text anchoring layer. The L2 decision graph construction unit takes the evidence unit library as input. Based on the modality type of each evidence unit, it uses the corresponding subgraph transformation strategy to convert each evidence unit into a decision subgraph and merge them into an initial L2 decision graph. Then, it sequentially performs semantic normalization deduplication (merging nodes from different modalities but semantically equivalent and taking the union of their reverse pointers) and loop elimination processing (prioritizing the explicit termination condition strategy and using the loop node encapsulation strategy for loop segments that cannot be directly eliminated). The result is a final L2 decision graph with node types including discriminant nodes, branch nodes, recommended action nodes, and loop nodes, and the overall structure is a directed acyclic structure. Each node maintains a reverse pointer to the source L1 evidence unit, evidence_refs, and an evidence level attribute field, node_evidence_grade. The L3 skill segmentation and encapsulation unit takes the final L2 decision graph as input, backtracks to the minimum complete set of predecessor nodes using recommended action nodes as segmentation anchors to obtain candidate skill subgraphs, completes subgraph segmentation through cross-boundary semantic similarity verification and hard constraints of clinical scenario ontology, and then performs static dependency analysis on each subgraph to establish a field-node dependency mapping table and pre-calculates the default influence of each field. The subgraph is encapsulated into an executable diagnostic and treatment skill unit containing input parameter mode, output parameter mode, internal decision structure, evidence pointer set, applicable population constraints and version number, and stored in the diagnostic and treatment skill unit library. All units together constitute the L3 executable skill layer.
[0082] The consistency verification module performs consistency verification on each executable diagnostic and treatment skill unit in the diagnostic and treatment skill unit library. The verification process consists of three steps: First, according to the node type, the internal decision structure is restored to natural language fragments by topological sorting using a predefined clinical language template and then concatenated to obtain the inverse rendered text R(s); Second, R(s) is concatenated with the original content of the L1 evidence unit referenced by the evidence pointer set of the skill unit to form the reference text O(s). The reference text O(s) is then calculated and weighted by three factors: semantic embedding similarity sim_emb, rule matching similarity sim_rule, and large model scoring similarity sim_llm, to obtain the comprehensive consistency score sim_total; Finally, tiered processing is performed based on sim_total: if sim_total ≥ 0.85 and the range of the three scores is less than 0.15, it is automatically passed; if 0.70 ≤ sim_total < 0.85 or the range of the three scores exceeds the limit, automatic correction is triggered, and the score is re-evaluated. If it still fails, it is pushed to the expert verification workbench; if sim_total < 0.70, it is directly pushed to the expert verification workbench. In addition, the consistency verification module also performs coverage cross-analysis, identifies explicit and implicit coverage blind spots through set difference operation and semantic embedding coverage retrieval, and pushes the skill units corresponding to the blind spots to the expert verification workbench.
[0083] The expert verification workbench provides three preset access points: an L2 to L3 conversion access point, an anti-rendering conflict access point, and a coverage blind spot access point. At each access point, the decision image segment or treatment skill unit to be verified, along with its anti-rendering text, reference text, and a summary of scoring differences, are presented to the experts. The expert verification workbench receives expert review comments and structures them into correction records containing change type, change fields, content before change, and content after change. On one hand, the correction records are directly applied to the corresponding skill units in the treatment skill unit library; on the other hand, the correction records are written to the sample library for subsequent incremental fine-tuning of the multimodal large model.
[0084] The version management module employs a three-tiered semantic versioning specification—major version number, minor version number, and revision number—to version-mark each executable clinical skill unit in the clinical skill unit library. When clinical guidelines are updated, the version management module performs a set difference operation and semantic similarity comparison on the SemanticIR of the old and new versions of the guidelines, using the original text anchor identifier as the smallest granularity. This identifies the set of anchor points with substantial changes, and then propagates the impact of the changes layer by layer along the reverse pointer link from L1 evidence unit → L2 decision graph node → L3 executable clinical skill unit, precisely locating the set of skill units that require version changes: the major version number increments when there are substantial changes to the cited content; the minor version number increments when the decision structure is adjusted but the core recommendations remain unchanged; and the revision number increments when only metadata or evidence pointers are adjusted. Each version change generates a difference record containing the change type, changed fields, and related guideline difference anchor points to support incremental synchronization and version tracking of executable clinical skill units in downstream systems. Skill units unaffected by guideline updates remain unchanged and are not rebuilt.
[0085] Furthermore, to verify the actual effectiveness of the technical solution of this invention, three clinical guidelines—the NCCN Clinical Practice Guidelines for Non-Small Cell Lung Cancer (3rd Edition, 2024), the CSCO Guidelines for the Diagnosis and Treatment of Breast Cancer (1st Edition, 2024), and the CSCO Guidelines for the Diagnosis and Treatment of Colorectal Cancer (1st Edition, 2024)—were used as inputs to conduct a multi-dimensional quantitative evaluation of the diagnostic and treatment skill unit set constructed by the method of this invention. All experiments were completed on a server cluster equipped with eight NVIDIA A100 80GB GPUs.
[0086] The baseline comparisons included: (a) a traditional artificial knowledge engineering method, collaboratively completed by three engineers with over five years of clinical knowledge engineering experience; (b) an ablation scheme that removed multimodal parsing capabilities and retained only the plain text large model processing; (c) an ablation scheme that removed the L1 original text anchoring layer and did not maintain the original text pointer in the structured product; (d) an ablation scheme that removed the anti-rendering comparison mechanism; and (e) an ablation scheme that removed the expert verification loop. Among these, (a) was used for construction efficiency comparison, and (b) through (e) were used for ablation experiments to quantify the independent contribution of each core module of the present invention.
[0087] (1) Construction efficiency Table 1 shows a comparison of the construction time of the method in this invention and the traditional artificial knowledge engineering method for the end-to-end structured construction task of the NCCN Clinical Practice Guidelines for Non-Small Cell Lung Cancer: Table 1 Both machine time and human time are directly compared in hours. Machine time is the calculation time of the automation module, and human time is the actual working time of the human. The sum of the two is the total time consumed in this stage.
[0088] The results show that the method of this invention reduces the end-to-end structured construction cycle of a single tumor diagnosis and treatment guideline from approximately 49 person-days (calculated at 8 hours / day) to approximately 1.5 person-days, significantly reducing the engineering cost of knowledge assets. In the consistency verification stage, the introduction of anti-rendering comparison and hierarchical processing mechanisms reduces the actual workload of expert intervention to 12 person-hours. In contrast, the automated part (0.6 machine-hours) requires entirely manual verification of scoring and hierarchical routing in traditional methods, lacking a corresponding automated time benchmark.
[0089] (2) Distribution of overall consistency scores The anti-rendering comparison and three-way similarity fusion verification described in step S5 were performed on the 186 diagnostic and treatment skill units constructed in the NCCN Clinical Practice Guidelines for Non-Small Cell Lung Cancer. The distribution of the overall consistency score sim_total was as follows: sim_total≥0.85 (automatic pass) accounted for 78.5% (146 units), 0.70≤sim_total<0.85 (triggered automatic repair) accounted for 16.1% (30 units), and sim_total<0.70 (directly transferred to expert verification) accounted for 5.4% (10 units). Among the 30 skill units that triggered automatic repair, 23 units reached the automatic pass threshold again after one round of automatic correction, with an automatic repair success rate of 76.7%. The mean values of the three-way similarity (sim_emb, sim_rule, sim_llm) on the automatically passed samples were 0.912, 0.879, and 0.928, respectively, with standard deviations all less than 0.06, indicating that the three-way scores corroborate each other and have good stability.
[0090] (3) Ablation test To verify the effectiveness of each core module of this invention, an ablation experiment was conducted using end-to-end accuracy (defined as the proportion of skill units that were deemed "semantically correct and with accurate evidence pointers" by three independent clinical experts in a blind review) and evidence pointer hit rate as evaluation indicators. The results are shown in Table 2. Table 2 Ablation results showed that the multimodal parsing capability described in step S2 contributed 32.9 percentage points to the end-to-end accuracy, indicating that the decision-making information in the form of tables and flowcharts in clinical guidelines has a decisive impact on the final skill quality, while pure text parsing suffers from serious information loss. The introduction of the L1 original text anchoring layer improved the evidence pointer hit rate by 59.6 percentage points, directly verifying the core role of the original text anchoring and reverse pointer mechanism in supporting original text tracing. The anti-rendering comparison mechanism described in step S5 contributed 14.7 percentage points to the end-to-end accuracy, indicating that this mechanism effectively intercepted logical deviations in the automated construction process. The expert verification closed loop contributed an additional 9.6 percentage points. Each module has a significant independent contribution, and there is a synergistic gain relationship between modules.
[0091] (4) Skill coverage The coverage cross-analysis described in step S5 was performed on the set of diagnostic and treatment skill units constructed in the NCCN Clinical Practice Guidelines for Non-Small Cell Lung Cancer. A total of 198 structured decision segments were extracted from the guidelines. The 186 diagnostic and treatment skill units constructed using the method of this invention covered 192 of these, achieving a coverage of 96.97%. The six uncovered segments were all determined to be prospective research statements at the end of the guidelines and did not constitute actionable diagnostic and treatment decisions. The average coverage across the three guidelines was 95.4%, significantly higher than the 81.6% achieved by traditional manual methods (manual methods are prone to omissions, especially in footnotes and appendices).
[0092] (5) Workload of expert verification After introducing the expert verification closed loop, when the cumulative verification workload of 3 oncology experts on the NCCN Clinical Practice Guidelines for Non-Small Cell Lung Cancer was 12 people, the corresponding number of skill units to be verified was 40 (7 that still failed after triggering automatic repair + 10 that were directly transferred for verification + 23 for coverage blind spot checks), with an average verification time of 18 minutes per skill unit. Compared with the traditional manual method requiring 390 people, the expert workload was reduced by 96.9%, indicating that the graded processing strategy in step S5 effectively focused the scope of expert intervention on skill units with genuine uncertainty.
[0093] (6) Incremental fine-tuning returns The 412 accumulated expert verification records were used as incremental fine-tuning samples to perform LoRA fine-tuning on the multimodal large model called in step S2. The fine-tuned model's initial comprehensive consistency score (sim_total) on the new guideline (CSCO Colorectal Cancer Diagnosis and Treatment Guidelines) increased from 0.823 to 0.881, the automatic pass rate increased from 71.2% to 83.7%, and the proportion of skill units requiring expert verification decreased from 8.4% to 3.1%. This indicates that the "construction-verification-optimization" closed loop described in this invention has a significant positive iterative effect, and the initial quality of subsequent guideline construction can be continuously improved as verification samples accumulate.
[0094] (7) Version management efficiency In the scenario of updating the NCCN Clinical Practice Guidelines for Non-Small Cell Lung Cancer from 2024.v3 to 2024.v4, the differential analysis described in step S5 identified 23 substantial changes. The incremental update mechanism of this invention only performs version changes on the 31 affected diagnostic and treatment skill units (with major version numbers increasing by 8, minor version numbers increasing by 14, and revision numbers increasing by 9), while the remaining 155 skill units remain unchanged. Compared to the traditional full reconstruction method, the incremental update time is reduced from approximately 240 person-hours to 0.9 machine-hours + 2.5 person-hours, an efficiency improvement of approximately 70 times. This verifies the practical value of the change impact propagation mechanism using the original text anchor as the smallest granularity unit in guideline update scenarios.
[0095] Based on the above experimental results, the technical solution described in this invention is significantly superior to existing methods in five dimensions: construction efficiency, consistency, coverage, traceability, and iteration capability. This verifies the effectiveness and synergy of the L1 / L2 / L3 three-layer architecture, the anti-rendering comparison mechanism, and the expert verification closed loop.
[0096] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for automatically constructing structured clinical treatment skills based on clinical guidelines, characterized in that, Includes the following steps: S1. The layout parsing engine is used to process the multimodal clinical guideline document page by page, dividing the document content into content blocks with different modal types, and generating original text anchor points carrying page number, position and modal type information for each content block; S2. Call the multimodal large model to perform semantic parsing on each content block obtained in S1. Extract clinical concepts, structured representations and inter-block relationships according to the modal features of each content block. Bind the parsing results with the corresponding original text anchor point identifiers to form a set of evidence units carrying semantic information and original text positioning information, which constitutes the L1 original text anchoring layer. S3. Based on the set of evidence units in the L1 original text anchoring layer, a combination of large model and rules is used to extract and organize the clinical decision conditions, decision results and their dependencies in the evidence units into a directed acyclic graph composed of nodes and directed edges. Each node maintains a reverse pointer to its source evidence unit, forming an L2 decision graph representing the clinical decision process. S4. The L2 decision graph is divided into subgraphs according to clinical application scenarios. Each subgraph is encapsulated into an executable diagnostic and treatment skill unit consisting of input parameter mode, output parameter mode and internal decision structure. All executable diagnostic and treatment skill units constitute the L3 executable skill layer. S5. Based on the L1 original text anchoring layer and the L2 decision graph, verify the consistency of each executable diagnosis and treatment skill unit in the L3 executable skill layer with the original clinical guidelines, correct the executable diagnosis and treatment skill units that fail the verification, and output a set of verified executable diagnosis and treatment skill units, which can be called by external systems. In step S1, the content blocks are divided into text blocks, table blocks, flowchart blocks, and image example blocks according to their modal types. The original text anchor point identifier is composed of the document identifier field, version field, page number field, block identifier field, block bounding box field, block type field, and reading order field. The block bounding box field describes the position of the content block using the coordinates of the top left corner and bottom right corner of the page, and the block type field records the modal type to which the content block belongs. Step S5 specifically includes: S51. For each executable diagnostic and treatment skill unit in the L3 executable skill layer, based on the node type of the subgraph corresponding to the skill unit in the L2 decision graph, restore the internal decision structure of the skill unit to a natural language fragment according to the predefined template, and splice them in topological order to obtain the inverse rendering text of the skill unit. S52. The anti-rendered text is concatenated with the original content of the L1 evidence unit referenced by the evidence pointer set of the executable diagnostic and treatment skill unit to form a reference text. A comprehensive consistency score is obtained by multi-dimensional similarity calculation and weighted fusion. S53. Perform tiered processing according to the comprehensive consistency score: if the score is higher than the high confidence threshold, it is determined to be consistent and automatically passes; if the score is in the middle range, a large model local correction is triggered and the score is re-scored. If it still fails, it is transferred to expert verification; if the score is lower than the low confidence threshold, it is directly transferred to expert verification; the correction records of the executable diagnostic and treatment skill units that have been corrected by expert verification are written back to the sample library in a structured form, and the set of verified executable diagnostic and treatment skill units is output.
2. The method for automatically constructing structured clinical treatment skills according to claim 1, characterized in that, Step S2 involves calling a multimodal large model to perform semantic parsing on each content block obtained in S1, specifically including: S21. For the text block, call the large model to extract the set of clinical concepts contained therein, and perform standard medical terminology mapping on each clinical concept to obtain the set of clinical concepts and standard terminology mapping results of the text block. S22. For the table block, call the multimodal large model to perform structural restoration of the table image to obtain a structured table representation containing the header hierarchy, cell content and merged cell information. S23. For the flowchart block, call the multimodal large model to identify nodes and edges in the flowchart image, and obtain a structured graph representation containing the node set, edge set and edge discrimination conditions; S24. For the image example block, call the multimodal large model to jointly understand the image and its captions, extract the description of clinical signs and the illustrated clinical situation, and establish cross-modal associations through the evidence units corresponding to the semantically related text blocks and flowchart blocks identified by the original anchor points. The above-obtained analysis results are bound to the original text anchor points of the corresponding content blocks to form a set of evidence units, constituting the L1 original text anchoring layer.
3. The method for automatically constructing structured clinical guidance diagnostic and treatment skills as described in claim 2, characterized in that, The process of identifying nodes and edges by calling a multimodal large model on the flowchart blocks includes the following sub-steps: S231. Call the multimodal large model to perform element detection on the flowchart image block and identify the set of nodes and the set of edges; S232. For each node, call the multimodal large model to classify the node as a candidate category, including the discriminant node, action node, start node, and end node. Combine the original text anchor point identifier to which the node belongs to perform context-aware text recognition on the text within the node, and extract the set of clinical concepts associated with the node from the recognized text. S233. For each edge, call the multimodal large model to extract the discrimination conditions on the edge from the label image of that edge; S234. Perform cycle checks on the obtained node set and edge set. For decision paths containing cycles, prioritize the explicit termination condition strategy for processing. That is, identify the exit judgment condition in the cycle and replace the back edges pointing to visited nodes in the cycle with edges pointing to the loop exit node caused by the exit judgment condition. For cycle segments that still cannot be resolved after explicit termination condition processing, adopt the cycle node encapsulation strategy to encapsulate the entire cycle segment into a cycle node. The cycle node retains the iterative decision logic inside and is connected to the structured graph representation as a single node. Record its iterative exit condition and maximum iteration number constraint. The output is a structured graph representation of the flowchart blocks formed by the node set, edge set, and edge discrimination conditions obtained from S231 to S234.
4. The method for automatically constructing structured clinical guidance diagnostic and treatment skills as described in claim 1, characterized in that, Constructing an L2 decision graph in S3 includes the following sub-steps: S31. Traverse all evidence units in the L1 original text anchoring layer, and use different subgraph transformation strategies to convert each evidence unit into a decision subgraph according to the modality type of each evidence unit. Establish a reverse pointer to the source L1 evidence unit for each node in each subgraph, and merge all subgraphs into the initial L2 decision graph. S32. Perform deduplication and merging on nodes that are determined to be semantically equivalent after normalization in the initial L2 decision graph. Take the union of the reverse pointers of each node to be merged and assign it to the merged node. Reconnect the inbound and outbound edges of the merged nodes to the merged node to obtain the normalized L2 decision graph. S33. Perform loop elimination processing on the loop-containing structures from the flowchart blocks in the normalized L2 decision graph. Prioritize the explicit termination condition strategy to eliminate loops. For loop segments that cannot be directly eliminated, use the loop node encapsulation strategy to encapsulate them as loop nodes to maintain the overall directed acyclicity of the final L2 decision graph; thus obtaining the final L2 decision graph.
5. The method for automatically constructing structured clinical guidance diagnostic and treatment skills as described in claim 4, characterized in that, The normalization process described in step S32 specifically includes: Clinical concepts in each node of the initial L2 decision graph are mapped to standard terms based on a standard medical terminology database. For the discriminant nodes and branch nodes in the initial L2 decision graph that express clinical discrimination conditions, their discrimination conditions are structured and standardized, and uniformly represented as standard triplet forms consisting of comparison objects, comparison operators, and comparison values. When two nodes are of the same type and the standard terminology set after the above normalization is consistent with the standard triplet discrimination conditions, they are determined to be semantically equivalent.
6. The method for automatically constructing structured clinical guidance diagnostic and treatment skills as described in claim 1, characterized in that, Step S4 includes the following sub-steps: S41. Traverse the L2 decision graph, using the recommended action node as the splitting anchor point. For each splitting anchor point, collect the smallest predecessor node set that makes it complete and decidable along the incoming edge direction. The subgraph induced by the smallest predecessor node set and the anchor point is taken as the candidate skill subgraph. Merge candidate skill subgraphs with strong semantic association across boundaries, and verify the segmentation results using the clinical scenario ontology as a constraint to obtain a set of candidate skill subgraphs with complete segmentation; S42. For each subgraph in the candidate skill subgraph set, construct its internal decision structure and perform static dependency analysis to establish a field-node dependency mapping table. Then, encapsulate the subgraph into an executable diagnostic and treatment skill unit. The executable diagnostic and treatment skill unit consists of an input parameter mode, an output parameter mode, an internal decision structure, and an evidence pointer set. The evidence pointer set records the L1 evidence unit identifier corresponding to each node. All executable diagnostic and treatment skill units constitute the L3 executable skill layer.
7. The method for automatically constructing structured clinical guidance diagnostic and treatment skills as described in claim 6, characterized in that, The static dependency analysis described in S42 includes: Traverse all the decision nodes in the internal decision structure, extract the input parameter fields on which the decision conditions of each decision node depend, establish the mapping relationship between each required input parameter field and the set of decision nodes it affects, and pre-calculate the default influence degree of each required field. The default influence degree comprehensively considers the proportion of decision paths affected after the field is omitted and the position sensitivity of the associated decision node in the internal decision structure. The field-node dependency mapping table and the default influence of each field are stored together with the internal decision structure in the executable diagnostic and treatment skill unit. This allows the skill unit to perform local reachability analysis based on the field-node dependency mapping table when there are default input parameters during runtime. It can output deterministic recommendations for reachable paths and conditional recommendations with default dependency fields marked for unreachable paths. Even when the input parameters are incomplete, it can still output local recommendation results.
8. The method for automatically constructing structured clinical guidance diagnostic and treatment skills as described in claim 1, characterized in that, S53 also includes: after outputting a set of validated executable diagnostic and treatment skill units, for each executable diagnostic and treatment skill unit in the set, traversing each complete decision path p in its internal decision structure, and calculating the comprehensive strength score of the path evidence level: ; in, The overall strength score for the path evidence level. Assign integrity weights to path evidence levels; It is the arithmetic mean of the internal scale values of the evidence level of each node on the path; The normalized information entropy for the path evidence level distribution; The entropy penalty coefficient; The comprehensive strength score of the path evidence level is extended and written into the output parameter mode of the corresponding executable diagnostic and treatment skill unit, which is used to provide clinical users with a graded and ranked recommendation result.
9. The method for automatically constructing structured clinical treatment skills according to claim 1, characterized in that, The construction method also includes version management of the executable diagnostic and treatment skill unit set output by S5, specifically including: A three-level semantic versioning specification—major version number, minor version number, and revision number—is used to version-mark each executable clinical skill unit. When the underlying clinical guidelines are updated, a difference analysis is performed on the old and new versions of the guidelines to identify newly added, modified, and obsolete guideline content fragments. Based on the difference analysis results, incremental version changes are performed on the affected executable clinical skill units: the major version number increments when the content of the L1 evidence unit referenced by the skill unit undergoes substantial changes; the minor version number increments when the internal decision-making structure of the skill unit is adjusted but the core recommendations remain unchanged; and the revision number increments when only metadata or evidence pointers are adjusted. Each version change generates a difference record, which includes the change type, change fields, content before the change, content after the change, and related guideline difference anchors to support incremental synchronization and version tracking of executable clinical skill units in downstream systems.
10. A clinical guideline-based structured diagnostic and treatment skills automatic construction system, characterized in that, The system is used to implement the method as described in any one of claims 1-9, comprising: The multimodal parsing module receives multimodal clinical guideline documents, performs layout parsing, identifies content units such as text blocks, table blocks, flowchart blocks, and image example blocks, and generates original text anchor point identifiers carrying page numbers, positions, and modal types; and calls the multimodal large model to perform semantic parsing on each content unit to obtain a multimodal semantic intermediate representation indexed by the original text anchor points; The L1 / L2 / L3 conversion module is used to sequentially construct the L1 original text anchoring layer, L2 decision graph, and L3 executable skill layer based on the multimodal semantic intermediate representation. Specifically, the L1 evidence unit construction unit encapsulates the original content, modal information, and semantic parsing results of each content unit into evidence units and stores them in the evidence unit library; the L2 decision graph construction unit extracts clinical decision conditions, decision results, and their dependencies from the evidence units, organizes them into a directed acyclic graph composed of nodes and directed edges, and maintains a reverse pointer to the source evidence unit in each node; the L3 skill segmentation and encapsulation unit segments the L2 decision graph into subgraphs according to clinical application scenarios and encapsulates the subgraphs into executable diagnostic and treatment skill units composed of input parameter modes, output parameter modes, and internal decision structures. The consistency verification module is used to perform anti-rendering comparison and multi-dimensional similarity calculation on the executable diagnostic and treatment skill unit, output a comprehensive consistency score, and trigger automatic correction or push the skill unit to be corrected to the expert verification workbench based on the score result. The expert verification workbench is used to present decision image segments and diagnostic and treatment skill units to be verified at preset access points, receive expert verification opinions, and write back structured correction records to the sample library and diagnostic and treatment skill unit library. The version management module is used to perform semantic version marking, change difference recording, and incremental release management of the executable diagnostic and treatment skill units to support incremental and traceable maintenance in clinical guideline update scenarios.
Citation Information
Patent Citations
Decision tree generation method based on large language model and related equipment
CN119227784A
Anesthesia virtual simulation training system fusing knowledge, skills and thinking closed loop
CN121747390A
Cross-guide recommendation suggestion consistency detection method and system based on structural features
CN122287632A