Clinical trial protocol generation system and method based on multi-source medical knowledge graph
Patent Information
- Application Number
- CN202610917348.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-22
AI Technical Summary
[0005]有鉴于此,本申请实施例提供了一种基于多源医学知识图谱的临床试验方案生成系统及方法,以解决现有技术存在的知识语义分散、方案要素缺乏联动、合规执行校验滞后的问题
通过知识图谱构建模块,用于获取多源医学知识数据,对多源医学知识数据进行术语归一、证据标注和规则结构化处理,构建包含医学实体节点、证据节点、规则节点和试验要素节点的多源医学知识图谱;试验意图解析模块,用于解析临床试验设计需求,生成试验目标向量、研究约束向量和方案骨架节点;证据子图生成模块,用于根据试验目标向量、研究约束向量和方案骨架节点,在多源医学知识图谱中提取目标证据子图,并基于目标证据子图生成候选方案路径;方案约束推理模块,用于对候选方案路径中的试验要素节点进行一致性校验、合规规则匹配和执行约束评价,生成方案推理结果;方案装配模块,用于根据方案推理结果确定目标方案路径,将目标方案路径转换为结构化临床试验方案对象,并基于结构化临床试验方案对象生成临床试验方案文本及对应的方案要素映射结果。本申请能够提高知识关联准确性、增强方案一致性、提升合规执行校验效率。
Smart Images

Figure CN122800282A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent clinical trial technology, and in particular to a clinical trial protocol generation system and method based on a multi-source medical knowledge graph. Background Technology
[0002] With the continuous improvement of informatization and intelligence in clinical research, the design of clinical trial protocols is gradually evolving from traditional manual writing to a generation-assisted approach based on medical knowledge bases, historical trial data, and intelligent generation models. Clinical trial protocols typically require the integration of disease diagnosis and treatment knowledge, drug intervention information, inclusion and exclusion criteria, endpoint indicators, visit arrangements, safety monitoring requirements, and regulatory compliance rules. There are strong medical logical relationships and execution constraints among their contents.
[0003] In existing technologies, some systems assist in generating clinical trial protocol texts through template filling, keyword retrieval, or large language model generation. While these methods can improve the efficiency of initial protocol draft generation, they typically focus on generating chapter content, lacking a unified semantic organization of multi-source medical knowledge. This makes it difficult to convert guidelines, literature, previous trials, real-world data, and compliance rules into computable and traceable structured knowledge relationships. Furthermore, in the current protocol generation process, elements such as research objectives, population conditions, intervention pathways, endpoint indicators, visit plans, statistical analysis, and safety monitoring are often generated in a decentralized manner, lacking cross-element consistency verification and constraint reasoning mechanisms. This easily leads to problems such as mismatched protocol elements, unclear evidence sources, and difficulty in identifying rule conflicts.
[0004] Furthermore, existing technologies typically require manual feasibility and compliance reviews after the protocol text is finalized. This makes it difficult to comprehensively assess participant recruitment accessibility, center execution capabilities, data collection availability, and regulatory compliance during the protocol generation stage. This leads to repeated protocol revisions and affects the structured integration of subsequent case report forms, visit plans, and statistical analysis plans. Therefore, it is necessary to propose a clinical trial protocol generation system capable of evidence association, constraint reasoning, and structured assembly based on multi-source medical knowledge graphs. Summary of the Invention
[0005] In view of this, embodiments of this application provide a clinical trial protocol generation system and method based on a multi-source medical knowledge graph to solve the problems of fragmented knowledge semantics, lack of linkage between protocol elements, and lagging compliance execution verification in the prior art.
[0006] The first aspect of this application provides a clinical trial protocol generation system based on a multi-source medical knowledge graph, comprising: a knowledge graph construction module for acquiring multi-source medical knowledge data, performing terminology normalization, evidence annotation, and rule structuring on the multi-source medical knowledge data, and constructing a multi-source medical knowledge graph containing medical entity nodes, evidence nodes, rule nodes, and trial element nodes; a trial intent parsing module for parsing clinical trial design requirements and generating trial target vectors, research constraint vectors, and protocol skeleton nodes; an evidence subgraph generation module for extracting target evidence subgraphs from the multi-source medical knowledge graph based on the trial target vectors, research constraint vectors, and protocol skeleton nodes, and generating candidate protocol paths based on the target evidence subgraphs; a protocol constraint reasoning module for performing consistency verification, compliance rule matching, and execution constraint evaluation on the trial element nodes in the candidate protocol paths, and generating protocol reasoning results; and a protocol assembly module for determining the target protocol path based on the protocol reasoning results, converting the target protocol path into a structured clinical trial protocol object, and generating clinical trial protocol text and corresponding protocol element mapping results based on the structured clinical trial protocol object.
[0007] The second aspect of this application provides a clinical trial protocol generation method based on a multi-source medical knowledge graph, using the system of the first aspect. The method includes: acquiring multi-source medical knowledge data; performing terminology normalization, evidence annotation, and rule structuring on the multi-source medical knowledge data to construct a multi-source medical knowledge graph; parsing clinical trial design requirements to generate trial target vectors, research constraint vectors, and protocol skeleton nodes; extracting a target evidence subgraph from the multi-source medical knowledge graph based on the trial target vectors, research constraint vectors, and protocol skeleton nodes, and generating candidate protocol paths based on the target evidence subgraphs; performing constraint reasoning on the trial element nodes in the candidate protocol paths to generate protocol reasoning results, and determining the target protocol path based on the protocol reasoning results; converting the target protocol path into a structured clinical trial protocol object, and generating clinical trial protocol text and corresponding protocol element mapping results based on the structured clinical trial protocol object.
[0008] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: The application employs a knowledge graph construction module to acquire multi-source medical knowledge data, performs terminology normalization, evidence annotation, and rule structuring on this data, and constructs a multi-source medical knowledge graph containing medical entity nodes, evidence nodes, rule nodes, and trial element nodes. A trial intent parsing module analyzes clinical trial design requirements, generating trial objective vectors, research constraint vectors, and protocol skeleton nodes. An evidence subgraph generation module extracts target evidence subgraphs from the multi-source medical knowledge graph based on the trial objective vectors, research constraint vectors, and protocol skeleton nodes, and generates candidate protocol paths based on these subgraphs. A protocol constraint reasoning module performs consistency verification, compliance rule matching, and execution constraint evaluation on the trial element nodes in the candidate protocol paths, generating protocol reasoning results. A protocol assembly module determines the target protocol path based on the protocol reasoning results, converts the target protocol path into a structured clinical trial protocol object, and generates the clinical trial protocol text and corresponding protocol element mapping results based on the structured clinical trial protocol object. This application can improve the accuracy of knowledge association, enhance protocol consistency, and improve the efficiency of compliance execution verification. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a schematic diagram of the structural composition of the clinical trial protocol generation system based on multi-source medical knowledge graph provided in the embodiments of this application; Figure 2 This is a flowchart illustrating the clinical trial protocol generation method based on a multi-source medical knowledge graph provided in this application embodiment. Detailed Implementation
[0011] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0012] In existing technologies, clinical trial protocols are typically generated through manual writing, template filling, keyword retrieval, or general intelligent generation models. While these methods can improve the efficiency of initial protocol draft generation to some extent, they rely heavily on fixed chapter structures and scattered data retrieval, making it difficult to unify disease diagnosis and treatment knowledge, drug intervention knowledge, previous clinical trial evidence, real-world data, medical terminology standards, and regulatory compliance rules into a calculable and traceable structured knowledge system. Furthermore, existing protocol generation processes usually focus on text content generation, lacking systematic modeling of the medical logical relationships and execution constraints between protocol elements.
[0013] The resulting technical problems include semantic dispersion, inconsistent terminology, and difficulty in tracing the sources of evidence among multiple sources of medical knowledge; a lack of interconnected reasoning among elements such as research objectives, population conditions, intervention pathways, endpoint indicators, visit plans, safety monitoring, and statistical analysis in clinical trial protocols, which easily leads to element mismatches, rule conflicts, and inconsistencies; in addition, existing technologies usually conduct manual compliance reviews and feasibility assessments after the protocol text is formed, making it difficult to simultaneously complete compliance rule matching and execution constraint evaluation during the protocol generation stage, resulting in repeated protocol revisions and affecting the structured connection of subsequent case report forms, visit plans, and statistical analysis plans.
[0014] To address the aforementioned issues, this application provides a clinical trial protocol generation system based on a multi-source medical knowledge graph. This system first acquires multi-source medical knowledge data and then performs terminology normalization, evidence annotation, and rule structuring on the data. This constructs a multi-source medical knowledge graph containing medical entity nodes, evidence nodes, rule nodes, and trial element nodes, thereby forming a unified semantic foundation for medical knowledge and an evidence constraint space.
[0015] Building upon this foundation, the system analyzes clinical trial design requirements, generating trial objective vectors, research constraint vectors, and protocol skeleton nodes. Then, based on these vectors, it extracts a target evidence subgraph from a multi-source medical knowledge graph and generates candidate protocol paths carrying evidence citation relationships. In this approach, the core elements of the clinical trial protocol are no longer generated by piecing together isolated text fragments, but rather driven by the evidence subgraph, preserving the citation relationships between protocol elements and medical evidence.
[0016] Furthermore, the system performs consistency verification, compliance rule matching, and execution constraint evaluation on the experimental element nodes in the candidate solution path, generating solution reasoning results. This process enables unified reasoning on the constraint relationships between research objectives and endpoint indicators, inclusion and exclusion criteria and target populations, visitation plans and data collection, statistical analysis and endpoint types, and safety monitoring and intervention risks during the solution generation stage, and combines rule nodes and execution constraints to screen and modify candidate solution paths.
[0017] Finally, the system determines the target protocol path based on the protocol reasoning results, converts the target protocol path into a structured clinical trial protocol object, and generates the clinical trial protocol text and corresponding protocol element mapping results based on the structured clinical trial protocol object. This structured clinical trial protocol object can simultaneously support protocol text generation, evidence tracing, visit plan configuration, data collection item generation, and statistical analysis item integration, so that the protocol generation results are not only presented as natural language text, but also serve as a structured data foundation for subsequent clinical trial execution and management system calls.
[0018] Through the above technical solutions, this application can improve the accuracy of multi-source medical knowledge association, enhance the consistency among clinical trial protocol elements, improve the efficiency of compliance rule matching and execution constraint verification, and reduce the semantic deviation between protocol generation results and subsequent trial execution objects.
[0019] The specific components and functions of the clinical trial protocol generation system based on multi-source medical knowledge graph provided in this application will be described in detail below with reference to the accompanying drawings and specific embodiments. Figure 1 This is a schematic diagram of the structural composition of the clinical trial protocol generation system based on a multi-source medical knowledge graph provided in the embodiments of this application, such as... Figure 1 As shown, the system may specifically include the following components: The knowledge graph construction module 101 is used to acquire multi-source medical knowledge data, perform terminology normalization, evidence annotation and rule structuring on the multi-source medical knowledge data, and construct a multi-source medical knowledge graph containing medical entity nodes, evidence nodes, rule nodes and experimental element nodes. The trial intent parsing module 102 is used to parse the clinical trial design requirements and generate trial objective vectors, research constraint vectors, and protocol skeleton nodes. The evidence subgraph generation module 103 is used to extract the target evidence subgraph from the multi-source medical knowledge graph based on the experimental target vector, the research constraint vector and the scheme skeleton node, and to generate candidate scheme paths based on the target evidence subgraph. The scheme constraint reasoning module 104 is used to perform consistency verification, compliance rule matching, and execution constraint evaluation on the test element nodes in the candidate scheme path, and generate scheme reasoning results. The protocol assembly module 105 is used to determine the target protocol path based on the protocol reasoning results, convert the target protocol path into a structured clinical trial protocol object, and generate the clinical trial protocol text and corresponding protocol element mapping results based on the structured clinical trial protocol object.
[0020] In some embodiments, terminology normalization, evidence annotation, and rule structuring are performed on multi-source medical knowledge data to construct a multi-source medical knowledge graph containing medical entity nodes, evidence nodes, rule nodes, and experimental element nodes, including: Source analysis and content segmentation are performed on multi-source medical knowledge data to generate standardized knowledge units that carry source version, scope of application and time status; Based on the medical terminology ontology and semantic embedding model, entity recognition, concept alignment and synonym merging are performed on medical concepts in standardized knowledge units to generate medical entity nodes. Standardized knowledge units are labeled according to evidence level, research scenario, and rule attributes to generate evidence nodes and rule nodes. Structured elements related to clinical trial design are extracted from standardized knowledge units to generate trial element nodes. Graph relationship edges are established based on the semantic dependencies between nodes to obtain a multi-source medical knowledge graph.
[0021] Specifically, when constructing a multi-source medical knowledge graph, the system first performs source analysis on the accessed medical knowledge data, identifying the data source type, publication time, applicable region, applicable disease scope, data version, and update status. Then, it generates content segments based on the minimum knowledge granularity required by the clinical trial protocol. Content segmentation can be based on disease definitions, diagnostic stratification, intervention measures, inclusion criteria, exclusion criteria, endpoint indicators, visit requirements, safety monitoring, statistical analysis rules, and data collection requirements.
[0022] For the same guideline document, the system can break down disease diagnostic criteria, staging criteria, recommended medication conditions, and monitoring requirements into multiple standardized knowledge units. For previous trial registration data, the system can break down study type, study stage, participant range, primary endpoint, observation period, and trial status into multiple standardized knowledge units. Each standardized knowledge unit is configured with source version, scope of application, and expiration status for subsequent evidence citation and version tracking.
[0023] After generating standardized knowledge units, the system identifies and normalizes medical concepts based on a medical terminology ontology and a semantic embedding model. The medical terminology ontology provides standardized hierarchical relationships between diseases, symptoms, examination items, drug components, intervention methods, trial endpoints, and safety events. The semantic embedding model calculates semantic similarity and contextual consistency between different expressions. When the same medical concept has abbreviations, aliases, translations, or expressions at different granularities, the system first identifies candidate medical concepts, then aligns them by considering the disease scope, research scenario, and relational semantics within the context. Finally, concepts that meet the merging criteria are merged into a unified medical entity node. For example, in the scenario of generating diabetes-related clinical trial protocols, the system can merge synonymous expressions in adult type 2 diabetes, type 2 diabetes patients, and the target disease population into the same disease entity node, and identify glycated hemoglobin, fasting blood glucose, and hypoglycemic events as examination indicator entity nodes or safety event entity nodes, respectively.
[0024] Furthermore, the system performs evidence annotation and rule structuring based on the evidence source, research design type, sample size, publication time, applicable population, and rule nature corresponding to the standardized knowledge units. Evidence annotation is used to generate evidence nodes, which at least include evidence level, source identifier, applicable stage, applicable population, and evidence credibility. Rule structuring is used to generate rule nodes, which can represent ethical restrictions, inclusion restrictions, contraindications, follow-up windows, reporting requirements for serious adverse events, and data retention requirements.
[0025] For recommendations in the guidelines, the system can generate evidence nodes based on the strength and scope of the recommendation. For contraindications and monitoring requirements in drug instructions, the system can convert them into rule nodes with triggering conditions and constraints. Constraints are established between rule nodes and medical entity nodes, enabling subsequent candidate pathways to invoke the corresponding rules when generating inclusion / exclusion criteria, safety monitoring, and visit plans.
[0026] Subsequently, the system extracts structured elements related to clinical trial design from standardized knowledge units and generates trial element nodes. These nodes express design content that can directly participate in protocol generation, including research objectives, population conditions, intervention pathways, control methods, endpoint indicators, visit events, data collection items, safety monitoring, and analysis items. The system establishes graph relationship edges based on the semantic dependencies between medical entity nodes, evidence nodes, rule nodes, and trial element nodes.
[0027] For example, in a Phase II trial of a new drug for type 2 diabetes, the system can associate the target disease entity node with the population condition node, the target drug entity node with the intervention path node, the change in glycated hemoglobin with the primary endpoint node, the observation period with the visit event node, and the hypoglycemic event with the safety monitoring node. If an exclusion condition originates from a drug contraindication rule, the system further establishes a rule reference edge between the exclusion condition node and the corresponding rule node.
[0028] Through the above processing, a multi-source medical knowledge graph can unify medical concepts, evidence, constraints, and protocol design elements from scattered sources into a computable graph structure. This graph structure can provide foundational data for evidence subgraph extraction, candidate protocol path generation, consistency verification, and compliance rule matching during subsequent clinical trial protocol generation. It can improve the accuracy of medical knowledge associations, enhance semantic consistency between protocol elements, and improve the evidence traceability of protocol generation results.
[0029] In some embodiments, the clinical trial design requirements are parsed to generate a trial objective vector, a study constraint vector, and protocol skeleton nodes, including: Intent identification and semantic slot extraction are performed on clinical trial design requirements to obtain core design elements; The core design elements are semantically aligned with the experimental element nodes in the multi-source medical knowledge graph, and the alignment results are vector-encoded to generate the experimental target vector. Based on the applicable scope, evidence constraints, and rule constraints corresponding to the core design elements, a research constraint vector is generated. Based on the experimental objective vector and research constraint vector, the chapter hierarchy and element occupancy relationships of the clinical trial protocol are determined, and the protocol skeleton nodes are generated.
[0030] Specifically, upon receiving clinical trial design requirements, the system first standardizes the input, converting natural language descriptions, structured form fields, and historical project configurations into a unified design requirement text. The system then uses an intent recognition model to identify the research task type, target disease, research phase, intervention subjects, control settings, observation objectives, and output format, and extracts core design elements using a semantic slot extraction model. These core design elements may include the target indication, target population, research objective, intervention measures, main observation directions, trial phase, study region, and expected protocol type. For input content that is omitted or incomplete, the system performs candidate completion based on trial element nodes associated with the target disease and intervention measures in a multi-source medical knowledge graph, and assigns low-confidence flags to the completed content for differentiated processing during subsequent evidence subgraph retrieval and rule validation.
[0031] After obtaining the core design elements, the system semantically aligns these elements with the experimental element nodes in a multi-source medical knowledge graph. During this alignment process, the system first determines the standard concepts corresponding to the core design elements based on a medical terminology ontology. Then, it calculates the semantic similarity between the core design elements and the experimental element nodes using a semantic embedding model, and adjusts the similarity based on the node context relationships. These node context relationships can include disease classification relationships, intervention attribution relationships, endpoint adaptation relationships, and research stage relationships.
[0032] For example, in a scenario where a user inputs a Phase II randomized controlled trial of a hypoglycemic drug for adult patients with type 2 diabetes to observe the improvement in glucose metabolism, the system aligns type 2 diabetes to the population condition node associated with the disease entity, aligns the hypoglycemic drug to the intervention path node, aligns the improvement in glucose metabolism to endpoint indicator nodes such as changes in glycated hemoglobin and fasting blood glucose, and aligns the Phase II randomized controlled trial to the study design node.
[0033] After semantic alignment, the system performs vector encoding on the alignment results to generate a trial target vector. The trial target vector represents the target direction of clinical trial design needs in the protocol generation space, and its encoded content includes semantic information such as study subjects, intervention categories, study phases, endpoint propensity, and protocol type. The system can use a graph neural network to aggregate and encode the aligned trial element nodes and their adjacency relationships, and then fuse the aggregation results with the text semantic encoding results to obtain a trial target vector that simultaneously expresses the user's design intent and the association relationships within the medical knowledge graph. For slots with low confidence, the system reduces the corresponding feature weights during vector encoding to avoid uncertain elements excessively affecting subsequent evidence retrieval.
[0034] Furthermore, the system generates research constraint vectors based on the applicable scope, evidence constraints, and rule constraints corresponding to the core design elements. The applicable scope defines the disease type, subject population, research phase, regional environment, and protocol suitability boundaries; evidence constraints represent the required level of evidence, source of evidence, and timeliness of evidence for candidate protocol elements; and rule constraints represent ethical restrictions, safety monitoring requirements, contraindications, data collection standards, and follow-up window limitations. Taking a Phase II trial of a new drug for type 2 diabetes as an example, the system can generate constraints on subject age and disease status based on the adult patient population, dose exploration and short-to-medium-term endpoint constraints based on the Phase II study phase, and hypoglycemia monitoring and liver and kidney function test constraints based on the safety rules of hypoglycemic drugs. These constraints are then encoded into a research constraint vector.
[0035] Subsequently, the system determines the chapter hierarchy and element placement relationships of the clinical trial protocol based on the trial objective vector and research constraint vector, generating protocol skeleton nodes. The system first matches the protocol document skeleton corresponding to the current research task type, then determines key chapters and necessary design elements based on the objective vector, and determines the reserved evidence citation locations, rule validation locations, and execution evaluation locations within each chapter based on the constraint vector. Protocol skeleton nodes can include research background nodes, research objective nodes, research design nodes, population criteria nodes, intervention protocol nodes, endpoint indicator nodes, visitation plan nodes, safety monitoring nodes, statistical analysis nodes, and data management nodes. Hierarchical, citation, and dependency relationships are established between these protocol skeleton nodes. For example, the primary endpoint node establishes an analytical dependency relationship with the statistical analysis node, the visitation plan node establishes a data collection dependency relationship with the endpoint indicator node, and the safety monitoring node establishes a risk constraint relationship with the intervention protocol node.
[0036] Through the above processing, the system can convert clinical trial design requirements into computable trial objective vectors, research constraint vectors, and protocol skeleton nodes, establishing a stable mapping relationship between natural language requirements and trial elements in a multi-source medical knowledge graph. This implementation method can improve the accuracy of requirement parsing, enhance the consistency of subsequent evidence subgraph extraction and candidate protocol path generation, and improve the automation level of clinical trial protocol chapter assembly and constraint reasoning.
[0037] In some embodiments, extracting a target evidence subgraph from a multi-source medical knowledge graph and generating candidate solution paths based on the target evidence subgraph includes: Using the experimental target vector and research constraint vector as retrieval conditions, semantic recall and relation expansion are performed on the multi-source medical knowledge graph to obtain a set of candidate evidence nodes that match the skeleton nodes of the scheme. Based on evidence fit, rule constraint relationship and node association strength, the candidate evidence node set is sorted and conflict pruning is performed to generate the target evidence subgraph. Based on the element occupancy relationship of the scheme skeleton node, the evidence path in the target evidence subgraph is converted into a combinable sequence of test elements, and candidate scheme paths carrying evidence citation relationships are generated.
[0038] Specifically, after generating the experimental target vector and research constraint vector, the system uses both as entry points for graph retrieval, performing semantic recall and relation expansion within a multi-source medical knowledge graph. Semantic recall first calculates the matching degree between the protocol skeleton nodes and medical entity nodes, evidence nodes, and experimental element nodes based on the experimental target vector, filtering candidate evidence nodes related to the target disease, intervention measures, research phase, and observation objectives. Relationship expansion then performs multi-hop traversal along the reference edges, adaptation edges, and constraint edges between evidence nodes and experimental element nodes, supplementing with associated nodes related to inclusion / exclusion criteria, endpoint indicators, visit windows, safety monitoring, and statistical analysis. For the recall results, the system filters based on the scope of application, evidence level requirements, and rule restrictions in the research constraint vector, ensuring that the set of candidate evidence nodes matches the element occupancy relationships in the protocol skeleton nodes.
[0039] In a specific example, the user inputs a Phase II randomized controlled trial of a hypoglycemic drug for adult patients with type 2 diabetes. The system recalls evidence nodes related to type 2 diabetes, hypoglycemic drug intervention, Phase II study, randomized controlled design, and glucose metabolism endpoints based on the trial target vector. It then excludes evidence nodes that are not applicable to adult patients, have a mismatched study phase, or are outdated based on the study constraint vector. The system subsequently expands along the disease entity nodes to target population conditions, along the intervention path nodes to dosing cycles and safety monitoring items, and along the endpoint indicator nodes to observation windows and data collection items, forming a set of candidate evidence nodes corresponding to the protocol skeleton nodes.
[0040] After obtaining the set of candidate evidence nodes, the system calculates the evidence fit of each candidate evidence node relative to the current scheme generation task. Evidence fit can be generated comprehensively based on evidence source level, overlap of research subjects, similarity of intervention measures, consistency of research stages, and publication time status. The system simultaneously reads the rule nodes connected to the candidate evidence nodes, identifies taboo conditions, ethical restrictions, visitation window restrictions, and security monitoring constraints, and determines the priority of each evidence node in the scheme path based on the node association strength. If two candidate evidence nodes support contradictory population conditions, or if the endpoint observation window corresponding to one evidence node does not match the visitation rate corresponding to another evidence node, the system will generate a conflict identifier and perform conflict pruning based on evidence fit and rule priority.
[0041] After path sorting and conflict pruning, the system generates a target evidence subgraph. This subgraph includes evidence nodes, trial element nodes, medical entity nodes, and rule nodes directly related to the current clinical trial design requirements, preserving the evidence citation relationships, semantic dependencies, and rule constraints between nodes. Following the element placement relationships in the protocol skeleton nodes, the system decomposes the evidence path in the target evidence subgraph into composable sequences of trial elements. These sequences are organized according to the dependency order between research objectives, population conditions, intervention pathways, endpoint indicators, visitation plans, safety monitoring, and statistical analysis, with each trial element associated with a corresponding evidence node and rule node.
[0042] Taking the aforementioned Phase II trial of a new drug for type 2 diabetes as an example, the system can generate a candidate protocol path from the target evidence subgraph. This candidate protocol path includes adult patients with type 2 diabetes as the target population, hypoglycemic drugs combined with lifestyle management as the intervention path, changes in glycated hemoglobin as the primary endpoint, changes in fasting blood glucose and hypoglycemic events as associated observational elements, fixed visit nodes within the pre-set observation period as the data collection arrangement, and hypoglycemic monitoring, liver and kidney function tests, and restrictions on concomitant medications as safety constraints. These elements are not generated in isolation but are formed through the transformation of the evidence path, and the source node, version identifier, and scope of application are recorded through evidence citation relationships.
[0043] Through the above processing, the system can extract target evidence subgraphs that match the clinical trial design requirements from a multi-source medical knowledge graph and convert these subgraphs into candidate solution paths carrying evidence citation relationships. This implementation method can improve the accuracy of evidence association in candidate solution generation, reduce semantic conflicts between solution elements, and enhance the reliability of subsequent consistency verification, compliance rule matching, and structured solution assembly.
[0044] In some embodiments, the evidence paths in the target evidence subgraph are converted into a composable sequence of test elements, and candidate solution paths carrying evidence reference relationships are generated, including: Based on the element type and hierarchical relationship corresponding to the skeleton node of the scheme, the evidence path in the target evidence sub-graph is decomposed to generate evidence fragments that match the different element positions. Based on the semantic dependencies and rule constraints corresponding to the evidence fragments, the combination order and reference boundaries between the experimental elements are determined, and a sequence of experimental elements is generated. The experimental element sequence is linked to the corresponding evidence node through a reference mapping, and candidate solution paths are generated by combining them according to the preset solution path generation rules.
[0045] Specifically, after obtaining the target evidence subgraph, the system first reads the type, level, prerequisite dependencies, and required attributes of each element's placeholder in the scheme skeleton node, and then decomposes the evidence path in the target evidence subgraph accordingly. During path decomposition, the system uses the evidence node as the starting point and identifies continuous paths that can support scheme generation along evidence reference edges, element adaptation edges, and rule constraint edges. It then segments these paths into multiple evidence fragments according to element types such as research design, population conditions, intervention path, endpoint indicators, visitation plans, safety monitoring, and statistical analysis. Each evidence fragment retains its source version, scope of application, evidence level, associated entities, and available boundaries to prevent the same evidence content from being incorrectly filled into mismatched scheme placeholders. For cases where the same path covers multiple element placeholders, the system splits the path according to the hierarchy of node relationships, allowing population condition evidence fragments, endpoint indicator evidence fragments, and visitation window evidence fragments to participate in subsequent combinations.
[0046] In a specific example, for a phase II randomized controlled trial of hypoglycemic drugs for type 2 diabetes in adults, the target evidence subplot may include evidence from previous similar trials, treatment guidelines, drug safety rules, and testing specifications. The system decomposes pathways related to the diagnostic scope of type 2 diabetes in adults into population-conditional evidence fragments, pathways related to changes in glycated hemoglobin and fasting blood glucose into endpoint indicator evidence fragments, pathways related to observation period and testing frequency into visitation plan evidence fragments, and pathways related to hypoglycemic events and liver and kidney function monitoring into safety monitoring evidence fragments. Each evidence fragment is assigned to a corresponding element placeholder in the protocol skeleton node and configured with a fragment identifier and a source node identifier.
[0047] Subsequently, the system determines the combination order of experimental elements based on the semantic dependencies and rule constraints between evidence fragments. Semantic dependencies are used to limit the basis of a certain experimental element to its preceding elements; for example, the primary endpoint needs to match the research objective and target population, the visitation plan needs to cover the observation window corresponding to the primary endpoint, and the statistical analysis items need to match the data type of the primary endpoint. Rule constraints are used to limit the combinability boundaries of a certain experimental element; for example, the intervention path triggers safety monitoring requirements, exclusion conditions trigger contraindication constraints, and data collection items trigger visitation window restrictions. The system generates citation boundaries based on the above relationships. Citation boundaries are used to indicate the range of elements that the evidence fragments can support, the range of chapters that cannot be crossed, and the rule nodes that need to be jointly cited.
[0048] When generating the experimental element sequence, the system follows a logical order: research objectives first, population conditions defined, intervention path implementation, endpoint indicator determination, visitation plan implementation, safety monitoring supplementation, and statistical analysis adaptation. This process converts evidence fragments into a sequence of experimental elements with dependent directions. If multiple endpoint indicators exist, the system prioritizes them based on evidence fit, research phase consistency, and rule conflict status. If a safety monitoring fragment has strong rule constraints with an intervention path, the system binds that safety monitoring element to the corresponding intervention element to prevent omission during subsequent scheme path combinations. After completing the experimental element sequence construction, the system establishes a reference mapping between each experimental element in the sequence and its corresponding evidence node, recording the evidence source, version status, and scope of application. It then generates candidate scheme paths according to preset scheme path generation rules.
[0049] Through the above processing, the system can transform complex evidence paths in the target evidence subgraph into a combinable, traceable, and verifiable sequence of experimental elements, forming candidate solution paths carrying evidence citation relationships. This implementation method can improve the evidence support integrity of candidate solution paths, reduce the risk of solution element mismatch and evidence citation overstepping, and improve the accuracy of subsequent solution constraint reasoning and structured solution assembly.
[0050] In some embodiments, consistency verification, compliance rule matching, and execution constraint evaluation are performed on the test element nodes in the candidate solution path to generate solution reasoning results, including: Based on the semantic dependencies and temporal constraints between the test element nodes in the candidate solution path, generate element consistency verification results; Based on the rule nodes in the multi-source medical knowledge graph, rule matching and conflict identification are performed on the candidate solution paths to generate compliant matching results. Based on the execution resource constraints and data acquisition constraints corresponding to the candidate solution paths, the feasibility of the candidate solution paths is evaluated, and execution evaluation results are generated. The results of consistency verification of integrated elements, compliance matching, and execution evaluation are combined to generate the solution reasoning results.
[0051] Specifically, after obtaining candidate solution paths, the system first reads the semantic dependencies and temporal constraints between the nodes of each experimental element in the candidate solution path, and constructs a solution element verification graph. The solution element verification graph uses research objectives, population conditions, intervention paths, endpoint indicators, visitation plans, safety monitoring, and statistical analysis as verification objects, and evidence citation relationships, data collection dependencies, analysis dependencies, and risk constraints as verification edges. The system traverses each verification edge according to the solution generation order, determining whether preceding experimental elements can provide a clear citation basis for subsequent experimental elements, and whether subsequent experimental elements exceed the limitations of preceding experimental elements, thereby generating element consistency verification results.
[0052] During consistency verification, the system can determine whether the primary endpoint covers the current research question based on the semantic matching degree between the research objective and the endpoint indicators; whether the target population can generate corresponding observation data based on the fit between population conditions and endpoint indicators; whether the visit time covers the key observation window based on the temporal constraints between the endpoint indicators and the visit plan; and whether the analysis method matches the data structure based on the analytical dependencies between the endpoint data type and the statistical analysis items. For example, in a phase II randomized controlled trial of hypoglycemic drugs for type 2 diabetes in adults, if the candidate protocol path sets the change in glycated hemoglobin as the primary endpoint, but the visit plan only includes examinations at weeks 2 and 4 after administration, the system will identify a temporal inconsistency between the primary endpoint observation window and the visit plan, and mark the visit node as needing to be supplemented or adjusted in the element consistency verification results.
[0053] Subsequently, the system performs rule matching and conflict identification on candidate protocol paths based on rule nodes in a multi-source medical knowledge graph. During rule matching, the system performs graph matching between the population conditions, intervention path, safety monitoring, data collection, and follow-up arrangements in the candidate protocol path and the rule nodes respectively, determining whether the candidate protocol path triggers ethical restrictions, contraindications, medication restrictions, adverse event monitoring requirements, data logging requirements, and follow-up window requirements. For trial element nodes that trigger rules, the system establishes rule reference edges and records the rule source, version status, and scope of application. If a candidate's inclusion condition overlaps with a contraindication rule, or if an intervention path triggers safety monitoring requirements but the corresponding monitoring item is not configured in the candidate protocol path, the system generates a compliance conflict identifier and writes the conflict location, conflict type, and associated rule nodes into the compliance matching results.
[0054] Furthermore, the system evaluates the feasibility of candidate study pathways based on execution resource constraints and data acquisition constraints. Execution resource constraints may include the research center's testing capacity, visit frequency capacity, subject recruitment scope, sample processing conditions, and personnel operation requirements; data acquisition constraints may include the availability of examination items, acquisition window width, data field completeness, consistency of acquisition methods, and availability for subsequent statistical analysis. The system matches these constraints with the trial element nodes in the candidate study pathways, calculating evaluation characteristics such as recruitment accessibility, visit feasibility, acquisition completeness, and center suitability. Using the aforementioned trial as an example, if the candidate study pathway requires all subjects to complete specific laboratory tests weekly, but the target research center only supports monthly centralized testing, the system will mark the visit plan node and data acquisition node as execution mismatch and generate corresponding adjustment flags in the execution evaluation results.
[0055] Finally, the system integrates the element consistency verification results, compliance matching results, and execution evaluation results to generate a solution reasoning result. During the integration, the system first categorizes conflicts into blocking, corrective, and suggestive types based on their severity. Then, it calculates the comprehensive evaluation value of candidate solution paths by combining evidence fit, rule priority, and execution evaluation weight. For blocking conflicts, the system marks candidate solution paths as unfittable; for corrective conflicts, the system generates a set of affected experimental element nodes and local correction directions; for suggestive conflicts, the system retains candidate solution paths and adds review prompts. The resulting solution reasoning result provides a basis for target solution path selection, local regeneration, and structured solution assembly.
[0056] Through the above processing, the system can perform joint reasoning on the medical logic, compliance rules, and execution conditions in candidate protocol paths before generating the clinical trial protocol text. This implementation method can improve the consistency of protocol elements, reduce the risk of rule conflicts and execution mismatch, and improve the accuracy of subsequent target protocol path determination and clinical trial protocol object assembly.
[0057] In some embodiments, determining the target protocol path based on the protocol reasoning results and converting the target protocol path into a structured clinical trial protocol object includes: Based on the consistency verification results, compliance matching results, and execution evaluation results of the scheme reasoning results, the candidate scheme paths are prioritized and the target scheme path that meets the preset assembly conditions is determined. Based on the experimental element nodes and the reference relationships between nodes in the target scheme path, generate the scheme chapter structure, element association structure, and evidence tracing structure; The structure of the protocol chapters, the structure of the elements, and the structure of the evidence traceability are encapsulated into objects to generate structured clinical trial protocol objects.
[0058] Specifically, after obtaining the reasoning results, the system first aggregates the verification status of each candidate path, converting the element consistency verification results, compliance matching results, and execution evaluation results into a unified path evaluation representation. The path evaluation representation includes consistency status, rule conflict status, execution adaptation status, evidence completeness, and correctable scope. The system performs tiered screening of candidate paths based on preset assembly conditions. These conditions may include the absence of blocking rule conflicts, the existence of evidentiary basis for key experimental elements, the temporal correspondence between the main endpoints and the visit plan, and the availability of execution conditions for core data collection items. For candidate paths with corrective conflicts, the system first calls the set of affected experimental element nodes in the reasoning results for local correction, and then recalculates the path evaluation representation. For candidate paths that still do not meet the assembly conditions, the system designates them as alternative paths or eliminates them.
[0059] During the prioritization process, the system generates a comprehensive ranking value based on evidence fit, element consistency, compliance matching degree, and execution evaluation results, and determines the target protocol path. Taking a phase II randomized controlled trial of hypoglycemic drugs for adult type 2 diabetes as an example, if the first candidate path uses changes in glycated hemoglobin as the primary endpoint, includes visits at weeks 12 and 24 after administration, and also includes monitoring of hypoglycemic events and liver and kidney function, and has few rule conflicts and the center's testing capacity can support it, then the system ranks this candidate path with a higher priority. If the second candidate path has a higher level of evidence, but the number of visits is too high and some testing items cannot be reliably completed at the target center, then the system lowers the priority of this candidate path. After the target protocol path is determined, the system locks the trial element nodes, evidence citation relationships, and rule constraint relationships in the target protocol path as the data foundation for structured assembly.
[0060] Subsequently, the system generates the protocol chapter structure according to the trial element nodes and their inter-node references in the target protocol path. This chapter structure represents the chapter hierarchy and content placeholders of the clinical trial protocol text. The system binds chapter nodes such as research background, research objectives, research design, population criteria, intervention protocol, endpoint indicators, visitation plan, safety monitoring, statistical analysis, and data management to their corresponding trial element nodes. The system also generates an element association structure to represent the dependencies and linkages between different trial elements. For example, a data collection dependency is established between the primary endpoint node and the visitation plan node; an analysis dependency is established between the primary endpoint node and the statistical analysis node; a risk constraint relationship is established between the intervention path node and the safety monitoring node; and a limiting relationship is established between the population criteria node and the inclusion / exclusion criteria node. This element association structure ensures logical consistency during subsequent chapter content generation, protocol modification, and partial regeneration.
[0061] Furthermore, the system generates an evidence traceability structure based on the evidence citation relationships retained in the target protocol path. This traceability structure records the evidence nodes, rule nodes, source versions, scope of application, and expiration status corresponding to each trial element node. For example, the primary endpoint node can be associated with evidence nodes from previous similar trials and guideline evidence nodes; the safety monitoring node can be associated with drug safety rule nodes; and the visit plan node can be associated with observation window evidence nodes and central execution constraint nodes. The system objectifies and encapsulates the protocol chapter structure, element association structure, and evidence traceability structure to form a structured clinical trial protocol object. This object can be uniformly managed using node identifiers, chapter identifiers, citation identifiers, version identifiers, and status identifiers, and supports subsequent clinical trial protocol text generation, protocol element mapping, evidence review, and version change recording.
[0062] Through the above processing, the system can convert candidate protocol path screening results into structured clinical trial protocol objects, establishing stable data associations between protocol chapters, trial elements, and evidence sources. This implementation method can improve the accuracy of target protocol path selection, enhance the consistency of structured assembly of clinical trial protocols, and improve the reliability of subsequent protocol text generation, evidence tracing, and execution object mapping.
[0063] In some embodiments, generating clinical trial protocol text and corresponding protocol element mapping results based on structured clinical trial protocol objects includes: Based on the preset protocol document skeleton, the content is filled and paragraphs are organized in the protocol chapter structure of the structured clinical trial protocol object to generate the clinical trial protocol text; Based on the element association structure, the test element nodes are mapped to the subsequent test execution objects to generate the scheme element mapping results; Based on the evidence tracing structure, configure evidence citation identifiers and version identifiers for clinical trial protocol texts and protocol element mapping results.
[0064] Specifically, after generating a structured clinical trial protocol object, the system first reads the protocol's chapter structure within the object and matches it with a pre-defined protocol document skeleton. The pre-defined protocol document skeleton defines the chapter hierarchy, paragraph order, content boundaries, and element placement positions of the clinical trial protocol text. The system determines the assembly order of chapters such as research background, research objectives, research design, population criteria, intervention protocol, endpoint indicators, visitation plan, safety monitoring, statistical analysis, and data management based on chapter identifiers, and extracts the corresponding structured content based on the trial element nodes bound to each chapter. During content filling, the system does not directly generate the clinical trial protocol freely; instead, it uses the node content, node relationships, and citation identifiers in the structured clinical trial protocol object as the basis for generation, converting the trial elements in the target protocol path into paragraph content that conforms to the expression habits of protocol documents.
[0065] During paragraph organization, the system determines the connection between paragraphs according to the hierarchical relationship of chapter nodes and the element association structure. For the research design chapter, the system extracts elements such as research stage, research type, control method, grouping method, and observation period from the target plan path, and generates continuous paragraphs according to the design logic in the plan skeleton; for the population criteria chapter, the system organizes the population condition nodes, inclusion condition nodes, and exclusion condition nodes according to the defined scope and rule reference relationship; for the endpoint indicator chapter, the system generates indicator descriptions based on the dependency relationship between primary endpoints, secondary endpoints, and observation windows; for the visit plan chapter, the system maps visit event nodes to data collection project nodes, so that each key observation time point can be associated with the corresponding collection content.
[0066] For example, for a phase II randomized controlled trial of hypoglycemic drugs for type 2 diabetes in adults, the system can fill in details such as randomization, parallel grouping, and pre-set observation period in the study design section; fill in adult type 2 diabetes patients, previous treatment status, basal metabolic index range, and contraindication restrictions in the population criteria section; fill in changes in glycated hemoglobin, changes in fasting blood glucose, and hypoglycemic event records in the endpoint indicators section; and fill in the data collection arrangements corresponding to the screening period, baseline period, post-dose observation period, and safety follow-up period in the visitation plan section. All of the above text content is generated driven by the protocol section structure and trial element nodes, and each paragraph retains the mapping relationship with the corresponding node, facilitating subsequent review, modification, and partial regeneration.
[0067] Furthermore, the system generates protocol element mapping results based on the element association structure. This structure records the conversion relationships between trial element nodes and subsequent trial execution objects. Based on this, the system maps population standard nodes to subject selection criteria, visit event nodes to visit plan objects, data collection item nodes to case report form fields, safety monitoring nodes to risk monitoring rules, and statistical analysis nodes to statistical analysis entries. During the mapping process, the system generates a unified mapping identifier based on node type, data attributes, collection time, and associated chapters to avoid semantic discrepancies for the same trial element across protocol text, case report forms, and statistical analysis plans. For example, the glycated hemoglobin change node is mapped to both the primary endpoint paragraph and the corresponding case report form field and statistical analysis entry; the week 12 collection node in the visit plan is simultaneously associated with the testing item, collection window, and data source for analysis.
[0068] The system also configures evidence citation identifiers and version identifiers for clinical trial protocol text and protocol element mapping results based on the evidence traceability structure. Evidence citation identifiers record which evidence node or rule node a protocol text paragraph, trial element node, or execution object originates from; version identifiers record the source data version, protocol object version, and mapping result version. When guidelines are updated, drug safety rules are adjusted, or experts modify protocol elements, the system can locate the affected chapters / paragraphs and execution objects based on the evidence citation identifiers and generate change records based on the version identifiers. For example, if the source version of the hypoglycemia monitoring rule is updated, the system can locate the content related to hypoglycemia events in the safety monitoring chapter, visit collection items, and risk monitoring rules, and trigger the verification and regeneration of the corresponding nodes.
[0069] Through the above processing, the system can simultaneously generate clinical trial protocol text and protocol element mapping results based on structured clinical trial protocol objects, forming a consistent structured association between protocol text content, implementation targets, and evidence sources. This implementation method can improve the standardization of protocol text generation, reduce semantic deviations between protocol content and subsequent implementation targets, and improve the efficiency of evidence tracing, version management, and local correction.
[0070] In some embodiments, a feedback update module is also included, which is used to obtain review feedback data on clinical trial protocol text, protocol element mapping results, and subsequent trial execution objects; The review feedback data is parsed into feedback event nodes, and the feedback event nodes are written back to the structured clinical trial protocol object and the multi-source medical knowledge graph according to the reference relationship between evidence citation identifiers, version identifiers and trial element nodes. Based on the deviation type and impact range corresponding to the feedback event node, update the weights of the evidence node, the weights of the rule node, and the candidate solution path generation parameters, and generate a set of affected solution nodes for local regeneration.
[0071] Specifically, after generating the clinical trial protocol text and protocol element mapping results, the feedback update module continuously receives review feedback data from medical experts, ethics reviewers, registration applicants, data managers, and research center executives. Review feedback data can come from online annotations, structured review forms, version comparison records, execution system feedback records, and rule verification reports. The system first identifies the source and object of the review feedback data, determining the corresponding paragraph position in the clinical trial protocol text, the execution object in the protocol element mapping results, and the specific configuration items in the subsequent trial execution object. For example, in a phase II randomized controlled trial of hypoglycemic drugs for adult type 2 diabetes, experts believe the observation period for the primary endpoint is too short, ethics reviewers believe some exclusion criteria do not cover high-risk drug users, and the research center reports a high execution burden for continuous liver and kidney function testing at weeks 4, 8, and 12. The system identifies these as endpoint window feedback, safety rule feedback, and visit execution feedback, respectively.
[0072] After obtaining review feedback data, the system converts natural language annotations and structured review items into feedback event nodes using a feedback parsing model. Feedback event nodes include the feedback source, feedback recipient, deviation type, scope of impact, processing priority, associated version, and suggested correction direction. Deviation types can include insufficient evidence deviation, rule conflict deviation, chapter consistency deviation, execution adaptation deviation, and mapping missing deviation. Based on the evidence citation identifier and version identifier corresponding to the feedback content, the system locates the protocol chapter structure, element association structure, and evidence traceability structure within the structured clinical trial protocol object, and writes the feedback event nodes back to the corresponding nodes based on the citation relationships between trial element nodes. For example, for feedback regarding a short observation period for the primary endpoint, the system associates the feedback event node with the primary endpoint node, visit plan node, statistical analysis node, and corresponding evidence node; for feedback regarding exclusion criteria not covering high-risk drug users, the system associates the feedback event node with the population standard node, safety monitoring node, and corresponding rule node.
[0073] Furthermore, the system synchronously writes feedback event nodes back to the multi-source medical knowledge graph. During the write-back, the system does not directly overwrite the original evidence nodes and rule nodes, but instead adds feedback relationship edges, feedback confidence, and version status identifiers to the original nodes. For feedback of insufficient evidence confirmed by experts, the system reduces the call weight of the corresponding evidence node in the same research scenario and increases the candidate priority of adjacent high-level evidence nodes; for rule conflict feedback confirmed by ethical review, the system increases the constraint weight of the corresponding rule node and sets the rule node as a strong constraint node in the subsequent candidate solution path generation process; for execution adaptation feedback returned by the research center's execution system, the system updates execution parameters related to visit rate, acquisition window, detection resources, and subject burden.
[0074] When updating generation parameters, the system generates a set of affected solution nodes based on the impact scope of feedback event nodes. This set includes not only nodes directly affected by feedback but also related nodes whose impact propagates through the element association structure. For example, if the primary endpoint observation period is adjusted from 12 weeks to 24 weeks, the affected solution node set includes endpoint indicator nodes, visit plan nodes, data collection project nodes, statistical analysis nodes, and sample size parameter nodes. If the hypoglycemia monitoring rule is upgraded to a strong constraint rule, the affected solution node set includes safety monitoring nodes, visit event nodes, case report form field nodes, and risk reporting rule nodes. The system performs partial regeneration based on the affected solution node set, updating only the affected chapters, mapping objects, and rule configurations; unaffected solution chapters and execution objects retain their original versions.
[0075] During the partial regeneration process, the system re-invokes the updated evidence nodes and rule nodes in the multi-source medical knowledge graph, recalculates the optional content of the corresponding trial element nodes based on the candidate protocol path generation parameters, and performs consistency verification, compliance rule matching, and execution constraint evaluation on the regenerated content. If the regenerated content passes the verification, the system generates a new version of the structured clinical trial protocol object and simultaneously updates the clinical trial protocol text, protocol element mapping results, and evidence traceability structure; if conflicts still exist, the system retains the unclosed-loop state of the feedback event node and generates a correction prompt awaiting manual confirmation.
[0076] Through the above processing, the system can transform review feedback and execution feedback into computable feedback event nodes, and write these feedback event nodes back to the structured clinical trial protocol object and the multi-source medical knowledge graph. This implementation method can improve the positioning accuracy of protocol updates, reduce content offset caused by full regeneration, enhance the adaptive correction capabilities of evidence nodes, rule nodes, and candidate protocol path generation parameters, and improve the efficiency of clinical trial protocol version iteration and local correction.
[0077] The above embodiments have described in detail the specific components and functions of the clinical trial protocol generation system based on multi-source medical knowledge graph of this application. The implementation process of the clinical trial protocol generation method based on multi-source medical knowledge graph of this application will be described in detail below with reference to specific embodiments. Figure 2 This is a flowchart illustrating the clinical trial protocol generation method based on a multi-source medical knowledge graph provided in this application embodiment, such as... Figure 2 As shown, the method may specifically include the following steps: S201: Acquire multi-source medical knowledge data, perform terminology normalization, evidence annotation, and rule structuring on the multi-source medical knowledge data, and construct a multi-source medical knowledge graph; S202 analyzes clinical trial design requirements and generates trial objective vectors, research constraint vectors, and protocol skeleton nodes; S203: Based on the experimental target vector, research constraint vector, and scheme skeleton nodes, extract the target evidence subgraph from the multi-source medical knowledge graph, and generate candidate scheme paths based on the target evidence subgraph; S204, perform constraint reasoning on the test element nodes in the candidate solution path, generate solution reasoning results, and determine the target solution path based on the solution reasoning results; S205, convert the target protocol path into a structured clinical trial protocol object, and generate the clinical trial protocol text and corresponding protocol element mapping results based on the structured clinical trial protocol object.
[0078] Specifically, in S201, the system first acquires multi-source medical knowledge data from disease diagnosis and treatment guidelines, drug instructions, previous clinical trial registration information, medical literature, real-world aggregated data, medical terminology standards, and regulatory rule bases. It then performs source analysis, version identification, and content segmentation on the data from different sources. The system generates the required knowledge granularity according to the clinical trial protocol, breaking down the original medical knowledge into standardized knowledge units, and configuring each standardized knowledge unit with source version, scope of application, level of evidence, timeliness, and rule attributes.
[0079] Subsequently, based on the medical terminology ontology and semantic embedding model, the system performs entity recognition, concept alignment, and synonym merging on diseases, intervention measures, examination items, endpoint indicators, safety events, and rule conditions in standardized knowledge units, generating medical entity nodes, evidence nodes, rule nodes, and test element nodes. Based on the semantic dependencies, evidence citation relationships, and rule constraint relationships between nodes, graph relationship edges are established to obtain a multi-source medical knowledge graph.
[0080] In S202, the system performs intent recognition and semantic slot extraction on the clinical trial design requirements input by the user, obtaining core design elements such as target disease, target population, research phase, intervention method, research objective, control method, and observation target. The system semantically aligns these core design elements with trial element nodes in a multi-source medical knowledge graph, and performs fusion encoding based on graph adjacency relationships and textual semantic features to generate a trial target vector. The system also generates a research constraint vector based on the applicable scope, evidence constraints, and rule constraints corresponding to the core design elements.
[0081] Then, the system determines the chapter hierarchy, element placement relationships, and node dependencies of the clinical trial protocol based on the trial objective vector and research constraint vector, generating protocol skeleton nodes. For example, for a phase II randomized controlled trial of hypoglycemic drugs for adults with type 2 diabetes, the system can generate protocol skeleton nodes including study design, population criteria, intervention protocol, endpoint indicators, visitation plan, safety monitoring, and statistical analysis.
[0082] In S203, the system uses the experimental target vector, research constraint vector, and scheme skeleton nodes as search criteria to perform semantic recall and relation expansion in a multi-source medical knowledge graph, obtaining a set of candidate evidence nodes matching the current experimental design task. The system further sorts the candidate evidence node set and prunes conflicts based on evidence fit, rule constraint relationships, and node association strength, generating a target evidence subgraph. The target evidence subgraph retains medical entity nodes, evidence nodes, rule nodes, and experimental element nodes related to the current scheme generation task. The system then decomposes the evidence paths in the target evidence subgraph into composable sequences of experimental elements according to the element occupancy relationships of the scheme skeleton nodes, and establishes corresponding evidence reference relationships for each experimental element, forming candidate scheme paths.
[0083] In S204, the system performs constraint reasoning on the experimental element nodes in the candidate protocol paths. First, based on the semantic dependencies and temporal constraints between the experimental element nodes, the system verifies the consistency between research objectives and endpoint indicators, target population and inclusion / exclusion criteria, endpoint indicators and visit plans, and endpoint data types and statistical analysis methods, generating element consistency verification results. Subsequently, based on rule nodes in a multi-source medical knowledge graph, the system performs compliance rule matching and conflict identification on the candidate protocol paths, generating compliance matching results. The system also evaluates the feasibility of the candidate protocol paths based on execution resource constraints such as research center testing capabilities, visit frequency, data acquisition window, availability of acquisition items, and subject burden, generating execution evaluation results. The system integrates the element consistency verification results, compliance matching results, and execution evaluation results to obtain the protocol reasoning results, and prioritizes the candidate protocol paths according to these results to determine the target protocol path.
[0084] In S205, the system generates a protocol chapter structure, an element association structure, and an evidence traceability structure based on the trial element nodes and their inter-node references in the target protocol path. These structures are then objectified and encapsulated to obtain a structured clinical trial protocol object. Subsequently, the system populates the protocol chapter structure within the structured clinical trial protocol object with content and organizes paragraphs according to a pre-defined protocol document skeleton, generating the clinical trial protocol text. Simultaneously, based on the element association structure, the system maps population criteria, visit events, data collection items, safety monitoring items, and statistical analysis entries to subsequent trial execution objects, generating protocol element mapping results. Furthermore, based on the evidence traceability structure, the system configures evidence citation identifiers and version identifiers for the clinical trial protocol text and the protocol element mapping results.
[0085] Through the above-described process, this application can transform dispersed, multi-source medical knowledge into a computable and traceable knowledge graph, and simultaneously complete evidence association, element linkage, rule verification, and execution evaluation during the clinical trial protocol generation process, thereby improving the accuracy of knowledge association, consistency of protocol elements, and efficiency of compliance execution verification in protocol generation.
[0086] It should be understood that the sequence number of each step in the above method embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0087] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although the technical solutions of this application have been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A clinical trial protocol generation system based on a multi-source medical knowledge graph, characterized in that, include: The knowledge graph construction module is used to acquire multi-source medical knowledge data, perform terminology normalization, evidence annotation and rule structuring on the multi-source medical knowledge data, and construct a multi-source medical knowledge graph containing medical entity nodes, evidence nodes, rule nodes and experimental element nodes. The trial intent parsing module is used to parse clinical trial design requirements and generate trial objective vectors, research constraint vectors, and protocol skeleton nodes; The evidence subgraph generation module is used to extract target evidence subgraphs from the multi-source medical knowledge graph based on the experimental target vector, research constraint vector, and scheme skeleton nodes, and to generate candidate scheme paths based on the target evidence subgraphs. The scheme constraint reasoning module is used to perform consistency verification, compliance rule matching, and execution constraint evaluation on the test element nodes in the candidate scheme path, and generate scheme reasoning results. The protocol assembly module is used to determine the target protocol path based on the protocol reasoning result, convert the target protocol path into a structured clinical trial protocol object, and generate clinical trial protocol text and corresponding protocol element mapping results based on the structured clinical trial protocol object.
2. The system according to claim 1, characterized in that, The process of performing terminology normalization, evidence annotation, and rule structuring on the multi-source medical knowledge data to construct a multi-source medical knowledge graph containing medical entity nodes, evidence nodes, rule nodes, and experimental element nodes includes: The multi-source medical knowledge data is analyzed for its source and segmented for content to generate standardized knowledge units that carry the source version, scope of application, and time status. Based on the medical terminology ontology and semantic embedding model, entity recognition, concept alignment and synonym merging are performed on the medical concepts in the standardized knowledge units to generate medical entity nodes. The standardized knowledge units are labeled according to the evidence level, research scenario, and rule attributes to generate evidence nodes and rule nodes. Structured elements related to clinical trial design are extracted from the standardized knowledge units to generate trial element nodes. Graph relationship edges are established based on the semantic dependencies between nodes to obtain the multi-source medical knowledge graph.
3. The system according to claim 1, characterized in that, The process of analyzing clinical trial design requirements generates trial objective vectors, research constraint vectors, and protocol skeleton nodes, including: The core design elements are obtained by performing intent identification and semantic slot extraction on the clinical trial design requirements. The core design elements are semantically aligned with the experimental element nodes in the multi-source medical knowledge graph, and the alignment results are vector-encoded to generate the experimental target vector. The research constraint vector is generated based on the applicable scope, evidence constraints, and rule constraints corresponding to the core design elements. Based on the experimental target vector and the research constraint vector, the chapter hierarchy and element placement relationships of the clinical trial protocol are determined, and the protocol skeleton nodes are generated.
4. The system according to claim 1, characterized in that, The step of extracting a target evidence subgraph from the multi-source medical knowledge graph and generating candidate solution paths based on the target evidence subgraph includes: Using the experimental target vector and the research constraint vector as retrieval conditions, semantic recall and relation expansion are performed on the multi-source medical knowledge graph to obtain a set of candidate evidence nodes that match the skeleton nodes of the scheme; Based on evidence fit, rule constraint relationship and node association strength, the candidate evidence node set is sorted and conflict pruning is performed to generate the target evidence subgraph; Based on the element occupancy relationship of the skeleton nodes of the scheme, the evidence path in the target evidence subgraph is converted into a combinable sequence of experimental elements, and candidate scheme paths carrying evidence reference relationships are generated.
5. The system according to claim 4, characterized in that, The step of converting the evidence paths in the target evidence subgraph into a composable sequence of experimental elements and generating candidate solution paths carrying evidence citation relationships includes: Based on the element type and hierarchical relationship corresponding to the skeleton node of the scheme, the evidence path in the target evidence subgraph is decomposed to generate evidence fragments that match the different element positions. Based on the semantic dependencies and rule constraints corresponding to the evidence fragments, the combination order and reference boundaries between the experimental elements are determined, and the sequence of experimental elements is generated. The test element sequence is linked to the corresponding evidence node through a reference mapping, and the candidate solution path is formed by combining them according to the preset solution path generation rules.
6. The system according to claim 1, characterized in that, The process of performing consistency verification, compliance rule matching, and execution constraint evaluation on the test element nodes in the candidate solution path to generate solution reasoning results includes: Based on the semantic dependencies and temporal constraints between the test element nodes in the candidate solution path, generate element consistency verification results; Based on the rule nodes in the multi-source medical knowledge graph, the candidate solution paths are matched by rules and conflict identified to generate compliant matching results. Based on the execution resource constraints and data acquisition constraints corresponding to the candidate solution paths, the feasibility of the candidate solution paths is evaluated, and execution evaluation results are generated. The consistency verification results, compliance matching results, and execution evaluation results of the aforementioned elements are combined to generate the reasoning results of the proposed solution.
7. The system according to claim 1, characterized in that, The step of determining the target protocol path based on the reasoning results of the protocol and converting the target protocol path into a structured clinical trial protocol object includes: Based on the consistency verification results, compliance matching results, and execution evaluation results of the scheme reasoning results, the candidate scheme paths are prioritized and the target scheme path that meets the preset assembly conditions is determined. Based on the experimental element nodes and the reference relationships between nodes in the target scheme path, generate the scheme chapter structure, element association structure, and evidence tracing structure; The structured clinical trial protocol object is generated by objectifying and encapsulating the chapter structure, element association structure, and evidence tracing structure of the protocol.
8. The system according to claim 7, characterized in that, The process of generating clinical trial protocol text and corresponding protocol element mapping results based on the structured clinical trial protocol object includes: According to the preset scheme document skeleton, the content is filled and paragraphs are organized in the scheme chapter structure of the structured clinical trial scheme object to generate the clinical trial scheme text; Based on the aforementioned element association structure, the experimental element nodes are mapped to subsequent experimental execution objects to generate the scheme element mapping result; Based on the evidence tracing structure, configure evidence citation identifiers and version identifiers for the clinical trial protocol text and protocol element mapping results.
9. The system according to claim 8, characterized in that, It also includes a feedback update module, used to obtain review feedback data on the clinical trial protocol text, the protocol element mapping results, and subsequent trial execution objects; The review feedback data is parsed into feedback event nodes, and the feedback event nodes are written back to the structured clinical trial protocol object and the multi-source medical knowledge graph according to the reference relationship between the evidence citation identifier, version identifier and trial element node; Based on the deviation type and impact range corresponding to the feedback event node, update the evidence node weight, rule node weight, and candidate solution path generation parameters, and generate a set of affected solution nodes for local regeneration.
10. A method for generating clinical trial protocols based on a multi-source medical knowledge graph, using the system described in any one of claims 1 to 9, characterized in that, include: Acquire multi-source medical knowledge data, perform terminology normalization, evidence annotation, and rule structuring on the multi-source medical knowledge data, and construct a multi-source medical knowledge graph; Analyze clinical trial design requirements to generate trial objective vectors, research constraint vectors, and protocol skeleton nodes; Based on the experimental target vector, research constraint vector, and scheme skeleton nodes, a target evidence subgraph is extracted from the multi-source medical knowledge graph, and candidate scheme paths are generated based on the target evidence subgraph. Constraint reasoning is performed on the test element nodes in the candidate solution path to generate solution reasoning results, and the target solution path is determined based on the solution reasoning results; The target protocol path is converted into a structured clinical trial protocol object, and the clinical trial protocol text and corresponding protocol element mapping results are generated based on the structured clinical trial protocol object.