Multi-source knowledge graph construction method and device, electronic equipment and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PENG CHENG LAB
- Filing Date
- 2026-06-02
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]但相关技术中,人体组织结构的电生理模型实验研究中,文献记载与对应的代码记载之间存在数据量多且记载可能存在混乱的情况或文献描述的实验协议与代码记载不完整,导致基于研究记载的内容难以复现
[0007]The multi-source knowledge graph construction method provided in this disclosure acquires a collection of literature related to human tissue electrophysiological model simulation experiments and corresponding open-source code. For any target literature in the collection, an executable reproducible experimental package is generated based on the specification data (including experimental protocol specification data and experimental indicator specification data), the code capability list of the open-source code, and the open-source code. Based on the execution results of the reproducible experimental package, the reproducible experimental results are obtained. The literature knowledge graph subgraph generated based on the target literature, the code knowledge graph subgraph generated based on the parsing results of the corresponding open-source code, and the reproducible knowledge graph subgraph generated based on the reproducible experimental results are written back into the target knowledge graph. For each target literature in the collection, corresponding literature knowledge graph subgraphs, code knowledge graph subgraphs, and reproducible knowledge graph subgraphs are generated and written back into the target knowledge graph. This allows for unified, intelligent, and structured management of experimental research content with a unified human tissue structure, facilitating the rapid retrieval of valuable experimental research content.
Smart Images

Figure CN122529032A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data analysis technology, specifically to a method, apparatus, electronic device, storage medium, and program product for constructing a multi-source knowledge graph. Background Technology
[0002] Multiscale electrophysiological modeling experiments on human tissue structures (such as the heart) can aid in the study of the operational mechanisms of related human tissue structures and provide data support for related research. Research on multiscale electrophysiological modeling simulation experiments of human tissue structures is presented in various forms, such as academic papers, journal articles, and open-source code used in the simulations.
[0003] However, in related technologies, in experimental studies using electrophysiological models of human tissue structures, there are often issues with the amount of data in the literature and the corresponding code records, such as large amounts of data and potential confusion in the records, or incomplete experimental protocols and code records described in the literature, making it difficult to reproduce the content based on the research records. Faced with a large amount of experimental research content, there is an urgent need for intelligent and structured management of the experimental research content. Summary of the Invention In view of this, this disclosure provides a method, apparatus, electronic device, storage medium and program product for constructing a multi-source knowledge graph, so as to realize intelligent and structured management of experimental research content.
[0004] In a first aspect, this disclosure provides a method for constructing a multi-source knowledge graph. The method includes: acquiring a collection of literature related to simulation experiments of human tissue electrophysiological models and corresponding open-source code; for any target literature in the collection, generating an executable reproducible experimental package based on the standard data corresponding to the target literature, the code capability list of the open-source code corresponding to the target literature, and the open-source code corresponding to the target literature, wherein the standard data includes experimental protocol standard data and experimental indicator standard data; obtaining reproducible experimental results based on the execution results of the reproducible experimental package; and writing back the literature knowledge graph subgraph generated based on the target literature, the code knowledge graph subgraph generated based on the parsing results of the corresponding open-source code, and the reproducible knowledge graph subgraph generated based on the reproducible experimental results to the target knowledge graph. Secondly, this disclosure provides a multi-source knowledge graph construction device, the device comprising: a first acquisition module, used to acquire a collection of literature related to human tissue electrophysiological model simulation experiments and corresponding open-source code; a first generation module, used to generate an executable reproducible experimental package for any target literature in the collection, based on the standard data corresponding to the target literature, the code capability list of the open-source code corresponding to the target literature, and the open-source code corresponding to the target literature, wherein the standard data includes experimental protocol standard data and experimental indicator standard data; a second acquisition module, used to obtain reproducible experimental results based on the execution results of the reproducible experimental package; and a write-back module, used to write back the literature knowledge graph subgraph generated based on the target literature, the code knowledge graph subgraph generated based on the parsing results of the corresponding open-source code, and the reproducible knowledge graph subgraph generated based on the reproducible experimental results to the target knowledge graph. Thirdly, this disclosure provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the multi-source knowledge graph construction method described in the first aspect or any corresponding embodiment.
[0005] Fourthly, this disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to execute the multi-source knowledge graph construction method described in the first aspect or any corresponding embodiment thereof.
[0006] Fifthly, this disclosure provides a computer program product, including computer instructions for causing a computer to execute the multi-source knowledge graph construction method described in the first aspect or any corresponding embodiment thereof.
[0007] The multi-source knowledge graph construction method provided in this disclosure acquires a collection of literature related to human tissue electrophysiological model simulation experiments and corresponding open-source code. For any target literature in the collection, an executable reproducible experimental package is generated based on the specification data (including experimental protocol specification data and experimental indicator specification data), the code capability list of the open-source code, and the open-source code. Based on the execution results of the reproducible experimental package, the reproducible experimental results are obtained. The literature knowledge graph subgraph generated based on the target literature, the code knowledge graph subgraph generated based on the parsing results of the corresponding open-source code, and the reproducible knowledge graph subgraph generated based on the reproducible experimental results are written back into the target knowledge graph. For each target literature in the collection, corresponding literature knowledge graph subgraphs, code knowledge graph subgraphs, and reproducible knowledge graph subgraphs are generated and written back into the target knowledge graph. This allows for unified, intelligent, and structured management of experimental research content with a unified human tissue structure, facilitating the rapid retrieval of valuable experimental research content. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a schematic diagram corresponding to the multi-source knowledge graph construction method according to the embodiments of this disclosure; Figure 2 This is a structural block diagram of a multi-source knowledge graph construction apparatus according to an embodiment of the present disclosure; Figure 3 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0011] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0012] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0013] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0014] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0015] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0016] According to an embodiment of this disclosure, a method for constructing a multi-source knowledge graph is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0017] This embodiment provides a method for constructing a multi-source knowledge graph, which can be used on any electronic device capable of executing the method, such as a terminal or server. Figure 1 This is a flowchart of a multi-source knowledge graph construction method according to embodiments of this disclosure, such as... Figure 1 As shown, the process includes the following steps: Step 101: Obtain the literature collection and corresponding open-source code related to the human tissue electrophysiological model simulation experiment.
[0018] For example, the human tissue electrophysiological model in this application embodiment takes the cardiac electrophysiological model as an example. This application embodiment does not limit the specific type of cardiac electrophysiological model. The corresponding literature collection (such as papers, journals, etc.) records research content related to cardiac electrophysiological models, which can be cardiac electrophysiological cell models, cardiac electrophysiological tissue models, or cardiac electrophysiological organ models; at the same time, the open source code repository records the model simulation open source code corresponding to the literature. This application embodiment does not limit the number of papers included in the literature collection. The trained model can be used to identify experimental research literature and corresponding open source code in the corresponding research direction from the corresponding public experimental research websites or platforms.
[0019] Step 102: For any target document in the document set, generate an executable reproducible experimental package based on the specification data corresponding to the target document, the code capability list of the open source code corresponding to the target document, and the open source code corresponding to the target document. The specification data includes experimental protocol specification data and experimental indicator specification data.
[0020] For example, for any target document in the acquired literature set, the corresponding standard data can be semantically recognized by a pre-trained natural language model. Based on the semantic recognition results, entity data related to the corresponding model simulation experiment, such as model entities, experimental entities, protocol entities, indicator entities, and conclusion entities recorded in the document, can be extracted. Then, structured processing under the corresponding domain mode can be performed to generate experimental protocol standard data and experimental indicator standard data corresponding to the experiment. This clarifies the experimental operation, modeling process, and simulation step standards, avoids ambiguity in the experimental process, and provides data required for experimental reproduction, such as experimental evaluation parameters, calculation methods, and judgment criteria. This allows for horizontal comparison and verification of experimental results from different models. The trained model is used to identify relevant materials of the open-source code corresponding to the target literature, resulting in a code capability list. This code capability list includes at least the following: a code entry point (characterizing where the code starts running, including the main program and startup script); parameter files (including model configuration, electrophysiological parameters, and initial condition files); a solver (a computational module for solving cardiac electrophysiological equations); output metrics (characterizing the results the code can produce, such as action potential duration, propagation velocity, and amplitude); scale capability classification (the simulation scale supported by the code, such as cellular, tissue, organ, or whole-heart level); and environment signature (characterizing the environment required to run the code, such as operating system, dependent libraries, GPU, and version number). Based on the obtained specification data, code capability list, and open-source code, a complete, directly executable, and reproducible experimental package is generated, integrating specification experimental parameters, running code, dependent environments, configuration files, calibration files, etc.
[0021] Step 103: Based on the execution results of the reproduction experiment package, obtain the reproduction experiment results.
[0022] For example, the results of the reproducible experiment obtained based on the execution results of the reproducible experiment package may include numerical results, such as action potentials, conduction velocities, electrophysiological waveforms, and values of various quantitative indicators; and graphical results, such as potential curves, electrocardiograms, regional excitation time-series visualization images, and operation log results.
[0023] Step 104: Write back the document knowledge graph subgraph generated based on the target document, the code knowledge graph subgraph generated based on the parsing results of the corresponding open source code, and the reproduction knowledge graph subgraph generated based on the reproduction experiment results to the target knowledge graph.
[0024] For example, entities extracted from target literature (such as models, experiments, indicators, and other experiment-related entity data) are used as nodes in the knowledge graph, and the relationships between entities are used as edges to construct a literature knowledge graph subgraph. Similarly, entities identified from open-source code that are related to experiments and the relationships between entities are used to construct a code knowledge graph subgraph, and entity data related to the results of the reproduced experiments obtained from running the reproduced experiment package and the relationships between entities are used to construct a reproduction knowledge graph subgraph. In this embodiment, the target knowledge graph can be a historically constructed knowledge graph of related human tissues. Writing the newly analyzed literature knowledge graph subgraph, code knowledge graph subgraph, and reproduction knowledge graph subgraph back to the historically constructed knowledge graph can achieve comprehensive management of the corresponding research, and the graph construction process can be versioned during the write-back process. The constructed target knowledge graph can help researchers quickly find corresponding literature research results during the research process, and can be visualized for easy searching and viewing of literature and code-related reproduction experiment results, facilitating the rapid selection of valuable reference literature.
[0025] The multi-source knowledge graph construction method provided in this embodiment generates corresponding literature knowledge graph subgraphs, code knowledge graph subgraphs, and reproduction knowledge graph subgraphs for each target document in the literature collection, and writes the generated subgraphs back into the target knowledge graph. This enables unified, intelligent, and structured management of experimental research content with a unified human tissue structure, facilitating the rapid retrieval of experimental research content with reference value.
[0026] In some optional implementations, step 102 includes: extracting multiple types of entity data from the target document, the multiple types of entity data including: simulation model entity data, experimental entity data, protocol entity data, indicator entity data, and experimental conclusion entity data; generating the specification data based on the extracted entity data; parsing the open-source code corresponding to the target document to generate the code capability list; determining the cross-source alignment degree between the target document and the corresponding open-source code based on the specification data and the code capability list; and responding to the generation operation of the executable reproducible package when the cross-source alignment degree meets the preset requirements.
[0027] For example, simulation model entity data may include model type (e.g., ventricular model, atrial model), scale (e.g., single-cell model, one-dimensional myocardial fiber model, two-dimensional myocardial tissue model, three-dimensional whole-heart model, etc.), type (e.g., atrial fibrillation pathology model, myocardial ischemia model, etc.), and model version, etc.; experimental entity data may be used to describe what kind of simulation experiment was conducted, the experimental object, the experimental method, etc.; protocol entity data may include entity data describing how to conduct the experiment, the experimental procedure, experimental parameters, experimental rules, etc.; index entity data may include entity data characterizing experimental measurement indicators (e.g., action potential duration, resting potential, peak potential, conduction velocity, etc.), experimental evaluation parameters, etc. The specific methods for extracting relevant entity data and generating the code capability list can be obtained through pre-trained models, and will not be elaborated here.
[0028] Cross-source alignment can be used to measure the degree of matching between the literature description specifications and the actual capabilities of the code. This can be quantified using matching scores or determined using multiple evaluation metrics. Specifically, determining the cross-source alignment between a target document and its corresponding open-source code based on the acquired specification data and code capability list can involve assessing whether the simulation model described in the document matches the code implementation model; whether their processes, parameters, and constraints correspond; whether the statistical standards for calculation metrics are consistent; and whether the solver, simulation scale, and operating environment are compatible. For each dimension, a corresponding score can be set based on the matching status. The degree of cross-source alignment can be determined based on the total score or a weighted score. When the preset score requirement is met, the corresponding executable reproducible package is generated. By determining the degree of cross-source alignment, discrepancies between the literature description and code functionality can be identified in advance, avoiding invalid reproduction and improving the efficiency of intelligent management of literature collections.
[0029] In some optional implementations, the method further includes: verifying the consistency between the reproduced experimental results and the original experimental results recorded in the target literature, and generating a verification report; when the consistency verification results between the reproduced experimental results and the original experimental results do not meet the requirements, responding to the differential cause analysis operation; recording the attribution tags and attribution basis corresponding to the analyzed differential causes in the verification report and writing the differential causes back into the target knowledge graph.
[0030] For example, the consistency verification between the reproduced experimental results and the original experimental results recorded in the literature can be performed. Specific verification dimensions may include indicator values, waveform characteristics, and whether the experimental conclusions match. Consistency verification can be completed by quantifying the magnitude of the deviation between the two and the size of a preset range, and a verification report is generated. When the consistency verification results do not meet the requirements, a pre-built preset rule base can be used to analyze the causes of differences. This preset rule base pre-records the causal tags corresponding to different differences, such as differences in parameter settings, deviations in algorithm implementation, different solver selections, differences in experimental procedures, differences in environmental configuration, omissions in literature descriptions, etc. After matching the causes of differences, the attribution tags are determined and the attribution basis is recorded in the verification report. The causes of differences are then associated with the subgraph of the reproduced knowledge graph in the target knowledge graph and written back.
[0031] In some optional implementations, the step of responding to the generation operation of the executable reproducible package when the cross-source alignment degree meets the preset requirements includes: recording the specification data, code capability list and corresponding open source code of the target document whose cross-source alignment degree meets the preset requirements; writing the recording results into the target review pool; and responding to the generation operation of the executable reproducible package when the review results meet the requirements.
[0032] For example, by recording the standard data, code capability list, and corresponding open-source code corresponding to the target documents whose cross-source alignment meets the preset requirements, document data with reproducibility value can be selected from the document collection, and the recorded results can be written into the target review pool for manual review or secondary model review. When the review results meet the reproducibility requirements, the generation operation of the executable reproducibility package is responded to. By reviewing the documents and code whose cross-source alignment meets the preset requirements, errors in the first model alignment determination can be avoided, which would affect the reproducibility results and thus the efficiency and credibility of intelligent management.
[0033] In some optional implementations, the method further includes: generating corresponding identification information for the acquired documents and open-source code, the identification information including source identifiers and version identifiers. By generating corresponding source identifiers and version information, and using the corresponding identification information as a node parameter in the corresponding knowledge graph, it can be used to distinguish data sources, and facilitate subsequent source tracing and data classification retrieval while avoiding confusion between old and new versions of data.
[0034] In some optional implementations, the step of extracting multiple types of entity data from the target document includes: performing page layout parsing on the target document; dividing the document into blocks based on the page layout parsing results to obtain multiple document block sets and generating corresponding evidence anchors for each document block; and performing natural language parsing on each document block to obtain the multiple types of entity data.
[0035] For example, natural language models can be used to perform semantic parsing on the multi-level headings of target documents, enabling the document to be segmented based on its headings. By performing natural language parsing on each document, multiple types of entity data can be obtained, avoiding parsing biases that can occur when directly parsing large documents. During the parsing process, evidence anchors can be generated based on information such as page numbers, paragraphs, start and end positions, and block numbers of document blocks in the original text, facilitating subsequent source tracing and entity association in the knowledge graph construction process.
[0036] As a specific implementation of this application, the literature in this embodiment is a paper. The multi-source knowledge graph construction method provided by this application is carried out through the following steps: (S1) Obtain the candidate paper set, the supplementary material set related to the paper for supplementing experimental data, and the open source code repository set related to the paper, and generate the corresponding source identifier and version identifier for each data source; (S2) Analyze and divide the paper into blocks to obtain a set of document blocks, and generate an evidence anchor for each document block. (S3) Extract structured information from the document block set, including at least model entities, experimental entities, protocol entities, indicator entities and conclusion entities, and generate protocol specification (ProtocolSpec) and indicator specification (MetricSpec). (S4) Perform syntax and semantic parsing on the open source code in the open source code repository to generate a code capability manifest, which includes at least the code entry point, parameter file, solver, output metrics, scale capability classification and environment signature; (S5) Calculate the cross-source alignment score based on the ProtocolSpec, MetricSpec and CapabilityManifest, and generate an alignment record (AlignmentRecord); when the alignment confidence is in the candidate interval, write the alignment record into the manual review pool; (S6) Based on the AlignmentRecord, the ProtocolSpec is compiled into an executable reproducible experimental package (ReproPackage), which includes at least environment signature, runtime plan, parameter mapping and result export specifications, etc. (S7) Execute the ReproPackage to obtain the ReproRun experimental run record and output the target results of the paper, such as the original waveform file, index file and run log file, etc. (S8) Perform a programmed consistency calculation on the target results of the paper and the results of the ReproRun experiment to generate a VerificationReport. (S9) When the verification results are inconsistent or unreproducible, the cause of the difference is attributed according to the preset rule tree, and the attribution label and supporting evidence are recorded in the verification report. (S10) Write back the paper knowledge graph subgraph, code knowledge graph subgraph, reproduction knowledge graph subgraph, and the associated ProtocolSpec, MetricSpec, CapabilityManifest, AlignmentRecord, ReproPackage, VerificationReport, DifferenceCause and EvidenceAnchor to the fused knowledge graph in a graph structure, and record the graph construction process in a versioned manner.
[0037] To ensure that the above embodiments can be stably executed by engineering systems, the embodiments of this application adopt a pattern constraint and data contract verification mechanism to solidify key intermediate products into machine-verifiable structured objects. Each object records schema_version (pattern / object structure version number) and content_hash (object content hash, used for caching and traceable writeback). The content_hash is used to identify the object content and supports caching and traceable writeback.
[0038] In this embodiment of the application, the protocol specification is preferably represented in YAML or JSON, and its fields include at least: protocol_id (protocol identifier (used for indexing and referencing)), pacing_cycle_length (pacing cycle length), stimulus_amplitude (stimulus amplitude), stimulus_duration (stimulus duration), temperature, ionic_conditions (ionic conditions / ionic concentration conditions), drug_condition (drug conditions / drug concentration conditions), solver_step (solver step size (time step)), steady_state_beats (number of beats required to reach steady state / number of warm-up beats), output_window (output time window (time range for calculating indicators)), units (field units and unit definitions), defaulting_policy (default / default policy (e.g., disallow default, allow default but require review)), missing_fields (missing field list (only records missing fields, does not complete)), and evidence_anchors (set of evidence anchors (pointing to the source location of papers / supplementary materials / figures, etc.)). Specifically, when a field is not explicitly specified in the paper, the field is left blank and written to missing_fields, without inferring it based on common sense; when leaving a field blank will affect reproducibility, the ProtocolSpec further records defaulting_policy to identify the default policy, such as "disable default" or "allow default but require manual review".
[0039] The MetricSpec, which can be represented in YAML or JSON, defines the set of metrics to be calculated and compared during the reproduction process, along with their tolerance standards. In this embodiment, its fields include at least: metric_spec_id (metric specification identifier (used for indexing and citation)), metrics (metric set (name, unit, calculation method identifier, threshold / error standard, etc.)), units (metric units and unit standard descriptions), tolerance_policy (tolerance / error standard policy (used for consistency judgment)), postprocess_policy (postprocessing strategy (waveform filtering, peak / valley detection, window truncation, etc. rules)), and evidence_anchors (set of evidence anchors (pointing to the source of the metric definition)). The metric set (metrics) must at least include the metric name, unit, calculation method identifier, and threshold standard. When a paper only provides a graphical or textual description without a clear numerical threshold, the metric specification records the missing item and enters the review pool to avoid supplementing it with common sense.
[0040] The CapabilityManifest is preferably represented in JSON format and its fields should include at least the following: repo_id (code repository identifier), commit_id (code commit identifier (version identifier)), model_name (model name / alias), cell_type (cell type), scale_support (supported scale range (e.g., cell, tissue, organ)), executable_level (executable capability level), entrypoint (running entry point (script / module / function)), parameter_files (parameter file list and path), solver (solver name and configuration), required_inputs (required inputs (protocol fields, external data, geometry, stimulus settings, etc.)), expected_outputs (expected outputs (waveforms, metrics, logs, etc.)), environment_signature (environment signature (language version, dependency list, operating system / architecture, etc.)), call_chain_hint (call chain hint (used to quickly locate the location of metrics and equation implementations)), and evidence_anchors (set of evidence anchors (pointing to the location of code files / functions / class definitions, etc.)). Among them, executable_level is used to distinguish the level of executable capability, and preferably can include at least: cell_only (executable only at the single cell scale), tissue_supported (supports tissue scale execution / assembly and simulation), organ_supported (supports organ scale execution / assembly and simulation) and non_executable_reference (represents that it cannot be directly executed and is only used as a reference implementation (e.g., containing only fragments or lacking entry points)).
[0041] In this embodiment of the application, the alignment record may include: paper_experiment_id (paper experiment identifier (unique identifier of the paper-side experiment)), candidate_code_id (candidate code implementation identifier (unique identifier of the code-side function / method / entry point)), score_breakdown (sub-item score (name / equation / protocol / parameter / scale / output / species, etc.)), score_total (total score), thresholds (threshold set (used to distinguish strong match / candidate / mismatch)), alignment_status (alignment status), evidence_anchors (evidence anchor set (used to support the alignment basis)), and manual_review_status (manual review status (whether it has been reviewed, conclusion, etc.)). The alignment_status preferably includes at least: strong_match (strong match (automatically approved and strong connection can be established)), candidate_match (candidate match (enters the manual review pool)) and unmatched (no match (no alignment connection is established)); when it is candidate_match, it is written to the manual review pool and the manual review status (manual_review_status) is recorded.
[0042] A ReproPackage should include at least the following: package_id (package identifier, a unique identifier for the reproducible package)), protocol_id (protocol identifier, referencing ProtocolSpec.protocol_id)), candidate_code_id (candidate code implementation identifier, referencing code-side entry point / function / method)), environment_signature (environment signature, used for environment reproduction and difference localization)), run_plan (run plan, such as commands, inputs / outputs, stopping conditions, resource constraints, etc.)), protocol_mapping (mapping rules from protocol fields to repository parameters), metrics_extractor_spec (metric extractor specification, how to calculate metrics from waveforms / logs)), artifact_layout (artifact directory structure specification, layout of waveforms, metrics, logs, reports, etc.)), cache_key (cache key, used for repeated runs for reuse and cost control)), and evidence_anchors (set of evidence anchors, supporting the run plan / mapping / extraction rules)). The runtime plan is used to declaratively describe the execution commands, input files, output files, and stopping conditions; the cache_key is preferably obtained by combining the commit_id (code commit identifier), protocol_hash (protocol hash), and key parameter hash to support repeated execution and cost control.
[0043] The Verification Report should include at least the following: ok (whether the write-back conclusion is allowed or not), status (verification status), repro_score (reproducibility score, between 0 and 1), metric_diff (details of numerical metric differences), curve_diff (details of waveform differences), method_check (method execution consistency check result), completeness_check (protocol / evidence / product completeness check result), causes (set of cause labels for differences), evidence_anchors (set of evidence anchors supporting differences and attributions), gate_decision (gateway decision (whether write-back of consistent / inconsistent conclusions is allowed)), and review_state (review status (whether manual review has been conducted and the conclusion is given)). Among these, gate_decision indicates whether write-back of consistent / inconsistent conclusions is allowed; when key fields are missing or evidence anchors are incomplete, gate_decision is preferably set to "fail write-back" and the report is entered into the review pool. The possible values for status include: consistent, partially consistent, inconsistent, or unreproducible.
[0044] Evidence anchors are used to trace the conclusions, alignments, and attributions in a graph back to their original carrier locations. Their fields should include at least: anchor_type (anchor type (paper text / figure captions / tables / code files / run logs, etc.)), source_id (source identifier (paper / supplementary materials / repository / files, etc.)), location (location information (page number, paragraph offset, table coordinates, code path and function signature, log line number, etc.)), span (segment range (start and end positions or range description)), content_hash (anchor segment content hash (used for consistency verification)), and created_at (anchor generation time). Preferably, the anchor type should include at least: paper text, figure captions, tables, supplementary materials, code files, function / class definitions, run logs, and result files. The location should preferably use verifiable location information such as page numbers and paragraph offsets, table cell coordinates, graph area coordinates (bbox), code file path and function signature, and log line number range.
[0045] In defining the domain model, this application takes the construction of a multi-scale ontology of cardiac electrophysiology as an example. The entity set E can be denoted as: E={Paper, Section, Figure, Table, CellType, CellModel, Experiment, Protocol, ProtocolSpec, Condition, Metric, MetricSpec, Conclusion, CodeRepo, Commit, CapabilityManifest, File, Class, Function, Parameter, Solver, TissueMethod, OrganMethod, ReproTask, ReproPackage, ReproRun} The `AlignmentRecord` (alignment record (structured object)), `VerificationReport` (verification report (structured object)), `DifferenceCause` (reason for difference), and `EvidenceAnchor` (evidence anchor (structured object)) are used, where the cell type (CellType) includes at least five core entities: VentricularCell, AtrialCell, PurkinjeCell, SANCell, and AVNCell. Preferably, the `Protocol` node is used to associate with or carry the `ProtocolSpec`, establishing a relationship of "Protocol -> has_spec -> ProtocolSpec"; the `Metric` node is used to associate with the `MetricSpec`, establishing a relationship of "Metric -> has_spec -> MetricSpec"; and the `Commit` node is used to associate with the `CapabilityManifest`, establishing a relationship of "Commit -> has_manifest -> CapabilityManifest".AlignmentRecords are used to record the score decomposition, thresholds, and evidence anchors for cross-source alignment. It is preferable to establish a relationship of "Experiment -> aligned_by -> AlignmentRecord" and "AlignmentRecord -> referers_to -> Function / TissueMethod / OrganMethod". Evidence Anchors are used to ensure evidence traceability. It is preferable to connect conclusions, alignment records, and validation reports to EvidenceAnchors via "supported_by".
[0046] In this embodiment of the application, the relation set R can be denoted as: R = {contains (contains files / sections, etc., such as repositories / papers)), describes (descriptions, such as paper description models)), has_cell_type (cell type), has_experiment (experiments), uses_protocol (protocol), has_condition (condition), reports_metric (report metrics), has_conclusion (conclusion), has_spec (specified specification, such as Protocol / Metric -> Spec)), has_manifest (manifested list, such as Commit -> CapabilityManifest)), implements (implementations, such as function / class implementation models)), defined_in (defined in, such as function / class definition file)), calls (calls, such as function call relationships)), reads_parameter (read parameters), uses_solver (uses solver), assembles_tissue (tissue assembly), assembles_organ (organ assembly), aligned_to (aligned to, such as strong / candidate alignment edges)), aligned_by (aligned by alignment records, such as Experiments)} ->AlignmentRecord), refers_to (referencing / pointing to (AlignmentRecord -> function / method, etc.)), reproduced_by (reproduced by the reproduction task), compiled_as (compiled to (ProtocolSpec -> ReproPackage)), executed_as (executed to (ReproPackage -> ReproRun)), validated_by (validated by the validation report), consistent_with (consistent with), inconsistent_with (inconsistent with), attributed_to (attributed to (VerificationReport -> DifferenceCause)), supported_by (supported by the evidence anchor)}; In this mode, the literature knowledge graph subgraph, code knowledge graph subgraph, and reproduction knowledge graph subgraph can form a cross-source closed loop through relationships such as "aligned_to", "compiled_as", "executed_as", "validated_by", and "attributed_to", thereby ensuring that any experimental conclusion can be traced back to the original evidence anchor, code implementation location, and reproduction experimental product along the graph path.
[0047] In this embodiment, the document knowledge graph subgraph takes PDFs, supplementary materials, and figure captions from the paper database as input. The system performs layout parsing and block division on the full text, and indexes the formula and figure areas; then, it generates a document block set D={d1} according to a multi-view format of "title—method—result—figure caption—supplementary material". ,d2 ,…,dn For each document block, a schema-constrained extraction engine is used to extract entities and relations under a strict schema, and evidence anchors are recorded in the extraction results to trace each triple back to its source location. To ensure that the extraction results closely match the cardiac electrophysiology scenario, the extraction template preferably explicitly constrains key attributes such as model name, cell type, experimental conditions, indicators, conclusions, units, version number, and species origin; fields are left empty when not explicitly given in the text to avoid guesswork.
[0048] The templates within the quotation marks below can be directly used as structured extraction templates in the document extraction stage in this application embodiment: "You are a cardiac electrophysiology knowledge graph extractor. Please extract entities, attributes, and relationships based solely on the input text; do not add common-sense guesses."
[0049] <schema>; Entity type: [Paper, CellType, CellModel, Experiment, Protocol, ProtocolSpec,Condition, Metric, MetricSpec, Conclusion, TissueMethod, OrganMethod, EvidenceAnchor]; Relationship type: [describes, has_cell_type, has_experiment, uses_protocol, has_condition, reports_metric, has_conclusion, has_spec, assembles_tissue,assembles_organ, supported_by]; constraint: 1. CellType only allows VentricularCell, AtrialCell, PurkinjeCell, SANCell, AVNCell; 2. The model version, units, frequency, drug concentration, temperature, step size, steady-state conditions, and species origin must be preserved; 3. If the text is not explicitly given, the field should be left blank and cannot be guessed; 4. Each triple must include an evidence_anchor (the fields must include at least anchor_type, source_id, location, and span); < / schema> ; <input> ; {chunk_text}; ; <output_format> ; Output a JSON array, each item containing: subject, subject_type, relation, object, object_type, properties, evidence_anchor, confidence; < / output_format> ".
[0050] In the above extraction template, "subject" represents the subject entity (name or identifier); "subject_type" represents the subject entity type; "relation" represents the relation type; "object" represents the object entity (name or identifier); "object_type" represents the object entity type; "properties" represents the triple attributes (unit, version, value, etc.); "evidence_anchor" represents the evidence anchor (pointing to the source location); and "confidence" represents the confidence level. Based on the extraction results, a typical path in the document knowledge graph subgraph can be represented as follows: Paper ->describes ->CellModel ->has_cell_type ->CellType; Paper ->has_experiment ->Experiment ->uses_protocol ->Protocol; Experiment ->has_condition ->Condition; Experiment ->reports_metric ->Metric; Experiment ->has_conclusion ->Conclusion; Conclusion ->supported_by ->EvidenceAnchor; Protocol ->has_spec ->ProtocolSpec; Metric ->has_spec ->MetricSpec; The Protocol node is associated with at least one ProtocolSpec. In this embodiment, the ProtocolSpec includes at least the following attributes: stimulation period, stimulation amplitude, stimulation duration, temperature, ion concentration, drug concentration, solution step size, steady-state criterion, geometric scale, and output time window. The Metric node is preferably associated with one MetricSpec, which includes at least the following indicators and thresholds: APD90 (action potential duration (90% repolarization)), APA (action potential amplitude), RMP (resting membrane potential), CV (conduction velocity), ERP (effective refractory period), EAD (early afterdepolarization), DAD (delayed afterdepolarization), alternans (alternation phenomena such as action potential / calcium alternation), EAD_flag (flag indicating whether EAD has occurred (Boolean)), alternans_flag (flag indicating whether alternans have occurred (Boolean)), and conduction block. For experimental methods in the figure captions and supplementary materials, evidence should be merged with the Experiment node in the main text, rather than creating new duplicate experiment nodes, to reduce the risk of the same experiment being extracted multiple times.
[0051] The code knowledge graph subgraph targets open-source code repositories, supplementary code packages, and example scripts. The system performs dual parsing at both the syntax and semantic levels on the code side: the syntax layer extracts repositories, directories, files, classes, functions, parameter files, configuration files, and dependent environments; the semantic layer extracts model equations, state variables, ion flow sets, solvers, input interfaces, output indicators, organization assembly logic, and organ-level structure construction logic. Core relationships in the code subgraph can include: CodeRepo -> contains -> File, Function / Class -> defined_in -> File, Function -> calls -> Function, Function -> implements -> CellModel, Function -> reads_parameter -> Parameter, Function -> uses_solver -> Solver, Function -> assembles_tissue -> TissueMethod, and Function -> assembles_organ -> OrganMethod. To enable code knowledge to be used for automatic reproduction, the system preferably generates a CapabilityManifest at the Commit level and writes it to Commit->has_manifest->CapabilityManifest. The CapabilityManifest records fields such as entrypoint, required_inputs, expected_outputs, executable_level, and environment_signature. In this embodiment, the CapabilityManifest can be documented as manifest.json.
[0052] During the code building and reproduction adaptation phases, the system can generate a unified runtime entry point and reproduction experiment scripts through the code adaptation unit. This code adaptation unit preferably adds only adaptation files, without modifying the core equations or original parameter files, and marks the source path and mapping relationship for the new files. An example of a structured instruction template that can be used in the code adaptation phase is as follows: "You are now in the root directory of a cardiac electrophysiology model repository. Please complete the following tasks, but do not change the numerical meaning of the original models:" 1. Identify the scales supported by this repository: cell (single cell), 1D tissue (one-dimensional tissue), 2D tissue (two-dimensional tissue), 3D ventricle (three-dimensional ventricle); 2. Locate the input function, parameter file, solver, and output metrics for each model; 3. Generate a machine-readable manifest.json for each model, including the following fields: model_name, cell_type, scale, entrypoint, parameter_files, solver, required_inputs, outputs, tissue_method, organ_method; 4. If the repository lacks a unified entry point, please add a run_repro.py file as an adaptation layer, which only handles parameter mapping, input assembly, and result export. 5. Generate tests / test_manifest.py and verify that the manifest matches the source code; 6. Output a summary.md file explaining which code implementations can be used to reproduce the experiments in the paper, and which can only be used as a reference for tissue / organ construction.
[0053] constraint: - Do not modify the core equations and original parameter files; - All new files must include the source file path; - If you cannot determine the field, write "null" and explain the reason in summary.md.
[0054] In the above extraction template, "manifest.json" represents the code capability manifest file (the disk-based form of CapabilityManifest); "run_repro.py" represents the unified run entry adaptation script (only performing parameter mapping, input assembly, and result export); "tests / test_manifest.py" represents the test script for verifying the consistency between the manifest and the source code; and "summary.md" represents the adaptation and reproducibility summary description file.
[0055] In this embodiment, during the cross-source evidence alignment stage, the system can match the "CellModel-Experiment-Protocol-Metric" path in the document knowledge graph subgraph with the "CellModel-Function-Parameter-Solver-TissueMethod / OrganMethod" path in the code knowledge graph subgraph and output an alignment record. Let the paper's experimental node be... The code implementation node is Then the cross-source alignment score ( ) can be defined as:
[0056] in, This indicates consistency in model aliases, author names, and years; This indicates the consistency of state variables, ion flow terms, and equation structure; This indicates consistency in stimulation method, frequency, temperature, and drug conditions; This indicates the consistency between the parameter set and the unit diameter; Indicates whether the scale capability classification matches; Indicates whether the code can be output or whether the paper's metrics can be calculated; Indicate whether the species or origin is consistent. The weights are indicated, and the specific values of the weights are not limited in this application embodiment. Those skilled in the art can determine them according to actual needs. Preferably, the calculation of each sub-score uses the comparison and unit normalization of the fields of ProtocolSpec, MetricSpec and CapabilityManifest as input, and the evidence anchor points are written into AlignmentRecord.
[0057] The system writes the AlignmentRecord back to the fused knowledge graph, following the paths: Experiment -> aligned_by -> AlignmentRecord and AlignmentRecord -> referers_to -> Function / TissueMethod / OrganMethod. The AlignmentRecord records score_breakdown, thresholds, and evidence_anchors. If... Then, a strong connection is established between Experiment -> aligned_to -> Function / TissueMethod / OrganMethod and marked as "strong_match"; if If so, a candidate edge is established and entered into the manual review pool, marked as "candidate_match"; if If the value is not specified, it is recorded as "unmatched". This application does not limit the specific threshold value. Through this process, single-cell experiments can be aligned with specific model functions, while tissue-scale or organ-scale experiments can be aligned with TissueMethod and OrganMethod nodes, thus ensuring that the three scales of "cell-tissue-organ" are continuous and traceable in the knowledge graph.
[0058] The automated reproduction and verification phase uses aligned paper experiments and code implementations as input to construct ReproTask task nodes and compiles ProtocolSpec into ReproPackage. The system reads ProtocolSpec and MetricSpec from the literature and information such as entrypoints, required_inputs, solver, and environment_signature from the CapabilityManifest in the code. It then generates or repairs runtime scripts, environment configuration files, and result extraction code to ensure the reproduction experiments adhere to a unified input / output specification. The protocol compilation process preferably includes: field unit normalization, parameter mapping generation, missing field gating, runtime plan generation, and metric extraction specification generation.
[0059] The environment signature includes at least the runtime language version, dependency list, operating system, and architecture information. The system can further generate container description files or environment lock files, and incorporate them as components of the executable reproducible package to reduce the non-reproducibility introduced by environment differences. In this embodiment, the execution artifacts of the executable reproducible package can be organized according to the artifact layout specification, and only URIs (Uniform Resource Identifiers) and hashes are stored in the graph to avoid directly storing large objects in the database. An example of a structured task template that can be used to generate the reproducible package is as follows: Read the following two files: 1. paper_protocol.yaml; 2. manifest.json; Task: Generate an executable script for reproducing experiments for the current repository based on the paper's license agreement.
[0060] Require: 1. Generate repro_runner.py, supporting command-line input of protocol_id; 2. Map the pacing_cycle_length, stimulus_amplitude, stimulus_duration, temperature, ionic_conditions, drug_condition, solver_step, and steady_state_beats in the protocol to repository parameters; 3. If the protocol field name does not match the repository parameter name, please create a mapping and write it to protocol_mapping.json; 4. After running, the output will be raw_trace.csv, metrics.json, and run_log.txt; 5. The metrics.json file must contain at least APD90, APA, RMP, CV, ERP, EAD_flag, and alternatives_flag; 6. Keep the core numerical values of the original model unchanged, and only add adaptation code and result export code; 7. If a protocol cannot be implemented, please write the specific reason in unreproducible_reason.txt.
[0061] In the above extraction template, "paper_protocol.yaml" represents the paper protocol file (a structured representation of the protocol extracted / organized from the paper); "manifest.json" represents the code capability manifest file (the disk-based version of CapabilityManifest); "repro_runner.py" represents the script for running the reproducible experiment (the entry point for the adaptation layer); "protocol_mapping.json" represents the mapping table from protocol fields to repository parameters; "raw_trace.csv" represents the raw waveform / time series output file; "metrics.json" represents the metric output file; "run_log.txt" represents the running log file; and "unreproducible_reason.txt" represents the explanation file for the reasons why the reproducibility is not reproducible.
[0062] In this embodiment, the result verification and discrepancy attribution stages can be performed by a verification attribution unit in the system. This verification attribution unit takes paper_results, metrics.json, raw_trace.csv, run_log.txt, and ReproPackage as input and outputs a VerificationReport. The verification results can be composed of "programmatic consistency calculation" and "evidence anchor verification": programmatic consistency calculation is responsible for numerical and curve indicators, while evidence anchor verification is responsible for method execution consistency and data completeness gatekeeping.
[0063] Algorithmically, let the paper's result be... The reproduction result is Then the consistency score is reproduced. Defined as:
[0064] in, This indicates the consistency of numerical indicators, which can be calculated using the thresholds and error calibers defined in MetricSpec; Waveform similarity can be represented by calculating the differences or similarity of waveform features using a program. The consistency of the execution of the representation method can be verified by running plan hash, key parameter overwrite records, key function call chain records and log anchors. This indicates consistency at the conclusion level, which can be inferred from waveforms and indicators using procedural rules and compared with the paper's conclusions. This indicates the agreement and evidence completeness correction items, used to penalize or restrict access for missing key fields, missing evidence anchors, or missing products. The weights are indicated by the symbols. In this application, the specific values of the weights are not limited, but can be determined by those skilled in the art according to actual needs.
[0065] when And when `gate_decision` passes, the "ReproRun -> consistent_with -> Experiment" sequence is written back to the knowledge graph; when... When it is partially consistent, it is marked as consistent and enters the review pool; when If the script fails to complete the runtime, it will be written to the inconsistent_with relationship or marked as "unreproducible" and enter the difference attribution module. This application embodiment does not limit the specific threshold value.
[0066] The set of possible causes for discrepancies can include: C = {MissingCode (e.g., missing executable code / entry script / key implementation), MissingCondition (e.g., missing key experimental conditions, such as incomplete protocol fields)), ParameterMismatch (parameter mismatch (inconsistent parameter values / units / calibers)), SolverMismatch (solver mismatch (e.g., inconsistent types or key configurations)), ScaleMismatch (scale mismatch (e.g., insufficient code capability grading or incompatible scales)), VersionMismatch (version inconsistency (e.g., differences in code / model / parameter versions)), InitStateMismatch (inconsistent initial states (e.g., differences in initial conditions / steady-state settings)), PostProcessMismatch (inconsistent post-processing (e.g., differences in metric calculation / filtering / windowing rules)), EnvironmentFailure (environmental failure (e.g., dependency / platform / runtime errors causing execution failure)), AmbiguousDescription (ambiguous description (e.g., insufficient information in the paper or code leading to indeterminacy)), MultipleCauses (multiple causes coexisting (cannot be uniquely attributed))} UnknownCause (Unknown cause (rule tree not covered and insufficient evidence))}.
[0067] During attribution, rule trees can be used to determine salient cases. For example, a missing entry script can be attributed to "MissingCode", a missing stimulus frequency or temperature to "MissingCondition", and insufficient code scale capability rating to "ScaleMismatch". For complex cases not covered by the rules, the system's verification attribution unit can infer and generate candidate attributions based on log anchors and code anchors, and output the attribution evidence set with EvidenceAnchors in the VerificationReport. When a candidate attribution cannot determine a unique cause, it can output "MultipleCauses" or "UnknownCause" and enter the review pool.
[0068] An example of a structured instruction template that can be used in the attribution verification stage is shown in the quotation marks below. This application embodiment uses an XML structure to enhance the stability of the structured output: “ <role>; You are the agent for the cardiac electrophysiology reproduction experiment verification.
[0069] < / role> ; <context>; The experimental protocol for the paper is located in paper_protocol.yaml; The target results for the paper are located in paper_results.json; The experimental results are reproduced in metrics.json and raw_trace.csv; The runtime log is located in run_log.txt; The code adaptation instructions are located in summary.md.
[0070] < / context> ; <task>; 1. Determine whether the current replication experiment truly adheres to the protocol outlined in the paper; 2. Compare the consistency between the results in the paper and the results of replication; 3. Provide the differences in numerical indicators, waveforms, and conclusions; 4. If there is an inconsistency or the problem cannot be reproduced, attribute it to one of the following sets: MissingCode, MissingCondition, ParameterMismatch, SolverMismatch, ScaleMismatch, VersionMismatch, InitStateMismatch, PostProcessMismatch, EnvironmentFailure, AmbiguousDescription; 5. Every conclusion must be accompanied by evidence.
[0071] < / task> ; <output>; Returns JSON: { "ok": true / false, "repro_score": 0-1, "status": "consistent|partially_consistent|inconsistent|unreproducible", "metric_diff": [...], "curve_diff": [...], "causes": [...], "evidence": [...], "reason": "If ok=false, write the reason for preventing passage"; } < / output> ".
[0072] The fusion write-back phase does not simply save a final conclusion, but rather writes the entire process back into the fusion knowledge graph in a graph structure, and records the construction process in a versioned manner. For each paper experiment Experiment_i and its aligned implementation Code_j, the system writes at least the following paths: Experiment_i ->aligned_to ->Code_j; Experiment_i ->reproduced_by ->ReproTask_k ->compiled_as ->ReproPackage_k ->executed_as ->ReproRun_k; ReproRun_k ->validated_by ->VerificationReport_k; VerificationReport_k ->consistent_with / inconsistent_with ->Experiment_i; VerificationReport_k ->attributed_to ->DifferenceCause_m; VerificationReport_k ->supported_by ->EvidenceAnchor_n.
[0073] The `EvidenceAnchor_n` can correspond to excerpts from the original paper, excerpts from figure captions, code function anchors, runtime log anchors, test output anchors, and result file anchors. Optionally, the original waveforms, logs, and result files are stored as external artifacts and referenced in the knowledge graph using URIs (Uniform Resource Identifiers) and `content_hash` to achieve lightweight graph placement and strong traceability. If the verification conclusion is "unreproducible," the path is not deleted from the knowledge graph; instead, the unreproducible state and the `DifferenceCause` node are retained, making "unreproducibility" itself formal knowledge.
[0074] Optionally, the system generates a build_id (build / writeback batch identifier) for each build and writeback, and associates the content_hash of ProtocolSpec, MetricSpec, CapabilityManifest, AlignmentRecord, ReproPackage, and VerificationReport with the build_id, so that the reproduction trajectory of the same paper experiment at different times, different code versions, or different environments can be distinguished and traced.
[0075] In this embodiment, the final fused knowledge graph can simultaneously contain three layers: first, a paper knowledge layer, recording models, experiments, protocols, and conclusions; second, a code knowledge layer, recording implementation, parameters, solvers, tissue assembly, and organ construction methods; and third, a reproduction knowledge layer, recording reproduction experimental packages, running records, verification results, and reasons for discrepancies. These three layers, through alignment edges, compilation edges, execution edges, and evidence anchor edges, constitute a unified multi-scale cardiac electrophysiological knowledge network.
[0076] From the perspective of implementation, the embodiments of this application are not "one-time summarization of literature," but rather place the large model or rule engine in three distinct and constrained positions: First, on the literature side, for entity relation extraction and protocol specification generation for pattern constraints; second, on the code side, for code semantic analysis, adaptation layer generation, reproduction experimental package generation, and test repair; third, on the verification side, for result verification, evidence anchor verification, difference attribution, and gated write-back review. Since both the extraction and verification sides are constrained by the domain schema, structured output, and evidence anchors, the drift caused by pure natural language generation can be reduced, and the final result can be solidified into a queryable, traceable, and verifiable fused knowledge graph.
[0077] In practical implementation, the schema constraint extraction engine, code parsing and adaptation unit, graph database writing unit, and verification and attribution unit on the document side are all replaceable components. Preferably, the document side can use Microsoft GraphRAG (a RAG framework for knowledge graphs), LlamaIndex's PropertyGraphIndex (a dedicated indexing component for attribute graphs built into LlamaIndex) / Schema constraint extractor, or the Neo4j KG Builder pipeline; graph storage can use Neo4j or other graph databases; code adaptation can use a code agent with file read / write and command execution capabilities; verification and attribution can be implemented using a verification agent that supports structured output and tool access (such as OpenClaw). Microsoft GraphRAG, LlamaIndex, Neo4j KG Builder, and OpenClaw are all existing framework / tool examples used to illustrate replaceable implementation paths for system components; their English names are inherent names or project names.
[0078] The multi-source knowledge graph construction method provided in this application constructs a domain ontology, transforms natural language experimental methods in papers into standardized experimental protocol objects, and performs semantic parsing on open-source code to form a structured representation of executable capabilities. Based on the protocol objects and code capability representation, the system performs cross-source evidence alignment, automatically generates reproducible experimental tasks and runs simulations, compares the reproducible experimental results with the paper results, and obtains a consistency score. When inconsistencies or inability to reproduce occur, the system attributes the differences and writes back the paper evidence, code evidence, reproducible verification evidence, and reasons for the differences to the fused knowledge graph, thereby constructing a traceable, verifiable, and iteratively updatable cardiac electrophysiological model knowledge graph covering cell, tissue, and organ scales. To enhance the feasibility of the project, this invention further solidifies at least five types of machine-verifiable structured products in the aforementioned closed loop: Protocol / MetricSpec, CapabilityManifest, AlignmentRecord, ReproPackage, and VerificationReport. These products, along with evidence anchors, are then written back to the fused knowledge graph in a versioned manner.
[0079] The multi-source knowledge graph construction method provided in this application aims to offer a technical solution for cross-source evidence alignment, automatic reproduction and verification, and differential attribution writing. It integrates scattered paper knowledge, experimental knowledge, and code knowledge in the field of cardiac electrophysiology into a searchable, traceable, verifiable, and iteratively updated multi-scale knowledge graph. This improves the efficiency of model knowledge organization, reproduction success rate, consistency analysis capability, and subsequent research reuse efficiency. It addresses the technical problems of inconsistent model names, experimental conditions, evaluation indicators, and conclusions in papers, making direct structuring difficult; the lack of a stable mapping relationship between paper methods and code implementations; significant differences in implementation versions, parameter specifications, and solver settings of the same model in different repositories; incomplete experimental protocols described in papers, making automatic generation and execution of reproducible experiments difficult; and the lack of a systematic attribution mechanism for discrepancies between paper results and reproducible results, even when the code can be run.
[0080] The multi-source knowledge graph construction method provided in this application does not focus on large-model text summarization or single-time knowledge extraction. Instead, it uses standardized experimental protocol representation as an intermediary to connect the natural language experimental descriptions in the paper, the executable implementations in the code, and the results of the reproducible experiments, forming the following closed loop: extraction of experimental methods from the paper → protocol standardization → code semantic parsing → cross-source alignment → automatic generation of reproducible experimental tasks → simulation execution → comparison with paper results → determination of the reasons for differences → writing the conclusions and evidence back into the knowledge graph. Through this closed loop, the knowledge graph can be upgraded from "static knowledge organization" to "dynamic evidence verification knowledge graph".
[0081] This embodiment also provides a multi-source knowledge graph construction apparatus, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0082] This embodiment provides a multi-source knowledge graph construction device, such as... Figure 2 As shown, it includes: The first acquisition module 201 is used to acquire a collection of literature related to the simulation experiment of human tissue electrophysiological model and the corresponding open source code. The first generation module 202 is used to generate an executable reproducible experimental package for any target document in the document set, based on the standard data corresponding to the target document, the code capability list of the open source code corresponding to the target document, and the open source code corresponding to the target document. The standard data includes experimental protocol standard data and experimental indicator standard data. The second acquisition module 203 is used to obtain the reproduction experiment results based on the execution results of the reproduction experiment package; The write-back module 204 is used to write back the literature knowledge graph subgraph generated based on the target literature, the code knowledge graph subgraph generated based on the parsing results of the corresponding open source code, and the reproduction knowledge graph subgraph generated based on the reproduction experiment results to the target knowledge graph.
[0083] The multi-source knowledge graph construction device provided in this embodiment generates corresponding literature knowledge graph subgraphs, code knowledge graph subgraphs, and reproduction knowledge graph subgraphs for each target document in the literature collection, and writes the generated subgraphs back into the target knowledge graph. This enables unified, intelligent, and structured management of experimental research content with a unified human tissue structure, facilitating the rapid retrieval of experimental research content with reference value.
[0084] In some optional implementations, the first generation module 202 includes: an extraction submodule, used to extract multiple types of entity data from the target document, the multiple types of entity data including: simulation model entity data, experimental entity data, protocol entity data, indicator entity data, and experimental conclusion entity data; a first generation submodule, used to generate the specification data based on the extracted entity data; a second generation submodule, used to parse the open-source code corresponding to the target document and generate the code capability list; a determination submodule, used to determine the cross-source alignment degree between the target document and the corresponding open-source code based on the specification data and the code capability list; and a response submodule, used to respond to the generation operation of the executable reproducible package when the cross-source alignment degree meets a preset requirement.
[0085] In some optional embodiments, the apparatus further includes: a second generation module, configured to verify the consistency between the reproduced experimental results and the original experimental results recorded in the target literature, and generate a verification report; an analysis module, configured to respond to a differential cause analysis operation when the consistency verification results between the reproduced experimental results and the original experimental results do not meet the requirements; and a recording module, configured to record the attribution tags and attribution basis corresponding to the analyzed differential causes in the verification report and write the differential causes back to the target knowledge graph.
[0086] In some optional implementations, the response submodule includes: a recording unit for recording the specification data, code capability list, and corresponding open-source code of the target document whose cross-source alignment meets preset requirements; a writing unit for writing the recording results into the target review pool; and a response unit for responding to the generation operation of the executable reproducible package when the review results meet the requirements.
[0087] In some optional embodiments, the apparatus further includes a third generation module, used to generate corresponding identification information for the acquired documents and open-source code, the identification information including a source identifier and a version identifier.
[0088] In some optional implementations, the extraction module includes: a parsing submodule for parsing the layout of the target document; a first acquisition submodule for dividing the document into blocks based on the layout parsing results, obtaining multiple document block sets, and generating corresponding evidence anchors for each document block; and a second acquisition submodule for obtaining the multiple types of entity data by parsing each document block using natural language.
[0089] The multi-source knowledge graph construction apparatus provided in this disclosure can execute the multi-source knowledge graph construction method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.
[0090] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.
[0091] The following is a detailed reference. Figure 3 The diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from memory 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device. The processor 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0092] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0093] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a memory 508, or installed from a ROM 502. When the computer program is executed by the processor 501, it performs the functions defined in the multi-source knowledge graph construction method of embodiments of this disclosure.
[0094] Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0095] This disclosure also provides a computer-readable storage medium in which the methods described in this disclosure can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium may also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the multi-source knowledge graph construction method shown in the above embodiments is implemented.
[0096] A portion of this disclosure can be applied to computer program products, such as computer program instructions, which, when executed by a computer, can invoke or provide methods and / or technical solutions according to this disclosure through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, and installation package files. Accordingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions; the computer compiling the instructions and then executing the corresponding compiled program; the computer reading and executing the instructions; or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0097] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for constructing a multi-source knowledge graph, characterized in that, The method includes: Acquire a collection of literature related to human tissue electrophysiological model simulation experiments and the corresponding open-source code; For any target document in the literature collection, an executable reproducible experimental package is generated based on the standard data corresponding to the target document, the code capability list of the open source code corresponding to the target document, and the open source code corresponding to the target document. The standard data includes experimental protocol standard data and experimental indicator standard data. Based on the execution results of the aforementioned reproducible experimental package, the reproducible experimental results are obtained; The document knowledge graph subgraph generated based on the target document, the code knowledge graph subgraph generated based on the parsing results of the corresponding open-source code, and the reproduction knowledge graph subgraph generated based on the reproduction experiment results are written back into the target knowledge graph.
2. The method according to claim 1, characterized in that, The step of generating an executable reproducibility test package for any target document in the document set, based on the specification data corresponding to the target document, the code capability list of the open-source code corresponding to the target document, and the open-source code corresponding to the target document, includes: Multiple types of entity data are extracted from the target documents, including: simulation model entity data, experimental entity data, protocol entity data, indicator entity data, and experimental conclusion entity data. The standardized data is generated based on the extracted entity data; The open-source code corresponding to the target document is parsed to generate the code capability list; Based on the specification data and the code capability list, the degree of cross-source alignment between the target document and the corresponding open-source code is determined; When the cross-source alignment meets the preset requirements, the system responds by generating the executable reproducible package.
3. The method according to claim 1, characterized in that, The method further includes: The consistency between the reproduced experimental results and the original experimental results recorded in the target literature is verified, and a verification report is generated. If the consistency verification results between the reproduced experimental results and the original experimental results do not meet the requirements, the response should perform a differential cause analysis operation. The attribution labels and attribution basis corresponding to the differentiated causes obtained from the analysis are recorded in the verification report, and the differentiated causes are written back into the target knowledge graph.
4. The method according to claim 2, characterized in that, When the cross-source alignment meets a preset requirement, the step of generating the executable reproducible package includes: Record the standard data, code capability list, and corresponding open source code for target documents that meet the preset requirements for cross-source alignment. Write the recorded results to the target review pool; When the verification results meet the requirements, the system will respond by generating the executable reproducible package.
5. The method according to claim 1, characterized in that, The method further includes generating corresponding identification information for the acquired documents and open-source code, wherein the identification information includes source identifier and version identifier.
6. The method according to claim 2, characterized in that, The extraction of multiple types of entity data from the target document includes: Perform page layout analysis on the target document; Based on the page layout analysis results, the document is divided into blocks to obtain multiple document block sets, and corresponding evidence anchors are generated for each document block. The entity data of the multiple types are obtained by parsing each document block using natural language.
7. A multi-source knowledge graph construction device, characterized in that, The device includes: The first acquisition module is used to acquire a collection of literature related to human tissue electrophysiological model simulation experiments and the corresponding open source code. The first generation module is used to generate an executable reproducible experimental package for any target document in the document set, based on the standard data corresponding to the target document, the code capability list of the open source code corresponding to the target document, and the open source code corresponding to the target document. The standard data includes experimental protocol standard data and experimental indicator standard data. The second acquisition module is used to obtain the reproduction experiment results based on the execution results of the reproduction experiment package; The write-back module is used to write back the literature knowledge graph subgraph generated based on the target literature, the code knowledge graph subgraph generated based on the parsing results of the corresponding open source code, and the reproduction knowledge graph subgraph generated based on the reproduction experiment results to the target knowledge graph.
8. An electronic device, characterized in that, include: A memory and a processor are interconnected, the memory stores computer instructions, and the processor executes the computer instructions to perform the multi-source knowledge graph construction method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to execute the multi-source knowledge graph construction method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, It includes computer instructions for causing a computer to execute the multi-source knowledge graph construction method according to any one of claims 1 to 6.