A multi-source heterogeneous disease data standardization collection method and system for clinical key specialties

CN122658697APending Publication Date: 2026-08-28SUZHOU HENGYIXIN INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610849866.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-12
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0007]有鉴于现有技术在临床重点专科多源异构疾病数据采集过程中存在数据分散、语义不一致、采集时间点错误、多源冲突处理不足以及标准化可追溯性差的问题,本发明的目的在于提供一种用于临床重点专科的多源异构疾病数据标准化采集方法及系统,以实现围绕特定诊疗事件的动态事件驱动采集、多源候选值冲突消解以及全流程标准化证据链追溯的技术效果

Benefits of technology

[0018] Based on the above technical solutions, the present invention provides a standardized acquisition method and system for multi-source heterogeneous disease data in key clinical specialties. By establishing a binding relationship between standard data elements, diagnostic and treatment event nodes, and event acquisition time windows, and dynamically judging acquisition trigger conditions in conjunction with the patient's diagnostic and treatment event timeline, event-driven acquisition tasks are generated, realizing disease data acquisition around specific diagnostic and treatment events and actual acquisition time windows. At the same time, through candidate value conflict graphs, logical weighted decision-making, and standardized evidence chains, conflict resolution and full-process traceability of multi-source candidate data are achieved. This solves the problems of inaccurate temporal semantics in multi-source heterogeneous disease data acquisition, reliance on manual maintenance for field mapping, insufficient handling of multi-source candidate value conflicts, and poor traceability of the standardization process in existing technologies, thereby improving the accuracy, completeness, maintainability, and traceability of disease data acquisition in key clinical specialties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122658697A_ABST
    Figure CN122658697A_ABST
Patent Text Reader

Abstract

The application discloses a multi-source heterogeneous disease data standardization collection method and system for clinical key specialties. By constructing a disease data standard model, a binding relationship is established among standard data elements, diagnosis and treatment event nodes and event collection time windows; in combination with a patient diagnosis and treatment event time axis and a collection trigger condition, an event-driven collection task is dynamically generated; multi-source candidate disease data is subjected to standardization conversion, and a standard main value is determined through a candidate value conflict graph and a logic weighted decision; and finally, a standardized evidence chain is generated, realizing the full-process traceability of data from collection to standardization. The method and system can improve the accuracy, integrity and maintainability of multi-source heterogeneous disease data, and provide a reliable data basis for clinical key specialty construction, scientific research analysis and quality control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical data processing technology, and in particular to a standardized acquisition method and system for multi-source heterogeneous disease data in key clinical specialties. Background Technology

[0002] With the rapid development of hospital information technology, medical institutions typically have multiple business systems, including Hospital Information System (HIS), Electronic Medical Record System (EMR), Laboratory Information System (LIS), Picture Archiving and Communication System (PACS), Pathology Information System, Anesthesia and Surgery System, Pharmacy System, Nursing System, Follow-up System, and disease-specific databases. These systems have accumulated a large amount of patient disease data during clinical diagnosis and treatment, scientific research and analysis, the development of key specialties, and medical quality management. However, because these systems were built by different vendors at different times, their data structures, field naming, coding systems, interface specifications, and storage methods vary significantly. This results in the disease data of the same patient exhibiting characteristics of being multi-sourced, heterogeneous, scattered, and semantically inconsistent.

[0003] Existing technologies typically aggregate data from different systems into a data center or disease-specific database through data interface integration, database extraction, field mapping, data cleaning, data transformation, and unified storage. These methods are centered on data sources or tables, relying on fixed field mappings and manually maintained data cleaning rules. When the fields, table structures, or business processes of the source system change, existing rules are prone to becoming invalid, requiring extensive manual reconfiguration and review. Furthermore, existing technologies lack data collection methods tailored to the temporal characteristics of events in key clinical specialties. For example, the same data element may have different clinical meanings at different stages of diagnosis and treatment; data such as preoperative examinations, postoperative complication records, pre-treatment staging, post-treatment efficacy evaluation, discharge, and follow-up outcomes are difficult to accurately match with corresponding events using existing methods.

[0004] Furthermore, in multi-source data environments, the same standard data element often has multiple candidate values. For example, the operation time may come from the anesthesia system, the medical record cover sheet, the progress notes, or the billing system; diagnostic information may come from the admission diagnosis, discharge diagnosis, pathology report, or a disease-specific database. Existing technologies typically determine the final value through fixed priorities, manual review, or simple coverage rules, lacking effective conflict identification, conflict resolution, and multi-source evidence preservation mechanisms, resulting in insufficient data accuracy and traceability.

[0005] Therefore, the existing technology has at least the following shortcomings: 1. Data collection methods are centered on data sources or fields, lacking a standard model-driven mechanism for key clinical specialties and diseases; 2. It cannot combine the patient's diagnosis and treatment event timeline with the event acquisition time window to achieve accurate event-driven data acquisition; 3. The mapping between source fields and standard data elements relies on manual configuration, making it difficult to cope with changes in source system fields or differences between multiple systems; 4. There is a lack of scientific conflict resolution and evidence preservation mechanisms for multiple candidate values ​​of the same standard data element; 5. Standardized data lacks full-process traceability, failing to meet the high-quality data needs for the development of key clinical specialties and scientific research analysis.

[0006] Based on the above shortcomings, there is an urgent need to provide a standardized acquisition method and system for multi-source heterogeneous disease data in key clinical specialties. This system should be able to combine disease data standard models, patient diagnosis and treatment event timelines, and event acquisition time windows to achieve event-driven data acquisition, dynamic logical judgment, multi-source conflict resolution, and standardized evidence chain tracing, thereby improving the accuracy, completeness, maintainability, and traceability of multi-source heterogeneous disease data. Summary of the Invention

[0007] In view of the problems existing in the collection of multi-source heterogeneous disease data in key clinical specialties, such as data dispersion, semantic inconsistency, incorrect collection time point, insufficient handling of multi-source conflicts, and poor standardization and traceability, the purpose of this invention is to provide a standardized collection method and system for multi-source heterogeneous disease data in key clinical specialties, so as to achieve the technical effects of dynamic event-driven collection around specific diagnosis and treatment events, resolution of conflicts of multi-source candidate values, and traceability of the entire process of standardized evidence chain.

[0008] To achieve the above objectives, the present invention provides the following technical solution: In one possible implementation, a method for standardized collection of multi-source heterogeneous disease data for key clinical specialties is provided, comprising: constructing a disease data standard model corresponding to the target key clinical specialty, wherein the disease data standard model includes a set of standard data elements, a set of diagnosis and treatment event nodes, and an event collection time window; establishing a binding relationship between the standard data elements, diagnosis and treatment event nodes, and the event collection time window, wherein the binding relationship is used to limit the collection stage and collection time range of the target standard data elements relative to the target diagnosis and treatment event nodes; accessing multiple heterogeneous medical data sources related to the target patient, and identifying key diagnosis and treatment events of the target patient from the heterogeneous medical data sources; constructing a patient diagnosis and treatment event timeline based on the event occurrence time of the key diagnosis and treatment events; and dynamically determining whether the target standard data elements meet the collection trigger conditions based on the binding relationship and the patient diagnosis and treatment event timeline, including whether the event has occurred and whether the fields are complete. The system checks the consistency of the data source, historical mapping, and whether the data collection time window has been reached. When the collection triggering conditions are met, an event-driven collection task is generated and executed to collect candidate disease data located within the actual collection time window from multiple heterogeneous medical data sources. The candidate disease data is standardized and converted, and standardized disease data is generated based on the results of encoding mapping, unit conversion, time format standardization, missing value annotation, value range verification, and structured extraction. When multiple candidate values ​​exist for the same standard data element, a candidate value conflict graph is constructed, and logical calculations and weighted decisions are performed based on the authority of the candidate value source, time accuracy, completeness, time proximity to the target diagnosis and treatment event node, and historical quality scores to determine the standard principal value. Candidate values ​​that are not selected as principal values ​​are saved as conflict values. A standardized evidence chain corresponding to the standardized disease data is generated, and the triggering logic, collection conditions, conflict decisions, and manual review information are recorded to achieve full-process traceability.

[0009] In one possible implementation, the binding relationship is also used to limit the data source range, collection priority, and data verification rules of different standard data elements under different diagnosis and treatment event nodes.

[0010] In one possible implementation, the event acquisition time window includes at least a number of the following: pre-event time window, post-event time window, event proximity time window, periodic acquisition time window, and follow-up time window, and can be dynamically adjusted based on the patient's medical event status.

[0011] In one possible implementation, after accessing heterogeneous medical data sources, a source field profile is generated, including the source field name, source field path, data type, unit distribution, value distribution, time attribute, source system identifier, field association, null value rate, historical mapping records, and historical quality score, which are used for subsequent logical judgment and trigger condition evaluation.

[0012] In one possible implementation, candidate mapping relationships are generated based on the source field profile and the standard data element set, and the mapping confidence is determined by logical weighting. When the mapping confidence is lower than a preset threshold, it automatically enters the manual review queue; when the mapping confidence meets the threshold, the candidate mapping relationship is automatically confirmed.

[0013] In one possible implementation, constructing a patient diagnosis and treatment event timeline includes: extracting the name of the diagnosis and treatment event, the time of occurrence of the event, the source of the event, and evidence of the event; merging the same or similar events; and dynamically updating the event timeline when the event time changes to trigger a re-collection task.

[0014] In one possible implementation, the event-driven acquisition task includes patient identifier, target standard data element, target diagnosis and treatment event node, actual acquisition time window, candidate data source, candidate source field, acquisition priority, acquisition triggering condition and logical judgment relationship, which are used to dynamically control the acquisition order and repeated acquisition.

[0015] In one possible implementation, the step of determining the standard principal value based on the candidate value conflict graph includes: constructing candidate value nodes and relation edges to represent the equality, approximation, inclusion or conflict relationships between candidate values; calculating the principal value score of each node using weighted logic; and dynamically selecting the principal value according to the score results and preset logic rules.

[0016] In one possible implementation, the standardized chain of evidence records the original data source, source field path, original value, standard value, standardized transformation rules, disease data standard model version, mapping confidence, target diagnosis and treatment event node, actual collection time window, conflict candidate value information, logical judgment results, and manual review records, so as to achieve dynamic traceability of the entire process.

[0017] In one possible implementation, a standardized acquisition system for multi-source heterogeneous disease data in key clinical specialties is provided, comprising: a disease data standard model construction module for constructing a disease data standard model and establishing binding relationships between standard data elements, diagnosis and treatment event nodes, and event acquisition time windows; a multi-source data access module for accessing multiple heterogeneous medical data sources; a diagnosis and treatment event timeline construction module for identifying key diagnosis and treatment events and constructing a patient diagnosis and treatment event timeline, dynamically updating the patient event status; an event-driven acquisition task generation module for dynamically generating event-driven acquisition tasks based on the binding relationships and logical judgments of the patient diagnosis and treatment event timeline; a data acquisition module for executing event-driven acquisition tasks and acquiring candidate disease data; a standardization conversion module for standardizing the candidate disease data; a multi-source conflict resolution module for constructing a candidate value conflict graph and selecting a standard principal value based on logical weighted judgments; and an evidence chain tracing module for generating a standardized evidence chain and recording logical judgments and triggering conditions.

[0018] Based on the above technical solutions, the present invention provides a standardized acquisition method and system for multi-source heterogeneous disease data in key clinical specialties. By establishing a binding relationship between standard data elements, diagnostic and treatment event nodes, and event acquisition time windows, and dynamically judging acquisition trigger conditions in conjunction with the patient's diagnostic and treatment event timeline, event-driven acquisition tasks are generated, realizing disease data acquisition around specific diagnostic and treatment events and actual acquisition time windows. At the same time, through candidate value conflict graphs, logical weighted decision-making, and standardized evidence chains, conflict resolution and full-process traceability of multi-source candidate data are achieved. This solves the problems of inaccurate temporal semantics in multi-source heterogeneous disease data acquisition, reliance on manual maintenance for field mapping, insufficient handling of multi-source candidate value conflicts, and poor traceability of the standardization process in existing technologies, thereby improving the accuracy, completeness, maintainability, and traceability of disease data acquisition in key clinical specialties. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of a standardized acquisition system for multi-source heterogeneous disease data in key clinical specialties, provided by an embodiment of the present invention.

[0020] like Figure 1 As shown, the system includes a multi-source data access module, a source field profile generation module, a disease data standard model construction module, a candidate mapping generation module, a diagnosis and treatment event timeline construction module, an event-driven acquisition task generation module, a standardization conversion module, a multi-source conflict resolution module, and an evidence chain tracing module. Through the collaborative work of these modules, the system achieves the acquisition, standardization processing, and conflict resolution of multi-source heterogeneous medical data, while simultaneously generating a standardized evidence chain throughout the entire process, ensuring the accuracy, completeness, and traceability of data collection for key clinical specialties.

[0021] Figure 2 This is a schematic diagram of the standardized collection method for multi-source heterogeneous disease data in key clinical specialties provided by an embodiment of the present invention.

[0022] like Figure 2 As shown, the method includes the following steps: constructing a standard model for disease data; accessing multiple heterogeneous medical data sources and generating source field profiles; generating candidate mapping relationships and calculating mapping confidence based on source field profiles and standard data element sets; identifying key diagnosis and treatment events of target patients and constructing a timeline of patient diagnosis and treatment events; generating event-driven acquisition tasks based on the standard model for disease data and the timeline of patient diagnosis and treatment events; executing event-driven acquisition tasks to obtain candidate disease data; standardizing and transforming the candidate disease data; constructing a candidate value conflict graph and determining the standard principal value; and generating standardized disease data and its standardized evidence chain. Detailed Implementation

[0023] To enable those skilled in the art to more clearly understand the technical solutions, principles, and improvements over the prior art, the embodiments of the present invention are described in detail below with reference to the accompanying drawings. It should be understood that the following embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0024] The core of this invention lies in: constructing a standard model of disease data based on the target key clinical specialties, and establishing a binding relationship between standard data elements, diagnosis and treatment event nodes, and event collection time windows; further combining the patient's diagnosis and treatment event time axis and collection trigger conditions to dynamically generate event-driven collection tasks; after obtaining multi-source candidate disease data, through standardized transformation, candidate value conflict resolution, and standardized evidence chain tracing, accurate collection, standardized processing, and full-process traceability of multi-source heterogeneous disease data are achieved.

[0025] The following will provide a detailed description of the invention from the aspects of overall system structure, construction of disease data standard model, generation of source field profiles, candidate mapping relationship and mapping confidence, construction of patient diagnosis and treatment event timeline, generation and execution of event-driven acquisition tasks, resolution of candidate value conflicts, standardization transformation and evidence chain generation, data quality evaluation and field drift monitoring, and application examples.

[0026] I. General Description The following combination Figure 1 and Figure 2 This invention provides a method and system for standardized acquisition of multi-source heterogeneous disease data in key clinical specialties. It should be understood that the following embodiments are only for explaining the technical solutions of this invention and are not intended to limit the scope of protection of this invention. Where there is no conflict, the technical features in this embodiment can be combined with each other.

[0027] The multi-source heterogeneous disease data standardization acquisition system provided in this invention is designed for application scenarios such as the construction of key clinical specialties, the development of disease-specific databases, the establishment of research cohorts, medical quality control, follow-up management, and specialty competency evaluation. The key clinical specialties may include cardiovascular medicine, oncology, neurology, respiratory medicine, orthopedics, obstetrics and gynecology, and critical care medicine. The target diseases can be determined according to actual operational needs, such as stroke, lung cancer, coronary heart disease, chronic obstructive pulmonary disease, and diabetic complications. This invention does not limit the specific specialty type or specific disease type.

[0028] like Figure 1As shown, the multi-source heterogeneous disease data standardization acquisition system of this invention includes a multi-source data access module, a source field profile generation module, a disease data standard model construction module, a candidate mapping generation module, a diagnosis and treatment event timeline construction module, an event-driven acquisition task generation module, a standardization conversion module, a multi-source conflict resolution module, and an evidence chain tracing module. Each module can be deployed on the same server, a hospital intranet server, a regional medical data platform, a private cloud platform, or other computing environments with data processing capabilities. Data interaction between modules can occur through databases, interface services, message queues, file exchange, or memory access.

[0029] The multi-source data access module is used to access multiple heterogeneous medical data sources. These heterogeneous medical data sources may include at least two of the following: hospital information systems, electronic medical record systems, laboratory information systems, medical imaging systems, pathology information systems, anesthesia and surgery systems, pharmacy systems, nursing systems, follow-up systems, disease-specific databases, and patient reporting systems. Because these systems differ in field naming, data structure, coding systems, time recording methods, and business meanings, directly collecting data based on fixed fields or fixed data tables can easily lead to errors in target data element collection, inconsistencies in collection stages, or conflicts in candidate values ​​from multiple sources.

[0030] The disease-specific data standard model construction module is used to build disease-specific data standard models corresponding to target key clinical specialties. This model is not simply a data dictionary; rather, it associates standard data elements, diagnostic and treatment event nodes, and event collection time windows. It describes which data elements should be collected around which diagnostic and treatment events, within what time frame, and from which data sources for a specific target disease. This model transforms traditional field-centric data extraction methods into data collection methods oriented towards clinical diagnostic and treatment events.

[0031] The treatment event timeline construction module is used to identify key treatment events for a target patient from multiple heterogeneous medical data sources, and construct a patient treatment event timeline based on the occurrence time of these key events. These key treatment events may include at least several of the following: initial visit, admission, diagnosis, examination, laboratory tests, surgery, medication, treatment initiation, discharge, follow-up visit, and re-examination. The patient treatment event timeline represents the actual treatment process of the target patient, providing a basis for subsequently determining the actual data collection time window and generating event-driven data collection tasks.

[0032] The event-driven acquisition task generation module is used to determine the target medical event node and actual acquisition time window corresponding to the target standard data element based on the binding relationship in the disease data standard model and the patient's diagnosis and treatment event time axis, and to generate event-driven acquisition tasks in combination with acquisition trigger conditions. The acquisition trigger conditions may include whether the diagnosis and treatment event has occurred, whether the acquisition time window has been reached, whether the candidate source field exists, whether the field integrity requirements are met, whether the historical mapping relationship is consistent, and whether there are any abnormal situations that require supplementary or re-acquisition. Therefore, this invention does not perform static extraction according to a preset fixed table structure, but dynamically generates acquisition tasks according to the actual patient diagnosis and treatment process.

[0033] The standardization transformation module is used to standardize candidate disease data obtained from event-driven acquisition tasks, resulting in standardized disease data. The standardization transformation may include encoding mapping, unit conversion, time format standardization, missing value annotation, value range verification, and structured extraction from unstructured or semi-structured medical text. Through standardization transformation, heterogeneous data from different systems can be converted into data formats and semantics that conform to the target disease data standard model requirements.

[0034] The multi-source conflict resolution module is used to identify the relationships between multiple candidate values ​​when multiple candidate values ​​exist for the same standard data element, and to determine the standard principal value. In one implementation, a candidate value conflict graph can be constructed, where candidate value nodes represent different candidate values ​​corresponding to the same standard data element, and candidate value relationship edges represent equality, approximation, inclusion, or conflict relationships between different candidate values. The system can calculate the principal value score of candidate value nodes based on factors such as the authority of the candidate value source, the time accuracy of the candidate value, the completeness of the candidate value, the time proximity of the candidate value to the target diagnostic event node, and the historical quality score of the data source, and determine the standard principal value based on the score results and preset logical rules.

[0035] The evidence chain tracing module is used to generate a standardized evidence chain corresponding to standardized disease data. This standardized evidence chain may include at least several of the following: original data source, source field path, original value, standard value, standardization transformation rules, disease data standard model version, mapping confidence level, target diagnosis and treatment event node, actual data collection time window, conflict candidate value information, logical judgment results, manual review records, and generation time. Through the standardized evidence chain, the processing path from original data to standardized results can be recorded, enabling the standardized disease data to be reviewed, located, traced, and interpreted.

[0036] like Figure 2As shown, the method provided in this embodiment of the invention generally includes: constructing a standard model for disease data; accessing multiple heterogeneous medical data sources and generating source field profiles; generating candidate mapping relationships and calculating mapping confidence based on the source field profiles and standard data element sets; identifying key diagnosis and treatment events of the target patient and constructing a patient diagnosis and treatment event timeline; determining the actual collection timeline based on the binding relationship between standard data elements, diagnosis and treatment event nodes, and event collection time windows, combined with the patient diagnosis and treatment event timeline; generating event-driven collection tasks according to collection triggering conditions; executing event-driven collection tasks to obtain candidate disease data; standardizing the candidate disease data; resolving conflicts and determining the standard principal value when multiple candidate values ​​exist; and generating standardized disease data and its standardized evidence chain.

[0037] Compared with existing data acquisition methods centered on data sources, data tables, or fixed fields, the embodiments of this invention have at least the following characteristics: First, by establishing a binding relationship between standard data elements, diagnosis and treatment event nodes, and event acquisition time windows through a disease-specific data standard model, the data acquisition has clear clinical event semantics; Second, by determining the actual diagnosis and treatment events and their time locations for the target patient through a patient diagnosis and treatment event timeline, the acquisition task can be dynamically generated around the patient's actual diagnosis and treatment process; Third, by controlling the generation, execution, supplementary acquisition, and re-acquisition of acquisition tasks through acquisition triggering conditions and logical judgment relationships, the risk of data from erroneous stages being acquired as target data is reduced; Fourth, by handling multi-source candidate value conflicts through candidate value conflict graphs and principal value scoring mechanisms, the accuracy of standardized disease data is improved; Fifth, by preserving the processing basis for acquisition, transformation, mapping, conflict resolution, and manual review through a standardized evidence chain, the traceability and interpretability of data results are improved.

[0038] Therefore, the embodiments of the present invention can solve the problems in the prior art of clinical key specialty disease data collection which is centered on data sources or fields, which makes it difficult to reflect the temporal semantics of diagnosis and treatment events, the collection rules are prone to failure after changes in source fields, the handling of conflicts between multiple candidate values ​​is insufficient, and the standardized results lack complete traceability basis, thereby improving the accuracy, completeness, maintainability and traceability of standardized collection of multi-source heterogeneous disease data.

[0039] II. Construction of Standard Model for Disease Data In this embodiment, the disease data standard model is used to describe the data collection rules for the target key clinical specialty or target disease. This model is not merely used to record data element names or field formats, but rather to establish the binding relationship between standard data elements, diagnosis and treatment event nodes, and event collection time windows, thereby enabling disease data collection to revolve around real diagnosis and treatment events.

[0040] In one possible implementation, the disease-specific data standard model includes a set of standard data elements, a set of diagnostic and treatment event nodes, an event collection time window, data source priority rules, standard mapping rules, unit conversion rules, value range verification rules, and conflict resolution rules. Specifically, the set of standard data elements defines the data items that need to be collected by the target key clinical specialty; the set of diagnostic and treatment event nodes defines the key events related to the diagnosis and treatment process of the target disease; and the event collection time window limits the collection time range of standard data elements relative to the diagnostic and treatment event nodes.

[0041] The standard data set can include basic patient information, diagnostic information, symptoms and signs, examination results, laboratory results, surgical information, medication information, pathological information, staging and grading information, scoring scales, complication information, discharge outcome, and follow-up outcome. For example, for a key specialty in stroke, standard data elements can include onset time, arrival time, first imaging examination time, NIHSS score, thrombolysis initiation time, thrombectomy information, discharge mRS score, and 90-day follow-up mRS score. For a key specialty in oncology, standard data elements can include pathological type, TNM stage, gene mutation status, initial treatment regimen, efficacy evaluation results, and survival outcome.

[0042] The set of diagnostic and treatment event nodes can include at least several of the following: initial visit event, hospital admission event, diagnosis event, examination event, laboratory test event, surgical event, medication event, treatment initiation event, discharge event, follow-up visit event, and follow-up event. Different diseases may correspond to different diagnostic and treatment event nodes. For example, the disease of stroke may include the onset event, hospital arrival event, first imaging examination event, thrombolysis initiation event, thrombectomy initiation event, discharge event, and follow-up event; the disease of oncology may include the initial visit event, imaging abnormality event, pathological diagnosis event, molecular testing event, first treatment event, efficacy evaluation event, and follow-up event.

[0043] An event acquisition time window is used to define the acquisition range of standard data elements relative to the diagnostic and treatment event node. The event acquisition time window may include at least one of the following: pre-event time window, post-event time window, event proximity time window, periodic acquisition time window, and follow-up time window. For example, preoperative laboratory indicators may be limited to the most recent laboratory results within 7 days prior to the surgical event; post-treatment efficacy evaluation may be limited to examination results within a preset period after the start of treatment; and follow-up outcomes may be limited to follow-up records within 30, 90, or 180 days after discharge.

[0044] In one possible implementation, the system establishes a binding relationship between standard data elements, diagnostic and treatment event nodes, and event collection time windows. This binding relationship is used to define which target diagnostic and treatment event node the target standard data element should be collected around, and within which time range before, after, or adjacent to that target diagnostic and treatment event node. This binding relationship prevents data from being mistakenly used as target standard data elements for non-target diagnostic and treatment stages. For example, the same test indicator, "hemoglobin," may have different clinical meanings in preoperative assessment, postoperative recovery, and follow-up examination stages; the system can use the binding relationship to define its corresponding diagnostic and treatment event node and event collection time window separately.

[0045] Furthermore, the binding relationship can also limit the data source range, collection priority, and data verification rules for different standard data elements under different diagnosis and treatment event nodes. For example, for "surgery start time", the priority data source can be set to the anesthesia and surgery system, followed by the electronic medical record system or the medical record cover page; for "test results", the priority data source can be set to the laboratory information system, followed by the medical record text extraction results; for "pathological classification", the priority data source can be set to the pathology information system, followed by the discharge record or disease-specific database.

[0046] In one possible implementation, the disease-specific data standard model also includes data source priority rules. These rules determine the initial collection order or scoring basis for candidate values ​​when multiple heterogeneous medical data sources contain such values. These rules can be configured based on the data source's business authority, data update timeliness, time accuracy, structure, and historical quality scores. For example, the surgery start time recorded in anesthesia systems typically has high time accuracy and can therefore be set as a high-priority source for surgical time-related data elements.

[0047] In one possible implementation, the disease-specific data standard model also includes standard mapping rules. Standard mapping rules describe the mapping relationship between source fields and standard data elements, or constrain the generation of subsequent candidate mapping relationships. Standard mapping rules may include field name correspondences, medical terminology correspondences, coding system correspondences, unit correspondences, and synonym or near-synonym correspondences. Standard mapping rules can reduce data element identification errors caused by differences in field naming in the source system.

[0048] In one possible implementation, the disease-specific data standard model also includes unit conversion rules and value range verification rules. Unit conversion rules address the issue of different units of measurement for the same indicator in different medical data sources. For example, the same test indicator might be recorded in mg / dL, mmol / L, or other units; the system can convert it to a unified standard unit according to the unit conversion rules. Value range verification rules determine whether the collected values ​​are within a reasonable range and can be used to detect outliers, data entry errors, or unit errors.

[0049] In one possible implementation, the disease-specific data standard model also includes conflict resolution rules. These rules provide a basis for constructing a conflict graph and determining the primary standard value when multiple candidate values ​​exist for the same standard data element. These rules may include weight settings for factors such as source authority, time accuracy, completeness, temporal proximity to the target diagnostic event node, and historical quality scores. They may also include criteria for determining equality, approximation, inclusion, or conflict relationships.

[0050] In one possible implementation, the disease-specific data standard model can be managed by version. Each version of the disease-specific data standard model can record version information for the standard data element set, the set of diagnostic and treatment event nodes, the event collection time window, binding relationships, mapping rules, verification rules, and conflict resolution rules. When the model content is adjusted, the system can generate a new model version and write the model version into the standardized evidence chain so that the model rules on which a certain standardized disease data is based can be traced later.

[0051] By constructing the aforementioned standard model for disease data, this invention transforms the traditional "field extraction" method for collecting clinical key specialty disease data into "event semantic-driven collection." This model binds standard data elements with diagnostic and treatment event nodes and event collection time windows, enabling the collection system to understand the clinical stage and temporal semantics corresponding to the data elements. This provides a unified rule foundation for subsequent patient diagnostic and treatment event timeline construction, event-driven collection task generation, multi-source candidate value conflict resolution, and evidence chain tracing.

[0052] III. Source Field Profile Generation In this embodiment, the source field profile generation module is used to parse and describe field information from multiple heterogeneous medical data sources, forming the input basis for subsequent candidate mapping generation, event-driven data acquisition task determination, and logical judgment. The source field profile not only records basic field information but also reflects the semantics, structure, and historical processing status of the fields in clinical data, providing dynamic judgment conditions for the system.

[0053] In one possible implementation, the source field profile includes, but is not limited to, the following information: 1. Basic Field Information Field name: The original name of the field in the source system; Field path: The storage path of the field in the database table, document template, or interface message; Data types: such as integer, floating-point, character, date, etc.; Unit distribution: Statistics on the possible units and frequencies of the field; Value distribution: Analysis of the field's value range and frequency of occurrence; Time attribute: The time dimension of the field record, such as the time of event occurrence or the time of record; Source system identifier: The data source system to which the field belongs, such as HIS, LIS, PACS, etc.

[0054] 2. Field Relationships and Semantic Information Field relationships: The relationships or dependencies between fields for the same patient or the same event; Synonyms or semantic mapping: the semantic correspondence between field names and standard data element names, supporting natural language processing or knowledge graph reasoning; Historical mapping records: the mapping results of fields to standard data elements and the results of manual review; Historical quality score: Scoring of the completeness, accuracy, and reliability of field data.

[0055] 3. Field Status and Processing Information Null value rate: The proportion of null values ​​in a field within historically collected data; Data anomaly characteristics: abnormal distribution of field values, outlier statistics, or coding anomalies; Update frequency and drift: The pattern of field changes over time or the degree of field drift, used to dynamically adjust the collection and mapping rules.

[0056] In one possible implementation, the source field profiling generation module performs unified modeling of the basic information, semantic information, relationships, and historical processing status of each field by parsing the database table structure, interface messages, document templates, or text extraction fields. The generated source field profiling can not only be used for candidate mapping generation, but also as conditional input for event-driven data acquisition task logic judgment, supporting the system in determining whether a field meets the data acquisition trigger conditions, whether there are conflicts, or whether re-acquisition is required.

[0057] Furthermore, the source field profile can support dynamic updates and version management. When the field structure, naming, or units of the source system change, the system periodically generates the latest field profile and compares it with historical field profiles to calculate the field drift. If the drift exceeds a preset threshold, it triggers event-driven data collection task adjustments or manual review, thereby ensuring the consistency and stability of data collection with the standard model.

[0058] By generating source field profiles, this invention achieves unified description and quantitative analysis of fields from heterogeneous medical data sources, providing a solid data foundation for event-driven data collection, candidate mapping relationship generation, multi-source conflict resolution, and standardized evidence chain construction, while also supporting the system's dynamic adaptability to field changes.

[0059] IV. Candidate Mapping Relationships and Mapping Confidence In this embodiment, the candidate mapping generation module is used to generate candidate mapping relationships between source fields and standard data elements based on the source field profile and the standard data element set in the disease data standard model, and to calculate the mapping confidence of the candidate mapping relationship. The candidate mapping relationship is used to represent the possible data semantic correspondence between source fields and target standard data elements in a heterogeneous medical data source, and the mapping confidence is used to represent the degree of credibility of the automatic adoption of the candidate mapping relationship.

[0060] In one possible implementation, the candidate mapping generation module first obtains the field names, source field paths, data types, unit distributions, value distributions, time attributes, source system identifiers, field relationships, historical mapping records, and historical quality scores from the source field profile. It also obtains the standard data element names, standard data element definitions, standard units, standard value ranges, associated treatment event nodes, and data source ranges from the disease data standard model. Subsequently, the candidate mapping generation module generates one or more candidate mapping relationships based on the above information.

[0061] In one possible implementation, candidate mapping relationships can be determined by at least several of the following: semantic similarity of field names, similarity of medical terminology, consistency of units, consistency of value ranges, consistency of time attributes, authority of the source system, historical manual review scores, and historical quality scores. For example, when the source field name is "postoperative Hb", the unit is "g / L", it comes from a laboratory information system, and its recording time is within a preset time window after the surgical event, the system can use this source field as a candidate mapping field for the "postoperative hemoglobin" standard data element.

[0062] In one possible implementation, the mapping confidence score can be obtained through a weighted calculation. Specifically, the system can assign weights to field name semantic similarity, medical terminology similarity, unit consistency, value range consistency, time attribute consistency, source system authority, historical manual review scores, and historical quality scores, and obtain the mapping confidence score based on the weighted result of each score. The weights can be configured according to the target specialty, target disease, standard data element type, or historical review results.

[0063] For example, for test-related standard data elements, unit consistency, value range consistency, and source system authority can have higher weights; for diagnostic standard data elements, medical terminology similarity, coding system consistency, and document source authority can have higher weights; and for time-related standard data elements, time attribute consistency and event time proximity can have higher weights. By configuring different weights according to data element type, mismapping caused by using uniform field matching rules can be avoided.

[0064] In one possible implementation, when the mapping confidence level is greater than or equal to a first confidence threshold, the system automatically confirms the candidate mapping relationship and uses it for subsequent event-driven data collection tasks; when the mapping confidence level is less than or equal to a second confidence threshold, the system rejects the candidate mapping relationship; when the mapping confidence level is between the first and second confidence thresholds, the system sends the candidate mapping relationship to a manual review queue. The manual review results can be written as historical manual review scores into the source field profile or mapping rule base for subsequent mapping confidence level calculations.

[0065] In one possible implementation, the first confidence threshold and the second confidence threshold can be dynamically configured based on the importance of the target specialty, target disease, or standard data element. For example, for data elements used as core indicators for specialty quality control, clinical outcome indicators, or research inclusion criteria, a higher first confidence threshold can be set to reduce the risk of automatic mapping errors; for general auxiliary data elements, a relatively lower threshold can be set to improve automatic data collection efficiency.

[0066] In one possible implementation, the candidate mapping generation module can also perform event semantic verification on the candidate mapping relationship by combining the binding relationship between standard data elements, diagnosis and treatment event nodes, and event collection time windows. That is, the system not only determines whether the source field name or unit matches the standard data element, but also whether the data record time corresponding to the source field falls within the event collection time window corresponding to the target diagnosis and treatment event node. If the source field name is similar, but its time attribute does not match the target diagnosis and treatment event node, the mapping confidence of the corresponding candidate mapping relationship is reduced or it is sent for manual review.

[0067] For example, if a source field name is similar to "post-treatment efficacy evaluation," but the corresponding record time is before the first treatment event, the system can determine that the field does not conform to the semantics of the collection stage of the target standard data element based on the event collection time window, thereby reducing its mapping confidence or rejecting the candidate mapping relationship. In this way, erroneous mappings caused solely by similar field names can be avoided.

[0068] In one possible implementation, the candidate mapping generation module can also utilize historical manual review results to form a feedback mechanism. When manual review confirms that a source field matches a standard data element, the system increases the historical manual review score between the source field profile and the corresponding standard data element; when manual review negates a candidate mapping relationship, the system decreases the corresponding score or records the exclusion rule. Subsequently, when the same or similar source fields appear again in the same hospital, the same hospital area, or a similar source system, the system can call upon historical manual review scores to improve the accuracy and consistency of candidate mapping relationship generation.

[0069] In one possible implementation, candidate mapping relationships can be associated and saved with the source field profile version, the disease data standard model version, and manual review records. If field drift occurs in the source system, such as changes in field name, field path, unit distribution, value distribution, or coding system, the system can recalculate the mapping confidence level and automatically confirm, reject, or trigger manual review based on the recalculation results. This can reduce the risk of erroneous data collection caused by source system upgrades or changes in business processes.

[0070] Through the aforementioned candidate mapping relationships and mapping confidence mechanism, this invention transforms the matching process between source fields and standard data elements from a fixed, manual configuration method to a dynamic judgment method based on field profiling, standard data element semantics, event time semantics, and historical review feedback. This approach not only improves the accuracy and maintainability of mapping relationship generation but also, when combined with the event-driven data collection task generation process, reduces data collection errors caused by semantically similar source fields but inconsistent treatment stages.

[0071] V. Construction of the Patient's Treatment Event Timeline In this embodiment, the patient diagnosis and treatment event timeline construction module is used to integrate and organize the diagnosis and treatment events of the target patient in multiple heterogeneous medical data sources in chronological order to form a complete patient event sequence, providing a basis for event-driven acquisition task generation and dynamic logic judgment.

[0072] In one possible implementation, the module first extracts key medical events for the patient from medical orders, medical records, examination reports, laboratory reports, surgical records, pathology reports, nursing records, expense records, and follow-up records. These key medical events include the event name, event occurrence time, event source, related fields, and event evidence. For example, the surgical name, start time, and end time recorded in the surgical record corresponding to the first surgical event constitute a key medical event node.

[0073] Subsequently, the system sorts the extracted medical events according to their occurrence time, generating a preliminary time series. During the sorting process, for events with overlapping or close times, the system can merge them based on the authority of the source system, the priority of the event type, or the completeness of the fields, forming a unified patient-level medical event node. For example, if the surgery start time is recorded in both the anesthesia system and the medical record homepage, the system can prioritize selecting the data with higher source authority as the event node time.

[0074] In one possible implementation, the constructed patient treatment event timeline can be dynamically updated. When a new data source is accessed or a source field shifts, the system re-parses the newly added or changed fields, updates the patient treatment event timeline, and determines whether to trigger a new event-driven data collection task. Through this dynamic update mechanism, the system can ensure that the data collection task always remains consistent with the actual patient treatment process.

[0075] Furthermore, the patient's treatment event timeline can be bound to standard data elements and event collection time windows in the disease-specific data standard model to achieve event-driven data collection. For example, the system can determine whether a certain standard data element is within the collection time window of the corresponding treatment event based on the patient's event timeline, thereby deciding whether to generate a collection task. This binding not only improves the accuracy of data collection but also reduces the possibility of data from erroneous stages being collected as target data.

[0076] By constructing a timeline of patient treatment events, this invention can organize scattered events in multi-source heterogeneous medical data into an ordered time series. Combined with event collection rules and logical judgments, it can achieve dynamic and accurate event-driven data collection, thereby enhancing the clinical semantic accuracy and traceability of the data.

[0077] VI. Event-Driven Data Acquisition Task Generation and Execution In this embodiment, the event-driven acquisition task generation module dynamically generates event-driven acquisition tasks oriented towards target standard data elements based on the binding relationships in the disease data standard model and the patient's diagnosis and treatment event timeline. Unlike the traditional method of extracting data periodically according to fixed data tables or fixed fields, the acquisition tasks in this embodiment are triggered by the actual diagnosis and treatment events of the patient, and logical judgments are made in combination with conditions such as event acquisition time window, field integrity, mapping confidence, and data source priority.

[0078] In one possible implementation, the system first determines the target diagnostic and treatment event node and the preset event collection time window corresponding to the target standard data element based on the disease data standard model. Then, the system searches for the actual diagnostic and treatment event matching the target diagnostic and treatment event node in the patient's diagnostic and treatment event timeline, and converts the preset event collection time window into an actual collection time window based on the event occurrence time of the actual diagnostic and treatment event. For example, if a standard data element is defined as "the most recent test result within 7 days before the surgical event," and the patient's actual surgical time is a specific point in time, the system generates the corresponding actual collection time window based on that actual surgical time.

[0079] In one possible implementation, the system performs a collection trigger condition judgment before generating the event-driven acquisition task. The collection trigger conditions may include: whether the target diagnostic event node has occurred; whether the current time has reached or exceeded the actual acquisition time window; whether there is collectable data from candidate data sources; whether the mapping confidence between candidate source fields and standard data elements meets preset requirements; whether the field integrity of candidate source fields meets a preset integrity threshold; and whether valid standardized results already exist for the target standard data elements. Only when the collection trigger conditions are met will the system generate and execute the corresponding event-driven acquisition task.

[0080] In one possible implementation, the event-driven data acquisition task includes patient identification, target standard data elements, target diagnostic and treatment event nodes, actual acquisition time window, candidate data sources, candidate source fields, acquisition priority, acquisition triggering conditions, and logical judgment results. By recording the above information in the acquisition task, the subsequent data acquisition process can have clear clinical event semantics and temporal semantics, and provide a basis for standardized evidence chain tracing.

[0081] In one possible implementation, the system determines the collection order based on data source priority rules and the mapping confidence of candidate source fields. For candidate source fields with high source authority, high mapping confidence, and good field completeness, the system prioritizes collection. For candidate source fields with low mapping confidence or insufficient field completeness, the system can postpone collection, mark them as alternative sources, or trigger manual review. This approach reduces the probability of erroneous fields or low-quality data entering the standardized processing flow.

[0082] In one possible implementation, when performing event-driven data acquisition tasks, the system collects candidate disease data from multiple heterogeneous medical data sources within the actual acquisition time window. If multiple records exist in the same candidate source field, the system can filter candidate records based on the time proximity of the record to the target diagnosis and treatment event node, record completeness, and data source priority. For example, for "preoperative test indicators," the system can select the test record closest to the surgical event and that has passed value range verification as candidate disease data within a preset time window before the surgical event.

[0083] In one possible implementation, the system can also dynamically generate supplementary or re-collection tasks based on the acquisition results. When no valid candidate disease data is obtained for the target standard data element within the actual acquisition time window, the system can generate a supplementary collection task based on the range of alternative data sources in the disease data standard model. When the patient's diagnosis and treatment event timeline is updated, the source field profile drifts, the candidate mapping relationship is manually reviewed and modified, or the standardization result fails the quality evaluation, the system can generate a re-collection task to re-execute data acquisition and standardization processing.

[0084] In one possible implementation, event-driven data acquisition tasks can be executed synchronously, asynchronously, or in batches. For standard data elements with high real-time requirements, such as emergency arrival time, thrombolysis start time, or surgery start time, the system can generate acquisition tasks immediately after identifying the relevant diagnostic and treatment events. For follow-up outcomes, efficacy evaluations, or periodic re-examination indicators, the system can generate acquisition tasks at regular intervals based on follow-up time windows or periodic acquisition time windows.

[0085] In one possible implementation, the system can also record the task status during the event-driven data acquisition process. The task status can include pending trigger, triggered, in progress, completed, no data acquired, requiring manual review, requiring supplementary acquisition, and requiring re-acquisition. The task status is associated with and saved with the acquisition triggering conditions, acquisition results, and subsequent processing results, so as to reflect the reason for the acquisition task's generation, execution process, and processing results in a standardized chain of evidence.

[0086] By employing the event-driven data acquisition task generation and execution method described above, this invention can transform the abstract acquisition rules in the standard model of disease data into executable acquisition tasks for specific patients, specific events, and specific time windows, combined with the patient's diagnosis and treatment event timeline. This approach avoids the problem of traditional fixed-field extraction methods failing to distinguish between diagnosis and treatment stages, and can reduce the occurrence of data from erroneous time points, non-target stages, or low-reliability data being mistakenly acquired as standard data, thereby improving the accuracy and stability of multi-source heterogeneous disease data acquisition in key clinical specialties.

[0087] VII. Candidate Value Conflict Resolution and Standard Principal Value Determination In this embodiment, the multi-source conflict resolution module is used to determine the final standard principal value when multiple candidate values ​​exist for the same standard data element. This is achieved by constructing a candidate value conflict graph and performing logical weighted decision-making, while retaining conflict value information to provide a basis for standardizing the evidence chain. Compared with traditional fixed priority or manual selection methods, this invention can realize dynamic processing, logical judgment, and traceable decision-making of multi-source candidate values, fully demonstrating its inventiveness.

[0088] In one possible implementation, the construction of the candidate value conflict graph includes the following steps: 1. Candidate value node construction Establish nodes for each candidate value of the same standard data element collected from various heterogeneous medical data sources.

[0089] Each candidate value node includes the field source, original value, record time, unit, mapping confidence, field completeness, and corresponding diagnosis and treatment event node information.

[0090] 2. Construction of candidate value relationship edges Relationship edges are established between nodes to represent equality, approximation, inclusion, or conflict relationships between candidate values.

[0091] The system can determine the relationship between candidate values ​​by numerical similarity, text matching, encoding consistency, and temporal proximity.

[0092] 3. Principal Value Score Calculation For each candidate value node, a principal value score is calculated. The scoring criteria include, but are not limited to: the authority of the source system, the accuracy of the record time, the completeness of the fields, the proximity of the time to the target diagnosis and treatment event, the historical quality score, and the mapping confidence.

[0093] Different factors can be assigned weights, which can be dynamically adjusted based on the type of standard data element and clinical needs. For example, for surgical time-related data elements, priority should be given to time accuracy, while for laboratory test-related data elements, priority should be given to the authority of the source and the consistency of the institution.

[0094] 4. Logically weighted decision making Based on the relationship between the principal value score and the candidate value, the candidate value with the highest score and that meets the preset confidence conditions is dynamically selected as the standard principal value.

[0095] For candidate values ​​that have approximate or inclusive relationships, a principal value that better conforms to clinical semantics can be selected through weighted logical judgment.

[0096] For candidate values ​​that have conflicting relationships, the system can trigger manual review or automatically mark them as conflicting values, while retaining their information in the standardized chain of evidence.

[0097] 5. Dynamic updates and iterations When the source field profile, patient treatment event timeline, or mapping relationship is updated, the candidate value conflict graph and principal value score can be recalculated.

[0098] The system can automatically generate supplementary or re-sampling tasks to ensure that the standard master values ​​are consistent with the patient's actual medical events and the latest data.

[0099] Through the above-described candidate value conflict resolution and standard principal value determination steps, the present invention can achieve the following effects: Automatic processing of multi-source candidate values: conflict resolution can be completed without relying entirely on manual intervention.

[0100] Logical judgment drives principal value selection: dynamic decision-making is formed by combining authority, time proximity, completeness and mapping confidence, reflecting technological innovation.

[0101] Conflict traceability: The candidate values ​​for which no principal value was selected and the basis for the decision are recorded in a standardized chain of evidence, enabling full-process traceability.

[0102] Dynamic adaptability: The system can adjust the main value judgment logic according to new data or field drift, ensuring long-term stability and reliability of the collected results.

[0103] This module is one of the core innovations of this invention. It directly solves the problems of insufficient handling of multi-source candidate value conflicts, untraceable standardized data, and loss of semantics of clinical events in the prior art, and provides solid technical support for the subsequent execution of event-driven data collection tasks and the generation of evidence chains.

[0104] VIII. Standardization and Chain of Evidence Generation In this embodiment, the standardization conversion module is used to standardize the candidate disease data obtained from the event-driven acquisition task, and the evidence chain tracing module is used to record the processing from the original data to the standardized disease data. Through standardization conversion and evidence chain generation, multi-source heterogeneous medical data can be kept consistent in terms of data format, coding system, unit of measurement, time expression, and clinical semantics, and the final results can be verified, interpreted, and traceable.

[0105] In one possible implementation, the standardization conversion module receives candidate disease data output from an event-driven acquisition task. The candidate disease data may include structured data, semi-structured data, and unstructured text data. Structured data may originate from laboratory systems, surgical anesthesia systems, pharmacy systems, or disease-specific databases; semi-structured data may originate from examination reports, pathology reports, or follow-up forms; and unstructured text data may originate from medical records, discharge summaries, nursing records, or outpatient medical records.

[0106] In one possible implementation, the standardization transformation includes at least several of the following: encoding mapping, unit conversion, time format standardization, missing value annotation, value range validation, and structured extraction. Encoding mapping is used to convert local codes from different medical data sources into unified standard codes; unit conversion is used to convert different units of measurement for the same indicator in different systems into standard units; time format standardization is used to convert different formats of record time, event time, or report time into a unified time format; missing value annotation is used to distinguish different reasons for missing values, such as not collected, not recorded, inapplicable, and pending review; value range validation is used to determine whether candidate values ​​are within a preset reasonable range; and structured extraction is used to extract information such as disease diagnosis, symptoms and signs, examination conclusions, test indicators, surgical names, medication regimens, pathological classifications, staging and grading, rating scales, and follow-up outcomes from medical texts.

[0107] In one possible implementation, the standardization conversion module, during the conversion process, also verifies the target diagnostic and treatment event node corresponding to the target standard data element and the actual collection time window. If candidate disease data satisfies the field mapping relationship, but its recording time is not within the actual collection time window, or its event semantics are inconsistent with the target diagnostic and treatment event node, the system can mark the candidate disease data as not meeting the collection conditions, an alternative value, or a value requiring manual review. This approach avoids misclassifying data from incorrect diagnostic and treatment stages as standardized disease data simply because of similar field names or consistent data formats.

[0108] In one possible implementation, the standardization conversion module can also perform standardization processing based on the version of the disease data standard model. Different target specialties or target diseases can correspond to different data element definitions, coding rules, unit conversion rules, and value range verification rules. When the disease data standard model is updated, the system can retain the standardization results corresponding to the old version model, while re-executing the conversion or generating a re-collection task based on the new version model, to ensure that historical results are traceable and new results are updatable.

[0109] In one possible implementation, after the standardization transformation is completed, the system generates standardized disease data. The standardized disease data includes at least a standard data element identifier, a standard data element name, a standard value, a standard unit, a target diagnostic event node, an actual data acquisition time window, a data source identifier, and a generation time. For standard data elements with multiple candidate values, the standard value in the standardized disease data can be the standard master value determined by the multi-source conflict resolution module.

[0110] In this embodiment, the evidence chain tracing module is used to generate a standardized evidence chain corresponding to the standardized disease data. The standardized evidence chain is used to record the entire process information of the standardized disease data from raw data acquisition, field mapping, event time window judgment, standardization transformation, conflict resolution to final output.

[0111] In one possible implementation, the standardized chain of evidence includes at least several of the following: original data source, source field path, original value, original unit, original record time, standard value, standard unit, standardized conversion rules, disease data standard model version, source field profile version, mapping confidence, target diagnosis and treatment event node, actual collection time window, collection triggering conditions, logical judgment results, conflict candidate value information, standard principal value determination basis, manual review records, and generation time.

[0112] Among them, the original data source is used to identify which heterogeneous medical data source the candidate disease data comes from; the source field path is used to identify the table field, interface field, document template field or text extraction location of the data in the source system; the original value and standard value are used to record the data content before and after the transformation; the standardization transformation rules are used to record the rules on which encoding mapping, unit conversion, time standardization or text extraction are based; the disease data standard model version is used to record the model version applicable during standardization processing; the mapping confidence is used to record the degree of credibility of the matching between the source field and the standard data element; and the target diagnosis and treatment event node and the actual collection time window are used to record the clinical event semantics and temporal semantics of data collection.

[0113] In one possible implementation, when multiple candidate values ​​exist for the same standard data element, the standardized evidence chain also records the candidate value conflict graph identifier, candidate value node information, candidate value relationship edge information, principal value score of each candidate value node, the basis for determining the standard principal value, and conflicting or alternative values ​​that were not selected as the standard principal value. Therefore, even if a single standard principal value is ultimately output, the system still retains other candidate values ​​and the reasons why they were not adopted, facilitating subsequent review and auditing.

[0114] In one possible implementation, the evidence chain tracing module also records the collection trigger conditions and logical judgment results. For example, the system can record whether the target diagnosis and treatment event node has occurred, whether the collection time window has been reached, whether the candidate source field exists, whether the field integrity meets the requirements, whether the mapping confidence level reaches the threshold, whether manual review is triggered, and whether a supplementary collection task or re-collection task is generated. By recording the above logical judgment results, it can be shown that standardized disease data is not simply extracted, but is generated based on event semantics, time windows, and multi-condition judgments.

[0115] In one possible implementation, the standardized chain of evidence can be stored together with the standardized disease data, or it can be stored independently in an evidence chain database, log database, or audit database. The standardized chain of evidence can be associated with the standardized disease data through standard data element identifiers, patient identifiers, event task identifiers, or standardized data identifiers. When a user views a piece of standardized disease data, the system can display its original source, processing rules, conflict resolution process, and manual review records based on the association relationships.

[0116] In one possible implementation, the standardized chain of evidence can also be used for data quality assessment, error localization, and review and tracing. When an anomaly is found in standardized disease data, the system can trace the original data source, source field path, original value, transformation rules, event collection time window, and principal value decision basis according to the chain of evidence, thereby locating the cause of the anomaly. If the cause of the anomaly is field drift, mapping error, or deterioration of source data quality, the system can further trigger remapping, manual review, supplementary data collection, or re-collection tasks.

[0117] Through the aforementioned standardization transformation and evidence chain generation methods, this invention can achieve both format and semantic unification of multi-source heterogeneous disease data while fully preserving the source, rules, time windows, mapping confidence levels, conflict resolution, and manual review basis from the standardization process. This approach enhances the credibility of standardized disease data in clinical key specialty quality control, scientific research analysis, follow-up management, and disease-specific database construction, and strengthens the system's interpretability in scenarios involving data verification, quality evaluation, and accountability.

[0118] IX. Data Quality Assessment and Field Drift Monitoring In this embodiment, data quality evaluation and field drift monitoring are used to continuously monitor the quality status of standardized disease data and the changes in fields of heterogeneous medical data sources, so as to ensure the stability and accuracy of event-driven acquisition tasks, candidate mapping relationships and standardization results in the long-term operation process.

[0119] In one possible implementation, the system performs a data quality assessment on standardized disease data. This data quality assessment can generate a data quality score based on at least several of the following: completeness rate, consistency rate, timeliness rate, outlier ratio, mapping confidence, traceability rate, and manual review pass rate. Specifically, the completeness rate represents the effective collection ratio of the target standardized data element within the target patient population; the consistency rate represents the degree of consistency among multiple candidate values ​​of the same standardized data element; the timeliness rate represents whether the data was collected within the actual collection time window corresponding to the target diagnostic event node; the outlier ratio represents the proportion of data that failed value range validation or logical validation; and the traceability rate represents whether the standardized disease data has a complete standardized evidence chain.

[0120] In one possible implementation, the system can assign different data quality evaluation weights to different standard data elements. For example, for core indicators of quality control in key clinical specialties, completeness, timeliness, and traceability can have higher weights; for laboratory indicators, consistency, outlier ratio, and unit conversion accuracy can have higher weights; and for follow-up outcome indicators, follow-up time window matching and evidence chain completeness can have higher weights. By setting evaluation weights according to data element type, data quality scoring can better meet the actual requirements of different clinical business scenarios.

[0121] In one possible implementation, when the data quality score is lower than a preset quality threshold, the system triggers corresponding processing tasks based on the type of quality anomaly. If the completeness rate is lower than the threshold, the system can trigger a supplementary data collection task; if the consistency rate is lower than the threshold, the system can trigger a candidate value conflict review task; if the timeliness rate is lower than the threshold, the system can re-verify the patient's treatment event timeline and the actual data collection time window; if the proportion of outliers is higher than the threshold, the system can trigger a value range rule review or a unit conversion rule review; if the traceability rate is lower than the threshold, the system can complete the standardized evidence chain or trigger manual review.

[0122] In one possible implementation, the data quality assessment results can be incorporated into a standardized chain of evidence. These results may include a data quality score, assessment time, assessment indicators, anomaly type, triggered processing tasks, and processing results. In this way, the system can not only output standardized disease data but also record the changes in the data's status during the quality assessment process, providing a basis for subsequent review, auditing, and accountability.

[0123] In this embodiment, field drift monitoring is used to detect whether changes occur in the source fields of heterogeneous medical data sources that affect the accuracy of data collection and mapping. Since hospital business systems may experience changes in field names, field paths, data types, units, value ranges, or encoding systems due to version upgrades, vendor modifications, table structure adjustments, interface changes, or business process modifications, continuing to use the original mapping relationships and collection rules may lead to erroneous or missed data collection. Therefore, this invention improves the system's adaptability to changes in the source system through field drift monitoring.

[0124] In one possible implementation, the system periodically acquires the current field profile of the source field and compares the current field profile with historical field profiles to obtain the field drift degree. The field drift degree can be determined based on at least several of the following: changes in source field name, changes in source field path, changes in data type, changes in unit distribution, changes in value distribution, changes in null value rate, changes in encoding system, changes in time attribute, and changes in field association relationship.

[0125] In one possible implementation, the system can assign different weights to different drift factors. For example, changes in field paths and data types may directly affect the data acquisition path and parsing method, so they can be assigned higher weights; changes in unit distribution and value distribution may affect standardization transformation and value range verification, so they can also be assigned higher weights; when field names change slightly but the field path, unit, and value distribution remain stable, a lower drift impact can be assigned. Through weighted calculation, frequent manual reviews can be avoided due to field name adjustments that have no substantial impact, while key changes affecting the accuracy of data acquisition can be identified in a timely manner.

[0126] In one possible implementation, when the field drift exceeds a preset drift threshold, the system pauses or marks the automatic data collection task related to that source field and triggers remapping or manual review. If the review confirms that the field can still match the original standard data element, the system updates the source field profile version and candidate mapping relationship; if the review confirms that the field semantics have changed, the system rejects the original mapping relationship and regenerates candidate mapping relationships or adjusts the data source rules in the disease data standard model.

[0127] In one possible implementation, field drift monitoring can also be linked to event-driven data acquisition tasks. When a source field drifts and is used in an incomplete or periodically triggered event-driven data acquisition task, the system can update the corresponding task status to pending review or paused execution. After remapping is completed, the system regenerates or continues executing the data acquisition task based on the new mapping relationship. This avoids erroneous data from entering the standardization transformation process during field changes.

[0128] In one possible implementation, field drift monitoring results can also be written into the standardized chain of evidence or field profile version record. The field drift monitoring results may include the current field profile version, historical field profile versions, drift factors, field drift degree, drift determination result, remapping result, manual review result, and update time. By recording the above information, the system can explain the field status and mapping version upon which standardized disease data is based, thereby enhancing the interpretability of the data results.

[0129] Through the aforementioned data quality evaluation and field drift monitoring, this invention enables closed-loop management of data quality and source field stability after data acquisition and standardization. This mechanism allows the system to not only complete one-time data extraction and transformation, but also to perform supplementary collection, re-collection, remapping, or manual review when the source system changes, data quality deteriorates, or the event timeline is updated, thereby improving the continuous maintainability and long-term reliability of standardized data acquisition for multi-source heterogeneous diseases in key clinical specialties.

[0130] 10. Application Examples and Implementation Results (I) Examples of Key Specialty Applications in Stroke In one possible implementation, the target key clinical specialty is stroke. The system constructs a standard data model for stroke, in which the diagnostic and treatment event nodes include onset events, hospitalization events, first imaging examination events, thrombolysis initiation events, thrombectomy initiation events, hospitalization events, discharge events, and follow-up events.

[0131] The corresponding standard data elements may include onset time, time of arrival at hospital, time of first imaging examination, NIHSS score, thrombolytic drugs, time of thrombolysis start, thrombectomy information, discharge mRS score, 90-day follow-up mRS score, mortality outcome, and readmission information.

[0132] The system identifies patient treatment events from emergency medical records, medical order systems, imaging examination reports, neurology progress notes, nursing records, and follow-up systems, and constructs a timeline of these events. For example, when the system identifies a patient's arrival event, it can generate an event-driven data collection task for "the first imaging examination result within a preset time after the arrival event" based on the disease data standard model; when the system identifies a thrombolysis initiation event, it can generate an event-driven data collection task for "the thrombolysis initiation event in relation to medication orders and administration time"; and when the system identifies a discharge event, it can generate an event-driven data collection task for "follow-up outcomes within 90±14 days after discharge".

[0133] When the NIHSS score exists simultaneously in emergency medical records, neurology progress notes, and disease follow-up forms, the system constructs a candidate value conflict graph and determines the standard principal value based on the authority of the candidate value source, the time proximity of the recording time to the target diagnosis and treatment event, field completeness, and historical quality score. At the same time, unselected candidate values ​​are written into the standardized evidence chain as alternative values ​​or conflict values.

[0134] This method can avoid mistaking scores from non-target stages as target scores, thereby improving the accuracy of quality control indicators and research analysis data for key stroke specialties.

[0135] (II) Examples of Application of Key Oncology Specialties In another possible implementation, the target clinical specialty is oncology, and the target disease is lung cancer. The system constructs a standard data model for lung cancer, where the diagnostic and treatment event nodes include initial diagnosis events, imaging abnormality events, pathological confirmation events, molecular testing events, initial treatment events, efficacy evaluation events, and follow-up events.

[0136] The corresponding standard data elements may include pathological type, TNM stage, gene mutation status, initial treatment regimen, surgical method, radiotherapy and chemotherapy regimen, targeted therapy drugs, immunotherapy drugs, efficacy evaluation results, adverse reactions, recurrence and metastasis status, and survival outcome.

[0137] The system collects pathological types and molecular test results for pathological diagnosis events, treatment plans for initial treatment events, imaging reports and efficacy evaluation results for efficacy evaluation events, and survival status and recurrence / metastasis for follow-up events.

[0138] When TNM staging exists simultaneously in medical records, imaging reports, pathology reports, and disease-specific databases, the system constructs a candidate value conflict graph and determines the standard principal value based on data source authority, event time proximity, text extraction confidence, field completeness, and historical quality score. For conflicting candidate values, the system records their source, original value, conflict relationship, and the reason for not being selected as the principal value, and writes this information into a standardized chain of evidence.

[0139] This approach can reduce staging errors caused by inconsistencies in multi-source records and improve the credibility of oncology data in research cohort construction, efficacy evaluation, and follow-up analysis.

[0140] (III) Implementation Results Through the above embodiments, the present invention can achieve at least the following technical effects: First, by establishing a binding relationship between standard data elements, diagnostic and treatment event nodes, and event collection time windows, the data collection process has clear clinical event semantics and temporal semantics, reducing the problem of data being mistakenly collected in non-target diagnostic and treatment stages.

[0141] Second, by using the patient's diagnosis and treatment event timeline and collection trigger conditions, event-driven collection tasks can be dynamically generated based on the patient's actual diagnosis and treatment process, thereby improving the accuracy and adaptability of disease data collection.

[0142] Third, by profiling source fields, candidate mapping relationships, and mapping confidence, the reliance on fixed field mapping and manual rule maintenance can be reduced, thereby improving the system's adaptability to multi-source heterogeneous data environments.

[0143] Fourth, through the candidate value conflict graph and principal value scoring mechanism, it is possible to identify relationships, resolve conflicts, and determine the principal value for multiple candidate values ​​of the same standard data element, thereby improving the consistency and reliability of standardized disease data.

[0144] Fifth, by recording the original source, source field path, transformation rules, event time window, mapping confidence, conflict decision and manual review results through a standardized chain of evidence, the entire process of standardization can be traced, which facilitates data verification, quality control and accountability.

[0145] Sixth, through data quality evaluation and field drift monitoring, it is possible to trigger supplementary data collection, re-collection, remapping, or manual review when the source system fields change, the mapping relationship changes, or the data quality deteriorates, thereby improving the long-term stability and maintainability of the system.

[0146] In summary, this invention does not simply extract, clean, and summarize multi-source medical data. Instead, it forms a closed-loop processing mechanism based on a disease-specific data standard model, a patient diagnosis and treatment event timeline, an event collection time window, dynamic logical judgment, multi-source conflict resolution, and a standardized evidence chain. This improves the accuracy, completeness, maintainability, and traceability of standardized collection of multi-source heterogeneous disease data in key clinical specialties.

[0147] In summary, this invention achieves dynamic standardized acquisition and full-process traceability of multi-source heterogeneous disease data by constructing a standard model for disease data oriented towards key clinical specialties, generating a timeline of patient diagnosis and treatment events, establishing binding relationships between event acquisition time windows, executing event-driven acquisition tasks, resolving conflicts between multi-source candidate values, and generating a standardized evidence chain. This method and system effectively improve the accuracy, completeness, maintainability, and traceability of disease data, providing a reliable data foundation for the construction of key clinical specialties, scientific research analysis, quality control, and follow-up management. It overcomes the shortcomings of existing technologies, such as unclear temporal semantics in multi-source heterogeneous data acquisition, insufficient conflict handling, and lack of traceability for standardized results.

[0148] The embodiments of the present invention are described exemplarily. Those skilled in the art will understand that various equivalent substitutions, combinations, or improvements can be made to the above embodiments without departing from the spirit and scope of the present invention, including adjusting the order of modules, decomposing functions, optimizing data processing methods, or changing system deployment methods, all of which should fall within the protection scope of the present invention. The scope of protection defined in the claims should be determined by the claims and their equivalent technical solutions. The content of this specification is only used to illustrate the principles, concepts, and implementation methods of the present invention and does not limit the scope of protection of the present invention.

Claims

1. A standardized collection method for multi-source heterogeneous disease data in key clinical specialties, characterized in that, include: Construct a standard model of disease data corresponding to the target key clinical specialty. The standard model of disease data includes a set of standard data elements, a set of diagnosis and treatment event nodes, and an event collection time window. Establish a binding relationship between standard data elements, diagnostic and treatment event nodes, and event collection time windows. The binding relationship is used to limit the collection stage and collection time range of the target standard data elements relative to the target diagnostic and treatment event nodes. Access multiple heterogeneous medical data sources related to the target patient, and identify key diagnostic and treatment events of the target patient from the heterogeneous medical data sources; Based on the occurrence time of the key diagnostic and treatment events, a timeline of patient diagnostic and treatment events is constructed; Based on the binding relationship and the patient diagnosis and treatment event timeline, dynamically determine whether the target standard data element meets the collection trigger conditions, including whether the event has occurred, field integrity, historical mapping consistency, and whether the collection time window has been reached; When the acquisition triggering condition is met, an event-driven acquisition task is generated and executed to acquire candidate disease data located within the actual acquisition time window from the multiple heterogeneous medical data sources. The candidate disease data is standardized and transformed, and standardized disease data is generated based on the results of encoding mapping, unit conversion, time format standardization, missing value annotation, value range verification, and structured extraction. When there are multiple candidate values ​​for the same standard data element, a candidate value conflict graph is constructed. Logical calculations and weighted decisions are performed based on the authority of the candidate value source, time accuracy, completeness, time proximity to the target diagnosis and treatment event node, and historical quality scores to determine the standard principal value. Candidate values ​​that are not selected as principal values ​​are saved as conflict values. A standardized evidence chain corresponding to the standardized disease data is generated, and the triggering logic, collection conditions, conflict decisions and manual review information are recorded to achieve full-process traceability.

2. The method according to claim 1, characterized in that, The binding relationship is also used to limit the data source range, collection priority and data verification rules of different standard data elements under different diagnosis and treatment event nodes.

3. The method according to claim 1, characterized in that, The event acquisition time window includes at least a number of the following: pre-event time window, post-event time window, event proximity time window, periodic acquisition time window, and follow-up time window, and can be dynamically adjusted based on the patient's medical event status.

4. The method according to claim 1, characterized in that, After accessing heterogeneous medical data sources, source field profiles are generated, including source field name, source field path, data type, unit distribution, value distribution, time attribute, source system identifier, field association, null value rate, historical mapping records, and historical quality score, which are used for subsequent logical judgment and trigger condition evaluation.

5. The method according to claim 4, characterized in that, Candidate mapping relationships are generated based on the source field profile and the standard data element set. The mapping confidence is judged by logical weighting. When the mapping confidence is lower than a preset threshold, it automatically enters the manual review queue. When the mapping confidence meets the threshold, the candidate mapping relationship is automatically confirmed.

6. The method according to claim 1, characterized in that, Constructing a patient treatment event timeline includes: extracting the event name, event occurrence time, event source, and event evidence; merging identical or similar events; and dynamically updating the event timeline when event times change to trigger a re-collection task.

7. The method according to claim 1, characterized in that, The event-driven acquisition task includes patient identifier, target standard data element, target diagnosis and treatment event node, actual acquisition time window, candidate data source, candidate source field, acquisition priority, acquisition triggering condition and logical judgment relationship, which are used to dynamically control the acquisition order and repeated acquisition.

8. The method according to claim 1, characterized in that, The steps for determining the standard principal value based on the candidate value conflict graph include: constructing candidate value nodes and relation edges to represent the equality, approximation, inclusion, or conflict relationships between candidate values; calculating the principal value score of each node using weighted logic; and dynamically selecting the principal value based on the score results and preset logic rules.

9. The method according to claim 1, characterized in that, The standardized evidence chain records the original data source, source field path, original value, standard value, standardization transformation rules, disease data standard model version, mapping confidence, target diagnosis and treatment event node, actual collection time window, conflict candidate value information, logical judgment results, and manual review records, realizing dynamic traceability of the entire process.

10. A standardized data acquisition system for multi-source heterogeneous diseases in key clinical specialties, characterized in that, include: The disease data standard model construction module is used to build a disease data standard model and establish the binding relationship between standard data elements, diagnosis and treatment event nodes and event collection time windows; Multi-source data access module, used to access multiple heterogeneous medical data sources; The treatment event timeline construction module is used to identify key treatment events and construct a patient treatment event timeline, dynamically updating the patient event status; The event-driven acquisition task generation module is used to dynamically generate event-driven acquisition tasks based on logical judgments of the binding relationship and the patient's diagnosis and treatment event timeline. The data acquisition module is used to execute event-driven acquisition tasks and obtain candidate disease data; The standardization transformation module is used to standardize candidate disease data; The multi-source conflict resolution module is used to construct a candidate value conflict graph and select the standard principal value based on logical weighting. The evidence chain tracing module is used to generate standardized evidence chains and record logical judgments and triggering conditions.