Multi-source data fusion management method and system for medical record information
By constructing a medical ontology knowledge graph for multi-source data fusion management, the problem of insufficient semantic relevance in medical record information is solved, and efficient integration and dynamic analysis of multi-source data are achieved, improving the quality and efficiency of medical services and supporting precision medicine and personalized treatment.
Patent Information
- Application Number
- CN202510612121.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Existing medical record management methods ignore the semantic relationships between data, limiting the ability to deeply understand and mine medical record information, especially in the process of diagnosing and treating complex diseases, making it difficult to provide comprehensive information support.
By constructing a multi-source data fusion management method based on medical ontology knowledge graph, including data collection, semantic mapping, concept association, temporal feature extraction and heterogeneous data fusion, a unified medical record data view is generated.
It enables efficient integration and dynamic analysis of multi-source data, provides comprehensive medical record information support, improves the quality and efficiency of medical services, and supports precision medicine and personalized treatment.
Smart Images

Figure CN120452824B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical technology, and in particular to a method and system for multi-source data fusion management of medical record information. Background Technology
[0002] In the modern medical environment, with the development and application of information technology, medical institutions have accumulated a large amount of medical record data. This data comes from different systems, such as electronic medical record systems, medical imaging workstations, and laboratory information systems, each with its unique data structure and format. This multi-source, heterogeneous data environment poses a significant challenge to the effective management and utilization of medical record information. On the one hand, due to the lack of unified standards and specifications, data from different sources is difficult to integrate effectively, leading to the emergence of data silos. On the other hand, traditional data processing methods often focus on static data analysis, failing to meet the needs of dynamic analysis of medical record data.
[0003] Furthermore, existing medical record management methods often neglect the semantic relationships between data, which limits the ability to deeply understand and mine medical record information. Especially when facing complex disease diagnosis and treatment processes, relying solely on a single type of data is insufficient to provide comprehensive information support. For example, when formulating a treatment plan, doctors may need to consider multiple aspects of information, such as the patient's clinical symptoms, test results, and imaging findings. This information is scattered across different systems, making efficient integration and sharing difficult. Therefore, establishing a method that can effectively integrate medical record data from multiple sources has become an important research topic.
[0004] To address the aforementioned issues, a multi-source data fusion management method for medical record information based on a medical ontology knowledge graph is proposed. This method not only focuses on data collection and integration but also emphasizes semantic relationships and dynamic feature analysis between data. By constructing a unified view of medical record data, it can help medical staff obtain relevant patient information more quickly and accurately, improving the quality and efficiency of medical services. Simultaneously, this method also provides strong data support for further research in cutting-edge fields such as precision medicine and personalized treatment. However, achieving this goal still requires overcoming challenges such as data privacy protection and cross-institutional data sharing. Summary of the Invention
[0005] The main objective of this invention is to provide a multi-source data fusion management method and system for medical record information, which solves the technical problem that existing medical record management methods usually ignore the semantic correlation between data, which limits the ability to deeply understand and mine medical record information.
[0006] To achieve the above objectives, the present invention provides a multi-source data fusion management method for medical record information, comprising the following steps:
[0007] Data is collected from the electronic medical record system, medical imaging workstation, and laboratory information system of medical institutions to obtain the original medical record dataset;
[0008] Based on a pre-defined medical ontology knowledge graph, semantic mapping and concept association are performed on the original medical record dataset to obtain a semantically associated medical record information network.
[0009] Dynamic temporal feature analysis is performed on the semantically associated medical record information network using a temporal feature extractor to obtain a temporal medical record feature sequence.
[0010] Heterogeneous data fusion is performed on the time-series medical record feature sequences to obtain a unified medical record data view.
[0011] Furthermore, the semantic mapping and concept association of the original medical record dataset based on a preset medical ontology knowledge graph to obtain a semantically associated medical record information network includes:
[0012] The original medical record dataset is segmented using a pre-defined medical terminology parser to obtain a sequence of medical terms. The sequence of medical terms is then semantically standardized and mapped based on a pre-defined medical ontology knowledge graph to obtain a standardized set of medical concepts.
[0013] The standardized medical concept set is analyzed for association using a semantic network builder to obtain a concept relationship graph. The concept relationship graph is then hierarchically organized to obtain a hierarchical semantic network.
[0014] The hierarchical semantic network is analyzed by a semantic reasoning engine to obtain a set of semantic reasoning rules. Based on the set of semantic reasoning rules, knowledge is expanded to obtain an extended semantic relationship network.
[0015] Based on a preset ontology fusion processor, the extended semantic relationship network is conceptually aligned and knowledge is integrated to obtain a semantically related medical record information network.
[0016] Furthermore, the association analysis of the standardized medical concept set based on the semantic network builder to obtain a concept relationship graph includes:
[0017] The standardized medical concept set is subjected to co-occurrence analysis by a concept relationship mining tool to obtain a concept co-occurrence matrix. The association degree of the concept co-occurrence matrix is then calculated to obtain a concept association strength map.
[0018] Path mining is performed on the concept association strength graph to obtain a set of concept semantic paths, and weights are assigned to the set of concept semantic paths to obtain a weighted semantic path network.
[0019] The weighted semantic path network is topologically optimized based on the semantic network builder to obtain an optimized semantic structure graph, and the optimized semantic structure graph is hierarchically divided to obtain a hierarchical semantic structure.
[0020] The hierarchical semantic structure is constructed using a knowledge graph generator to obtain a concept relationship graph.
[0021] Furthermore, the step of performing dynamic temporal feature analysis on the semantically associated medical record information network using a temporal feature extractor to obtain a temporal medical record feature sequence includes:
[0022] The semantically associated medical record information network is subjected to trend component extraction by a temporal feature extractor to obtain disease progression trend features, and the disease progression trend features are divided into disease stages to obtain staged disease progression features.
[0023] Based on the phased disease progression features, periodic pattern mining is performed on the semantically associated medical record information network to obtain disease recurrence cycle features, and the fluctuation amplitude of the disease recurrence cycle features is quantified to obtain quantified disease recurrence features.
[0024] Based on the quantitative disease recurrence characteristics, outlier detection is performed on the semantically associated medical record information network to obtain a sequence of abnormal disease events. Then, event correlation analysis is performed on the sequence of abnormal disease events to obtain key event characteristics of the disease.
[0025] Based on the key event characteristics of the disease, temporal logical reasoning and disease recording are performed on the abnormal event sequence of the disease to obtain a temporal medical record feature sequence.
[0026] Furthermore, the periodic pattern mining of the semantically associated medical record information network based on the phased disease progression features to obtain disease recurrence cycle features includes:
[0027] Based on the staged disease progression features, the semantically associated medical record information network is segmented into subsequences to obtain disease progression subsequences, and the similarity of the disease progression subsequences is measured to obtain a subsequence similarity matrix.
[0028] Cluster analysis is performed on the disease progression subsequences based on the subsequence similarity matrix to obtain disease progression pattern clusters, and pattern features are extracted from the disease progression pattern clusters to obtain disease progression pattern feature vectors.
[0029] Based on the feature vector of the disease progression pattern, periodic pattern matching is performed on the disease progression pattern cluster to obtain a candidate relapse cycle set, and statistical significance test is performed on the candidate relapse cycle set to obtain a significant relapse cycle set.
[0030] Based on the significant recurrence cycle set, the cycle length of the candidate recurrence cycle set is estimated to obtain the recurrence cycle length value, and the confidence interval of the recurrence cycle length value is calculated to obtain the recurrence cycle confidence interval.
[0031] Based on the confidence interval of the recurrence cycle, the set of significant recurrence cycles is periodically integrated to obtain the disease recurrence cycle characteristics.
[0032] Furthermore, the step of performing temporal logical reasoning and disease recording on the abnormal event sequence based on the key event characteristics of the disease to obtain a temporal medical record feature sequence includes:
[0033] Based on the key event characteristics of the disease, an event causal association analysis is performed on the abnormal event sequence of the disease to obtain an event causal association graph. Then, the causal strength of the event causal association graph is calculated to obtain a weighted causal association network.
[0034] Based on the weighted causal association network, temporal constraint mining is performed on the abnormal event sequence of the disease to obtain a temporal constraint rule set, and the constraint strength of the temporal constraint rule set is evaluated to obtain an effective temporal constraint network.
[0035] Based on the effective temporal constraint network, temporal path reasoning is performed on the weighted causal association network to obtain a set of candidate disease development paths, and path probability evaluation is performed on the set of candidate disease development paths to obtain a probabilistic development path graph.
[0036] Based on the probabilistic development path diagram, the clinical trajectory of the abnormal disease event sequence is reconstructed and the disease is recorded to obtain a time-series medical record feature sequence.
[0037] Furthermore, the step of fusing heterogeneous data from the time-series medical record feature sequences to obtain a unified medical record data view includes:
[0038] The time-series medical record feature sequence is aligned in the spatiotemporal dimension by a spatiotemporal feature aligner to obtain a spatiotemporal aligned feature set, and features are extracted from the spatiotemporal aligned feature set to obtain a multi-dimensional feature vector;
[0039] The multi-dimensional feature vectors are transformed into a unified feature space, and feature association calculations are performed on the unified feature space to obtain a feature association network.
[0040] The feature association network is fused using a multimodal feature fusion processor to obtain a fused feature map. The fused feature map is then optimized to obtain an optimized feature map. A knowledge graph is constructed based on the optimized feature map to obtain a unified medical record data view.
[0041] This invention also provides a multi-source data fusion management system for medical record information, comprising:
[0042] The data acquisition module is used to collect data from the electronic medical record system, medical imaging workstation and laboratory information system of medical institutions to obtain the original medical record dataset;
[0043] The association module is used to perform semantic mapping and concept association on the original medical record dataset based on a preset medical ontology knowledge graph to obtain a semantically associated medical record information network.
[0044] The analysis module is used to perform dynamic temporal feature analysis on the semantically associated medical record information network through a temporal feature extractor to obtain a temporal medical record feature sequence;
[0045] The fusion module is used to perform heterogeneous data fusion on the time-series medical record feature sequences to obtain a unified medical record data view.
[0046] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the methods described above.
[0047] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the methods described above.
[0048] This invention provides a multi-source data fusion management method for medical record information, comprising the following steps: collecting data from the electronic medical record system, medical imaging workstation, and laboratory information system of a medical institution to obtain an original medical record dataset; performing semantic mapping and concept association on the original medical record dataset based on a preset medical ontology knowledge graph to obtain a semantically related medical record information network; performing dynamic temporal feature analysis on the semantically related medical record information network using a temporal feature extractor to obtain a temporal medical record feature sequence; and fusing heterogeneous data on the temporal medical record feature sequence to obtain a unified medical record data view. This method solves the technical problem that existing medical record management methods often ignore the semantic relationships between data, which limits the ability to deeply understand and mine medical record information. It achieves a unified medical record data view by fusing heterogeneous data from temporal medical record feature sequences, enabling data from different sources and with different formats to be effectively organized and displayed within the same framework. This not only facilitates access and use by medical staff but also improves the performance of decision support systems. Attached Figure Description
[0049] Figure 1 This is a schematic diagram illustrating the steps of a multi-source data fusion management method for medical record information in one embodiment of the present invention;
[0050] Figure 2 This is a structural block diagram of a multi-source data fusion management system for medical record information according to an embodiment of the present invention;
[0051] Figure 3 This is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.
[0052] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0054] like Figure 1 As shown, Figure 1 This is a schematic diagram illustrating the steps of a multi-source data fusion management method for medical record information in one embodiment of the present invention;
[0055] One embodiment of the present invention provides a method for multi-source data fusion management of medical record information, comprising the following steps:
[0056] Step S1: Collect data from the electronic medical record system, medical imaging workstation, and laboratory information system of the medical institution to obtain the original medical record dataset.
[0057] Specifically, in the aforementioned scheme, data is collected from the electronic medical record system, medical imaging workstation, and laboratory information system of medical institutions to obtain the original medical record dataset. This process forms the foundation of the entire multi-source data fusion management method. Specifically, the electronic medical record system stores structured or unstructured text data such as patients' clinical records, diagnostic information, and treatment plans; the medical imaging workstation stores patients' imaging examination results, such as X-ray films, CT scans, and MRI data, which are typically in image format; and the laboratory information system records patients' laboratory test results, such as numerical data like blood indicators and urine analysis. To achieve this data collection step, it is first necessary to interface with each system through standardized interfaces or protocols (such as HL7, FHIR, etc.) to ensure data can be obtained from different sources while guaranteeing data integrity and consistency. For example, in an application scenario of a large general hospital, a doctor completes a comprehensive physical examination for a patient, including medical history records, imaging examinations, and laboratory tests. At this time, the data collection module will extract the patient's medical history and diagnostic information from the electronic medical record system, obtain relevant imaging data from the medical imaging workstation, and read the laboratory test results from the laboratory information system. After initial processing, these data were integrated into a raw medical record dataset, serving as the foundation for subsequent semantic mapping and dynamic analysis. However, in practice, issues such as inconsistent data formats and varying data quality may arise. Therefore, data cleaning techniques are needed to preprocess the collected data, such as removing duplicate records, filling in missing values, and standardizing timestamp formats, to ensure that the quality of the raw medical record dataset meets the needs of subsequent analysis. This process not only achieves efficient collection of multi-source data but also lays a solid foundation for subsequent semantic mapping and concept association based on medical ontology knowledge graphs.
[0058] Step S2: Based on the preset medical ontology knowledge graph, perform semantic mapping and concept association on the original medical record dataset to obtain a semantically associated medical record information network.
[0059] Specifically, in the above scheme, semantic mapping and conceptual association are performed on the original medical record dataset based on a pre-defined medical ontology knowledge graph to obtain a semantically associated medical record information network. This process is a key step in achieving deep integration of multi-source data. Specifically, a medical ontology knowledge graph is a structured form of knowledge representation that includes core concepts, terms, and relationships within the medical field, such as the association between diseases and symptoms, and the connection between laboratory indicators and pathological states. To achieve semantic mapping, it is first necessary to match various types of data in the original medical record dataset (including text records in electronic medical record systems, image descriptions in medical imaging workstations, and numerical results in laboratory information systems) with standardized terms in the medical ontology knowledge graph. For example, in a large hospital application scenario, a patient's electronic medical record shows that they have "hypertension," while the laboratory information system records that the patient has "elevated serum creatinine levels." The medical ontology knowledge graph can map this scattered information into a unified conceptual system, identifying the potential association between "hypertension" and "abnormal kidney function." Simultaneously, semantic mapping also needs to address the heterogeneity issues between different data sources, such as transforming unstructured text data into structured representations or mapping descriptive language in imaging reports to specific medical terms. Furthermore, the concept association process further uncovers deeper connections between data; for example, through relational reasoning in knowledge graphs, it can be discovered that an abnormality in a specific test indicator may be an early signal of a certain disease. This process not only enhances the semantics of the original medical record dataset but also provides a richer and more meaningful data foundation for subsequent dynamic temporal feature analysis. The resulting semantically linked medical record information network helps medical personnel understand the progression of a patient's condition from a holistic perspective, while also providing crucial support for precision medicine and personalized treatment.
[0060] Step S3: Perform dynamic temporal feature analysis on the semantically associated medical record information network using a temporal feature extractor to obtain a temporal medical record feature sequence.
[0061] Specifically, in the above scheme, a temporal feature extractor is used to perform dynamic temporal feature analysis on the semantically associated medical record information network to obtain a temporal medical record feature sequence. This process aims to uncover the changing patterns and development trends of medical record data over time. Specifically, the semantically associated medical record information network has integrated multi-source patient data into a unified, semantically related network structure. However, this data is still scattered across different time points and cannot directly reflect the dynamic evolution of the patient's condition. To achieve dynamic temporal feature analysis, the temporal feature extractor first needs to identify and extract time-related features from the network, such as diagnostic timestamps in electronic medical record systems, the time sequence of imaging examinations in medical imaging workstations, and the change curves of laboratory indicators in laboratory information systems. For example, in a large hospital application scenario, a patient's medical record data may show that their blood pressure has been continuously rising over the past three months, while renal function-related indicators such as serum creatinine levels also show a gradual upward trend. This temporal change pattern can be captured and transformed into a set of feature sequences by the temporal feature extractor. Meanwhile, the dynamic analysis process also needs to incorporate conceptual relationships within the semantically linked medical record information network to further reveal the potential connections between these features. For example, analysis revealed a temporal causal relationship between "elevated blood pressure" and "deteriorating kidney function." Furthermore, the temporal feature extractor can handle data at different time scales, such as short-term fluctuations and long-term patterns, thus providing medical personnel with a more comprehensive perspective on disease progression. The resulting temporal medical record feature sequence provides crucial input for subsequent heterogeneous data fusion and also offers a scientific basis for early disease warning, diagnostic decisions, and treatment plan adjustments, helping doctors more accurately grasp the dynamics of patients' conditions and develop more personalized medical plans.
[0062] Step S4: Perform heterogeneous data fusion on the time-series medical record feature sequence to obtain a unified medical record data view.
[0063] Specifically, in the above scheme, heterogeneous data fusion of time-series medical record feature sequences to obtain a unified medical record data view is a key step in achieving efficient integration and utilization of multi-source data. Specifically, while the time-series medical record feature sequences have already extracted dynamic features in the time dimension from the semantically related medical record information network, these features may still exist in different data formats and structures, such as text descriptions in electronic medical record systems, image analysis results from medical imaging workstations, and numerical change curves in laboratory information systems. To achieve heterogeneous data fusion, it is first necessary to design a fusion framework compatible with multiple data types. This framework can transform time-series features from different sources into a unified representation while preserving their original semantics and temporal attributes. For example, in a large hospital application scenario, a patient's time-series medical record feature sequence may include continuous changes in their blood pressure, trend analysis of renal function indicators, and records of changes in kidney morphology in imaging examinations. These features exist in numerical, text, and image forms, respectively. Through heterogeneous data fusion technology, this scattered information can be integrated into a unified medical record data view, allowing doctors to simultaneously view the patient's blood pressure trend, the process of renal function deterioration, and the temporal evolution of imaging findings on a single interface. Furthermore, this process also needs to address data conflicts and redundancy. For example, when different systems record the same indicator differently, the fusion algorithm will select the optimal result based on credibility weights or consistency verification mechanisms. The resulting unified medical record data view not only provides medical staff with a comprehensive and intuitive overview of patient conditions but also supports further data mining and decision support applications, thereby helping doctors develop treatment plans more quickly and improving the accuracy of medical services.
[0064] In a specific embodiment, the step of semantically mapping and conceptually associating the original medical record dataset based on a preset medical ontology knowledge graph to obtain a semantically associated medical record information network includes:
[0065] The original medical record dataset is segmented using a pre-defined medical terminology parser to obtain a sequence of medical terms. The sequence of medical terms is then semantically standardized and mapped based on a pre-defined medical ontology knowledge graph to obtain a standardized set of medical concepts.
[0066] The standardized medical concept set is analyzed for association using a semantic network builder to obtain a concept relationship graph. The concept relationship graph is then hierarchically organized to obtain a hierarchical semantic network.
[0067] The hierarchical semantic network is analyzed by a semantic reasoning engine to obtain a set of semantic reasoning rules. Based on the set of semantic reasoning rules, knowledge is expanded to obtain an extended semantic relationship network.
[0068] Based on a preset ontology fusion processor, the extended semantic relationship network is conceptually aligned and knowledge is integrated to obtain a semantically related medical record information network.
[0069] Specifically, in the above scheme, the original medical record dataset is first segmented using a pre-defined medical terminology parser to extract the sequence of medical terms. For example, in a large hospital application scenario, a patient with a chronic disease has visited multiple times, and their electronic medical record contains descriptive terms such as "hypertension," "proteinuria," "left ventricular hypertrophy," and "elevated fasting blood glucose." The medical terminology parser uses natural language processing techniques (such as named entity recognition) to identify specific medical terms from this text and arranges them into a sequence of medical terms according to their order of appearance. Assuming the patient has 10 outpatient records, a total of 52 medical terms are extracted, forming a medical terminology sequence of length 52.
[0070] Subsequently, the extracted medical terminology sequences are semantically standardized and mapped based on a pre-defined medical ontology knowledge graph, generating a standardized medical concept set. Medical ontology knowledge graphs typically include standard systems such as ICD-10, SNOMED CT, and LOINC. For example, the term "hypertension" might correspond to ICD-10 code I10.001, while "proteinuria" might be mapped to LOINC code 28496-7. During this process, terms that might have been ambiguous or inconsistent in expression (such as "hyperglycemia" and "elevated fasting blood glucose") are uniformly mapped to standardized concept identifiers. Continuing with the patient example, after semantic standardization, the original 52 terms are merged into 37 standardized medical concepts, forming a standardized medical concept set. Next, a semantic network builder performs association analysis on the standardized medical concept set, generating a concept relationship graph, which is further hierarchically organized to form a layered semantic network. For example, analysis of the patient's standardized medical concept set revealed that "hypertension" and "left ventricular hypertrophy" appeared simultaneously at multiple time points, with a frequency of 85%; a strong co-occurrence relationship also existed between "proteinuria" and "abnormal renal function." The semantic network builder established edge connections based on this, representing the potential causal or correlational relationships between them. This ultimately formed a conceptual relationship graph containing 37 nodes (medical concepts) and 58 edges (conceptual relationships). Based on this, according to the hierarchical structure in the medical ontology knowledge graph (such as the disease classification in ICD-10), the graph was hierarchically divided, for example, "hypertension" was classified under "circulatory system diseases," and "proteinuria" under "urinary system diseases," thus forming a hierarchical semantic network with a clear semantic structure. Furthermore, a semantic reasoning engine performed semantic reasoning analysis on the hierarchical semantic network, generating a set of semantic reasoning rules, and based on these rules, knowledge was expanded to obtain an extended semantic relationship network. For example, if the knowledge graph already contains two rules, "hypertension leads to left ventricular hypertrophy" and "left ventricular hypertrophy increases the risk of heart failure," then a new conclusion, "hypertension increases the risk of heart failure," can be derived through chain reasoning. The semantic reasoning engine automatically extracted 12 similar rules and supplemented them with new association paths by combining external medical literature, such as "microalbuminuria is an early marker of diabetic nephropathy." After reasoning and expansion, the number of nodes in the original graph increased from 37 to 45, and the number of edges increased from 58 to 82, forming a richer extended semantic relationship network. Finally, based on the preset ontology fusion processor, the extended semantic relationship network is used for concept alignment and knowledge integration to generate a semantically related medical record information network. Since different medical institutions may use different terminology systems, such as Hospital A using ICD-10 and Hospital B using ICD-9, the ontology fusion processor completes cross-system concept alignment by looking up terminology mapping tables and calculating similarity (e.g., cosine similarity > 0.85 is considered a match).During the integration process, 13 duplicate relationships were removed and 7 conflicting relationships were merged, ultimately forming a semantically linked medical record information network containing 43 nodes and 75 edges. This network not only reflects the individual patient's disease progression path but also incorporates authoritative medical knowledge, possessing good interpretability and scalability.
[0071] In a specific embodiment, the step of performing association analysis on the standardized medical concept set based on a semantic network builder to obtain a concept relationship graph includes:
[0072] The standardized medical concept set is subjected to co-occurrence analysis by a concept relationship mining tool to obtain a concept co-occurrence matrix. The association degree of the concept co-occurrence matrix is then calculated to obtain a concept association strength map.
[0073] Path mining is performed on the concept association strength graph to obtain a set of concept semantic paths, and weights are assigned to the set of concept semantic paths to obtain a weighted semantic path network.
[0074] The weighted semantic path network is topologically optimized based on the semantic network builder to obtain an optimized semantic structure graph, and the optimized semantic structure graph is hierarchically divided to obtain a hierarchical semantic structure.
[0075] The hierarchical semantic structure is constructed using a knowledge graph generator to obtain a concept relationship graph.
[0076] Specifically, in the above scheme, the process of performing association analysis on the standardized medical concept set based on a semantic network builder to obtain a concept relationship graph is a complex but crucial step. First, a co-occurrence analysis is performed on the standardized medical concept set using a concept relationship miner to generate a concept co-occurrence matrix. This matrix is then further analyzed for association strength to obtain a concept association strength graph. Specifically, in a large hospital application scenario, suppose a patient undergoes a comprehensive examination for multiple health issues. Their electronic medical record system records terms such as "hypertension" and "diabetes," the medical imaging workstation may contain heart-related image descriptions, and the laboratory information system records various blood indicator values. The concept relationship miner can identify the co-occurrence of these terms in different documents or data records. For example, "hypertension" and "abnormal kidney function" may frequently appear together in the same patient's medical record. Statistical analysis of these co-occurrences generates a concept co-occurrence matrix, showing the frequency of each pair of concepts appearing together. Then, using association strength calculation methods, such as point mutual information (PMI) or correlation coefficients, the association strength between each pair of concepts is extracted from the co-occurrence matrix to form a concept association strength graph. Next, path mining is performed on the concept association strength graph to discover sets of semantic paths, and weights are assigned to these paths to construct a weighted semantic path network. This process aims to reveal the potential connections and importance between concepts. Continuing with the patient example above, if a certain association is known between "hypertension" and "abnormal kidney function," then path mining techniques can identify all possible paths connecting these two concepts from the concept association strength graph. For example, "hypertension" → "cardiovascular disease" → "abnormal kidney function" is one path. Then, weights are assigned to paths based on the association strength between nodes on each path; paths with higher weights indicate a stronger connection between these concepts or higher clinical significance. The resulting weighted semantic path network not only demonstrates the relationships between concepts but also quantifies the importance of these relationships, facilitating subsequent in-depth analysis and application. Subsequently, the weighted semantic path network is topologically optimized using a semantic network builder to generate an optimized semantic structure graph, which is then hierarchically partitioned to obtain a layered semantic structure. The purpose of topology optimization is to improve the interpretability and usability of the network. In this process, algorithms may be employed to simplify the network structure, remove redundant connections, or merge similar concept nodes, resulting in a more concise and clear semantic structure. For example, when dealing with complex networks containing numerous medical concepts, the connections between nodes can be adjusted to ensure that the most important paths and relationships are preserved, while secondary information is reasonably simplified. Furthermore, hierarchical partitioning divides the entire network into different levels based on the importance, category, or other criteria of the concepts.For example, it can be divided according to disease type (such as cardiovascular disease, endocrine disease) or disease stage (early symptoms, mid-stage progression, late-stage manifestations) to help doctors more intuitively understand the overall development of the patient's condition. Finally, a knowledge graph generator is used to construct a graph of the hierarchical semantic structure, ultimately resulting in a concept relationship graph. The role of the knowledge graph generator is to transform the hierarchical semantic structure obtained after a series of processing steps into a visualized knowledge graph form. This step not only needs to consider how to effectively display each concept and its interrelationships, but also to ensure the readability and interactivity of the graph. For example, in practical applications, doctors can quickly understand the connection between a disease and other related factors by viewing the concept relationship graph, such as how "hypertension" affects the function of other organs, and which changes in test indicators may be manifestations of this influence. In this way, doctors can not only obtain more comprehensive patient information, but also formulate more precise and effective treatment plans based on the latest medical research results. In summary, this series of steps, starting from the original medical terminology, and through multiple rounds of data processing and analysis, ultimately forms a structured, hierarchical concept relationship graph, greatly improving the ability of medical information management and clinical decision support.
[0077] To better understand the above technical solutions, the following explanation is provided. Medical Concept Relationship Mining Technical Solution 1: Co-occurrence Analysis and Concept Association Strength Map Construction of Concept Co-occurrence Matrix Calculation of Concept Relationship Miner on Standardized Medical Concept Set Co-occurrence analysis was performed. In the original medical record dataset, all text records were traversed, and for each text segment, co-occurrence analysis was conducted between any two concepts. and The number of times they appear simultaneously. Let the concept co-occurrence matrix be... Its elements Representing concepts and The number of times it appears in all texts, that is: in, It is the total number of medical record texts. Indicates the first Individual medical record text, It is an indicator function; when the condition inside the parentheses is true, The value is 1 if it is positive and 0 otherwise. Association degree calculation is based on the concept co-occurrence matrix. The strength of the association between concepts is calculated using the Pointwise Mutual Information (PMI) algorithm. Pointwise Mutual Information measures the degree of association between two concepts, and its formula is: in, It is a concept The probability of appearing in all medical record texts It is a concept The probability of appearing in all medical record texts It is a concept and The probability of simultaneous occurrence. In actual calculations, frequency is used to approximate probability: Substituting the above probabilities into the point mutual information formula, we obtain the correlation degree between each concept pair, and then construct a concept correlation strength map. ,in It is a set of edges, and the weight of each edge is the corresponding concept association degree. II. Path Mining and Weighted Semantic Path Network Construction (I) Concept Semantic Path Set Mining for Concept Association Strength Graph Path discovery is performed using the Depth-First Search (DFS) algorithm. From each node in the graph... Departure, set path length limit The DFS algorithm is used to search for nodes starting from this node with a length not exceeding [a certain value]. All paths. During the search process, the sequence of nodes traversed is recorded, forming a conceptual semantic path. Let the set of conceptual semantic paths obtained by mining be... Each path It is a node sequence , It is a path The length of the path. Weight allocation assigns weights to each conceptual semantic path, taking into account both the association strength of edges in the path and the path length. The weight calculation formula is: in, It is a path Middle node and The weight of the edges between them (i.e., the degree of association). It is a path The length of the path is calculated using this formula. Dividing the sum of the edge association strengths in the path by the path length yields the average weight of each path, forming a weighted semantic path network. ,in It is a set of path weights. III. Topology Optimization and Hierarchical Semantic Structure Construction, (I) Topology Optimization Semantic Network Builder for Weighted Semantic Path Networks Topology optimization is performed using a graph shrinking algorithm from graph theory. This algorithm simplifies the network structure by merging strongly related and redundant paths and nodes. The specific steps are as follows: Calculate the density between nodes, which can be measured using both the node's degree and the edge weights. Let the nodes... The degree is , with nodes The sum of the weights of the connected edges is Then the node tightness It can be represented as: in, These are parameters related to the degree of adjustment and the influence of edge weights, with a value range of [value missing]. For node pairs that are highly closely related and semantically similar, merge them into one node and update the relevant paths and weights. Repeat this process until the network structure reaches the preset optimization criteria, resulting in an optimized semantic structure graph. Hierarchical partitioning optimizes the semantic structure graph. Hierarchical partitioning is performed using a method based on conceptual semantic similarity and node importance. First, the semantic similarity between nodes is calculated using the cosine similarity method based on embedding vectors mentioned earlier. Then, the importance of a node is determined based on its degree, weight, and semantic similarity to other nodes. Let the node... Importance The calculation formula is: in, , , It is a parameter that adjusts the degree of influence of various factors, and Nodes are divided into different levels according to their importance, with higher-importance nodes placed at the top level and lower-importance nodes at the bottom level, resulting in a hierarchical semantic structure. ,in Indicates the first The set of nodes in each layer. IV. Concept Relationship Graph Construction: Knowledge Graph Generator Based on Hierarchical Semantic Structure Constructing a concept relationship graph The nodes in the hierarchical semantic structure are used as nodes in the graph, the connections between nodes are preserved as edges, and the edge weights and node hierarchical information are used as attributes of the graph. The resulting concept relationship graph... ,in It is a set of nodes, corresponding to nodes in a hierarchical semantic structure; It is a set of edges, corresponding to the connection relationships between nodes; It is a set of attributes, containing information such as edge weights and node hierarchy.
[0078] In a specific embodiment, the step of performing dynamic temporal feature analysis on the semantically associated medical record information network using a temporal feature extractor to obtain a temporal medical record feature sequence includes:
[0079] The semantically associated medical record information network is subjected to trend component extraction by a temporal feature extractor to obtain disease progression trend features, and the disease progression trend features are divided into disease stages to obtain staged disease progression features.
[0080] Based on the phased disease progression features, periodic pattern mining is performed on the semantically associated medical record information network to obtain disease recurrence cycle features, and the fluctuation amplitude of the disease recurrence cycle features is quantified to obtain quantified disease recurrence features.
[0081] Based on the quantitative disease recurrence characteristics, outlier detection is performed on the semantically associated medical record information network to obtain a sequence of abnormal disease events. Then, event correlation analysis is performed on the sequence of abnormal disease events to obtain key event characteristics of the disease.
[0082] Based on the key event characteristics of the disease, temporal logical reasoning and disease recording are performed on the abnormal event sequence of the disease to obtain a temporal medical record feature sequence.
[0083] Specifically, in the above scheme, the process of performing dynamic temporal feature analysis on the semantically associated medical record information network using a temporal feature extractor to obtain a temporal medical record feature sequence is a complex but crucial step. First, the temporal feature extractor extracts trend components from the semantically associated medical record information network to obtain disease progression trend features. These features are then further divided into disease stages to form staged disease progression features. Specifically, in a large hospital application scenario, suppose a patient undergoes a comprehensive examination for multiple health problems. Their electronic medical record system records diagnostic information such as "hypertension" and "diabetes," the medical imaging workstation may contain heart-related image descriptions, and the laboratory information system records various blood indicator values. The temporal feature extractor can identify patterns in these data over time. For example, if a patient's blood pressure has been consistently rising over the past few years, this can be considered a trend component of disease progression. Based on this trend component, the algorithm automatically divides the disease progression into different stages, such as early, middle, and late-stage hypertension, to more accurately understand the development of the disease. Next, based on the phased disease progression characteristics, periodic pattern mining is performed on the semantically associated medical record information network to discover disease recurrence cycle characteristics. The fluctuation amplitude of these characteristics is then quantified to obtain quantitative disease recurrence characteristics. In this process, the algorithm not only focuses on the long-term trend of disease development but also seeks out periodic change patterns. Continuing with the above patient example, if the patient's blood glucose level exhibits seasonal fluctuations—that is, blood glucose levels rise every winter and stabilize in summer—this periodic pattern can be identified using periodic pattern mining techniques. Furthermore, by quantifying the fluctuation amplitude of these periodic patterns, such as calculating the difference between each blood glucose peak and trough, more detailed quantitative disease recurrence characteristics can be obtained. This process helps doctors better understand the natural progression of the disease and the degree to which it is affected by external factors, providing a basis for developing personalized treatment plans. Subsequently, based on the quantitative disease recurrence characteristics, outlier detection is performed on the semantically associated medical record information network to screen out disease-related event sequences. Event correlation analysis is then performed on these events to obtain key disease event characteristics. In practical applications, certain physiological indicators or symptoms of patients may suddenly fluctuate abnormally, such as a sudden heart attack or a severe hypoglycemic event. Outlier detection technology can accurately identify these abnormal events from large amounts of time-series data and mark them as part of a disease-related event sequence. Then, by analyzing the correlations of these abnormal events—for example, exploring whether a heart attack is related to strenuous exercise in the preceding days, or whether a hypoglycemic event is related to changes in dietary habits—key factors leading to the abnormal events can be revealed, forming key event signatures for the disease. This process not only helps doctors quickly pinpoint potential sources of problems but also provides guidance for preventing similar events from recurring.Finally, based on the key event features of the disease, temporal logical reasoning and disease recording are performed on the abnormal event sequence, ultimately yielding a temporal medical record feature sequence. In this step, the algorithm utilizes known key event features, combined with the patient's detailed medical history and clinical manifestations, to predict potential future health problems through logical reasoning. For example, if a patient previously experienced an acute exacerbation of hypertension due to a high-salt diet, the system can issue an early warning under similar dietary conditions in the future, suggesting adjustments to the diet to prevent the condition from worsening. Simultaneously, these reasoning results and disease records are integrated into a complete temporal medical record feature sequence, allowing doctors to view the patient's entire medical history, disease progression trends, potential risks, and recommended interventions from a unified perspective. This approach not only improves the quality and efficiency of medical services but also promotes the development of personalized medicine, enabling each patient to receive the most suitable health management plan. In summary, this series of steps, starting from raw temporal data and undergoing multiple rounds of data processing and analysis, ultimately forms a structured and ordered temporal medical record feature sequence, significantly enhancing medical information management and clinical decision support capabilities. The detailed explanation of the temporal feature analysis technology based on semantic association medical record information networks, including disease progression trend feature extraction and stage division, is presented in this paper. Trend component extraction utilizes a temporal feature extractor to analyze the semantically correlated medical record information network. Disease-related data within this network changes over time, forming a time series. To extract disease progression trend features, the Empirical Mode Decomposition (EMD) algorithm is employed. Based on the local feature scale of the signal, the EMD algorithm decomposes complex time series data into a finite number of intrinsic mode functions (IMFs) and a residual component. The residual component represents the trend component of the time series. The selection process begins with the original time series as follows: Find The upper envelope of all maxima and minima is obtained by cubic spline interpolation. and lower envelope Calculate the mean The difference is obtained. .examine Does it satisfy the IMF conditions (the number of extreme points is equal to or differs from the number of zero crossings by at most 1; at any given time, the upper envelope determined by the local maxima and the lower envelope determined by the local minima are locally symmetric about the time axis)? If not, As a new Repeat the above steps until the condition is met. This is the first IMF component. Iterative decomposition: computation ,Will As a new Repeat the above screening process to obtain the second IMF component. And so on, until the residual component. It becomes a monotonic function or satisfies the stopping condition. Final residual component. This refers to the disease progression trend characteristics. 1.2 Disease Stage Classification Based on Extracted Disease Progression Trend Characteristics The Dynamic Time Warping (DTW) algorithm is used to match the disease stage templates with a pre-defined set of disease stage templates to achieve disease stage segmentation. Let the pre-defined set of disease stage templates be... Each template It is a time series. Regarding the characteristics of disease progression trends... The DTW algorithm is used to calculate its relationship with each template. distance The DTW algorithm calculates the cumulative distance along the optimal matching path by determining the non-linear alignment of two time series on the time axis. Specifically, in the calculation, let... The length is , The length is , build a Distance matrix ,in Then search from arrive Optimal path This makes the cumulative distance Minimum. The disease stage corresponding to the template with the smallest distance is selected as the current stage of the disease, thus obtaining the staged disease progression characteristics. Periodic pattern mining and recurrence feature quantification: Periodic pattern mining is based on the staged disease progression characteristics and uses Fourier Transform to mine the disease recurrence cycle characteristics. Fourier Transform can convert the time domain signal into a frequency domain signal, and by analyzing the peak values in the frequency domain signal, the periodic component of the signal is determined. For the time series composed of staged disease progression characteristics... Its Discrete Fourier Transform (DFT) formula is: in, It is the length of the time series. It is a frequency domain representation. It is the imaginary unit. After obtaining the frequency domain results, identify the frequency components with larger amplitudes, and then determine the frequency components based on the relationship between frequency and period. ( For a period of time, The frequency of disease recurrence is used to determine the recurrence cycle, thus obtaining the characteristics of the disease recurrence cycle. The fluctuation amplitude is quantified by using the standard deviation to measure the degree of fluctuation of the disease recurrence characteristics within the cycle. Let the time series corresponding to the disease recurrence cycle characteristics be... Divide it into multiple subsequences according to its period. , For each subsequence Calculate its mean ( (where is the length of the subsequence), then the standard deviation of the subsequence is: By combining the standard deviations of all subsequences, a quantitative disease recurrence characteristic is obtained, reflecting the fluctuation range of disease recurrence. Outlier detection and key event analysis are performed. Outlier detection, based on the quantitative disease recurrence characteristic, uses the Isolation Forest algorithm to detect outliers in the semantically associated medical record information network, obtaining disease-related abnormal event sequences. The Isolation Forest algorithm partitions the data by constructing multiple isolated trees, where the path length of each data point in the isolated tree represents its degree of isolation. For a sample... In the The path length in the isolated tree is The average path length is ( (Number of isolated trees). Samples The formula for calculating the abnormal score is: in, The sample size is The theoretical value of the average path length at that time. When anomaly scores... When the threshold is exceeded, the sample is judged. For outliers, all events corresponding to outliers are grouped into a disease-related abnormal event sequence. Event correlation analysis is then performed on the disease-related abnormal event sequence, using the Pearson correlation coefficient to calculate the correlation between events and obtain the key event characteristics of the disease. Let's assume there are two events in the disease-related abnormal event sequence corresponding to time series data... and The formula for calculating its Pearson correlation coefficient is: in, and They are time series and The mean, This refers to the length of the time series. Events with larger absolute values of correlation coefficients are selected as key disease event features. 4. Temporal Logic Reasoning and Medical Record Feature Sequence Generation: Based on the key disease event features, temporal logic reasoning algorithms (such as temporal logic-based reasoning methods) are used to reason about the abnormal disease event sequences. Temporal logic describes and reasons about the temporal relationships of events by introducing time operators (such as "until" and "next"). For example, linear temporal logic (LTL) is used to define logical formulas to describe the temporal and causal relationships between events. By performing logical judgment and reasoning on the abnormal disease event sequences, combined with disease record information, related events and information are integrated in chronological order to generate a temporal medical record feature sequence, providing more comprehensive and orderly information for disease diagnosis, treatment, and research.
[0084] In a specific embodiment, the step of performing periodic pattern mining on the semantically associated medical record information network based on the phased disease progression characteristics to obtain disease recurrence cycle characteristics includes:
[0085] Based on the staged disease progression features, the semantically associated medical record information network is segmented into subsequences to obtain disease progression subsequences, and the similarity of the disease progression subsequences is measured to obtain a subsequence similarity matrix.
[0086] Cluster analysis is performed on the disease progression subsequences based on the subsequence similarity matrix to obtain disease progression pattern clusters, and pattern features are extracted from the disease progression pattern clusters to obtain disease progression pattern feature vectors.
[0087] Based on the feature vector of the disease progression pattern, periodic pattern matching is performed on the disease progression pattern cluster to obtain a candidate relapse cycle set, and statistical significance test is performed on the candidate relapse cycle set to obtain a significant relapse cycle set.
[0088] Based on the significant recurrence cycle set, the cycle length of the candidate recurrence cycle set is estimated to obtain the recurrence cycle length value, and the confidence interval of the recurrence cycle length value is calculated to obtain the recurrence cycle confidence interval.
[0089] Based on the confidence interval of the recurrence cycle, the set of significant recurrence cycles is periodically integrated to obtain the disease recurrence cycle characteristics.
[0090] Specifically, in the above scheme, the process of mining periodic patterns in the semantically associated medical record information network based on the phased disease progression characteristics to obtain the disease recurrence cycle characteristics is a systematic and multi-level analysis process. First, the semantically associated medical record information network is segmented into subsequences based on the phased disease progression characteristics to obtain disease progression subsequences. These subsequences are then further measured for similarity to generate a subsequence similarity matrix. Specifically, in a large hospital application scenario, suppose a patient undergoes long-term follow-up examinations for a chronic disease, and their electronic medical record system records data such as blood pressure and blood sugar levels changing over time. Based on the disease progression stage (e.g., early hypertension, mid-stage hypertension), this data can be divided into multiple subsequences. For example, if the patient's blood pressure values have fluctuated multiple times over the past few years, this blood pressure data can be segmented into several subsequences according to different time periods or disease development stages. Then, by measuring the similarity between these subsequences, such as using the Dynamic Time Warping (DTW) algorithm to calculate their distances, a subsequence similarity matrix can be constructed. This matrix quantifies the degree of similarity between each pair of subsequences. Next, cluster analysis is performed on the disease progression subsequences based on the subsequence similarity matrix to identify disease progression pattern clusters, and disease progression pattern feature vectors are extracted from these clusters. In this process, cluster analysis helps identify subsequences with similar trends, thus forming different disease progression pattern clusters. Continuing with the patient example above, suppose cluster analysis reveals several main blood pressure change patterns: one is a continuously rising trend, and the other is a fluctuating but generally stable trend within a certain range. For each such pattern cluster, its core features can be extracted to form a disease progression pattern feature vector, such as the average rate of change, maximum peak value, and minimum trough value. These feature vectors not only summarize the main characteristics of each pattern cluster but also provide a foundation for subsequent periodic pattern matching. Subsequently, periodic pattern matching is performed on the disease progression pattern clusters based on the disease progression pattern feature vectors to screen out candidate recurrence cycle sets, and statistical significance tests are conducted on these candidate sets to finally determine the significant recurrence cycle set. In this step, using the existing disease progression pattern feature vectors, the algorithm attempts to find pattern clusters that exhibit periodic changes. For example, if a patient's blood sugar levels exhibit a pattern of rising in winter and falling in summer each year, this cyclical change can be identified as part of a candidate relapse cycle. Then, by performing statistical significance tests on these candidate relapse cycles, such as chi-square tests or t-tests, it can be verified which cycles are indeed significant and not simply random fluctuations. This step ensures the authenticity and reliability of the selected cycles.Next, the cycle length of the candidate recurrence cycle set is estimated based on the significant recurrence cycle set, yielding recurrence cycle length values. Confidence intervals are then calculated for these length values to obtain the recurrence cycle confidence intervals. In this process, to more accurately understand the disease's recurrence cycle, a detailed length estimate is needed for each cycle in the initially screened significant recurrence cycle set. For example, for the aforementioned patient, if seasonal fluctuations in blood glucose levels are confirmed, the specific cycle length of these fluctuations can be further calculated, such as approximately 12 months. Simultaneously, to assess the accuracy of this estimate, the corresponding confidence interval needs to be calculated, allowing doctors to understand the possible range and reliability of the cycle length estimate. Finally, the significant recurrence cycle set is integrated based on the recurrence cycle confidence intervals to ultimately obtain the disease recurrence cycle characteristics. This step aims to integrate all validated and estimated significant recurrence cycles to form a comprehensive characteristic description reflecting the disease's recurrence pattern. For example, after comprehensively considering the patient's blood glucose levels, blood pressure changes, and other relevant physiological indicators, a characteristic set containing multiple disease recurrence cycles can be constructed. These cycle characteristics not only reveal the inherent laws of disease development but also provide important reference for prevention and treatment. This approach not only helps doctors better understand the natural progression of diseases but also guides the development of more scientific and rational treatment plans, improving the quality and effectiveness of medical services. In short, this series of steps, starting from raw time-series data and undergoing multiple rounds of data processing and analysis, ultimately forms a structured and orderly pattern of disease recurrence cycles, significantly enhancing medical information management and clinical decision support capabilities.
[0091] In a specific embodiment, the step of performing temporal logical reasoning and disease recording on the abnormal event sequence based on the key disease event characteristics to obtain a temporal medical record feature sequence includes:
[0092] Based on the key event characteristics of the disease, an event causal association analysis is performed on the abnormal event sequence of the disease to obtain an event causal association graph. Then, the causal strength of the event causal association graph is calculated to obtain a weighted causal association network.
[0093] Based on the weighted causal association network, temporal constraint mining is performed on the abnormal event sequence of the disease to obtain a temporal constraint rule set, and the constraint strength of the temporal constraint rule set is evaluated to obtain an effective temporal constraint network.
[0094] Based on the effective temporal constraint network, temporal path reasoning is performed on the weighted causal association network to obtain a set of candidate disease development paths, and path probability evaluation is performed on the set of candidate disease development paths to obtain a probabilistic development path graph.
[0095] Based on the probabilistic development path diagram, the clinical trajectory of the abnormal disease event sequence is reconstructed and the disease is recorded to obtain a time-series medical record feature sequence.
[0096] Specifically, in the above scheme, the process of obtaining a time-series medical record feature sequence by performing temporal logical reasoning and disease recording based on the characteristics of key disease events is a complex and multi-layered analytical process. First, causal association analysis is performed on the sequence of abnormal disease events based on the characteristics of key disease events, thereby generating an event causal association graph. This graph is then further analyzed for causal strength to form a weighted causal association network. Specifically, in a large hospital application scenario, suppose a patient undergoes a comprehensive examination for multiple health problems. Their electronic medical record system records diagnostic information such as "hypertension" and "diabetes," the medical imaging workstation may contain heart-related image descriptions, and the laboratory information system records various blood indicator values. When certain physiological indicators or symptoms of the patient fluctuate abnormally, such as a sudden heart attack or a severe hypoglycemic event, these abnormal events are marked as part of the disease abnormal event sequence. By performing causal association analysis on these events, potential causal relationships can be identified, such as whether the heart attack is related to strenuous exercise in the preceding days, or whether the hypoglycemic event is related to changes in dietary habits. The event causal relationship graph constructed in this way not only reveals the potential connections between different events but also provides a foundation for further calculation of causal strength. Using statistical methods such as the Granger Causality Test, the causal strength between each pair of events can be quantified, thus forming a weighted causal relationship network. Next, based on the weighted causal relationship network, temporal constraint mining is performed on the abnormal event sequence of the disease to discover a set of temporal constraint rules. The constraint strength of these rules is then evaluated to ultimately determine the effective temporal constraint network. In this process, the algorithm not only identifies the causal relationships between events but also considers their temporal order and time intervals. Continuing with the patient example above, if a record of strenuous exercise is found before each heart attack, this pattern can be identified through temporal constraint mining techniques, and corresponding constraint rules can be set, such as heart attacks typically occurring within 24 hours after strenuous exercise. Then, by evaluating the constraint strength of these temporal constraint rule sets, such as using indicators like support and confidence to measure the effectiveness of the rules, rules with high reliability can be selected to form an effective temporal constraint network. This process helps to more accurately predict potential future health problems and their time windows of occurrence. Subsequently, based on the effective temporal constraint network, temporal path reasoning is performed on the weighted causal association network to filter out a set of candidate disease development paths. The path probabilities of these paths are then evaluated to generate a probabilistic development path graph. In this step, the algorithm utilizes the existing effective temporal constraint network and weighted causal association network to attempt to infer all possible disease development paths.For example, for the aforementioned patient, if a significant causal relationship has been established between strenuous exercise and a heart attack, and a corresponding time constraint has been set, a typical disease progression path can be deduced: strenuous exercise → elevated blood pressure → heart attack. By probabilistically assessing all possible paths, such as using Bayesian networks or Markov models to calculate the probability of each path occurring, a probabilistic progression path diagram containing multiple possible paths can be generated. This step not only helps doctors understand the potential trajectory of the disease but also provides a basis for developing personalized preventive measures. Finally, based on the probabilistic progression path diagram, the clinical trajectory of the abnormal event sequence is reconstructed and the disease is recorded, ultimately yielding a time-series medical record feature sequence. This step aims to integrate all the analysis results to form a complete and ordered description of the clinical trajectory. For example, after comprehensively considering the patient's blood pressure changes, heart attack, and other relevant physiological indicators, a detailed clinical trajectory can be constructed. This trajectory not only shows the patient's past health status and abnormal events but also predicts future potential risks and development trends. Simultaneously, to facilitate subsequent medical decision support, these analysis results need to be recorded to form a structured time-series medical record feature sequence. In this way, doctors can not only view a patient's current health status but also understand their past medical history, family genetic risk factors, and other multi-dimensional information, which helps in developing more precise and effective treatment plans. This approach not only improves the quality and efficiency of medical services but also promotes the development of personalized medicine, allowing each patient to receive the most suitable health management plan. In short, this series of steps, starting from the original abnormal disease events and undergoing multiple rounds of data processing and analysis, ultimately forms a structured and ordered temporal sequence of medical record characteristics, greatly enhancing the capabilities of medical information management and clinical decision support.
[0097] In a specific embodiment, the step of fusing heterogeneous data on the time-series medical record feature sequences to obtain a unified medical record data view includes:
[0098] The time-series medical record feature sequence is aligned in the spatiotemporal dimension by a spatiotemporal feature aligner to obtain a spatiotemporal aligned feature set, and features are extracted from the spatiotemporal aligned feature set to obtain a multi-dimensional feature vector;
[0099] The multi-dimensional feature vectors are transformed into a unified feature space, and feature association calculations are performed on the unified feature space to obtain a feature association network.
[0100] The feature association network is fused using a multimodal feature fusion processor to obtain a fused feature map. The fused feature map is then optimized to obtain an optimized feature map. A knowledge graph is constructed based on the optimized feature map to obtain a unified medical record data view.
[0101] Specifically, in the above scheme, the process of fusing heterogeneous data from time-series medical record feature sequences to obtain a unified medical record data view is a multi-layered and complex analytical process. First, a spatiotemporal feature aligner is used to align the time-series medical record feature sequences in terms of spatiotemporal dimensions, thereby generating a spatiotemporally aligned feature set. Further, multi-dimensional feature vectors are extracted from this set. Specifically, in a large hospital application scenario, suppose a patient undergoes a comprehensive examination for multiple health issues. Their electronic medical record system records diagnostic information such as "hypertension" and "diabetes," the medical imaging workstation may contain cardiac-related image descriptions, and the laboratory information system records various blood indicator values. This data not only comes from different systems but also includes temporal trends and spatial distribution characteristics. For example, the patient's blood pressure changes over time, the imaging manifestations of cardiac function, and the spatial distribution of different laboratory indicators. The spatiotemporal feature aligner can identify and adjust the timestamps and spatial coordinates of this data, ensuring that they can be compared and analyzed within the same temporal and spatial framework. By extracting features from these spatiotemporally aligned data, such as calculating statistical features like mean, variance, and peak value, a multi-dimensional feature vector can be generated. This vector comprehensively reflects various aspects of the patient's health status. Next, feature space transformation is performed on the multi-dimensional feature vector to generate a unified feature space. Furthermore, feature association calculations are performed on this space to form a feature association network. In this process, the algorithm not only integrates data from different sources into a unified representation but also explores the potential relationships between these data. For example, continuing with the patient example above, assuming a multi-dimensional feature vector containing blood pressure values, cardiac imaging descriptions, and laboratory indicators has been obtained, feature space transformation techniques can be used to map these features into a common feature space. This step typically involves dimensionality reduction methods such as principal component analysis (PCA), t-SNE, or autoencoders, aiming to remove redundant information and retain the most important features. Then, by performing association calculations on the features in this unified feature space, such as using correlation coefficients or mutual information to measure the relationship between each pair of features, a feature association network can be constructed. This network reveals the intrinsic connections between different features, contributing to a deeper understanding of the patient's condition. Subsequently, a multimodal feature fusion processor is used to fuse features in the feature association network, selecting a fused feature map. This map is then further optimized to generate an optimized feature map. In this step, the algorithm utilizes the existing feature association network to attempt to integrate all relevant information to form a comprehensive feature description reflecting the patient's health status. For example, for the aforementioned patient, if a feature association network showing the relationship between blood pressure and cardiac function has already been constructed, the multimodal feature fusion processor can combine this information with data from other sources (such as laboratory indicators) to generate a fused feature map.This graph not only illustrates the interactions between different features but also lays the foundation for subsequent feature optimization. The feature optimization process aims to improve the quality of the fused feature map, removing noise and unnecessary complexity, making the final result more concise and clear. Common optimization methods include regularization and sparse representation, which can help reduce the risk of overfitting and improve the model's generalization ability. Finally, a knowledge graph is constructed based on the optimized feature map, ultimately resulting in a unified medical record data view. This step aims to integrate all the analysis results into a structured, easy-to-understand, and easy-to-use representation of clinical data. For example, after comprehensively considering a patient's blood pressure changes, cardiac function, and other relevant physiological indicators, a detailed medical record data view can be constructed. This view not only shows the patient's past health status and abnormal events but also predicts future potential risks and trends. To facilitate subsequent medical decision support, these analysis results need to be recorded to form a structured, unified medical record data view. In this way, doctors can not only view the patient's current health status but also understand their past medical history, family genetic risk factors, and other multi-dimensional information, which helps to develop more precise and effective treatment plans. This approach not only improves the quality and efficiency of healthcare services but also promotes the development of personalized medicine, enabling each patient to receive the most suitable health management plan. In short, this series of steps, starting from raw, multi-source data and undergoing multiple rounds of data processing and analysis, ultimately forms a structured, orderly, and unified view of medical record data, significantly enhancing the capabilities of medical information management and clinical decision support.
[0102] The above describes the multi-source data fusion management method for medical record information in the embodiments of the present invention. The following describes the multi-source data fusion management system for medical record information in the embodiments of the present invention. Please refer to [link / reference]. Figure 2 One embodiment of the multi-source data fusion management system for medical record information in this invention includes:
[0103] The data acquisition module 21 is used to collect data from the electronic medical record system, medical imaging workstation and laboratory information system of medical institutions to obtain the original medical record dataset;
[0104] The association module 22 is used to perform semantic mapping and concept association on the original medical record dataset based on a preset medical ontology knowledge graph to obtain a semantically associated medical record information network.
[0105] Analysis module 23 is used to perform dynamic temporal feature analysis on the semantically associated medical record information network through a temporal feature extractor to obtain a temporal medical record feature sequence;
[0106] The fusion module 24 is used to perform heterogeneous data fusion on the time-series medical record feature sequence to obtain a unified medical record data view.
[0107] In this embodiment, the specific implementation of each unit in the above system embodiment is described in the above method embodiment, and will not be repeated here.
[0108] Reference Figure 3 This invention also provides a computer device whose internal structure can be as follows: Figure 3 As shown, the computer device includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores the data corresponding to this embodiment. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method.
[0109] Those skilled in the art will understand that Figure 3 The structures shown are merely block diagrams of some structures related to the present invention and do not constitute a limitation on the computer devices on which the present invention is applied.
[0110] An embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0111] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the present invention and embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0112] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0113] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A multi-source data fusion management method of medical record information, characterized in that, The method comprises the following steps: Data collection is performed on an electronic medical record system, a medical image workstation and a test information system of a medical institution to obtain an original medical record data set; Semantic mapping and concept association are performed on the original medical record data set based on a preset medical ontology knowledge graph to obtain a semantic association medical record information network; Dynamic time sequence feature analysis is performed on the semantic association medical record information network by a time sequence feature extractor to obtain a time sequence medical record feature sequence; Heterogeneous data fusion is performed on the time sequence medical record feature sequence to obtain a unified medical record data view; The semantic mapping and concept association on the original medical record data set based on the preset medical ontology knowledge graph to obtain the semantic association medical record information network comprises: Word segmentation processing is performed on the original medical record data set by a preset medical term analyzer to obtain a medical term sequence, and semantic standardized mapping is performed on the medical term sequence based on the preset medical ontology knowledge graph to obtain a standardized medical concept set; Association analysis is performed on the standardized medical concept set based on a semantic network constructor to obtain a concept relationship graph, and the concept relationship graph is hierarchically organized to obtain a hierarchical semantic network; Semantic reasoning analysis is performed on the hierarchical semantic network by a semantic reasoning engine to obtain a semantic reasoning rule set, and knowledge expansion is performed based on the semantic reasoning rule set to obtain an expanded semantic relationship network; Concept alignment and knowledge integration are performed on the expanded semantic relationship network based on a preset ontology fusion processor to obtain the semantic association medical record information network; The dynamic time sequence feature analysis on the semantic association medical record information network by the time sequence feature extractor to obtain the time sequence medical record feature sequence comprises: Trend component extraction is performed on the semantic association medical record information network by the time sequence feature extractor to obtain a disease progression trend feature, and the disease progression trend feature is divided into stages to obtain a staged disease progression feature; Periodic pattern mining is performed on the semantic association medical record information network based on the staged disease progression feature to obtain a disease recurrence cycle feature, and the disease recurrence cycle feature is quantified in terms of fluctuation amplitude to obtain a quantified disease recurrence feature; Anomaly value detection is performed on the semantic association medical record information network based on the quantified disease recurrence feature to obtain a disease anomaly event sequence, and event correlation analysis is performed on the disease anomaly event sequence to obtain a disease key event feature; Time sequence logic reasoning and disease recording are performed on the disease anomaly event sequence based on the disease key event feature to obtain the time sequence medical record feature sequence; The heterogeneous data fusion on the time sequence medical record feature sequence to obtain the unified medical record data view comprises: Temporal and spatial dimension alignment is performed on the time sequence medical record feature sequence by a time-space feature aligner to obtain a time-space aligned feature set, and feature extraction is performed on the time-space aligned feature set to obtain a multi-dimensional feature vector; Feature space conversion is performed on the multi-dimensional feature vector to obtain a unified feature space, and feature association calculation is performed on the unified feature space to obtain a feature association network; The feature correlation network is fused by a multi-modal feature fusioner to obtain a fused feature map, the fused feature map is optimized to obtain an optimized feature map, and a knowledge graph is constructed based on the optimized feature map to obtain a unified medical record data view.
2. The medical record information multi-source data fusion management method according to claim 1, characterized in that, The standardized medical concept set is analyzed for correlation by the semantic network constructor to obtain a concept relationship graph, including: The standardized medical concept set is analyzed for co-occurrence by a concept relationship miner to obtain a concept co-occurrence matrix, and the concept co-occurrence matrix is calculated for correlation to obtain a concept correlation strength graph; The concept correlation strength graph is mined for paths to obtain a concept semantic path set, and the concept semantic path set is assigned weights to obtain a weighted semantic path network; The weighted semantic path network is topologically optimized by the semantic network constructor to obtain an optimized semantic structure graph, and the optimized semantic structure graph is hierarchically divided to obtain a hierarchical semantic structure. The hierarchical semantic structure is graphically constructed by the knowledge graph generator to obtain a concept relationship graph.
3. The medical record information multi-source data fusion management method according to claim 1, characterized in that, The semantic correlation medical record information network is periodically mined for patterns based on the staged disease progression features to obtain disease recurrence cycle features, including: The semantic correlation medical record information network is segmented for sub-sequences based on the staged disease progression features to obtain disease progression sub-sequences, and the disease progression sub-sequences are measured for similarity to obtain a sub-sequence similarity matrix; The disease progression sub-sequences are analyzed for clustering based on the sub-sequence similarity matrix to obtain disease progression pattern clusters, and the disease progression pattern clusters are extracted for pattern features to obtain a disease progression pattern feature vector; The disease progression pattern clusters are periodically matched for patterns based on the disease progression pattern feature vector to obtain a candidate recurrence cycle set, and the candidate recurrence cycle set is statistically tested for significance to obtain a significant recurrence cycle set; The candidate recurrence cycle set is estimated for cycle length based on the significant recurrence cycle set to obtain a recurrence cycle length value, and the recurrence cycle length value is calculated for a confidence interval to obtain a recurrence cycle confidence interval; The significant recurrence cycle set is integrated for cycles based on the recurrence cycle confidence interval to obtain disease recurrence cycle features.
4. The medical record information multi-source data fusion management method according to claim 1, characterized in that, The disease abnormal event sequence is temporally logically reasoned and recorded for diseases based on the disease key event features to obtain a temporal medical record feature sequence, including: The disease abnormal event sequence is analyzed for event causal correlation based on the disease key event features to obtain an event causal correlation graph, and the event causal correlation graph is calculated for causal strength to obtain a weighted causal correlation network; The disease abnormal event sequence is temporally constrained mined based on the weighted causal correlation network to obtain a temporal constraint rule set, and the temporal constraint rule set is evaluated for constraint strength to obtain an effective temporal constraint network; performing temporal path reasoning on the weighted causal association network based on the effective temporal constraint network to obtain a candidate disease development path set, and performing path probability evaluation on the candidate disease development path set to obtain a probabilistic development path graph; performing clinical trajectory reconstruction and disease record on the disease abnormal event sequence based on the probabilistic development path graph to obtain a temporal medical record feature sequence.
5. A medical record information multi-source data fusion management system, characterized in that, The method comprises: a collection module configured to collect data from an electronic medical record system, a medical image workstation, and a laboratory information system of a medical institution to obtain an original medical record data set; an association module configured to perform semantic mapping and concept association on the original medical record data set based on a preset medical ontology knowledge graph to obtain a semantic association medical record information network; an analysis module configured to perform dynamic temporal feature analysis on the semantic association medical record information network by a temporal feature extractor to obtain a temporal medical record feature sequence; a fusion module configured to perform heterogeneous data fusion on the temporal medical record feature sequence to obtain a unified medical record data view. The method comprises: performing trend component extraction on the semantic association medical record information network by the temporal feature extractor to obtain disease progression trend features, and performing disease stage division on the disease progression trend features to obtain staged disease progression features; performing periodic pattern mining on the semantic association medical record information network based on the staged disease progression features to obtain disease recurrence cycle features, and performing fluctuation amplitude quantification on the disease recurrence cycle features to obtain quantified disease recurrence features; performing abnormal value detection on the semantic association medical record information network based on the quantified disease recurrence features to obtain a disease abnormal event sequence, and performing event correlation analysis on the disease abnormal event sequence to obtain disease key event features; performing temporal logic reasoning and disease record on the disease abnormal event sequence based on the disease key event features to obtain a temporal medical record feature sequence; The method comprises: performing spatio-temporal dimension alignment on the temporal medical record feature sequence by a spatio-temporal feature aligner to obtain a spatio-temporal alignment feature set, and performing feature extraction on the spatio-temporal alignment feature set to obtain a multi-dimensional feature vector; performing feature space conversion on the multi-dimensional feature vector to obtain a unified feature space, and performing feature association calculation on the unified feature space to obtain a feature association network; performing feature fusion on the feature association network by a multi-modal feature fusioner to obtain a fused feature map, and performing feature optimization on the fused feature map to obtain an optimized feature map, and constructing a knowledge graph based on the optimized feature map to obtain a unified medical record data view. 6.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-5. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 4.
7. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 4.
Citation Information
Patent Citations
Multi-modal medical data fusion and analysis platform
CN119622621A