Multi-source data fusion management method and system for medical record information
By conducting semantic mapping and timing feature analysis of multi-source data in medical institutions, a unified medical record data view is generated, the problem of semantic correlation neglect in medical record management is solved, efficient integration and deep understanding of multi-source data is achieved, and the quality of medical services and decision-making support capabilities are improved.
Patent Information
- Application Number
- CN202510612121.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Existing medical record management methods ignore the semantic correlation between data, limiting the ability to deeply understand and explore medical record information, especially in the diagnosis and treatment of complex diseases, it is difficult to provide comprehensive information support.
By collecting data from the electronic medical record system, medical imaging workstation and inspection information system of medical institutions, semantic mapping and conceptual association are performed based on the medical ontology knowledge graph, dynamic timing feature analysis is performed using the timing feature extractor, and finally heterogeneous data fusion is performed to generate a unified medical record data view.
It has achieved effective integration and display of medical record data from different sources under the same framework, improved the quality and efficiency of medical services, supported precise medical treatment and personalized treatment, and improved the performance of decision support systems.
Smart Images

Figure CN120452824A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical technology, and in particular to a multi-source data fusion management method and system for medical record information. Background Art
[0002] In the modern healthcare environment, with the development and application of information technology, medical institutions have accumulated vast amounts of medical record data. This data originates from diverse systems, such as electronic medical records, medical imaging workstations, and laboratory information systems, each with its own unique data structure and format. This multi-source, heterogeneous data environment presents significant challenges for the effective management and utilization of medical record information. On the one hand, the lack of unified standards and specifications makes it difficult to effectively integrate data from different sources, leading to the emergence of data silos. On the other hand, traditional data processing methods often focus on static data analysis and are unable to meet the needs of dynamic analysis of medical record data.
[0003] In addition, existing medical record management methods often ignore the semantic correlation between data, which limits the ability to deeply understand and mine medical record information. Especially when faced with complex disease diagnosis and treatment processes, relying solely on a single type of data is difficult to provide comprehensive information support. For example, when formulating a treatment plan, a doctor may need to comprehensively consider multiple aspects of information such as the patient's clinical symptoms, test results, and imaging manifestations. However, this information is stored in different systems and is difficult to integrate and share efficiently. Therefore, how to establish a method that can effectively integrate medical record data from multiple sources has become an important topic of current research.
[0004] To address these challenges, a multi-source data fusion management method for medical record information based on a medical ontology knowledge graph was proposed. This method not only focuses on data collection and integration but also emphasizes semantic associations and dynamic feature analysis between data. By constructing a unified view of medical record data, medical personnel can more quickly and accurately obtain relevant patient information, improving the quality and efficiency of medical services. Furthermore, this method provides strong data support for further research in cutting-edge fields such as precision medicine and personalized treatment. However, achieving this goal requires overcoming challenges such as data privacy protection and cross-institutional data sharing. Summary of the Invention
[0005] The main purpose of the present invention is to provide a multi-source data fusion management method and system for medical record information, which solves the technical problem that existing medical record management methods usually ignore the semantic correlation between data, which limits the ability to deeply understand and mine medical record information.
[0006] To achieve the above object, the present invention provides a multi-source data fusion management method for medical record information, comprising the following steps: Collect data from the electronic medical record system, medical imaging workstation and laboratory information system of medical institutions to obtain the original medical record data set; Performing semantic mapping and concept association on the original medical record data set based on a preset medical ontology knowledge graph to obtain a semantically associated medical record information network; Performing dynamic time series feature analysis on the semantically associated medical record information network through a time series feature extractor to obtain a time series medical record feature sequence; Heterogeneous data fusion is performed on the time series medical record feature sequence to obtain a unified medical record data view.
[0007] Furthermore, the original medical record data set is semantically mapped and conceptually associated based on the preset medical ontology knowledge graph to obtain a semantically associated medical record information network, including: Performing word segmentation processing on the original medical record dataset through a preset medical term parser to obtain a medical term sequence, and performing semantic standardization mapping on the medical term sequence based on a preset medical ontology knowledge graph to obtain a standardized medical concept set; Performing a correlation analysis on the standardized medical concept set based on a semantic network builder to obtain a concept relationship map, and hierarchically organizing the concept relationship map to obtain a hierarchical semantic network; Performing semantic reasoning analysis on the hierarchical semantic network through a semantic reasoning engine to obtain a semantic reasoning rule set, and performing knowledge expansion based on the semantic reasoning rule set to obtain an extended semantic relationship network; Based on a preset ontology fusion processor, concept alignment and knowledge integration are performed on the extended semantic relationship network to obtain a semantically associated medical record information network.
[0008] Furthermore, the semantic network builder performs a correlation analysis on the standardized medical concept set to obtain a concept relationship map, including: Performing co-occurrence analysis on the standardized medical concept set using a concept relationship miner to obtain a concept co-occurrence matrix, and performing correlation calculation on the concept co-occurrence matrix to obtain a concept correlation strength graph; Performing path mining on the concept association strength graph to obtain a concept semantic path set, and performing weight assignment on the concept semantic path set to obtain a weighted semantic path network; Performing topological optimization on the weighted semantic path network based on a semantic network builder to obtain an optimized semantic structure graph, and hierarchically dividing the optimized semantic structure graph to obtain a hierarchical semantic structure; The hierarchical semantic structure is graphed based on a knowledge graph generator to obtain a concept relationship graph.
[0009] Furthermore, the dynamic time series feature analysis of the semantically associated medical record information network is performed by a time series feature extractor to obtain a time series medical record feature sequence, including: Extracting trend components from the semantically associated medical record information network using a temporal feature extractor to obtain disease progression trend features, and dividing the disease progression trend features into disease stages to obtain staged disease progression features; Based on the staged disease progression characteristics, the semantically associated medical record information network is mined for periodic patterns to obtain disease recurrence period characteristics, and the disease recurrence period characteristics are quantified for fluctuation amplitude to obtain quantified disease recurrence characteristics; Performing outlier detection on the semantically associated medical record information network based on the quantified disease recurrence features to obtain a disease abnormal event sequence, and performing event correlation analysis on the disease abnormal event sequence to obtain a disease key event feature; Based on the key event characteristics of the disease, the abnormal disease event sequence is subjected to temporal logic reasoning and disease records to obtain a temporal medical record feature sequence.
[0010] Furthermore, the periodic pattern mining of the semantically associated medical record information network based on the staged disease progression characteristics to obtain disease recurrence cycle characteristics includes: Performing subsequence segmentation on the semantically associated medical record information network based on the staged disease progression features to obtain disease progression subsequences, and performing similarity measurement on the disease progression subsequences to obtain a subsequence similarity matrix; performing cluster analysis on the disease progression subsequences based on the subsequence similarity matrix to obtain disease progression pattern clusters, and performing pattern feature extraction on the disease progression pattern clusters to obtain disease progression pattern feature vectors; performing periodic pattern matching on the disease progression pattern cluster based on the disease progression pattern feature vector to obtain a candidate recurrence period set, and performing a statistical significance test on the candidate recurrence period set to obtain a significant recurrence period set; Estimating the cycle length of the candidate recurrence cycle set based on the significant recurrence cycle set to obtain a recurrence cycle length value, and performing confidence interval calculation on the recurrence cycle length value to obtain a recurrence cycle confidence interval; The significant recurrence cycle set is cycle-integrated based on the recurrence cycle confidence interval to obtain the disease recurrence cycle characteristics.
[0011] Furthermore, the time-series logical reasoning and disease recording of the abnormal disease event sequence based on the disease key event characteristics to obtain a time-series medical record feature sequence includes: Performing event causal association analysis on the abnormal disease event sequence based on the key event characteristics of the disease to obtain an event causal association graph, and performing causal strength calculation on the event causal association graph to obtain a weighted causal association network; Performing temporal constraint mining on the disease abnormal event sequence based on the weighted causal association network to obtain a temporal constraint rule set, and performing constraint strength evaluation on the temporal constraint rule set to obtain an effective temporal constraint network; Performing temporal path reasoning on the weighted causal association network based on the effective temporal constraint network to obtain a set of candidate disease development pathways, and performing path probability evaluation on the set of candidate disease development pathways to obtain a probabilistic development pathway diagram; Based on the probabilistic development path diagram, the clinical trajectory of the abnormal disease event sequence is reconstructed and the disease record is recorded to obtain a time-series medical record feature sequence.
[0012] Furthermore, the heterogeneous data fusion of the time series medical record feature sequence is performed to obtain a unified medical record data view, including: Performing spatiotemporal alignment on the time series medical record feature sequence using a spatiotemporal feature aligner to obtain a spatiotemporal alignment feature set, and performing feature extraction on the spatiotemporal alignment feature set to obtain a multi-dimensional feature vector; Performing feature space conversion on the multi-dimensional feature vector to obtain a unified feature space, and performing feature association calculation on the unified feature space to obtain a feature association network; The feature association network is subjected to feature fusion by a multimodal feature fusion device to obtain a fused feature graph, and the fused feature graph is subjected to feature optimization to obtain an optimized feature graph. A knowledge graph is constructed based on the optimized feature graph to obtain a unified medical record data view.
[0013] The present invention also provides a multi-source data fusion management system for medical record information, comprising: The acquisition module is used to collect data from the electronic medical record system, medical imaging workstation and laboratory information system of the medical institution to obtain the original medical record data set; An association module is used to perform semantic mapping and concept association on the original medical record data set based on a preset medical ontology knowledge graph to obtain a semantically associated medical record information network; An analysis module, configured to perform dynamic time series feature analysis on the semantically associated medical record information network through a time series feature extractor to obtain a time series medical record feature sequence; The fusion module is used to fuse heterogeneous data of the time series medical record feature sequence to obtain a unified medical record data view.
[0014] The present invention also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0015] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0016] The present invention provides a multi-source data fusion management method for medical record information, comprising the following steps: collecting data from the electronic medical record system, medical imaging workstation, and laboratory information system of a medical institution to obtain an original medical record data set; performing semantic mapping and concept association on the original medical record data set based on a preset medical ontology knowledge graph to obtain a semantically associated medical record information network; performing dynamic temporal feature analysis on the semantically associated medical record information network through a temporal feature extractor to obtain a temporal medical record feature sequence; and performing heterogeneous data fusion on the temporal medical record feature sequence to obtain a unified medical record data view. This method solves the technical problem that existing medical record management methods generally ignore the semantic relevance between data, which limits the ability to deeply understand and mine medical record information. It achieves the goal of obtaining a unified medical record data view by performing heterogeneous data fusion on the temporal medical record feature sequence, so that data from different sources and in different formats can be effectively organized and displayed under the same framework. This not only facilitates the review and use of medical personnel, but also improves the performance of the decision support system. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a schematic diagram of the steps of a multi-source data fusion management method for medical record information in one embodiment of the present invention; Figure 2 This is a structural block diagram of a multi-source data fusion management system for medical record information in one embodiment of the present invention; Figure 3 It is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.
[0018] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0020] like Figure 1 As shown, Figure 1 This is a schematic diagram of the steps of a multi-source data fusion management method for medical record information in one embodiment of the present invention; In one embodiment of the present invention, a multi-source data fusion management method for medical record information is provided, comprising the following steps: Step S1: collect data from the electronic medical record system, medical imaging workstation and laboratory information system of the medical institution to obtain the original medical record data set.
[0021] Specifically, in the above solution, data is collected from a medical institution's electronic medical record system, medical imaging workstation, and laboratory information system to obtain the original medical record dataset. This process forms the foundation of the entire multi-source data fusion management method. Specifically, the electronic medical record system stores structured or unstructured text data such as a patient's clinical history, diagnosis information, and treatment plan. The medical imaging workstation stores the patient's imaging examination results, such as X-rays, CT scans, and MRI data, which are typically in image format. The laboratory information system records the patient's laboratory test results, such as numerical data such as blood tests and urinalysis. To implement this data collection step, it is first necessary to connect with each system through standardized interfaces or protocols (such as HL7 and FHIR) to ensure data access from different sources while ensuring data integrity and consistency. For example, in an application scenario at a large general hospital, a doctor completes a comprehensive physical examination for a patient, including medical history, imaging examinations, and laboratory tests. The data collection module extracts the patient's medical history and diagnosis information from the electronic medical record system, obtains relevant imaging data from the medical imaging workstation, and reads laboratory test results from the laboratory information system. After preliminary processing, this data is integrated into a raw medical record dataset, which serves as the basis for subsequent semantic mapping and dynamic analysis. However, in actual operations, issues such as inconsistent data formats or uneven data quality may arise. Therefore, data cleaning techniques are required to preprocess the collected data, such as removing duplicate records, filling missing values, and unifying timestamp formats, to ensure that the quality of the raw medical record dataset meets the needs of subsequent analysis. This process not only enables efficient multi-source data collection but also lays a solid foundation for subsequent semantic mapping and concept association based on the medical ontology knowledge graph.
[0022] Step S2: semantic mapping and concept association are performed on the original medical record data set based on a preset medical ontology knowledge graph to obtain a semantically associated medical record information network.
[0023] Specifically, in the above solution, semantic mapping and concept association are performed on the original medical record dataset based on a pre-set medical ontology knowledge graph to generate a semantically linked medical record information network. This process is a key step in achieving deep integration of multi-source data. Specifically, a medical ontology knowledge graph is a structured knowledge representation that contains core medical concepts, terms, and their relationships, such as the association between diseases and symptoms, and the connection between test indicators and pathological conditions. To achieve semantic mapping, it is first necessary to match various data types in the original medical record dataset (including text records in the electronic medical record system, image descriptions in the medical imaging workstation, and numerical results in the laboratory information system) with standardized terms in the medical ontology knowledge graph. For example, in a large hospital application scenario, a patient's electronic medical record indicates "hypertension," while the laboratory information system records the patient's "elevated serum creatinine level." The medical ontology knowledge graph can map this disparate information into a unified conceptual system, identifying potential connections between "hypertension" and "abnormal renal function." At the same time, semantic mapping also needs to deal with heterogeneity between different data sources, such as converting unstructured text data into structured representations, or matching descriptive language in imaging reports with specific medical terms. In addition, the process of concept association further explores the deep connections between data. For example, through relational reasoning in the knowledge graph, it is found that abnormalities in a certain specific test indicator may be an early signal of a certain disease. This process not only achieves semantic enhancement of the original medical record data set, but also provides a richer and more meaningful data foundation for subsequent dynamic time series feature analysis. The resulting semantically associated medical record information network can help medical staff understand the patient's disease progression from a global perspective, and also provide important support for precision medicine and personalized treatment.
[0024] Step S3: Performing dynamic time series feature analysis on the semantically associated medical record information network through a time series feature extractor to obtain a time series medical record feature sequence.
[0025] Specifically, in the above scheme, a temporal feature extractor performs dynamic temporal feature analysis on a semantically linked medical record information network to obtain temporal medical record feature sequences. This process aims to uncover the temporal variations and trends of medical record data. Specifically, the semantically linked medical record information network integrates multiple patient data sources into a unified, semantically linked network structure. However, this data is still scattered across different time nodes and cannot directly reflect the dynamic evolution of the disease. To achieve dynamic temporal feature analysis, the temporal feature extractor first needs to identify and extract time-related features in the network, such as diagnosis timestamps in electronic medical record systems, the chronological order of imaging examinations in medical imaging workstations, and the changing curves of laboratory parameters in laboratory information systems. For example, in a large hospital scenario, a patient's medical record data may show a continuous increase in blood pressure over the past three months, while renal function-related indicators such as serum creatinine levels also show a gradual upward trend. This temporal variation pattern can be captured by the temporal feature extractor and converted into a set of feature sequences. At the same time, the dynamic analysis process also needs to combine the conceptual relationships in the semantically associated medical record information network to further reveal the potential connections between these features. For example, through analysis, it was found that there is a temporal causal relationship between "elevated blood pressure" and "worsening renal function." In addition, the time series feature extractor can also process data at different time scales, such as short-term fluctuation trends and long-term change patterns, thereby providing medical staff with a more comprehensive perspective on the development of the disease. The resulting time series medical record feature sequence provides important input for subsequent heterogeneous data fusion, and also provides a scientific basis for early warning of the disease, diagnostic decision-making, and treatment plan adjustment, helping doctors to more accurately grasp the patient's condition dynamics and develop more personalized medical plans.
[0026] Step S4: performing heterogeneous data fusion on the time series medical record feature sequence to obtain a unified medical record data view.
[0027] Specifically, in the above solution, heterogeneous data fusion is performed on time-series medical record feature sequences to obtain a unified medical record data view. This process is a key step in achieving efficient integration and utilization of multi-source data. Specifically, time-series medical record feature sequences have extracted dynamic features along the temporal dimension from the semantically linked medical record information network. However, these features may still exist in different data formats and structures, such as text descriptions in electronic medical record systems, image analysis results from medical imaging workstations, and numerical change curves in laboratory information systems. To achieve heterogeneous data fusion, it is first necessary to design a fusion framework that is compatible with multiple data types. This framework can transform time-series features from different sources into a unified representation while preserving their original semantic and temporal attributes. For example, in a large hospital application scenario, a patient's time-series medical record feature sequence may include continuous changes in blood pressure values, trend analysis of renal function indicators, and records of renal morphological changes during imaging examinations. These features exist in the form of numerical values, text, and images. Through heterogeneous data fusion technology, this dispersed information can be integrated into a unified medical record data view, allowing doctors to simultaneously view the patient's blood pressure trend, renal function deterioration, and the temporal evolution of imaging findings on a single interface. This process also requires resolving data conflicts and redundancies. For example, when different systems record the same metric differently, the fusion algorithm selects the optimal result based on credibility weights or consistency verification mechanisms. The resulting unified medical record data view not only provides medical staff with a comprehensive and intuitive overview of the patient's condition, but also supports further data mining and decision support applications, helping doctors develop treatment plans more quickly and improving the precision of medical services.
[0028] In a specific embodiment, the semantic mapping and concept association of the original medical record data set based on the preset medical ontology knowledge graph to obtain a semantically associated medical record information network includes: Performing word segmentation processing on the original medical record dataset through a preset medical term parser to obtain a medical term sequence, and performing semantic standardization mapping on the medical term sequence based on a preset medical ontology knowledge graph to obtain a standardized medical concept set; Performing a correlation analysis on the standardized medical concept set based on a semantic network builder to obtain a concept relationship map, and hierarchically organizing the concept relationship map to obtain a hierarchical semantic network; Performing semantic reasoning analysis on the hierarchical semantic network through a semantic reasoning engine to obtain a semantic reasoning rule set, and performing knowledge expansion based on the semantic reasoning rule set to obtain an extended semantic relationship network; Based on a preset ontology fusion processor, concept alignment and knowledge integration are performed on the extended semantic relationship network to obtain a semantically associated medical record information network.
[0029] Specifically, in the above solution, a pre-set medical term parser first performs word segmentation on the original medical record dataset to extract medical term sequences. For example, in a large hospital scenario, a patient has visited the doctor multiple times for a chronic condition, and their electronic medical record contains descriptive terms such as "hypertension," "proteinuria," "left ventricular hypertrophy," and "elevated fasting blood glucose." The medical term parser uses natural language processing techniques (such as named entity recognition) to identify specific medical terms from this text and arranges them into a medical term sequence in order of appearance. Assuming the patient has a total of 10 outpatient visits, 52 medical terms are extracted, forming a medical term sequence of length 52.
[0030] Subsequently, the extracted medical term sequences are semantically normalized and mapped based on a pre-set medical ontology knowledge graph to generate a standardized medical concept set. Medical ontology knowledge graphs typically include standard systems such as ICD-10, SNOMED CT, and LOINC. For example, the term "hypertension" may correspond to ICD-10 code I10.001, while "proteinuria" may be mapped to LOINC code 28496-7. This process uniformly maps potentially ambiguous or inconsistent terms (such as "hyperglycemia" and "elevated fasting blood glucose") to standardized concept identifiers. Continuing with this patient example, after semantic normalization, the original 52 terms were merged into 37 standardized medical concepts, forming the standardized medical concept set. Next, the standardized medical concept set is subjected to correlation analysis using a semantic network builder to generate a concept relationship graph, which is then further organized hierarchically to form a hierarchical semantic network. For example, analysis of the patient's standardized medical concept set revealed that "hypertension" and "left ventricular hypertrophy" co-occurred at multiple time points, with a frequency of 85%. A strong co-occurrence relationship also existed between "proteinuria" and "abnormal renal function." The semantic network builder used these connections to establish edges, representing potential causal or correlational relationships between them. This ultimately resulted in a concept relationship graph consisting of 37 nodes (medical concepts) and 58 edges (concept relationships). Based on this, the graph was hierarchically partitioned according to the hierarchical structure of the medical ontology knowledge graph (e.g., the ICD-10 disease classification system). For example, "hypertension" was classified under "circulatory system diseases" and "proteinuria" under "urinary system diseases," thereby forming a hierarchical semantic network with a clear semantic structure. Furthermore, a semantic reasoning engine performed semantic reasoning analysis on the hierarchical semantic network, generating a set of semantic reasoning rules. Based on these rules, knowledge was expanded to form an extended semantic relationship network. For example, if the knowledge graph already contains the rules "hypertension causes left ventricular hypertrophy" and "left ventricular hypertrophy increases the risk of heart failure," chain reasoning can lead to the new conclusion "hypertension increases the risk of heart failure." The semantic reasoning engine automatically extracts 12 similar rules and, based on external medical literature, adds new association paths, such as "microalbuminuria is an early sign of diabetic nephropathy." After reasoning and expansion, the number of nodes in the original graph increases from 37 to 45, and the number of edges from 58 to 82, forming a richer extended semantic relationship network. Finally, the extended semantic relationship network is subjected to concept alignment and knowledge integration based on a pre-defined ontology fusion processor, generating a semantically connected medical record information network. Because different medical institutions may use different terminology systems (for example, Hospital A uses ICD-10 and Hospital B uses ICD-9), the ontology fusion processor achieves cross-system concept alignment by searching term mapping tables and calculating similarity (for example, a cosine similarity greater than 0.85 is considered a match).During the integration process, 13 duplicate relationships were deleted and 7 conflicting relationships were merged, ultimately forming a semantically linked medical record information network consisting of 43 nodes and 75 edges. This network not only reflects the individual patient's disease progression path but also incorporates authoritative medical knowledge, demonstrating good interpretability and scalability.
[0031] In a specific embodiment, the semantic network builder is used to perform correlation analysis on the standardized medical concept set to obtain a concept relationship map, including: Performing co-occurrence analysis on the standardized medical concept set using a concept relationship miner to obtain a concept co-occurrence matrix, and performing correlation calculation on the concept co-occurrence matrix to obtain a concept correlation strength graph; Performing path mining on the concept association strength graph to obtain a concept semantic path set, and performing weight assignment on the concept semantic path set to obtain a weighted semantic path network; Performing topological optimization on the weighted semantic path network based on a semantic network builder to obtain an optimized semantic structure graph, and hierarchically dividing the optimized semantic structure graph to obtain a hierarchical semantic structure; The hierarchical semantic structure is graphed based on a knowledge graph generator to obtain a concept relationship graph.
[0032] Specifically, in the above solution, the process of performing association analysis on the standardized medical concept set based on a semantic network builder to generate a concept relationship graph is a complex but crucial step. First, a concept relationship miner performs co-occurrence analysis on the standardized medical concept set to generate a concept co-occurrence matrix. This matrix is then used to calculate associations to generate a concept relationship strength graph. Specifically, in a large hospital scenario, assume a patient undergoes a comprehensive examination for multiple health issues. Their electronic medical record system records terms such as "hypertension" and "diabetes." A medical imaging workstation may have descriptions of heart-related images, and a laboratory information system records various blood test values. The concept relationship miner can identify instances where these terms co-occur across different documents or data records. For example, "hypertension" and "abnormal renal function" may frequently appear together in the medical records of the same patient. Statistical analysis of these co-occurrences generates a concept co-occurrence matrix, which shows the frequency of co-occurrence of each pair of concepts. Then, using association calculation methods such as pointwise mutual information (PMI) or correlation coefficients, the strength of the association between each pair of concepts is extracted from the co-occurrence matrix to form a concept relationship strength graph. Next, path mining is performed on the concept association strength graph to discover sets of concept semantic paths. These paths are then assigned weights to construct a weighted semantic path network. This process aims to reveal the potential connections between concepts and their importance. Continuing with the patient example mentioned above, if it is known that "hypertension" and "abnormal renal function" are associated, path mining techniques can identify all possible paths connecting these two concepts from the concept association strength graph. For example, "hypertension" → "cardiovascular disease" → "abnormal renal function" is a path. Next, weights are assigned to each path based on the strength of the association between the nodes. Paths with higher weights indicate a closer connection between the concepts or greater clinical significance. The resulting weighted semantic path network not only demonstrates the relationships between concepts but also quantifies the importance of these relationships, facilitating subsequent in-depth analysis and application. Subsequently, topology optimization is performed on the weighted semantic path network using a semantic network builder to generate an optimized semantic structure graph. This graph is then hierarchically partitioned to obtain a hierarchical semantic structure. The purpose of topology optimization is to improve the interpretability and practicality of the network. During this process, algorithms may be employed to simplify the network structure, removing redundant connections or merging similar concept nodes, resulting in a more concise and clear semantic structure. For example, when dealing with a complex network containing numerous medical concepts, the connections between nodes can be adjusted to ensure that the most important paths and relationships are preserved while less important information is appropriately simplified. Furthermore, hierarchical partitioning involves dividing the entire network into different levels based on concept importance, category, or other criteria.For example, classification can be done by disease type (e.g., cardiovascular disease, endocrine disease) or stage of disease progression (early symptoms, mid-stage progression, late-stage manifestations), allowing doctors to more intuitively understand the overall course of a patient's condition. Finally, a knowledge graph generator constructs a graph of the hierarchical semantic structure, ultimately generating a concept relationship graph. The knowledge graph generator transforms the hierarchical semantic structure generated through a series of processing steps into a visual knowledge graph. This step requires not only effective presentation of concepts and their interrelationships, but also ensuring the graph's readability and interactivity. For example, in practical applications, doctors can quickly understand the connections between a disease and other related factors by viewing the concept relationship graph. For example, by viewing the concept relationship graph, doctors can understand how high blood pressure affects the function of other organs and which changes in laboratory indicators may indicate this effect. This approach not only provides doctors with more comprehensive patient information but also allows them to develop more precise and effective treatment plans based on the latest medical research findings. In summary, this series of steps, starting from raw medical terminology and undergoing multiple rounds of data processing and analysis, ultimately creates a structured, hierarchical concept relationship graph, significantly enhancing medical information management and clinical decision support capabilities. In order to better understand the above technical solutions, the following explanations are made. Medical concept relationship mining technical solution 1. Co-occurrence analysis and concept association strength graph construction concept co-occurrence matrix calculation concept relationship mining tool for standardized medical concept set Perform co-occurrence analysis. In the original medical record data set, traverse all text records and count any two concepts for each text segment. and The number of times the concept appears at the same time. Let the concept co-occurrence matrix be , whose elements Representation Concept and The number of times they appear together in all texts, that is: in, is the total number of medical record texts, Indicates the Medical record text, Is an indicator function. When the condition in the brackets is met, The value of is 1, otherwise it is 0. The correlation calculation is based on the concept co-occurrence matrix To calculate the strength of the association between concepts, the Pointwise Mutual Information (PMI) algorithm is used. Pointwise Mutual Information is used to measure the degree of association between two concepts. The formula is: in, It's a concept The probability of appearing in all medical record texts, It's a concept The probability of appearing in all medical record texts, It's a concept and The probability of simultaneous occurrence. In actual calculations, frequency is used to approximate probability: Substitute the above probability into the point mutual information formula to obtain the correlation between each concept pair, and then construct the concept correlation strength graph ,in It is a set of edges, and the weight of the edge is the corresponding concept association degree. 2. Path mining and weighted semantic path network construction (I) Concept semantic path set mining on concept association strength graph To perform path mining, we use the Depth-First Search (DFS) algorithm. Start by setting a path length limit , search from this node through DFS algorithm, the length is no more than During the search process, the node sequence passed is recorded to form the concept semantic path. Let the concept semantic path set mined be , where each path is a node sequence , Is the path The weight assignment assigns a weight to each concept semantic path, taking into account the association strength of the edges in the path and the path length. The weight calculation formula is: in, Is the path midpoint and The weight of the edge between them (i.e., the degree of association), Is the path By using this formula, the sum of the association strengths of the edges in the path is divided by the path length to obtain the average weight of each path, forming a weighted semantic path network. ,in is a set of path weights. 3. Topology optimization and hierarchical semantic structure construction, (I) Topology optimization semantic network builder for weighted semantic path network To perform topology optimization, we use the graph contraction algorithm in graph theory. This algorithm simplifies the network structure by merging paths and nodes with strong associations and redundancy. The specific steps are as follows: Calculate the closeness between nodes, which can be measured by the node degree and edge weight. Assume that the nodes The degree is , and the node The sum of the weights of the connected edges is , then the node The tightness It can be expressed as: in, It is a parameter that regulates the degree of influence of the edge weight, and its value range is For pairs of nodes with high closeness and similar semantics, merge them into one node and update the relevant paths and weights. Repeat this process until the network structure reaches the preset optimization standard and obtains the optimized semantic structure graph. Hierarchical division, to optimize the semantic structure diagram To perform hierarchical division, a method based on concept semantic similarity and node importance is used. First, the semantic similarity between nodes is calculated. The cosine similarity method based on embedding vectors mentioned above can be used. Then, the importance of the node is determined based on the node's degree, weight, and semantic similarity with other nodes. Let node Importance The calculation formula is: in, 、 、 is a parameter that adjusts the degree of influence of various factors, and Nodes are divided into different levels according to their importance, with nodes of high importance placed in the upper level and nodes of low importance placed in the lower level, thus obtaining a hierarchical semantic structure. ,in Indicates the 4. Conceptual Relationship Graph Construction Knowledge Graph Generator is based on hierarchical semantic structure Build a concept relationship map The nodes in the hierarchical semantic structure are used as the nodes of the graph, the connection relationship between the nodes is retained as the edge of the graph, and the weight of the edge and the hierarchical information of the nodes are used as the attributes of the graph. The final concept relationship graph is ,in It is a set of nodes, corresponding to the nodes in the hierarchical semantic structure; is a set of edges, corresponding to the connection relationship between nodes; It is a set of attributes, including information such as edge weights and node levels.
[0033] In a specific embodiment, the dynamic time series feature analysis of the semantically associated medical record information network is performed by a time series feature extractor to obtain a time series medical record feature sequence, including: Extracting trend components from the semantically associated medical record information network using a temporal feature extractor to obtain disease progression trend features, and dividing the disease progression trend features into disease stages to obtain staged disease progression features; Based on the staged disease progression characteristics, the semantically associated medical record information network is mined for periodic patterns to obtain disease recurrence period characteristics, and the disease recurrence period characteristics are quantified for fluctuation amplitude to obtain quantified disease recurrence characteristics; Performing outlier detection on the semantically associated medical record information network based on the quantified disease recurrence features to obtain a disease abnormal event sequence, and performing event correlation analysis on the disease abnormal event sequence to obtain a disease key event feature; Based on the key event characteristics of the disease, the abnormal disease event sequence is subjected to temporal logic reasoning and disease records to obtain a temporal medical record feature sequence.
[0034] Specifically, in the above solution, the process of using a temporal feature extractor to perform dynamic temporal feature analysis on the semantically linked medical record information network to obtain a temporal medical record feature sequence is a complex but crucial step. First, the temporal feature extractor extracts trend components from the semantically linked medical record information network to obtain disease progression trend features. These features are then further classified into disease stages to form staged disease progression features. Specifically, in a large hospital application scenario, suppose a patient undergoes a comprehensive examination for multiple health issues. Their electronic medical record system records diagnoses such as "hypertension" and "diabetes." The medical imaging workstation may have heart-related image descriptions, and the laboratory information system records various blood test values. The temporal feature extractor can identify patterns in these data over time. For example, a patient's blood pressure values have consistently increased over the past few years, which can be considered a trend component of disease progression. Based on this trend component, the algorithm automatically divides the disease process into different stages, such as early, middle, and late hypertension, to more accurately understand the disease's progression. Next, based on the staged disease progression characteristics, the semantically linked medical record information network is mined for cyclical patterns, identifying disease recurrence cycle characteristics and quantifying the fluctuation amplitude of these characteristics to obtain quantitative disease recurrence signatures. In this process, the algorithm not only focuses on long-term disease trends but also looks for cyclical patterns within them. Continuing with the aforementioned patient example, if the patient's blood sugar levels exhibit seasonal fluctuations—rising in winter and stabilizing in summer—this cyclical pattern can be identified using cyclical pattern mining techniques. Furthermore, by quantifying the fluctuation amplitude of these cyclical patterns, such as calculating the difference between peak and trough blood sugar levels, more detailed quantitative disease recurrence signatures can be obtained. This process helps doctors better understand the natural progression of the disease and its influence by external factors, providing a basis for developing personalized treatment plans. Subsequently, based on the quantified disease recurrence signatures, the semantically linked medical record information network is subjected to outlier detection, filtering out abnormal disease event sequences. These events are then subjected to event correlation analysis to obtain key disease event signatures. In real-world applications, certain physiological indicators or symptoms of patients may suddenly experience abnormal fluctuations, such as a sudden heart attack or severe hypoglycemia. Outlier detection technology can accurately identify these abnormal events from large amounts of time series data and mark them as part of a sequence of abnormal disease events. Next, by analyzing the correlations between these abnormal events—for example, exploring whether a heart attack is related to strenuous exercise in the preceding days, or whether hypoglycemia is associated with changes in dietary habits—the key factors leading to the abnormal events can be revealed, forming a key disease event signature. This process not only helps doctors quickly locate the possible source of the problem but also provides guidance for preventing similar events from recurring.Finally, based on the key disease event features, temporal logical reasoning is performed on the abnormal disease event sequence and the disease record, ultimately generating a temporal medical record feature sequence. In this step, the algorithm uses known key event features, combined with the patient's detailed medical history and clinical manifestations, to predict potential future health issues through logical reasoning. For example, if a patient has experienced an acute exacerbation of hypertension due to a high-salt diet, the system can issue an early warning and recommend dietary adjustments to prevent further disease progression under similar dietary conditions in the future. Simultaneously, these reasoning results and disease records are integrated into a complete temporal medical record feature sequence, allowing doctors to view the patient's entire medical history, disease progression trends, potential risks, and recommended interventions in a unified view. This approach not only improves the quality and efficiency of medical services but also promotes the development of personalized medicine, ensuring that each patient receives the most appropriate health management plan. In summary, this series of steps, starting from raw time series data and undergoing multiple rounds of data processing and analysis, ultimately forms a structured, ordered temporal medical record feature sequence, significantly enhancing the capabilities of medical information management and clinical decision support. A detailed technical solution for temporal feature analysis based on a semantically linked medical record information network is provided, along with disease progression trend feature extraction and stage classification. Trend component extraction uses a time series feature extractor to analyze the semantically linked medical record information network. Disease-related data in the semantically linked medical record information network changes over time, forming a time series. To extract disease progression trend characteristics, the Empirical Mode Decomposition (EMD) algorithm is employed. Based on the local characteristic scale of the signal, the EMD algorithm decomposes complex time series data into a finite number of intrinsic mode functions (IMFs) and a residual component. The residual component represents the trend component of the time series. Screening process: Assume the original time series is . , find out All the maximum and minimum points of are interpolated by cubic spline to get the upper envelope and lower envelope , calculate the mean , and get the difference .examine Whether the IMF condition is met (the number of extreme points is equal to the number of zero-crossing points or the difference is at most 1; at any time, the upper envelope determined by the local maximum point and the lower envelope determined by the local minimum point are locally symmetric about the time axis). If not, As a new Repeat the above steps until the conditions are met. , which is the first IMF component . Iterative decomposition: Compute ,Will As a new Repeat the above screening process to obtain the second IMF component , and so on, until the residual component Become a monotonic function or meet the stopping condition. The final residual component That is the disease progression trend feature. 1.2 Disease stage division based on the extracted disease progression trend feature , Dynamic Time Warping (DTW) algorithm is used to match the preset disease stage template to achieve disease stage division. Assume that the preset disease stage template set is , each template is a time series. For the disease progression trend characteristics , calculate its difference with each template through DTW algorithm distance The DTW algorithm calculates the nonlinear alignment of two time series on the time axis and obtains the cumulative distance under the optimal matching path. The length is , The length is , build a The distance matrix ,in Then search for arrive The optimal path , so that the cumulative distance Minimum. The disease stage corresponding to the template with the smallest distance is selected as the current stage of the disease, thereby obtaining the staged disease progression characteristics. Periodic pattern mining and recurrence feature quantification, periodic pattern mining is based on the staged disease progression characteristics, using Fourier transform (Fourier Transform) to mine the disease recurrence period characteristics. Fourier transform can convert time domain signals into frequency domain signals, and by analyzing the peaks in the frequency domain signals, the periodic components of the signals are determined. For the time series composed of staged disease progression characteristics , its discrete Fourier transform (DFT) formula is: in, is the length of the time series, is the frequency domain representation, Is an imaginary unit. After calculating the frequency domain results, find the frequency component with larger amplitude, and according to the relationship between frequency and period ( For the cycle, As frequency), determine the cycle of disease recurrence and obtain the disease recurrence cycle characteristics. Fluctuation amplitude quantification quantifies the fluctuation amplitude of disease recurrence cycle characteristics, and uses standard deviation to measure the degree of fluctuation of disease recurrence characteristics within the cycle. Suppose the time series corresponding to the disease recurrence cycle characteristics is , which is divided into multiple subsequences according to the period , For each subsequence , calculate its mean ( is the subsequence length), then the standard deviation of the subsequence is: The standard deviation of all subsequences is combined to obtain the quantitative disease recurrence characteristics, reflecting the fluctuation range of disease recurrence. Outlier detection and key event analysis, outlier detection is based on the quantitative disease recurrence characteristics. The isolation forest algorithm is used to detect outliers on the semantically associated medical record information network to obtain the disease abnormal event sequence. The isolation forest algorithm divides the data by constructing multiple isolation trees. The path length of each data point in the isolation tree represents its degree of isolation. For a sample , in The path length in an isolated tree is , the average path length is ( is the number of isolated trees). Sample The anomaly score calculation formula is: in, The sample size is The theoretical value of the average path length when the anomaly score When the set threshold is exceeded, the sample is judged As an outlier, all the events corresponding to the outliers are combined into a disease abnormal event sequence. Event correlation analysis is performed on the disease abnormal event sequence. The Pearson Correlation Coefficient is used to calculate the correlation between events and obtain the key event characteristics of the disease. Suppose there are two events in the disease abnormal event sequence corresponding to the time series data and , and its Pearson correlation coefficient calculation formula is: in, and They are time series and The mean of is the length of the time series. Events with larger absolute values of correlation coefficients are selected as key disease event features. 4. Temporal logic reasoning and medical record feature sequence generation Based on the key disease event features, a temporal logic reasoning algorithm (such as a reasoning method based on temporal logic) is used to reason about the abnormal disease event sequence. Temporal logic describes and reasons about the temporal relationship of events by introducing time operators (such as "until", "next", etc.). For example, using Linear Temporal Logic (LTL), logical formulas are defined to describe the temporal relationship and causal relationship between events. By performing logical judgment and reasoning on the abnormal disease event sequence, combined with disease record information, related events and information are integrated in chronological order to generate a temporal medical record feature sequence, providing more comprehensive and orderly information for disease diagnosis, treatment and research.
[0035] In a specific embodiment, the periodic pattern mining of the semantically associated medical record information network based on the staged disease progression characteristics to obtain disease recurrence period characteristics includes: Performing subsequence segmentation on the semantically associated medical record information network based on the staged disease progression features to obtain disease progression subsequences, and performing similarity measurement on the disease progression subsequences to obtain a subsequence similarity matrix; performing cluster analysis on the disease progression subsequences based on the subsequence similarity matrix to obtain disease progression pattern clusters, and performing pattern feature extraction on the disease progression pattern clusters to obtain disease progression pattern feature vectors; performing periodic pattern matching on the disease progression pattern cluster based on the disease progression pattern feature vector to obtain a candidate recurrence period set, and performing a statistical significance test on the candidate recurrence period set to obtain a significant recurrence period set; Estimating the cycle length of the candidate recurrence cycle set based on the significant recurrence cycle set to obtain a recurrence cycle length value, and performing confidence interval calculation on the recurrence cycle length value to obtain a recurrence cycle confidence interval; The significant recurrence cycle set is cycle-integrated based on the recurrence cycle confidence interval to obtain the disease recurrence cycle characteristics.
[0036] Specifically, in the above scheme, mining periodic patterns in a semantically linked medical record information network based on staged disease progression features to identify disease recurrence cycle characteristics is a systematic and multi-layered analytical process. First, the semantically linked medical record information network is segmented into subsequences based on the staged disease progression features to obtain disease progression subsequences. These subsequences are then similarly measured to generate a subsequence similarity matrix. Specifically, in a large hospital scenario, assume a patient undergoes long-term follow-up for a chronic disease. Their electronic medical record system records time-varying data such as blood pressure and blood sugar levels. This data can be divided into multiple subsequences based on the disease progression stage (e.g., early-stage hypertension, mid-stage hypertension). For example, if a patient's blood pressure has fluctuated multiple times over the past few years, these blood pressure data can be segmented into several subsequences according to different time periods or stages of disease progression. Then, by measuring the similarity between these subsequences, such as by calculating the distance between them using the dynamic time warping (DTW) algorithm, a subsequence similarity matrix can be constructed. This matrix quantifies the degree of similarity between each pair of subsequences. Next, cluster analysis is performed on the disease progression subsequences based on the subsequence similarity matrix to identify disease progression pattern clusters. Disease progression pattern feature vectors are then extracted from these pattern clusters. In this process, cluster analysis can help identify subsequences with similar change trends, thereby forming distinct disease progression pattern clusters. Continuing with the aforementioned patient example, suppose cluster analysis reveals several key blood pressure change patterns: one characterized by a continuous upward trend, and another characterized by fluctuations within a certain range but overall stability. For each of these pattern clusters, core features, such as the average rate of change, maximum peak value, and minimum trough value, are extracted to form a disease progression pattern feature vector. These feature vectors not only summarize the key characteristics of each pattern cluster but also provide a basis for subsequent periodic pattern matching. Subsequently, periodic pattern matching is performed on the disease progression pattern clusters based on the disease progression pattern feature vectors to screen candidate recurrence period sets. These candidate sets are then statistically tested for significance, ultimately identifying significant recurrence period sets. In this step, using the existing disease progression pattern feature vectors, the algorithm attempts to identify pattern clusters that exhibit periodic variation patterns. For example, if a patient's blood sugar levels exhibit a pattern of increasing in the winter and decreasing in the summer each year, this cyclical variation can be identified as part of a candidate recurring cycle. Next, statistical significance tests, such as chi-square tests or t-tests, can be performed on this set of candidate recurring cycles to verify which cycles are truly significant and not the result of random fluctuations. This step ensures the authenticity and reliability of the selected cycles.Next, based on the set of significant recurrence cycles, the cycle lengths of the candidate recurrence cycles are estimated, resulting in recurrence cycle length values. Confidence intervals are then calculated for these lengths to obtain recurrence cycle confidence intervals. To more accurately understand the recurrence cycle of a disease, detailed length estimates are required for each cycle in the initially selected set of significant recurrence cycles. For example, if seasonal fluctuations in the patient's blood sugar levels are confirmed, the specific cycle length of these fluctuations can be further calculated, such as approximately 12 months. Furthermore, to assess the accuracy of this estimate, corresponding confidence intervals are calculated, providing the physician with an understanding of the possible range and reliability of the cycle length estimate. Finally, based on the recurrence cycle confidence intervals, the set of significant recurrence cycles is integrated to ultimately generate disease recurrence cycle signatures. This step aims to integrate all verified and estimated significant recurrence cycles into a comprehensive signature that reflects the recurrence patterns of the disease. For example, by comprehensively considering the patient's blood sugar levels, blood pressure fluctuations, and other relevant physiological indicators, a signature set encompassing the recurrence cycles of multiple diseases can be constructed. These cycle signatures not only reveal the underlying patterns of disease development but also provide important insights for prevention and treatment. This approach not only helps doctors better understand the natural course of the disease but also guides the development of more scientific and rational treatment plans, improving the quality and effectiveness of medical services. In short, this series of steps, starting from raw time series data and undergoing multiple rounds of data processing and analysis, ultimately forms a structured and orderly disease recurrence cycle, greatly enhancing the capabilities of medical information management and clinical decision support.
[0037] In a specific embodiment, the process of performing temporal logic reasoning and disease recording on the abnormal disease event sequence based on the key disease event characteristics to obtain a temporal medical record feature sequence includes: Performing event causal correlation analysis on the abnormal disease event sequence based on the key event characteristics of the disease to obtain an event causal correlation graph, and performing causal strength calculation on the event causal correlation graph to obtain a weighted causal correlation network; Performing temporal constraint mining on the disease abnormal event sequence based on the weighted causal association network to obtain a temporal constraint rule set, and performing constraint strength evaluation on the temporal constraint rule set to obtain an effective temporal constraint network; Performing temporal path reasoning on the weighted causal association network based on the effective temporal constraint network to obtain a set of candidate disease development pathways, and performing path probability evaluation on the set of candidate disease development pathways to obtain a probabilistic development pathway diagram; Based on the probabilistic development path diagram, the clinical trajectory of the abnormal disease event sequence is reconstructed and the disease record is recorded to obtain a time-series medical record feature sequence.
[0038] Specifically, in the above solution, the process of performing temporal logic reasoning and disease records on the abnormal disease event sequence based on key disease event characteristics to obtain a temporal medical record feature sequence is a complex and multi-layered analytical process. First, causal association analysis is performed on the abnormal disease event sequence based on the key disease event characteristics to generate an event causal association graph. This graph is then further subjected to causal strength calculation to form a weighted causal association network. Specifically, in a large hospital application scenario, suppose a patient undergoes a comprehensive examination for multiple health issues. Their electronic medical record system records diagnostic information such as "hypertension" and "diabetes." The medical imaging workstation may have heart-related image descriptions, and the laboratory information system records various blood test values. When a patient's physiological indicators or symptoms show abnormal fluctuations, such as a sudden heart attack or severe hypoglycemia, these abnormal events are marked as part of the abnormal disease event sequence. By performing causal association analysis on these events, potential causal relationships can be identified, such as whether a heart attack is related to strenuous exercise in the previous few days, or whether hypoglycemia is related to changes in dietary habits. The event causal association graph constructed in this way not only reveals potential connections between different events but also provides a basis for further causal strength calculations. Using statistical methods such as the Granger Causality Test, the causal strength between each pair of events can be quantified, thereby forming a weighted causal association network. Next, temporal constraint mining is performed on the disease anomaly event sequence based on the weighted causal association network to discover temporal constraint rule sets. These rules are then evaluated for constraint strength, ultimately determining a valid temporal constraint network. In this process, the algorithm not only identifies causal relationships between events but also considers the temporal order and time intervals between their occurrences. Continuing with the patient example mentioned above, if each heart attack is preceded by a history of strenuous exercise, temporal constraint mining techniques can identify this pattern and establish corresponding constraint rules, such as the assumption that heart attacks typically occur within 24 hours of strenuous exercise. Next, by evaluating the strength of these temporal constraint rule sets, using metrics such as support and confidence to measure rule effectiveness, highly reliable rules can be identified, forming a valid temporal constraint network. This process helps more accurately predict potential future health issues and their potential time windows. Subsequently, the algorithm uses the valid temporal constraint network to perform temporal path inference on the weighted causal association network, screening a set of candidate disease progression pathways. These pathways are then evaluated for path probability to generate a probabilistic progression pathway graph. In this step, the algorithm utilizes the existing valid temporal constraint network and the weighted causal association network to attempt to infer all possible disease progression pathways.For example, for the aforementioned patient, if a significant causal relationship between strenuous exercise and heart attack has been established and corresponding time constraints have been set, a typical disease progression path can be inferred: strenuous exercise → elevated blood pressure → heart attack. By probabilistically assessing all possible pathways, such as using a Bayesian network or Markov model to calculate the likelihood of each path occurring, a probabilistic progression diagram can be generated that encompasses multiple possible progression paths. This step not only helps doctors understand the potential trajectory of the disease but also provides a basis for developing personalized preventive measures. Finally, based on the probabilistic progression diagram, the clinical trajectory and disease record of the sequence of abnormal disease events are reconstructed, ultimately generating a time-series medical record feature sequence. This step aims to integrate all analysis results into a complete and organized clinical trajectory description. For example, by comprehensively considering the patient's blood pressure changes, heart attack, and other relevant physiological indicators, a detailed clinical trajectory can be constructed. This trajectory not only reflects the patient's past health status and abnormal events but also predicts potential future risks and development trends. Furthermore, to facilitate subsequent medical decision support, these analysis results need to be recorded to form a structured time-series medical record feature sequence. This allows doctors to not only review a patient's current health status but also understand multiple dimensions of information, such as their past medical history and family genetic risk factors, helping to develop more precise and effective treatment plans. This approach not only improves the quality and efficiency of medical services but also promotes the development of personalized medicine, enabling each patient to receive the health management plan that best suits them. In short, this series of steps, starting from the original abnormal disease event and undergoing multiple rounds of data processing and analysis, ultimately forms a structured, ordered, time-series medical record feature sequence, greatly enhancing the capabilities of medical information management and clinical decision support.
[0039] In a specific embodiment, the heterogeneous data fusion of the time series medical record feature sequence to obtain a unified medical record data view includes: Performing spatiotemporal alignment on the time series medical record feature sequence using a spatiotemporal feature aligner to obtain a spatiotemporal alignment feature set, and performing feature extraction on the spatiotemporal alignment feature set to obtain a multi-dimensional feature vector; Performing feature space conversion on the multi-dimensional feature vector to obtain a unified feature space, and performing feature association calculation on the unified feature space to obtain a feature association network; The feature association network is subjected to feature fusion by a multimodal feature fusion device to obtain a fused feature graph, and the fused feature graph is subjected to feature optimization to obtain an optimized feature graph. A knowledge graph is constructed based on the optimized feature graph to obtain a unified medical record data view.
[0040] Specifically, in the above solution, the process of fusing heterogeneous data from time-series medical record feature sequences to obtain a unified view of the medical record data is a multi-layered and complex analytical process. First, the time-series medical record feature sequences are aligned in their spatiotemporal dimensions using a spatiotemporal feature aligner to generate a spatiotemporally aligned feature set. Multidimensional feature vectors are then extracted from this set. Specifically, in a large hospital scenario, suppose a patient undergoes a comprehensive examination for multiple health issues. Their electronic medical record system records diagnoses such as "hypertension" and "diabetes." A medical imaging workstation may have cardiac-related image descriptions, and a laboratory information system records various blood test values. These data not only come from different systems but also contain temporal trends and spatial distribution characteristics. For example, the patient's blood pressure changes over time, imaging findings of cardiac function, and the spatial distribution of various laboratory parameters. The spatiotemporal feature aligner identifies and adjusts the timestamps and spatial coordinates of these data to ensure that they can be compared and analyzed within the same temporal and spatial framework. By extracting features from these spatiotemporally aligned data, such as calculating statistical features like mean, variance, and peak, a multidimensional feature vector can be generated that comprehensively reflects various aspects of the patient's health status. Next, the multidimensional feature vector undergoes feature space transformation to generate a unified feature space. Feature correlations are then calculated within this space to form a feature correlation network. In this process, the algorithm not only integrates data from different sources into a unified representation but also explores potential connections between these data. For example, continuing with the patient example above, assuming a multidimensional feature vector containing blood pressure values, cardiac imaging descriptions, and laboratory parameters has been generated, feature space transformation techniques can be used to map these features into a common feature space. This step typically involves dimensionality reduction methods such as principal component analysis (PCA), t-SNE, or autoencoders to remove redundant information and retain the most important features. Next, by performing correlation calculations on the features in this unified feature space, such as using correlation coefficients or mutual information to measure the relationship between each pair of features, a feature correlation network can be constructed. This network reveals the inherent connections between different features, providing a deeper understanding of the patient's condition. The multimodal feature fusion engine then fuses the features of the feature association network, extracts a fused feature graph, and further optimizes the features of this graph to generate an optimized feature graph. In this step, the algorithm utilizes the existing feature association network and attempts to integrate all relevant information to form a comprehensive feature description that reflects the patient's health status. For example, for the aforementioned patient, if a feature association network has already been constructed that shows the relationship between blood pressure and cardiac function, the multimodal feature fusion engine can combine this information with data from other sources (such as laboratory indicators) to generate a fused feature graph.This graph not only illustrates the interactions between different features but also provides a foundation for subsequent feature optimization. The feature optimization process aims to improve the quality of the fused feature graph, removing noise and unnecessary complexity, resulting in a more concise and clear final result. Common optimization methods include regularization and sparse representation, which can help reduce the risk of overfitting and improve model generalization. Finally, a knowledge graph is constructed based on the optimized feature graph, ultimately resulting in a unified medical record data view. This step aims to integrate all analysis results into a structured, easy-to-understand, and user-friendly representation of clinical data. For example, by comprehensively considering a patient's blood pressure, cardiac function, and other relevant physiological indicators, a detailed medical record data view can be constructed. This view not only displays the patient's past health status and abnormal events, but also predicts potential future risks and trends. To facilitate subsequent medical decision support, these analysis results need to be recorded to form a structured, unified medical record data view. This allows doctors to not only review the patient's current health status but also understand multiple dimensions of information, such as their medical history and family genetic risk factors, helping to formulate more precise and effective treatment plans. This approach not only improves the quality and efficiency of medical services but also promotes the development of personalized medicine, enabling each patient to receive the health management plan that best suits them. In short, this series of steps, starting with raw multi-source data and undergoing multiple rounds of data processing and analysis, ultimately forms a structured, organized, and unified medical record data view, significantly enhancing medical information management and clinical decision support capabilities.
[0041] The above describes the multi-source data fusion management method of medical record information in the embodiment of the present invention. The following describes the multi-source data fusion management system of medical record information in the embodiment of the present invention. Figure 2 An embodiment of the multi-source data fusion management system for medical record information in the embodiment of the present invention includes: The acquisition module 21 is used to collect data from the electronic medical record system, medical imaging workstation and laboratory information system of the medical institution to obtain the original medical record data set; An association module 22 is configured to perform semantic mapping and concept association on the original medical record data set based on a preset medical ontology knowledge graph to obtain a semantically associated medical record information network; An analysis module 23 is configured to perform dynamic time series feature analysis on the semantically associated medical record information network through a time series feature extractor to obtain a time series medical record feature sequence; The fusion module 24 is used to perform heterogeneous data fusion on the time series medical record feature sequence to obtain a unified medical record data view.
[0042] In this embodiment, for the specific implementation of each unit in the above system embodiment, please refer to the above method embodiment, which will not be repeated here.
[0043] Reference Figure 3 The embodiment of the present invention further provides a computer device, the internal structure of which can be as follows Figure 3 As shown. The computer device includes a processor, memory, display screen, input device, network interface and database connected via a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the above method is implemented.
[0044] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0045] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the above-described method when executed by a processor. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0046] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media provided herein and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM.
[0047] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.
[0048] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A multi-source data fusion management method for medical record information, characterized in that: The following steps are involved: Collect data from the electronic medical record system, medical imaging workstation and laboratory information system of medical institutions to obtain the original medical record data set; Performing semantic mapping and concept association on the original medical record data set based on a preset medical ontology knowledge graph to obtain a semantically associated medical record information network; Performing dynamic time series feature analysis on the semantically associated medical record information network through a time series feature extractor to obtain a time series medical record feature sequence; Heterogeneous data fusion is performed on the time series medical record feature sequence to obtain a unified medical record data view.
2. The multi-source data fusion management method for medical record information according to claim 1 is characterized in that: The original medical record data set is semantically mapped and conceptually associated based on a preset medical ontology knowledge graph to obtain a semantically associated medical record information network, including: Performing word segmentation processing on the original medical record dataset through a preset medical term parser to obtain a medical term sequence, and performing semantic standardization mapping on the medical term sequence based on a preset medical ontology knowledge graph to obtain a standardized medical concept set; Performing a correlation analysis on the standardized medical concept set based on a semantic network builder to obtain a concept relationship map, and hierarchically organizing the concept relationship map to obtain a hierarchical semantic network; Performing semantic reasoning analysis on the hierarchical semantic network through a semantic reasoning engine to obtain a semantic reasoning rule set, and performing knowledge expansion based on the semantic reasoning rule set to obtain an extended semantic relationship network; Based on a preset ontology fusion processor, concept alignment and knowledge integration are performed on the extended semantic relationship network to obtain a semantically associated medical record information network.
3. The multi-source data fusion management method for medical record information according to claim 2 is characterized in that: The semantic network builder is used to perform correlation analysis on the standardized medical concept set to obtain a concept relationship map, including: Performing co-occurrence analysis on the standardized medical concept set using a concept relationship miner to obtain a concept co-occurrence matrix, and performing correlation calculation on the concept co-occurrence matrix to obtain a concept correlation strength graph; Performing path mining on the concept association strength graph to obtain a concept semantic path set, and performing weight assignment on the concept semantic path set to obtain a weighted semantic path network; Performing topological optimization on the weighted semantic path network based on a semantic network builder to obtain an optimized semantic structure graph, and hierarchically dividing the optimized semantic structure graph to obtain a hierarchical semantic structure; The hierarchical semantic structure is graphed based on a knowledge graph generator to obtain a concept relationship graph.
4. The multi-source data fusion management method for medical record information according to claim 1, characterized in that: The method of performing dynamic time series feature analysis on the semantically associated medical record information network by using a time series feature extractor to obtain a time series medical record feature sequence includes: Extracting trend components from the semantically associated medical record information network using a temporal feature extractor to obtain disease progression trend features, and dividing the disease progression trend features into disease stages to obtain staged disease progression features; Based on the staged disease progression characteristics, the semantically associated medical record information network is mined for periodic patterns to obtain disease recurrence period characteristics, and the disease recurrence period characteristics are quantified for fluctuation amplitude to obtain quantified disease recurrence characteristics; Performing outlier detection on the semantically associated medical record information network based on the quantified disease recurrence features to obtain a disease abnormal event sequence, and performing event correlation analysis on the disease abnormal event sequence to obtain a disease key event feature; Based on the key event characteristics of the disease, the abnormal disease event sequence is subjected to temporal logic reasoning and disease records to obtain a temporal medical record feature sequence.
5. The multi-source data fusion management method for medical record information according to claim 4 is characterized in that: The periodic pattern mining of the semantically associated medical record information network based on the staged disease progression characteristics to obtain disease recurrence period characteristics includes: Performing subsequence segmentation on the semantically associated medical record information network based on the staged disease progression features to obtain disease progression subsequences, and performing similarity measurement on the disease progression subsequences to obtain a subsequence similarity matrix; performing cluster analysis on the disease progression subsequences based on the subsequence similarity matrix to obtain a disease progression pattern cluster, and performing pattern feature extraction on the disease progression pattern cluster to obtain a disease progression pattern feature vector; performing periodic pattern matching on the disease progression pattern cluster based on the disease progression pattern feature vector to obtain a candidate recurrence period set, and performing a statistical significance test on the candidate recurrence period set to obtain a significant recurrence period set; Estimating the cycle length of the candidate recurrence cycle set based on the significant recurrence cycle set to obtain a recurrence cycle length value, and performing confidence interval calculation on the recurrence cycle length value to obtain a recurrence cycle confidence interval; The significant recurrence cycle set is cycle-integrated based on the recurrence cycle confidence interval to obtain the disease recurrence cycle characteristics.
6. The multi-source data fusion management method for medical record information according to claim 5 is characterized in that: The method of performing temporal logic reasoning and disease recording on the abnormal disease event sequence based on the key disease event characteristics to obtain a temporal medical record feature sequence includes: Performing event causal association analysis on the abnormal disease event sequence based on the key event characteristics of the disease to obtain an event causal association graph, and performing causal strength calculation on the event causal association graph to obtain a weighted causal association network; Performing temporal constraint mining on the disease abnormal event sequence based on the weighted causal association network to obtain a temporal constraint rule set, and performing constraint strength evaluation on the temporal constraint rule set to obtain an effective temporal constraint network; Performing temporal path reasoning on the weighted causal association network based on the effective temporal constraint network to obtain a set of candidate disease development pathways, and performing path probability evaluation on the set of candidate disease development pathways to obtain a probabilistic development pathway diagram; Based on the probabilistic development path diagram, the clinical trajectory of the abnormal disease event sequence is reconstructed and the disease record is recorded to obtain a time-series medical record feature sequence.
7. The multi-source data fusion management method for medical record information according to claim 1, characterized in that: The heterogeneous data fusion of the time series medical record feature sequence to obtain a unified medical record data view includes: Performing spatiotemporal alignment on the time series medical record feature sequence using a spatiotemporal feature aligner to obtain a spatiotemporal alignment feature set, and performing feature extraction on the spatiotemporal alignment feature set to obtain a multi-dimensional feature vector; Performing feature space conversion on the multi-dimensional feature vector to obtain a unified feature space, and performing feature association calculation on the unified feature space to obtain a feature association network; The feature association network is subjected to feature fusion by a multimodal feature fusion device to obtain a fused feature graph, and the fused feature graph is subjected to feature optimization to obtain an optimized feature graph. A knowledge graph is constructed based on the optimized feature graph to obtain a unified medical record data view.
8. A multi-source data fusion management system for medical record information, characterized in that: include: The acquisition module is used to collect data from the electronic medical record system, medical imaging workstation and laboratory information system of the medical institution to obtain the original medical record data set; An association module is used to perform semantic mapping and concept association on the original medical record data set based on a preset medical ontology knowledge graph to obtain a semantically associated medical record information network; An analysis module, configured to perform dynamic time series feature analysis on the semantically associated medical record information network through a time series feature extractor to obtain a time series medical record feature sequence; The fusion module is used to fuse heterogeneous data of the time series medical record feature sequence to obtain a unified medical record data view.
9. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Medical data management method and system based on artificial intelligence and storage medium
CN118609743A
Medical diagnosis intelligent decision-making system based on multi-modal data fusion
CN119495423A
Multi-modal medical data fusion and analysis platform
CN119622621A
Cited By
Intelligent pathology review method and system based on case history comparison
CN120913888A
A pathological intelligent review method and system based on case history comparison
CN120913888B
Chronic disease data dynamic desensitization method, device and equipment and storage medium
CN121211500A
Industrial system modeling method, device and equipment based on dynamic view
CN121411190A
Puerpera health condition index monitoring and analyzing system based on multi-data fusion
CN121439229A