Medical record AI intelligent integration and analysis system based on big data

By constructing a co-occurrence matrix to analyze the relationship between subcategories and major categories, the problem of ignoring the semantic differences between subcategories and major categories in existing technologies is solved. This enables hierarchical classification and personalized diagnosis and treatment analysis of medical data, improving the interpretability of the data and clinical decision support.

CN121545656AActive Publication Date: 2026-02-17HUNAN RENJI BIOTECHNOLOGY CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202610076759.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-21
Publication Date
2026-02-17
Estimated Expiration
2046-01-21

AI Technical Summary

Technical Problem

Existing medical data clustering methods ignore the clinical semantic differences between small and large categories, resulting in inaccurate clustering results and making it difficult to provide doctors with reliable clinical decision support.

Method used

By constructing a co-occurrence matrix, we analyze the keyword combinations in medical records and their degree of improvement in condition over a time period, determine the importance of smaller clusters to larger clusters, and determine the splitting threshold based on this to achieve refined classification.

Benefits of technology

It improves the accuracy of medical data clustering, meets the personalized needs of clinical scenarios, and enhances the interpretability of data organization and its ability to support clinical decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545656A_ABST
    Figure CN121545656A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric digital data processing, in particular to a medical record AI intelligent integration and analysis system based on big data, and the system corresponds to the steps: determining a first class cluster and a second class cluster in medical data; determining each keyword group between the first category cluster and the second category cluster, and combining the medical records in the first category cluster to construct a co-occurrence matrix; determining the confidence of the target medical record by using the patient condition improvement degree obtained based on the diagnosis and treatment data, the examination data and the medication data in the co-occurrence matrix; and by utilizing the confidence coefficient and each element of the target keyword group in the co-occurrence matrix, determining the importance degree of the second-class cluster to the first-class cluster, and determining a splitting threshold value of the first-class cluster and each sub-class cluster in the first-class cluster. Through the technical scheme of the invention, the accuracy of medical data clustering is improved, and the interpretability of data organization and the support capability of clinical decision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing technology, specifically to an AI-powered intelligent integration and analysis system for medical records based on big data. Background Technology

[0002] In existing research and applications of medical data integration, mainstream methods are based on clustering or classification techniques using input features. This involves extracting and vectorizing features from patient medical records, laboratory indicators, and medication information, and then using traditional similarity metrics (such as Euclidean distance and cosine similarity) to cluster and integrate different records. These methods can, to some extent, group medical data according to the similarity of surface features, thereby achieving preliminary integration of multi-source information.

[0003] However, this process often overlooks the differences and hierarchical relationships in clinical semantics between minor categories (minor clinical symptom labels) and major categories (major clinical symptom labels), leading to significant bias and inaccuracy in the clustering results. This makes it difficult to provide reliable clinical decision support for physicians. For example, some minor symptom conditions may appear in different major symptom conditions, but their clinical importance and mechanisms of action may differ significantly. Existing methods often struggle to effectively model these differences. Summary of the Invention

[0004] To address the significant bias and inaccuracy in clustering results obtained from existing methods for clustering medical data, this invention aims to provide an AI-powered intelligent integration and analysis system for medical records based on big data. The specific technical solution adopted is as follows: This invention provides an AI-powered intelligent integration and analysis system for medical records based on big data, the system comprising: The matrix generation module is used to determine the first-class clusters and second-class clusters related to the disease in medical data, wherein the first-class clusters include multiple second-class clusters; determine each keyword group between the first-class clusters and the second-class clusters, and construct a co-occurrence matrix based on the keyword groups and medical records in the first-class clusters; The cluster splitting module is used to determine the confidence level of a target medical record by utilizing the degree of improvement in the patient's condition based on diagnosis, examination, and medication data within the time period of the target medical record in the co-occurrence matrix; it uses the confidence level and each element of the target keyword group in the co-occurrence matrix to determine the importance of the second-category cluster to the first-category cluster; and it uses the importance level to determine the splitting threshold of the first-category cluster and obtain each sub-category cluster in the first-category cluster.

[0005] Further, determining the keyword groups between the first category cluster and the second category cluster includes: Obtain keywords from each medical record in the first category cluster and determine the similarity between keywords; The filtered keywords are obtained by selecting one of the keywords between two keywords whose similarity is greater than a preset similarity threshold. By combining the filtering keywords in the first category cluster with the text information in the second category cluster, multiple keyword groups are obtained.

[0006] Furthermore, the construction of the co-occurrence matrix based on keyword groups and medical records in the first category cluster includes: Each keyword group is used as a column attribute of the co-occurrence matrix, and the medical records in the first category cluster are used as row attributes of the co-occurrence matrix. If the target element's row attribute data contains corresponding column attribute data, then the target element's data value is recorded as 1. If the target element's row attribute data does not contain corresponding column attribute data, the target element's data value is recorded as 0.

[0007] Furthermore, the determination of the confidence level of the target medical record by utilizing the degree of improvement in the patient's condition based on diagnosis data, examination data, and medication data within the time period of the target medical record in the co-occurrence matrix includes: Determine the diagnostic and treatment indicators corresponding to the diagnostic and treatment data of the target medical record within the time period in the co-occurrence matrix; Determine the examination indicators corresponding to the examination data within the time period of the target medical record in the co-occurrence matrix; Determine the medication indicators corresponding to the medication data of the target medical record within the time period in the co-occurrence matrix; The confidence level of the patient's target medical record is determined by using the degree of improvement in the patient's condition represented by the diagnostic and treatment indicators, the examination indicators, and the medication indicators, respectively.

[0008] Furthermore, determining the diagnostic and treatment indicators corresponding to the diagnostic and treatment data within the time period of the target medical record in the co-occurrence matrix includes: Determine the differences in disease severity between adjacent treatment stages and the frequency of occurrence of keywords related to disease improvement in the treatment data; By using the differences in the severity of the illness and the frequency of occurrence, the corresponding diagnostic and treatment indicators for the diagnostic and treatment data are determined.

[0009] Furthermore, determining the examination indicators corresponding to the examination data within the time period of the target medical record in the co-occurrence matrix includes: Determine the normal range of the patient's target vital signs and the target characteristic values ​​at the target examination stage from the examination data; Determine the difference in lesion volume between adjacent examination stages in the examination data; Using the normal range, the target features, and the lesion volume difference value, the examination indicators corresponding to the examination data are determined.

[0010] Furthermore, determining the medication indicators corresponding to the medication data within the time period of the target medical record in the co-occurrence matrix includes: Determine the difference in the dosage of the target drug between adjacent medication phases in the medication data; By utilizing the differences in dosage of various drugs, the corresponding drug use indicators can be determined based on the drug use data.

[0011] Furthermore, determining the importance of the second-category cluster to the first-category cluster using the confidence score and each element of the target keyword group in the co-occurrence matrix includes: The frequency of medical records corresponding to the target keyword group in the co-occurrence matrix is ​​determined by using the confidence score and the data values ​​of each element in the co-occurrence matrix. The importance of the second cluster to the first cluster is determined by the frequency of occurrence of the medical records.

[0012] Further, the step of determining the splitting threshold of the first category cluster using importance and obtaining each sub-category cluster in the first category cluster includes: The split threshold between nodes within the first category of clusters is determined by using the importance level and a preset initial split threshold; Using the aforementioned splitting threshold, each sub-category cluster in the first category cluster is obtained; wherein, a node represents a medical record in the first category cluster.

[0013] Further, the step of using the splitting threshold to obtain each sub-category cluster in the first category cluster includes: The reciprocal of the cosine similarity of the vectors between nodes within the first category cluster is used as the original edge value of the first category cluster; If the original edge value is less than the splitting threshold, disconnecting the relationships between nodes yields the individual sub-clusters within the first-category cluster.

[0014] The present invention has the following beneficial effects: This invention aims to construct a medical data integration framework with hierarchical classification capabilities to meet the needs of two-tiered diagnosis and treatment analysis in clinical scenarios, focusing on "major illnesses" and "minor illnesses." Based on the results of coarse clustering of major illnesses, an initial classification structure centered on major illnesses is established. Building upon this, the system further grants doctors flexible integration permissions, enabling them to propose refined classification requirements based on specific clinical goals, i.e., reclassifying minor illness categories within major illnesses. In this way, doctors can extract minor category data that meets specific diagnostic and treatment objectives from the existing classification results, achieving a transition from coarse to fine classification. This classification method aligns with the top-down diagnostic logic of medical data in clinical practice, laying the foundation for subsequent personalized diagnosis and treatment analysis. Furthermore, by establishing a co-occurrence matrix between minor and major categories, the system analyzes the closeness of the relationship between minor and target major categories. This allows for targeted and refined screening and reorganization within the existing classification structure, ensuring that medical data not only possesses hierarchical clustering characteristics but also meets the personalized needs of clinical applications. Ultimately, this invention achieves intelligent integration and semantic association of multi-source heterogeneous medical records, improves the accuracy of medical data clustering, and enhances the interpretability of data organization and its ability to support clinical decision-making. Attached Figure Description

[0015] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 The flowchart below shows the steps of an AI-powered intelligent integration and analysis system for medical records based on big data, provided as an embodiment of the present invention. Figure 2 A detailed flowchart of step S2 in a big data-based AI intelligent integration and analysis system for medical records, provided as an embodiment of the present invention; Figure 3 A detailed flowchart of step S2 in a big data-based AI intelligent integration and analysis system for medical records, provided as another embodiment of the present invention; Figure 4 A detailed flowchart of step S3 in a big data-based AI intelligent integration and analysis system for medical records, provided as an embodiment of the present invention; Figure 5 A detailed flowchart of step S4 in a big data-based AI intelligent integration and analysis system for medical records, provided as an embodiment of the present invention; Figure 6 A detailed flowchart of step S5 in a big data-based AI intelligent integration and analysis system for medical records, provided as an embodiment of the present invention; Figure 7 This is a schematic diagram of the hardware operating environment of the AI-powered intelligent integration and analysis device for medical records based on big data, which is involved in the embodiments of the present invention. Figure 8 This is a schematic diagram of the framework structure of the AI-powered intelligent integration and analysis system for medical records based on big data, which is involved in the embodiments of the present invention. Figure 9 This is a schematic diagram illustrating the set relationship between major and minor categories in the AI-powered intelligent integration and analysis system for medical records based on big data, which is part of an embodiment of the present invention. Detailed Implementation

[0017] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a big data-based AI intelligent integration and analysis system for medical records proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0019] The following description, in conjunction with the accompanying drawings, details the specific solution of the AI-powered intelligent integration and analysis system for medical records based on big data provided by this invention.

[0020] Example 1: For the AI-powered intelligent integration and analysis system for medical records based on big data provided in this invention, please refer to [link to relevant documentation]. Figure 8 , Figure 8 This is a schematic diagram of the framework structure of the AI-powered intelligent integration and analysis system for medical records based on big data, which is involved in the embodiments of the present invention.

[0021] The big data-based AI-powered intelligent integration and analysis system for medical records (hereinafter referred to as the "AI-powered intelligent integration and analysis system for medical records" or the "system") includes: The matrix generation module A10 is used to determine the first-class clusters and second-class clusters related to the disease in medical data, wherein the first-class clusters include multiple second-class clusters; determine each keyword group between the first-class clusters and the second-class clusters; and construct a co-occurrence matrix based on the keyword groups and medical records in the first-class clusters. The cluster splitting module A20 is used to determine the confidence level of the target medical record by utilizing the degree of improvement of the patient's condition based on diagnosis data, examination data, and medication data within the time period of the target medical record in the co-occurrence matrix; using the confidence level and each element of the target keyword group in the co-occurrence matrix, it determines the importance of the second category cluster to the first category cluster; using the importance level, it determines the splitting threshold of the first category cluster and obtains each sub-category cluster in the first category cluster.

[0022] Please see Figure 1 , Figure 1 The diagram illustrates the steps of an AI-powered intelligent integration and analysis system for medical records based on big data, provided in one embodiment of the present invention.

[0023] The methods and steps corresponding to the AI-powered intelligent integration and analysis system for medical records based on big data include: Step S1: Determine the first-category clusters and second-category clusters related to the medical condition in the medical data, wherein the first-category clusters include multiple second-category clusters; The medical data in this embodiment comes from the hospital information system and the electronic medical record system, covering various types of information about patients from initial diagnosis to follow-up.

[0024] Medical data (multiple medical records) specifically includes: diagnosis and treatment data (outpatient medical records, inpatient medical records, multiple follow-up medical records, etc.), which reflect the patient's condition description, clinical diagnosis and medical orders at different time points; medication data, which characterizes the patient's actual medication use and compliance; and examination data, such as laboratory tests (complete blood count, biochemical indicators), imaging examinations (CT, MRI) and functional examinations (electrocardiogram, pulmonary function), which can provide objective quantitative indicators. The introduction of multi-source medical data enables the patient's disease course to be characterized from multiple dimensions, providing a solid data foundation for subsequent modeling of the relationship between minor and major illnesses.

[0025] It should be noted that in this embodiment, minor illnesses, or minor categories, correspond to the first category cluster; major illnesses, or major categories, correspond to the second category cluster. Here, major and minor illnesses can be artificially classified according to multiple dimensions in clinical medical practice, such as the degree of life-threatening danger, the complexity of the illness, the difficulty of treatment, and the consumption of resources.

[0026] Each medical record obtained above is treated as an independent node. First, it is directly categorized based on the presence of a clear major illness label. If no label is found, entity extraction and semantic matching techniques are used to identify key symptoms, medications, or laboratory indicators related to major illnesses. A hierarchical clustering algorithm is then used to form a candidate set of major illnesses. This coarse clustering step effectively reduces cross-category confounding and establishes an initial data framework centered on major illnesses.

[0027] For minor illnesses, first clarify the clinical purpose of the subcategory the doctor wishes to define, and then gradually collect and formalize the definition of that subcategory in the system using a structured process: 1. Doctors enter the name of the subcategory and a brief clinical description on the interface; 2. The doctor should provide clear criteria and exclusion criteria (such as comorbidities, history of specific medications), and may specify time constraints (such as the first occurrence within 30 days before diagnosis). 3. Doctors can select and filter data modalities (medical record text, test values, imaging reports, medication records); Finally, the structured process data provided by the doctors above will serve as the conditions for the clustering splitting process and classification screening in the following embodiments.

[0028] Please refer to Figure 9 , Figure 9 This is a schematic diagram illustrating the set relationship between major and minor categories in the AI-powered intelligent integration and analysis system for medical records based on big data, which is part of an embodiment of the present invention.

[0029] In this embodiment, the clustering result is calculated based on the relationships between nodes. However, in real diagnosis and treatment processes, some subcategories may exhibit similar association patterns with multiple larger categories simultaneously, for example... Figure 9 Subcategories 'a' belong to both major categories A and B, yet possess differentiated clinical significance and weight within each category. For instance, the same subcategory might be a core symptom in one major category but merely an accompanying symptom in another. Ignoring these differences and relying solely on superficial relationships for clustering can easily lead to confusion in clinical semantics and biased results. Therefore, before clustering, it is necessary to clearly calculate the strength and importance of the relationships between the current subcategory and each major category to ensure that subsequent clustering not only reflects data-level correlations but also embodies discriminative value and clinical rationale within the medical context.

[0030] Step S2: Determine each keyword group between the first category cluster and the second category cluster, and construct a co-occurrence matrix based on the keyword groups and medical records in the first category cluster; In the analysis of clinical medical records, the co-occurrence patterns of specific secondary symptoms or signs with certain key terms often reflect their intrinsic association with the primary condition. For example, in the description of pneumonia, the symptom of "high fever" is frequently observed to co-occur with terms such as "pulmonary rales" and "chest X-ray showing infiltrates"; this high co-occurrence indicates that high fever may play an important role in the pathological manifestations and clinical diagnosis of pneumonia. Conversely, if a secondary symptom such as "rash" is only occasionally recorded in the diagnosis of pneumonia along with a few terms such as "adverse drug reaction" or "viral infection," it suggests that this manifestation has a weak association with the core clinical manifestations of pneumonia, and its importance is relatively limited.

[0031] Based on these co-occurrence patterns, a co-occurrence matrix can be constructed. By analyzing the probability distribution of minor categories (minor illnesses) within the co-occurrence matrix, the relationship between minor categories and each major category can be analyzed, thereby systematically evaluating the contribution of minor illnesses to the diagnostic decision-making of major illnesses. The process of calculating the importance of minor categories within major categories is shown below. The analytical objective is the relationship between the minor category currently input by the doctor and any (target) major category.

[0032] Specifically, in one embodiment, please refer to Figure 2 Step S2, determining each keyword group between the first category cluster and the second category cluster, includes: Step S21: Obtain keywords from each medical record in the first category cluster and determine the similarity between keywords; Step S22: Select one of the keywords between two keywords whose similarity is greater than a preset similarity threshold to obtain the selected keywords; Step S23: Combine the selected keywords in the first category cluster with the text information in the second category cluster to obtain multiple keyword groups.

[0033] In this embodiment, before constructing a co-occurrence matrix to analyze disease associations, it is first necessary to extract associated words that co-occur with minor diseases from large-scale diagnostic texts. Specifically: 1. Within the target major category cluster (any first-category cluster), for each medical record, use the Jieba word segmentation tool to segment the patient symptoms and medical orders into words or phrases. 2. Use the NER (Named Entity Recognition) model to identify entities such as symptoms and drugs in the text; 3. The TF-IDF (Term Frequency–Inverse Document Frequency) extraction algorithm is used to extract keywords from entities, i.e., important words in the text; 4. There are a large number of synonyms, near-synonyms, and colloquial expressions in medical texts. For example, "fever" and "have a fever", "pulmonary rales" and "lung rales", etc. If these words with different surface forms but the same semantics are not processed, they will disperse the co-occurrence frequency and reduce the statistical reliability, thus affecting the accuracy of subsequent importance calculation. Therefore, it is necessary to perform semantic homogenization on the extracted related words to ensure the consistency and comparability of subsequent matrix construction and probability calculation.

[0034] Specifically, using a word embedding model, map the above keywords into a low-dimensional vector space, and measure the semantic similarity of keywords by calculating the cosine similarity between vectors. Denote the cosine similarity of any two keyword vectors as , set a preset similarity threshold H (such as 0.7, which can be adjusted specifically). When the similarity of two keywords is greater, it means that the semantics of the two keywords are more similar. Take all pairs of keywords, and represent these two keywords with one of the keywords as the screening keyword.

[0035] Construct a co-occurrence matrix: Combine all the keywords (screening keywords) in the first category cluster obtained above with the small-category text (text information in the second category cluster) to form N keyword groups.

[0036] Specifically, in another embodiment, please refer to Figure 3 , the step S2 of constructing a co-occurrence matrix based on the keyword groups and the medical records in the first category cluster includes: Step S201, take each keyword group as the column attribute of the co-occurrence matrix, and take the medical records in the first category cluster as the row attribute of the co-occurrence matrix; Step S202, when there is corresponding column attribute data in the row attribute data of the target element, record the data value of the target element as 1; Step S203, when there is no corresponding column attribute data in the row attribute data of the target element, record the data value of the target element as 0.

[0037] Based on the above embodiments, for the N keyword groups formed, denote the small-category text as S, take all the constructed keyword groups as the column attributes of the co-occurrence matrix, and take a medical record as the row attribute of the co-occurrence matrix.

[0038] Generate a co-occurrence matrix: If in a certain medical record, there is a certain keyword combination (when there is corresponding column attribute data in the row attribute data of the target element), then record the data value at the target element as 1, and the others as 0 (when there is no corresponding column attribute data in the row attribute data of the target element). Finally, a co-occurrence matrix of whether the key small category and the keyword exist in the medical record can be formed, and the co-occurrence matrix is shown as follows: The keyword phrase can be represented as: … k represents the kth keyword that appears. This indicates the h-th medical record within the target category.

[0039] Step S3: Determine the confidence level of the target medical record by using the degree of improvement of the patient's condition based on diagnosis data, examination data and medication data within the time period of the target medical record in the co-occurrence matrix. In clinical text mining, relying solely on the co-occurrence frequency of symptoms or laboratory tests to measure the importance of minor ailments within major illnesses often leads to serious biases. Disease progression is dynamic; for some patients, the absence of certain symptoms does not indicate their unrelatedness to the major illness but rather reflects the patient's gradual improvement during treatment. If these symptoms do not reappear in subsequent records and are treated as "negative examples," it underestimates the symptom's role and representativeness in the early stages of the disease. This means that simply counting "occurrences" ignores the temporal evolution and fluctuations in disease progression, making it difficult to accurately depict the importance of symptoms.

[0040] Therefore, for any medical record within a broad category, its confidence level needs to be calculated. This confidence level represents the true representation of the keyword combination (symptoms) in that record. The confidence level of the medical record is indicated by the degree of improvement in the patient's condition as shown by the timestamp of that record. If the patient's condition gradually improved in medical records prior to that timestamp, even if the keyword combination in those records gradually decreased, the keyword combination in those records still better represents the co-occurrence of the subcategory and the keyword combination; that is, the medical record has a higher confidence level. The process for determining the confidence level of the current medical record is as follows.

[0041] By constructing a longitudinal time series with the patient at the center, medical records, including diagnosis and treatment records, examination records, and medication records, are placed on a unified time axis to ensure that data from different sources can be cross-referenced in the time dimension.

[0042] Specifically, please refer to Figure 4 Step S3 includes: Step S31: Determine the diagnostic and treatment indicators corresponding to the diagnostic and treatment data within the time period of the target medical record in the co-occurrence matrix; Specifically, step S31 includes: Determine the differences in disease severity between adjacent treatment stages and the frequency of occurrence of keywords related to disease improvement in the treatment data; By using the differences in the severity of the illness and the frequency of occurrence, the corresponding diagnostic and treatment indicators for the diagnostic and treatment data are determined.

[0043] In this embodiment, the construction of multi-dimensional indicators for the quantitative analysis of the degree of improvement of the condition requires full integration of different types of medical data.

[0044] The quantification of diagnostic and treatment data dimensions is represented by the sentiment (improvement instructions) and symptom trend scores in medical records. Specifically, the severity of the illness in the medical records can be obtained through the existing SOFA (Sequential Organ Failure Assessment) score, and the diagnostic and treatment indicators of the current medical data can be calculated. The formula is shown below: In the formula: This represents the severity value of the patient's condition recorded in the t-th medical record (corresponding to the t-th stage of treatment). This represents the severity value of the patient's condition in the (t-1)th medical record adjacent to t. That is, the difference in disease severity between adjacent stages of diagnosis and treatment; for a value of 1, it leads to... If the value is 0, it can be ignored and not included in the calculation. Alternatively, 0.1 can be used as a replacement to ensure that the formula can be performed effectively. This represents the total number of medical records corresponding to the patient's specified timestamp t. A lower severity value indicates a greater degree of improvement in the patient's condition. The frequency of textual keywords indicating improvement in the condition (such as "improvement," "recovery," "disappearance," etc.). This indicates the total number of the patient's medical records. , This represents the number of records containing text keywords indicating improvement (improvement, better, disappearance). The higher the frequency of these keywords, the greater the improvement in the patient's condition, i.e., the higher the diagnostic indicator. The Sigmoid normalization function maps the real numbers to the interval (0, 1).

[0045] Step S32: Determine the examination indicators corresponding to the examination data within the time period of the target medical record in the co-occurrence matrix; Specifically, step S32 includes: Determine the normal range of the patient's target vital signs and the target characteristic values ​​at the target examination stage from the examination data; Determine the difference in lesion volume between adjacent examination stages in the examination data; Using the normal range, the target features, and the lesion volume difference value, the examination indicators corresponding to the examination data are determined.

[0046] In this embodiment, the quantification of examination data dimensions is achieved by quantifying changes in the patient's vital signs and imaging results in the examination records, and then calculating the examination indicators of the current examination data. The formula is shown below: In the formula: e represents the patient's e-th vital sign (target vital sign, referring to any vital sign) indicator. This represents the data value (target feature value) of vital sign indicator e in the t-th examination record (corresponding to the target examination stage). This represents the normal range for vital sign e. Since the normal range for vital signs is a range of values, that is... , This represents the median value within the interval. The difference between vital sign e and the median of the normal range in the t-th examination record is indicated. The smaller the difference, the greater the improvement of the patient at the t-th timestamp of that record. This indicates the volume of the lesion in the t-th examination record. This represents the volume of the lesion in the (t-1)th examination record adjacent to t, where the difference in lesion volume is... The greater the difference, the greater the improvement in the patient's condition as indicated in the imaging report. This represents the rate of lesion volume reduction between adjacent examination stages; a higher rate indicates a greater degree of improvement in the condition. Since the data for all dimensions are on the same time axis, This can also be used to represent the total number of examination records corresponding to the patient's timestamp t.

[0047] Step S33: Determine the medication indicators corresponding to the medication data of the target medical record within the time period in the co-occurrence matrix; Specifically, step S33 includes: Determine the difference in the dosage of the target drug between adjacent medication phases in the medication data; By utilizing the differences in dosage of various drugs, the corresponding drug use indicators can be determined based on the drug use data.

[0048] In this embodiment, the quantification of medication data is achieved by quantifying the changes in medication intensity in the medication records, and calculating the medication index of the current medication data. The formula is shown below: In the formula: This represents the dosage of drug m (target drug, referring to any one drug) used by the patient in the t-th medication record (corresponding to the t-th medication stage). This represents the dosage of drug m in the (t-1)th medication record adjacent to t, where the dosage difference value is... The larger the value, the less medication is needed, meaning the greater the improvement in the patient's condition, and the higher the medication index. This represents the total number of drug types in the t-th medication record. Since the data for each dimension is on the same time axis, This can also represent the total number of medication records corresponding to that time stamp for the patient. In this embodiment of the invention, only the medications used are analyzed; therefore, Greater than 0.

[0049] Step S34: Using the degree of improvement of the patient's condition represented by the diagnostic and treatment indicators, the examination indicators, and the medication indicators, respectively, determine the confidence level of the patient's target medical record; Specifically, there are significant differences across different diseases and disease stages. For example, in acute illnesses, rapid changes in laboratory test indicators may reflect disease fluctuations better than long-term medication adherence; while in the maintenance phase of chronic diseases, long-term medication regularity and the frequency of follow-up visits may be more explanatory for disease control.

[0050] Therefore, the entropy weight method (existing technology) is used to determine the weight of each indicator based on its performance in different diseases and at different stages. This essentially involves calculating the weight of each indicator under different medical records. Information entropy, followed by... Information entropy and The information entropy of these three factors is then normalized to obtain their respective weights.

[0051] Based on the multi-dimensional indicators of the current medical records obtained above, calculate the medical record indicators. Detection and recording indicators and medication records The weighted sum between these values ​​yields the confidence level of the current (target) medical record. This ensures clinical validity and comparability across different treatment scenarios.

[0052] Step S4: Using the confidence score and each element of the target keyword group in the co-occurrence matrix, determine the importance of the second-category cluster to the first-category cluster; Specifically, please refer to Figure 5 Step S4 includes: Step S41: Using the confidence level and the data values ​​of each element of the target keyword group in the co-occurrence matrix, determine the frequency of occurrence of medical records corresponding to the target keyword group in the co-occurrence matrix; Step S42: Using the frequency of occurrence of the medical records, determine the importance of the second cluster to the first cluster.

[0053] Based on the above embodiments, the co-occurrence matrix and the confidence level of each medical record in the co-occurrence matrix can be obtained according to the above calculations. The higher the frequency of records appearing in each keyword group, the closer the relationship between the subcategory and the target major category, i.e., the stronger the importance of the subcategory to the target major category. The relationship between the subcategory and the major category is represented by the frequency of all medical records appearing in each keyword group. The formula for calculating the frequency of medical records appearing in each keyword group is as follows: In the formula: This represents the total number of medical records within the target category. This represents the confidence level of the d-th medical record. This indicates the location of the target keyword group in the co-occurrence matrix. The data value of the element in row d. That is, 0 or 1. Keyword phrase The frequency of occurrence of the corresponding medical records.

[0054] Finally, the product of all keyword groups represents the importance of the smaller category to the larger category, denoted as importance. .

[0055] Step S5: Determine the splitting threshold of the first category cluster using the importance level and obtain each sub-category cluster in the first category cluster.

[0056] Specifically, please refer to Figure 6 Step S5 includes: Step S51: Determine the split threshold between nodes within the first category cluster using the importance level and the preset initial split threshold; Step S52: Using the splitting threshold, obtain each sub-category cluster in the first category cluster; In this context, nodes represent medical records within the first category of clusters.

[0057] More specifically, step S52 includes: The reciprocal of the cosine similarity of the vectors between nodes within the first category cluster is used as the original edge value of the first category cluster; If the original edge value is less than the splitting threshold, disconnecting the relationships between nodes yields the individual sub-clusters within the first-category cluster.

[0058] At this point, the relationship between the current subcategory and each major category w can be obtained. The closer the relationship between a minor category and a major category, that is, the more important the minor category is to the major category. A larger value indicates that smaller categories are less likely to be split, resulting in a smaller splitting threshold. The formula for the splitting threshold for each larger category is: In the formula: This represents the threshold for splitting a connected graph, also known as the preset initial split threshold. (Specific details can be adjusted) This indicates the importance of the smaller category to the larger category w. This represents the splitting threshold between nodes within a large category w.

[0059] At this point, based on the splitting threshold of each major category... And the similarity of vectors between nodes (which can be cosine similarity), can be used to split each large category cluster, when the original edge value When the original edge value is used, the relationship between nodes is preserved; when the original edge value is used... At this point, the relationships between nodes are broken; ultimately, we can obtain the sub-category clusters (sub-category clusters) in each major category, and the medical records in these sub-category clusters are strongly correlated with the corresponding major category.

[0060] Building upon this foundation, and combining the sub-category filtering conditions input by doctors, the system can achieve targeted and refined filtering and recombination within the existing classification structure. This ensures that medical data not only possesses hierarchical clustering characteristics but also meets the personalized needs of clinical applications. Ultimately, the aforementioned embodiments achieve intelligent integration and semantic association of multi-source heterogeneous medical records, enhancing the interpretability of data organization and its support capabilities for clinical decision-making.

[0061] This invention aims to construct a medical data integration framework with hierarchical classification capabilities to meet the needs of two-tiered diagnosis and treatment analysis in clinical scenarios, focusing on "major illnesses" and "minor illnesses." Based on the results of coarse clustering of major illnesses, an initial classification structure centered on major illnesses is established. Building upon this, the system further grants doctors flexible integration permissions, enabling them to propose refined classification requirements based on specific clinical goals, i.e., reclassifying minor illness categories within major illnesses. In this way, doctors can extract minor category data that meets specific diagnostic and treatment objectives from the existing classification results, achieving a transition from coarse to fine classification. This classification method aligns with the top-down diagnostic logic of medical data in clinical practice, laying the foundation for subsequent personalized diagnosis and treatment analysis. Furthermore, by establishing a co-occurrence matrix between minor and major categories, the system analyzes the closeness of the relationship between minor and target major categories. This allows for targeted and refined screening and reorganization within the existing classification structure, ensuring that medical data not only possesses hierarchical clustering characteristics but also meets the personalized needs of clinical applications. Ultimately, this invention achieves intelligent integration and semantic association of multi-source heterogeneous medical records, improves the accuracy of medical data clustering, and enhances the interpretability of data organization and its ability to support clinical decision-making.

[0062] Example 2: This invention also proposes an AI-powered intelligent integration and analysis device for medical records based on big data. The device can be a data processing device such as a computer or a server, or a combination of multiple devices.

[0063] like Figure 7 As shown, Figure 7 This is a schematic diagram of the hardware operating environment of the AI-powered intelligent integration and analysis device for medical records based on big data, which is involved in the embodiments of the present invention.

[0064] like Figure 7 As shown, this AI-powered medical record integration and analysis device based on big data may include: a processor 1001, such as a CPU; a network interface 1004; a user interface 1003; a memory 1005; and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display or an input unit such as a control panel; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001. The memory 1005, as a computer storage medium, may include a medical record AI integration and analysis program.

[0065] Those skilled in the art will understand that Figure 7 The hardware structure shown does not constitute a limitation on the device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0066] Continue to refer to Figure 7 , Figure 7 The memory 1005, which is a computer-readable storage medium, may include an operating device, a user interface module, a network communication module, and a medical record AI intelligent integration and analysis program.

[0067] exist Figure 7 In this embodiment, the network communication module is mainly used to connect to the server and can communicate with the server for data; while the processor 1001 can call the medical record AI intelligent integration and analysis program stored in the memory 1005 and execute the steps in the above embodiments.

[0068] Based on the hardware structure of the above-mentioned AI-powered intelligent integration and analysis device for medical records based on big data, various embodiments of the AI-powered intelligent integration and analysis system for medical records based on big data of the present invention are implemented.

[0069] Furthermore, the present invention also provides a computer-readable storage medium. This computer-readable storage medium stores a medical record AI intelligent integration and analysis program, wherein, when executed by a processor, the medical record AI intelligent integration and analysis program implements the steps of the method corresponding to the big data-based medical record AI intelligent integration and analysis system described above.

[0070] The method implemented when the medical record AI intelligent integration and analysis program is executed can be referred to in various embodiments of the medical record AI intelligent integration and analysis system based on big data of the present invention, and will not be repeated here.

[0071] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0072] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0073] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0074] The above description is only a preferred embodiment of the present invention and does not limit the scope of protection of the present invention. All equivalent structural / method transformations made under the inventive concept of the present invention using the contents of the present invention specification and drawings, or direct / indirect applications in other related technical fields, are included within the scope of protection of the present invention.

Claims

1. A medical record AI intelligent integration and analysis system based on big data, characterized in that, The system includes: The matrix generation module is used to determine the first-class clusters and second-class clusters related to the disease in medical data, wherein the first-class clusters include multiple second-class clusters; determine each keyword group between the first-class clusters and the second-class clusters, and construct a co-occurrence matrix based on the keyword groups and medical records in the first-class clusters; The cluster splitting module is used to determine the confidence level of a target medical record by utilizing the degree of improvement in the patient's condition based on diagnosis, examination, and medication data within the time period of the target medical record in the co-occurrence matrix; it uses the confidence level and each element of the target keyword group in the co-occurrence matrix to determine the importance of the second-category cluster to the first-category cluster; and it uses the importance level to determine the splitting threshold of the first-category cluster and obtain each sub-category cluster in the first-category cluster. Methods for determining the confidence level of target medical records by utilizing the degree of patient improvement obtained from the target medical records in the co-occurrence matrix include: Determine the diagnostic and treatment indicators corresponding to the diagnostic and treatment data within the time period of the target medical record in the co-occurrence matrix; Determine the examination indicators corresponding to the examination data within the time period of the target medical record in the co-occurrence matrix; Determine the medication indicators corresponding to the medication data of the target medical record within the time period in the co-occurrence matrix; The confidence level of the patient's target medical record is determined by using the degree of improvement of the patient's condition represented by the diagnostic and treatment indicators, the examination indicators, and the medication indicators, respectively; wherein, the entropy weight method is used to determine the weights of the diagnostic and treatment indicators, the examination indicators, and the medication indicators, and the confidence level is obtained by weighted summation. Methods for determining diagnostic and treatment indicators include: Determine the differences in disease severity between adjacent treatment stages and the frequency of occurrence of keywords related to disease improvement in the medical data; Using the differences in the severity of the illness and the frequency of occurrence, the corresponding diagnostic and treatment indicators for the diagnostic and treatment data are determined; The methods for determining the inspection indicators include: Determine the normal range of the patient's target vital signs and the target characteristic values ​​at the target examination stage from the examination data; Determine the difference in lesion volume between adjacent examination stages in the examination data; Using the normal range, the target features, and the lesion volume difference value, the examination indicators corresponding to the examination data are determined; Methods for determining medication indicators include: Determine the difference in the dosage of the target drug between adjacent medication phases in the medication data; By utilizing the differences in dosage of various drugs, the corresponding drug use indicators can be determined based on the drug use data.

2. The AI-powered intelligent integration and analysis system for medical records based on big data as described in claim 1, characterized in that, The determination of each keyword group between the first category cluster and the second category cluster includes: Obtain keywords from each medical record in the first category cluster and determine the similarity between keywords; The filtered keywords are obtained by selecting one of the keywords between two keywords whose similarity is greater than a preset similarity threshold. By combining the filtering keywords in the first category cluster with the text information in the second category cluster, multiple keyword groups are obtained.

3. The AI-powered intelligent integration and analysis system for medical records based on big data as described in claim 1, characterized in that, The co-occurrence matrix constructed based on keyword groups and medical records in the first category cluster includes: Each keyword group is used as a column attribute of the co-occurrence matrix, and the medical records in the first category cluster are used as row attributes of the co-occurrence matrix. If the target element's row attribute data contains corresponding column attribute data, then the target element's data value is recorded as 1. If the target element's row attribute data does not contain corresponding column attribute data, the target element's data value is recorded as 0.

4. The AI-powered intelligent integration and analysis system for medical records based on big data as described in claim 1, characterized in that, The determination of the importance of the second-category cluster to the first-category cluster using confidence scores and elements of the target keyword groups in the co-occurrence matrix includes: Using the confidence score and the data values ​​of each element of the target keyword group in the co-occurrence matrix, the frequency of occurrence of medical records corresponding to the target keyword group in the co-occurrence matrix is ​​determined. The corresponding calculation formula is as follows: In the formula: This represents the total number of medical records in the first-category cluster. This represents the confidence level of the d-th medical record. This indicates the location of the target keyword group in the co-occurrence matrix. The data value of the element in row d. , Keyword phrase The frequency of occurrence of the corresponding medical records; The importance of the second cluster to the first cluster is determined by the frequency of occurrence of the medical records. In the formula, Indicates the degree of importance. Represents the normalization function. Indicates all Perform cumulative multiplication.

5. The AI-powered intelligent integration and analysis system for medical records based on big data as described in claim 1, characterized in that, The step of determining the splitting threshold of the first category cluster based on importance and obtaining each sub-category cluster in the first category cluster includes: The split threshold between nodes within the first category of clusters is determined by using the importance level and a preset initial split threshold; Using the aforementioned splitting threshold, each sub-category cluster in the first category cluster is obtained; wherein, a node represents a medical record in the first category cluster.

6. The AI-powered intelligent integration and analysis system for medical records based on big data as described in claim 5, characterized in that, The process of obtaining each sub-cluster within the first category cluster using the splitting threshold includes: The reciprocal of the cosine similarity of the vectors between nodes within the first category cluster is used as the original edge value of the first category cluster; If the original edge value is less than the splitting threshold, disconnecting the relationships between nodes yields the individual sub-clusters within the first-category cluster.

Citation Information

Patent Citations

  • Intelligent pediatric disease diagnosis auxiliary system

    CN118538399A

  • Clinical medication data recording method and system

    CN119580916A

  • Medical image segmentation method and system based on knowledge migration and attention mechanism

    CN120339192A

  • Multi-modal medical data intelligent association analysis system based on deep learning

    CN121034512A

  • Medical diagnosis auxiliary method and system based on thinking chain visualization

    CN121096595A