Method and apparatus for identifying electronic medical record data
By calculating the electronic frailty index in electronic medical record data and using machine learning models, the problem of low accuracy in identifying electronic medical record data in existing technologies has been solved, the accuracy of calculating the correlation of chronic obstructive pulmonary disease has been improved, the rate of missed diagnosis has been reduced, and early treatment and prevention have been promoted.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU NAT LAB
- Filing Date
- 2024-03-22
- Publication Date
- 2026-07-21
AI Technical Summary
Existing methods for identifying chronic obstructive pulmonary disease (COPD) based on electronic medical record data rely solely on common features, resulting in low identification accuracy.
By acquiring the electronic medical record data of the users to be tested, the electronic weakness index is calculated, and a probabilistic recognition model is used. Based on the electronic weakness index, demographic characteristics, physical examination characteristics, and biochemical indicators, a machine learning model is trained to improve the accuracy of the recognition probability calculation.
It improves the accuracy of correlation calculation between electronic medical record data and chronic obstructive pulmonary disease (COPD), reduces the rate of missed diagnosis of COPD, helps doctors identify potential patients early, prevents the occurrence of complications, and improves patients' quality of life and life expectancy.
Smart Images

Figure CN118352057B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for identifying and processing electronic medical record data. Background Technology
[0002] This section is intended to provide background or context for the embodiments of the invention set forth in the claims. The description herein is not an admission that it is prior art simply because it is included in this section.
[0003] Chronic obstructive pulmonary disease (COPD) has become a major disease that seriously affects health. COPD can lead to various serious complications such as respiratory failure and cardiovascular disease, severely impacting patients' quality of life and life expectancy. COPD has an insidious onset; therefore, early identification and intervention are crucial for preventing and slowing the disease's progression.
[0004] Regarding the problem of chronic obstructive pulmonary disease (COPD), some scholars have proposed identifying the correlation probability between electronic medical record (EMR) data and COPD based on EMR systems. However, existing methods for identifying and processing EMR data only rely on some common user characteristics, resulting in low accuracy. Summary of the Invention
[0005] This invention provides a method for identifying and processing electronic medical record data to improve the accuracy of calculating the identification probability of electronic medical record data, thereby improving the accuracy of calculating the correlation between electronic medical record data and chronic obstructive pulmonary disease. The method includes:
[0006] Obtain the electronic medical record data of the user to be tested;
[0007] Based on the electronic medical record data of the user to be tested, the electronic frailty index of the user to be tested is determined; the electronic frailty index is used to characterize the aging characteristics of the user to be tested.
[0008] Based on the electronic deterioration index of the user to be tested and the probability recognition model, the recognition probability corresponding to the electronic medical record data of the user to be tested is obtained. The recognition probability is used to characterize the correlation between the electronic medical record data of the user to be tested and chronic obstructive pulmonary disease. The probability recognition model is generated by pre-training a machine learning model based on the sample data of the relationship between the historical electronic deterioration index and the historical recognition probability determined from the historical electronic medical record data of multiple users.
[0009] This invention also provides an electronic medical record data recognition and processing device to improve the accuracy and efficiency of electronic medical record data recognition and processing, thereby helping doctors improve the accuracy and efficiency of identifying users' chronic obstructive pulmonary disease. The device includes:
[0010] The acquisition unit is used to acquire the electronic medical record data of the user to be tested;
[0011] The electronic decay index determination unit is used to determine the electronic decay index of the user to be tested based on the user's electronic medical record data; the electronic decay index is used to characterize the aging characteristics of the user to be tested.
[0012] The probability determination unit is used to obtain the recognition probability corresponding to the electronic medical record data of the user to be detected based on the electronic weakness index of the user to be detected and the probability recognition model. The recognition probability is used to characterize the correlation between the electronic medical record data of the user to be detected and chronic obstructive pulmonary disease. The probability recognition model is generated by pre-training a machine learning model based on the sample data of the relationship between the historical electronic weakness index and the historical recognition probability determined from the historical electronic medical record data of multiple users.
[0013] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-described method for identifying and processing electronic medical record data.
[0014] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for identifying and processing electronic medical record data.
[0015] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described method for identifying and processing electronic medical record data.
[0016] Compared with existing technologies that rely solely on common user characteristics to identify electronic medical record (EMR) data, resulting in low accuracy, the EMR data identification processing scheme provided in this invention determines the EMR index of the user to be tested based on the EMR data. This EMR index characterizes the aging characteristics of the user. By incorporating these aging characteristics into probabilistic identification, the identification probability corresponding to the user's EMR data is obtained based on the EMR index and the probabilistic identification model. This identification probability characterizes the correlation between the user's EMR data and chronic obstructive pulmonary disease (COPD). The probabilistic identification model is pre-trained using sample data of the relationship between historical EMR indices and historical identification probabilities determined from historical EMR data of multiple users. Therefore, this EMR data identification processing scheme can improve the calculation accuracy of the EMR data identification probability, thereby improving the accuracy of the correlation calculation between EMR data and COPD. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0018] Figure 1 This is a flowchart illustrating the method for identifying and processing electronic medical record data in an embodiment of the present invention.
[0019] Figure 2 This is a schematic diagram illustrating the principle of electronic medical record data recognition and processing in an embodiment of the present invention;
[0020] Figure 3 This is a schematic diagram of the AUC of patients with chronic obstructive pulmonary disease detected in a centralized dataset of City A in an embodiment of the present invention;
[0021] Figure 4 This is a schematic diagram of the AUC of patients with chronic obstructive pulmonary disease detected in the data of City B in an embodiment of the present invention;
[0022] Figure 5 This is a schematic diagram of the structure of the electronic medical record data recognition and processing device in an embodiment of the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0024] The information collected in this application's technical solution is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure, and application of this data all comply with relevant national and regional laws, regulations, and standards, and necessary confidentiality measures have been taken. This does not violate public order and good morals, and a corresponding operation entry point is provided for the user to choose to authorize or refuse. This application provides users with a corresponding operation entry point to choose whether to agree to or refuse the automated decision-making results; if the user chooses to refuse, the process proceeds to the expert decision-making stage.
[0025] The method for identifying and processing electronic medical record data claimed in this application is an information processing method in which all steps are performed by a computer or other device.
[0026] The inventors discovered that previous methods for identifying chronic obstructive pulmonary disease (COPD) based on electronic medical record (EMR) systems only incorporated routine demographic characteristics, physical examinations, and biochemical indicators, neglecting aging-related characteristics beyond age. Frailty is a progressive, multidimensional, and complex geriatric syndrome associated with adverse outcomes such as hospitalization and death. Therefore, the inventors proposed an EMR data processing scheme that helps physicians determine the probability of a user having COPD based on an electronic frailty index, thereby identifying potential COPD patients. This frailty index is a commonly used indicator for assessing the degree of frailty and, compared to age, more accurately reflects the degree of physiological aging. Studies have shown that the proportion of frail individuals is significantly higher among COPD patients than among non-COPD patients. Therefore, the frailty index is closely related to COPD. The electronic frailty index is an indicator calculated based on an EMR system that can accurately assess the degree of frailty. A probabilistic identification model based on the electronic decay index can improve the accuracy of calculating the identification probability of electronic medical record data, thereby improving the accuracy of calculating the correlation between electronic medical record data and chronic obstructive pulmonary disease. The identification and processing scheme for this electronic medical record data is described in detail below.
[0027] Figure 1 This is a flowchart illustrating the electronic medical record data identification and processing method in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes the following steps:
[0028] Step 101: Obtain the electronic medical record data of the user to be tested;
[0029] Step 102: Determine the electronic frailty index of the user to be tested based on the user's electronic medical record data; the electronic frailty index is used to characterize the aging characteristics of the user to be tested.
[0030] Step 103: Based on the electronic weakness index of the user to be tested and the probability recognition model, obtain the recognition probability corresponding to the electronic medical record data of the user to be tested. The recognition probability is used to characterize the correlation between the electronic medical record data of the user to be tested and chronic obstructive pulmonary disease. The probability recognition model is generated by pre-training the machine learning model based on the sample data of the relationship between the historical electronic weakness index and the historical recognition probability determined from the historical electronic medical record data of multiple users.
[0031] The electronic medical record data identification and processing method provided in this embodiment of the invention operates as follows: An electronic frailty index (EFI) of the user to be tested is determined based on the user's EPI data. This EPI characterizes the user's aging characteristics. Incorporating these aging characteristics into the probabilistic identification process, the identification probability corresponding to the user's EPI and the probabilistic identification model is obtained. This identification probability characterizes the correlation between the user's EPI data and chronic obstructive pulmonary disease (COPD). The probabilistic identification model is pre-trained using sample data representing the relationship between historical EPI and historical identification probabilities determined from historical EPI data of multiple users. Therefore, compared to existing methods that identify EPI data based solely on some common user characteristics, resulting in low accuracy, the electronic medical record data identification and processing method provided in this embodiment of the invention can improve the calculation accuracy of the electronic medical record data identification probability, thereby improving the accuracy of the correlation calculation between electronic medical record data and COPD. The following is a detailed description of this electronic medical record data identification and processing method.
[0032] The electronic medical record (EMR) data identification and processing method provided in this invention is an EMR data identification and processing method based on the electronic weakening index. This method can improve the calculation accuracy of the EMR data identification probability, thereby improving the accuracy of the correlation calculation between EMR data and chronic obstructive pulmonary disease (COPD). It mainly includes three parts: EMR data preprocessing; calculation of the electronic weakening index based on the EMR system; and determination of the identification probability corresponding to the EMR data of the user to be detected based on the electronic weakening index and a probability identification model. The following is a detailed explanation... Figure 2 A detailed introduction will be provided.
[0033] First, we will introduce step 101 above and the steps for electronic medical record data preprocessing.
[0034] 1) Multi-source data association and anonymization
[0035] Electronic medical record (EMR) data primarily originates from various data sources, including community health checkups, hospital laboratory tests, disease diagnoses (test results), and medication prescriptions. Each data source possesses a unique identifier, necessitating data linking between these different sources before using EMRs. This module uses an individual's national identification number as the unique identifier to link data from different sources, connecting annual health checkups, laboratory tests, disease diagnoses, and medication prescriptions stored in different systems. This data is then used for subsequent calculations of the Electronic Frailty Index (EFI) and detection of Chronic Obstructive Pulmonary Disease (COPD). After this multi-source data linking, personal information such as the individual's national identification number, name, address, and phone number is deleted to achieve anonymization and protect personal privacy.
[0036] 2) Outlier filtering
[0037] Human errors frequently occur when medical staff enter electronic medical record data, resulting in data that exceeds the biologically reasonable range; this data is called outliers. Outliers reduce the accuracy of the model in detecting chronic obstructive pulmonary disease and therefore need to be removed. This module calculates the standard deviation and mean for each continuous feature, and finally filters feature values that are greater than the mean plus three times the standard deviation or less than the mean minus three times the standard deviation. The calculation methods for the standard deviation and mean are shown in formulas (1) and (2), where X represents the mean, n represents the sample size, and x... i Let Y represent the characteristic value of the i-th individual, and Y represent the standard deviation.
[0038]
[0039]
[0040] 3) Electronic medical record data normalization
[0041] This section normalizes continuous features. The normalization method is shown in Formula (3), where x represents the feature value, X represents the mean (calculation method is shown in Formula (1)), Y represents the standard deviation (calculation method is shown in Formula (2)), and Z represents the normalized feature value.
[0042]
[0043] As can be seen from the above, in one embodiment, the above-mentioned electronic medical record data identification and processing method may further include the following preprocessing operations on the electronic medical record data:
[0044] The electronic medical record data from different data sources are associated using the unique identifier of the user to be tested. The personal data in the associated electronic medical record data is then anonymized to obtain the multi-source data association and anonymized electronic medical record data; that is, the above steps of multi-source data association and anonymization.
[0045] The electronic medical record data after multi-source data association and anonymization is subjected to outlier filtering to obtain outlier-filtered electronic medical record data; that is, the outlier filtering steps mentioned above.
[0046] The continuous features in the outlier-filtered electronic medical record data are normalized to obtain preprocessed electronic medical record data, which is then used as the electronic medical record data of the user to be detected; that is, the above-mentioned electronic medical record data normalization step.
[0047] In practice, the above-mentioned implementation plan for preprocessing electronic medical record data can improve the accuracy of calculating the recognition probability of electronic medical record data, and further improve the accuracy of calculating the correlation between electronic medical record data and chronic obstructive pulmonary disease.
[0048] As can be seen from the above, in one embodiment, outlier filtering processing is performed on the electronic medical record data after multi-source data association and anonymization to obtain outlier-filtered electronic medical record data, which may include:
[0049] For each continuous feature in the electronic medical record data, the standard deviation and mean are calculated. Feature values that are greater than the mean plus a preset multiple (e.g., three times) or less than the mean minus the preset multiple are filtered to obtain the electronic medical record data after outlier filtering, which is the detailed implementation scheme of the above outlier filtering.
[0050] In practice, the above-mentioned implementation scheme for filtering outliers in electronic medical record data can improve the accuracy of calculating the probability of electronic medical record data identification, and thus further improve the accuracy of calculating the correlation between electronic medical record data and chronic obstructive pulmonary disease.
[0051] As can be seen from the above, in one embodiment, normalizing the continuous features in the outlier-filtered electronic medical record data to obtain preprocessed electronic medical record data as the electronic medical record data of the user to be detected may include:
[0052] Based on the mean and standard deviation of each continuous feature in the electronic medical record data, each continuous feature in the electronic medical record data is normalized to obtain the preprocessed electronic medical record data as the electronic medical record data of the user to be tested, which is a detailed implementation plan for electronic medical record data normalization.
[0053] In practice, the above-mentioned implementation scheme for normalizing electronic medical record data can improve the calculation accuracy of the recognition probability of electronic medical record data, and thus further improve the accuracy of the calculation of the correlation between electronic medical record data and chronic obstructive pulmonary disease.
[0054] Secondly, the steps for calculating the electron decay index are introduced, namely in step 102 above.
[0055] 1) Disease score calculation
[0056] Disease diagnosis data (disease examination result data) is extracted from electronic medical records. Individuals are scored based on whether they have a specific disease. Disease names and scoring criteria are shown in Table 1 below. The scores for each disease type are summed to obtain the disease score (i.e., the comprehensive value for each disease type). The calculation method is shown in formula (4), where n represents the total number of items used to calculate the disease score, and D... k This represents the score of the k-th item.
[0057]
[0058] Table 1. Disease names and scoring criteria used for calculating the Electroweakness Index.
[0059] 1 diabetes Diseased: D=1, Unaffected: D=0 2 asthma Diseased: D=1, Unaffected: D=0 3 peptic ulcer Diseased: D=1, Unaffected: D=0 4 gallstones Diseased: D=1, Unaffected: D=0 5 rheumatoid arthritis Diseased: D=1, Unaffected: D=0 6 fracture Diseased: D=1, Unaffected: D=0 7 neurasthenia Diseased: D=1, Unaffected: D=0 8 Chronic kidney disease Diseased: D=1, Unaffected: D=0 9 hypertension Diseased: D=1, Unaffected: D=0 10 heart disease Diseased: D=1, Unaffected: D=0 11 Stroke Diseased: D=1, Unaffected: D=0 12 cancer Diseased: D=1, Unaffected: D=0 13 Osteoporosis Diseased: D=1, Unaffected: D=0 14 goiter Diseased: D=1, Unaffected: D=0 15 arrhythmia Diseased: D=1, Unaffected: D=0
[0060] In practice, the disease types mentioned above can include any combination of the disease types in Table 1. The number and types of diseases can be flexibly selected according to the actual needs of the user, which improves the flexibility and efficiency of calculating the comprehensive value of each type of disease, thereby improving the flexibility and efficiency of calculating the electronic frailty index. Ultimately, it improves the flexibility and efficiency of calculating the correlation between the electronic medical record data of the user to be tested and chronic obstructive pulmonary disease. For example, when it is necessary to efficiently determine the comprehensive value of each type of disease and thus efficiently determine the electronic frailty index, the top 8 items with the highest weights of the disease types in Table 1 can be selected for calculation.
[0061] 2) Calculation of inspection and testing scores
[0062] Extract the laboratory test data from the electronic medical record data, and score the indicators according to whether they exceed a certain range. The scoring table for this module is shown in Table 2 below. Sum the scores of each indicator type to obtain the laboratory test score (i.e., the comprehensive value of each laboratory test item type). The calculation method is shown in formula (5), where I represents the i-th laboratory test indicator, n represents the total number of items used to calculate the score, and C l This indicates the score for item l.
[0063]
[0064] Table 2. Inspection and scoring criteria for calculating the electron decay index.
[0065]
[0066]
[0067] In practice, the above-mentioned indicator types can include any combination of the indicator types in Table 2 above. The number and type of indicators can be flexibly selected according to the actual needs of the user, which improves the flexibility and efficiency of calculating the comprehensive value of each test and examination item type, thereby improving the flexibility and efficiency of calculating the electronic frailty index, and ultimately improving the flexibility and efficiency of calculating the correlation between the electronic medical record data of the user to be tested and chronic obstructive pulmonary disease. For example, when it is necessary to efficiently determine the comprehensive value of each test and examination item type, and thus efficiently determine the electronic frailty index, the top 7 items with the highest indicator weights in Table 2 above can be selected for calculation.
[0068] 3) Electron Weakness Index Score
[0069] After calculating the disease score and the laboratory test score, the sum of these two scores yields the Electronic Frailty Index score. The calculation method is shown in formula (6), where m represents the m-th person. m This represents the electronic deterioration index score of the m-th person, or Disease score. m This represents the disease score of the m-th person. (Check score) m This represents the test score for the m-th person.
[0070] Electronic frailty score m =Disease score m +Check score m (6)
[0071] As can be seen from the above, in one embodiment, determining the electronic deterioration index of the user to be tested based on the user's electronic medical record data may include:
[0072] Extract the disease examination results data of the user to be tested from the electronic medical record data, and determine the comprehensive value of the user's disease for each type of disease based on the disease examination results data; that is, the steps of calculating the disease score mentioned above.
[0073] Extract the test and examination data of the user to be tested from the electronic medical record data, and determine the comprehensive value of each test and examination item type of the user to be tested based on the test and examination data; that is, the above steps for calculating the test and examination score.
[0074] The electronic weakness index of the user to be tested is determined based on the comprehensive values of various types of diseases and the comprehensive values of various types of test items; that is, the steps of scoring the electronic weakness index mentioned above.
[0075] In practice, the above-mentioned implementation scheme for determining the electronic frailty index of the user to be tested can improve the calculation accuracy of the electronic medical record data recognition probability, and thus further improve the accuracy of the calculation of the correlation between electronic medical record data and chronic obstructive pulmonary disease.
[0076] Next, let's introduce step 103 above.
[0077] In this step, a logistic regression model is used to calculate the recognition probability based on 12 features: electronic frailty index (representing aging characteristics), age, gender, smoking status, fasting blood glucose, exercise status, waist circumference, body mass index, low-density lipoprotein, high-density lipoprotein, total cholesterol, and triglycerides. This recognition probability can be used to characterize the correlation between the electronic medical record data of the user being tested and chronic obstructive pulmonary disease. The specific calculation method is shown in formulas (7) and (8), where a represents the intercept (which can be the ordinate), j represents the j-th feature, and β... j t represents the coefficient of the j-th feature (which can be obtained through training on sample data). j Let represent the value of the j-th feature, q represent the total number of features used for probability calculation (e.g., 12 as mentioned above), and g(z(x)) represent the recognition probability corresponding to the electronic medical record data.
[0078]
[0079]
[0080] In specific implementation, the samples used to train the probability recognition model in this embodiment of the invention may include positive samples and negative samples. Positive samples may be electronic medical record data related to chronic obstructive pulmonary disease (i.e., sample data with high recognition probability corresponding to the electronic weakness index), while negative samples may be electronic medical record data unrelated to chronic obstructive pulmonary disease (i.e., sample data with low recognition probability corresponding to the electronic weakness index).
[0081] As can be seen from the above, in one embodiment, the identification probability corresponding to the electronic medical record data of the user to be detected is obtained based on the electronic decay index of the user to be detected and the probability identification model, including: obtaining the identification probability corresponding to the electronic medical record data of the user to be detected based on the electronic decay index of the user to be detected and the logistic regression probability identification model.
[0082] In practice, using a logistic regression probability recognition model for probability recognition can further improve the accuracy of calculating the recognition probability of electronic medical record data, thereby further improving the accuracy of calculating the correlation between electronic medical record data and chronic obstructive pulmonary disease.
[0083] As can be seen from the above, in one embodiment, the identification probability corresponding to the electronic medical record data of the user to be detected is obtained based on the electronic deterioration index of the user to be detected and the probability identification model, including:
[0084] Based on the electronic debility index, demographic characteristics, physical examination characteristics, and biochemical indicators of the user to be tested, as well as the probability recognition model, the recognition probability corresponding to the electronic medical record data of the user to be tested is obtained.
[0085] In practice, by comprehensively considering the electronic frailty index, demographic characteristics, physical examination characteristics, and biochemical indicators, and performing probability identification, the accuracy of the electronic medical record data identification probability calculation can be further improved, thereby further improving the accuracy of the correlation calculation between electronic medical record data and chronic obstructive pulmonary disease.
[0086] As described above, in one embodiment, the identification probability corresponding to the electronic medical record data of the user to be detected is obtained based on the user's electronic deterioration index, demographic characteristics, physical examination characteristics, and biochemical indicators, as well as a probability recognition model. This includes obtaining the identification probability corresponding to the electronic medical record data of the user to be detected according to the following probability recognition model:
[0087]
[0088]
[0089] Where: a represents the intercept, j represents the j-th feature, and β j t represents the coefficient of the j-th characteristic. j Let represent the value of the j-th feature, q represent the total number of features used for probability calculation, and g(z(x)) represent the recognition probability corresponding to the electronic medical record data.
[0090] In practice, the above-mentioned probability recognition model can further improve the calculation accuracy of the recognition probability of electronic medical record data, thereby further improving the accuracy of the calculation of the correlation between electronic medical record data and chronic obstructive pulmonary disease.
[0091] In summary, the electronic medical record data identification and processing method provided in this embodiment of the invention can improve the calculation accuracy of the electronic medical record data identification probability, thereby improving the accuracy of the correlation calculation between electronic medical record data and chronic obstructive pulmonary disease. The beneficial technical effects achieved are as follows:
[0092] (1) Chronic obstructive pulmonary disease (COPD) has an insidious onset, and many COPD patients are missed in diagnosis. This invention can help doctors accurately determine the correlation between electronic medical record data and COPD at low cost by using the identification probability as an intermediate parameter obtained by the electronic medical record data identification and processing method implemented by computers and other devices in all steps. This can reduce the missed diagnosis rate of COPD.
[0093] (2) Chronic obstructive pulmonary disease can lead to a variety of serious complications such as respiratory failure and cardiovascular disease. This invention uses the identification probability as an intermediate parameter obtained by the electronic medical record data identification and processing method implemented by computers and other devices in all steps to help remind people whose electronic medical record data is highly related to chronic obstructive pulmonary disease to undergo medical testing as early as possible to determine whether they have chronic obstructive pulmonary disease. For patients who are diagnosed, treatment measures can be taken as early as possible to prevent the occurrence of serious complications and improve the quality of life and life expectancy of patients.
[0094] To test the accuracy of this invention in detecting patients with chronic obstructive pulmonary disease, the accuracy was validated based on two datasets from cities A and B. The indicators for evaluating the accuracy were area under the curves (AUC), sensitivity, and specificity.
[0095] Implementation Results
[0096] (1) Results of data aggregation in City A
[0097] The dataset for City A contains 42,995 individuals, including 927 patients with chronic obstructive pulmonary disease (COPD). This invention identifies the probability of electronic medical record (EMR) data recognition within this dataset, and then uses an intermediate parameter of this recognition probability to identify the correlation between EMR data and COPD. This dataset involves the AUC (Area Under the Receiver Operating Characteristic Curve, defined as the area under the ROC curve and the coordinate axes) of COPD patients. ROC stands for Receiver Operating Characteristic Curve. (See...) Figure 3 The horizontal axis represents specificity, and the vertical axis represents sensitivity. Figure 3 The curves represent the sensitivity and specificity for detecting patients with chronic obstructive pulmonary disease (COPD) at different discrimination thresholds. In this dataset, the AUC for detecting COPD patients using this invention is 81.69%, with optimal sensitivity and specificity of 75.71% and 74.00%, respectively.
[0098] (2) Results of data collection in City B
[0099] This invention was also tested on a dataset from City B, which contains 156,429 individuals, including 1,037 patients with chronic obstructive pulmonary disease (COPD). This invention identifies the probability of electronic medical record (EMR) data recognition in this dataset, and then uses an intermediate parameter of this probability to identify the correlation between EMR data and COPD. The AUC of COPD patients in this dataset is shown below. Figure 4 The horizontal axis represents specificity, and the vertical axis represents sensitivity. Figure 4 The curves represent the sensitivity and specificity for detecting patients with chronic obstructive pulmonary disease (COPD) at different discrimination thresholds. In this dataset, the AUC for detecting COPD patients in City B using this invention is 84.35%, with optimal sensitivity and specificity of 78.56% and 74.25%, respectively.
[0100] Conclusion: Experimental results demonstrate that the embodiments of the present invention improve the calculation accuracy of the recognition probability of electronic medical record data, thereby improving the accuracy of the calculation of the correlation between electronic medical record data and chronic obstructive pulmonary disease, and have good generalizability.
[0101] This invention also provides an electronic medical record data identification and processing device, as described in the following embodiments. Since the principle by which this device solves the problem is similar to that of the electronic medical record data identification and processing method, the implementation of this device can refer to the implementation of the electronic medical record data identification and processing method; repeated details will not be elaborated further.
[0102] Figure 5 This is a schematic diagram of the structure of the electronic medical record data recognition and processing device in an embodiment of the present invention, as shown below. Figure 5 As shown, the device includes:
[0103] Acquisition unit 01 is used to acquire the electronic medical record data of the user to be tested;
[0104] The electronic decay index determination unit 02 is used to determine the electronic decay index of the user to be tested based on the electronic medical record data of the user to be tested; the electronic decay index is used to characterize the aging characteristics of the user to be tested.
[0105] The probability determination unit 03 is used to obtain the recognition probability corresponding to the electronic medical record data of the user to be detected based on the electronic weakness index of the user to be detected and the probability recognition model. The recognition probability is used to characterize the correlation between the electronic medical record data of the user to be detected and chronic obstructive pulmonary disease. The probability recognition model is generated by pre-training a machine learning model based on the sample data of the relationship between the historical electronic weakness index and the historical recognition probability determined from the historical electronic medical record data of multiple users.
[0106] In one embodiment, the above-mentioned electronic medical record data identification and processing device may further include a preprocessing unit for performing the following preprocessing operations on the electronic medical record data:
[0107] The electronic medical record data from different data sources are associated using the unique identifier of the user to be tested. The personal data in the associated electronic medical record data is then anonymized to obtain multi-source data association and anonymized electronic medical record data.
[0108] Outlier filtering is performed on the electronic medical record data after multi-source data association and anonymization to obtain outlier-filtered electronic medical record data.
[0109] The continuous features in the outlier-filtered electronic medical record data are normalized to obtain preprocessed electronic medical record data, which is then used as the electronic medical record data of the user to be detected.
[0110] In one embodiment, outlier filtering is performed on the electronic medical record data after multi-source data association and anonymization to obtain outlier-filtered electronic medical record data, including:
[0111] For each continuous feature in the electronic medical record data, calculate the standard deviation and mean. Filter feature values that are greater than the mean plus a preset multiple or less than the mean minus the preset multiple to obtain the outlier-filtered electronic medical record data.
[0112] In one embodiment, the continuous features in the outlier-filtered electronic medical record data are normalized to obtain preprocessed electronic medical record data, which is then used as the electronic medical record data of the user to be detected. This includes:
[0113] Based on the mean and standard deviation of each continuous feature in the electronic medical record data, each continuous feature in the electronic medical record data is normalized to obtain the preprocessed electronic medical record data as the electronic medical record data of the user to be tested.
[0114] In one embodiment, the electronic decay index determination unit is specifically used for:
[0115] Extract the disease examination results data of the user to be tested from the electronic medical record data, and determine the comprehensive value of the user to be tested for various types of diseases based on the disease examination results data of the user to be tested.
[0116] Extract the test and examination data of the user to be tested from the electronic medical record data, and determine the comprehensive value of each test and examination item type of the user to be tested based on the test and examination data.
[0117] The electronic deterioration index of the user to be tested is determined based on the comprehensive values of various types of diseases and the comprehensive values of various types of test items.
[0118] In one embodiment, the probability determination unit is specifically used to: obtain the recognition probability corresponding to the electronic medical record data of the user to be detected based on the electronic deterioration index of the user to be detected and the logistic regression probability recognition model.
[0119] In one embodiment, the probability determination unit is specifically used to: obtain the recognition probability corresponding to the electronic medical record data of the user to be detected based on the user's electronic deterioration index, demographic characteristics, physical examination characteristics and biochemical index characteristics, as well as the probability recognition model.
[0120] In one embodiment, the probability determination unit is specifically used to: obtain the recognition probability corresponding to the electronic medical record data of the user to be detected according to the following probability recognition model:
[0121]
[0122]
[0123] Where: a represents the intercept, j represents the j-th feature, and β j t represents the coefficient of the j-th characteristic. j Let represent the value of the j-th feature, q represent the total number of features used for probability calculation, and g(z(x)) represent the recognition probability corresponding to the electronic medical record data.
[0124] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-described method for identifying and processing electronic medical record data.
[0125] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for identifying and processing electronic medical record data.
[0126] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described method for identifying and processing electronic medical record data.
[0127] Compared with existing technologies that rely solely on common user characteristics to identify electronic medical record (EMR) data, resulting in low accuracy, the EMR data identification processing scheme provided in this invention determines the EMR index of the user to be tested based on the EMR data. This EMR index characterizes the aging characteristics of the user. By incorporating these aging characteristics into probabilistic identification, the identification probability corresponding to the user's EMR data is obtained based on the EMR index and the probabilistic identification model. This identification probability characterizes the correlation between the user's EMR data and chronic obstructive pulmonary disease (COPD). The probabilistic identification model is pre-trained using sample data of the relationship between historical EMR indices and historical identification probabilities determined from historical EMR data of multiple users. Therefore, this EMR data identification processing scheme can improve the calculation accuracy of the EMR data identification probability, thereby improving the accuracy of the correlation calculation between EMR data and COPD.
[0128] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0129] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0130] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0131] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0132] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for identifying and processing electronic medical record data, characterized in that, include: Obtain the electronic medical record data of the user to be tested; Disease examination results data of the user to be tested are extracted from the electronic medical record data. Based on the disease examination results data, the comprehensive value of the user's various diseases is determined. Laboratory test data of the user to be tested are extracted from the electronic medical record data. Based on the laboratory test data, the comprehensive value of each type of laboratory test is determined. Based on the comprehensive value of the user's various diseases and the comprehensive value of each type of laboratory test, the electronic frailty index of the user is determined. The electronic frailty index is used to characterize the aging characteristics of the user to be tested. The recognition probability of the electronic medical record data of the user to be detected is obtained according to the following probability recognition model: in: The intercept is represented by j, which represents the j-th feature. The coefficient of the j-th characteristic is... Let q represent the value of the j-th feature, and q represent the total number of features used for probability calculation. This indicates the recognition probability corresponding to the electronic medical record data; The identification probability is used to characterize the correlation between the electronic medical record data of the user to be detected and chronic obstructive pulmonary disease; The probability recognition model is generated by pre-training a machine learning model based on sample data of the relationship between historical electronic decay index and historical recognition probability determined from historical electronic medical record data of multiple users.
2. The method as described in claim 1, characterized in that, This also includes the following preprocessing operations on electronic medical record data: Electronic medical record data from different data sources are linked using the unique identifier of the user to be tested. Personal data in the linked electronic medical record data is then anonymized to obtain multi-source data linked and anonymized electronic medical record data. Outlier filtering is performed on the electronic medical record data after multi-source data association and anonymization to obtain outlier-filtered electronic medical record data. The continuous features in the outlier-filtered electronic medical record data are normalized to obtain preprocessed electronic medical record data, which is then used as the electronic medical record data of the user to be detected.
3. The method as described in claim 2, characterized in that, Outlier filtering is performed on the electronic medical record data after multi-source data association and anonymization to obtain outlier-filtered electronic medical record data, including: For each continuous feature in the electronic medical record data, calculate the standard deviation and mean. Filter feature values that are greater than the mean plus a preset multiple or less than the mean minus the preset multiple to obtain the outlier-filtered electronic medical record data.
4. The method as described in claim 2, characterized in that, The continuous features in the outlier-filtered electronic medical record data are normalized to obtain preprocessed electronic medical record data, which serves as the electronic medical record data of the user to be detected. This includes: Based on the mean and standard deviation of each continuous feature in the electronic medical record data, each continuous feature in the electronic medical record data is normalized to obtain the preprocessed electronic medical record data as the electronic medical record data of the user to be tested.
5. An electronic medical record data processing device, characterized in that, include: The acquisition unit is used to acquire the electronic medical record data of the user to be tested; The electronic frailty index determination unit is used to extract disease examination results data of the user to be tested from electronic medical record data, and determine the comprehensive value of the user's various types of diseases based on the disease examination results data; extract the user's laboratory test data from electronic medical record data, and determine the comprehensive value of each type of laboratory test item based on the laboratory test data; and determine the user's electronic frailty index based on the comprehensive value of the user's various types of diseases and the comprehensive value of each type of laboratory test item; the electronic frailty index is used to characterize the aging characteristics of the user to be tested. The probability determination unit is used to obtain the recognition probability corresponding to the electronic medical record data of the user to be detected according to the following probability recognition model: in: The intercept is represented by j, which represents the j-th feature. The coefficient of the j-th characteristic is... Let q represent the value of the j-th feature, and q represent the total number of features used for probability calculation. This indicates the recognition probability corresponding to the electronic medical record data; The identification probability is used to characterize the correlation between the electronic medical record data of the user to be detected and chronic obstructive pulmonary disease; The probability recognition model is generated by pre-training a machine learning model based on sample data of the relationship between historical electronic decay index and historical recognition probability determined from historical electronic medical record data of multiple users.
6. The apparatus as claimed in claim 5, characterized in that, It also includes a preprocessing unit for performing the following preprocessing operations on the electronic medical record data: The electronic medical record data from different data sources are associated using the unique identifier of the user to be tested. The personal data in the associated electronic medical record data is then anonymized to obtain multi-source data association and anonymized electronic medical record data. Outlier filtering is performed on the electronic medical record data after multi-source data association and anonymization to obtain outlier-filtered electronic medical record data. The continuous features in the outlier-filtered electronic medical record data are normalized to obtain preprocessed electronic medical record data, which is then used as the electronic medical record data of the user to be detected.
7. The apparatus as claimed in claim 6, characterized in that, The preprocessing unit is used for: For each continuous feature in the electronic medical record data, calculate the standard deviation and mean. Filter feature values that are greater than the mean plus a preset multiple or less than the mean minus the preset multiple to obtain the outlier-filtered electronic medical record data.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 4.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 4.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 4.