Method and system for predicting patient auxiliary examination items based on supervised learning

The patient's electronic medical record data is preprocessed and characterized by random forest classification method, which solves the complexity and efficiency problems in predicting patient assisted examination items in the prior art, and achieves a more efficient, accurate and interpretable prediction effect.

CN119943427AActive Publication Date: 2025-05-06HUNAN UNIV

Patent Information

Application Number
CN202510012182.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The prior art has problems such as complex training, relying on a large amount of historical time series data, ignoring the correlation between features, high computational complexity, noise sensitivity, high resource consumption and lack of interpretability when predicting patient-assisted examination items.

Method used

The random forest classification method is used to predict, and the electronic medical record data of hospitalized patients is preprocessed, feature item data is extracted and feature coded, and the pre-trained random forest classification prediction model is used to predict.

Benefits of technology

The problems of insufficient prediction ability of patients' early disease auxiliary examination items, insufficient processing of feature correlations, high computational complexity, poor model robustness, large resource consumption and lack of explanatory nature are solved, and more efficient, accurate and interpretable auxiliary examination items are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943427A_ABST
    Figure CN119943427A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting an auxiliary examination item of a patient based on supervised learning, and the method comprises the following steps: obtaining an electronic medical record of a hospitalized patient, storing the electronic medical record and a corresponding hospitalization number as source data of the hospitalized patient in the form of an Excel table, carrying out the preprocessing of the obtained source data of the hospitalized patient, so as to obtain the preprocessed source data, and carrying out the prediction of the auxiliary examination item of the patient. And inputting the preprocessed source data into a pre-trained auxiliary examination item prediction model to obtain a preliminary prediction result of the auxiliary examination item of the inpatient, and performing inverse coding on the preliminary prediction result by using a pre-established auxiliary examination coding table to obtain a final prediction result of the auxiliary examination item of the inpatient. The technical problem that an existing LSTM model is complex in training and depends on a large amount of historical time sequence data, so that the early-stage auxiliary examination items of the disease of the patient cannot be predicted can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of supervised learning in machine learning, and more specifically, relates to a method and system for predicting patient auxiliary examination items based on supervised learning. Background Art

[0002] With the growth of population and the intensification of aging problems, medical needs are rising, and the shortage of medical resources has affected the efficiency and quality of medical services. In this context, how to rationally plan medical resources and improve medical efficiency has become an important issue that medical institutions need to solve urgently. Predicting the auxiliary examination items that patients may need provides strong support for improving the efficiency of medical resource utilization and optimizing medical service processes. By accurately predicting the auxiliary examination needs of patients, it is possible to arrange examination time and equipment in advance, reduce patient waiting time, avoid unnecessary repeated examinations, and effectively save medical costs. In addition, accurate auxiliary examination predictions can help doctors better formulate treatment plans and improve the timeliness and accuracy of diagnosis. Ultimately, this will not only help improve the quality of medical services, but also promote the advancement of personalized medicine and intelligent medical decision-making, thereby promoting the efficient operation of the entire medical system.

[0003] Currently, there is little research on predicting patient auxiliary examination items, and existing methods mainly tend to the fields of deep learning and machine learning. The first method is to implement it by training a long-short term memory (LSTM) model. This method mainly uses the patient's historical medical record data (such as past examination records, treatment history, etc.) as time series input to predict the auxiliary examination items that the patient may need in the future. The second method is to implement it through a Bayesian classifier (Naive Bayes). This method is based on the Bayesian theorem, assuming that the various characteristics of the patient are independent of each other, and uses prior probability and conditional probability to calculate the posterior probability for classification, calculate the probability of each examination item, and select the one with the largest probability as the prediction result. The third method is to use the K-nearest neighborhood (KNN) algorithm, which calculates the distance between the test patient and the historical patient in various characteristics (such as age, symptoms, past medical history, etc.), selects the most similar K historical patients' examination results, and performs voting or weighted averaging to predict the examination items required for the patient. The fourth method is a reinforcement learning (RL) method, which regards each patient's medical treatment process as a "state", auxiliary examinations as "actions", and the final treatment effect as a "reward". Through multiple training sessions, the reinforcement learning model can learn which examination items are most likely to affect the patient's diagnosis or treatment outcomes, thereby making reasonable predictions.

[0004] However, the above-mentioned existing methods for predicting patient auxiliary examination items all have some defects that cannot be ignored:

[0005] First, the LSTM model can effectively process time series data, but its training process is very complicated and requires a large amount of labeled data. For long-term dependencies in medical data, the LSTM model may rely on a large amount of patient historical data, resulting in the inability to predict the auxiliary examination items for patients in the early stages of the disease.

[0006] Second, the Bayesian classifier assumes that patient characteristics are independent of each other. However, in reality, various characteristics of patients are often highly correlated (such as the correlation between medical history and symptoms). The model does not adequately handle the dependencies between features, which may lead to biased prediction results. In particular, the performance of the Bayesian classifier is poor when dealing with complex medical data.

[0007] Third, the computational complexity of the KNN method is high, especially when the sample size is large. Calculating the distance between each test sample and all training samples may result in a long computation time. In addition, the KNN method is sensitive to noise and outliers and may be affected by a small number of abnormal cases, resulting in poor robustness of the model. At the same time, the method relies on an appropriate distance metric. If the dimension of the feature space is high, the performance of the KNN method will also drop significantly.

[0008] Fourth, reinforcement learning faces many challenges when processing medical data: First, medical decisions often involve long-term chain reactions, and model training requires a large amount of interactive data and high demand for computing resources; second, the reward function of reinforcement learning is complex to design, and it is difficult to determine the appropriate reward standard in medical applications. Over-reliance on the reward mechanism may cause the model to favor certain decisions while ignoring the individual differences of patients and the complexity of treatment; in addition, reinforcement learning has poor interpretability in practical applications, and it is not easy to understand how the model comes up with a specific prediction result, which affects its widespread application in medical scenarios. Summary of the invention

[0009] In response to the above defects or improvement needs of the prior art, the present invention provides a method and system for predicting patient auxiliary examination items through random forest classification, with the aim of solving the technical problems that the existing LSTM model cannot predict the auxiliary examination items of patients in the early stage of their disease due to complex training and reliance on a large amount of historical time series data, as well as the technical problems that the existing Bayesian classifier ignores the correlation between features and has poor ability to process complex data, resulting in deviations in prediction results, and the technical problems that the existing KNN method has high algorithm calculation complexity under large-scale data and is sensitive to outliers, resulting in poor model robustness, and the existing reinforcement learning method has high requirements on data type and data set size and complex reward mechanism design, resulting in high resource consumption and lack of explainability in the model training process.

[0010] To achieve the above object, according to one aspect of the present invention, a method for predicting patient auxiliary examination items based on supervised learning is provided, comprising the following steps:

[0011] (1) Obtain the electronic medical records of inpatients and store them and the corresponding hospitalization numbers in the form of an Excel spreadsheet as the source data of inpatients.

[0012] (2) Preprocessing the source data of the hospitalized patients obtained in step (1) to obtain preprocessed source data.

[0013] (3) Inputting the source data preprocessed in step (2) into a pre-trained auxiliary examination item prediction model to obtain preliminary prediction results of the auxiliary examination items of hospitalized patients, and using a pre-established auxiliary examination coding table to reverse-encode the preliminary prediction results to obtain the final prediction results of the auxiliary examination items of hospitalized patients.

[0014] Preferably, step (2) specifically comprises: first, cleaning and deduplication processing is performed on the obtained source data of hospitalized patients to obtain the source data after preliminary processing; then, keyword segmentation and data desensitization processing are performed on the source data after preliminary processing to obtain the source data after secondary processing; then, the required feature item data is extracted from the source data after secondary processing, and feature encoding is performed on the feature item data by the encoding method corresponding to the feature item data, and the source data finally obtained is the preprocessed source data;

[0015] The encoding method of feature item data is:

[0016] For the feature item data of gender characteristics, one-hot encoding is used to convert it into a binary feature encoding result;.

[0017] For the feature item data of age characteristics, it is necessary to classify them according to the distribution of hospitalized patients at different age stages, and then perform one-hot encoding on each age stage to obtain the feature encoding results;

[0018] For feature item data including admission reason features and preliminary diagnosis features (which are multi-label), first extract all labels of the feature item data to obtain a label list, then traverse the labels in the label list one by one, and obtain the corresponding code of each label in the pre-established feature coding table, and finally summarize the corresponding codes of all labels to obtain the feature coding result of the feature item data.

[0019] Preferably, the feature coding table is established according to the following steps: first, all labels of each feature item data are extracted and summarized from a large amount of sample data to obtain a corresponding label list, and then the label list is deduplicated and normalized to obtain a standardized label list, and finally the standardized label list is custom-encoded to obtain a feature coding table corresponding to the feature item data.

[0020] The auxiliary examination coding table is established according to the following steps: first, all different labels of each auxiliary examination data are extracted and summarized from a large amount of sample data to obtain the corresponding auxiliary examination label list (the auxiliary examination label list summarizes all different auxiliary examination item names), and then the auxiliary examination label list is deduplicated and normalized to obtain a standardized auxiliary examination label list, and finally, each label in the standardized auxiliary examination label list is assigned a custom code starting with the letter "E" followed by three digits to obtain the auxiliary examination coding table corresponding to the auxiliary examination data.

[0021] Preferably, the auxiliary inspection item prediction model includes a data processing module and a random forest classification prediction model;

[0022] The input of the data processing module is the source data of hospitalized patients, and the output is the preprocessed source data, which includes a three-layer structure:

[0023] The input of the first layer is the source data of hospitalized patients, which is cleaned and deduplicated, and the output is the source data after preliminary processing.

[0024] The second layer inputs the source data after preliminary processing output by the first layer. It performs content segmentation and data desensitization on the source data, and outputs the source data after secondary processing.

[0025] The input of the third layer is the secondary processed source data output by the second layer. It encodes the feature data of the source data and outputs the preprocessed source data.

[0026] Preferably, the random forest classification prediction model includes multiple individual decision trees, the input of the random forest classification prediction model is the preprocessed source data output by the data processing module, and the output is the prediction result of the auxiliary inspection item.

[0027] The random forest classification prediction model consists of the following three layers:

[0028] The input of the first layer is the preprocessed source data output by the data processing module, which performs feature extraction, data set division, data separation and other processing on the source data, and outputs data sets such as training set feature data, training set target variables, test set and validation set.

[0029] The input of the second layer is the training set feature data and training set target variables output by the first layer. Based on the classification and regression tree CART method, multiple individual decision trees are constructed and output according to the training set feature data and training set target variables. Each individual decision tree includes its own split structure and prediction results.

[0030] The input of the third layer is the multiple individual decision trees constructed as the output of the second layer, which are integrated and the prediction results are combined to output the prediction results of the auxiliary inspection items.

[0031] Preferably, the first layer of the data processing module first sorts the source data of hospitalized patients according to the hospitalization number to form an Excel list, then uses the dropna() method of the pandas library of Python to delete the missing item data and blank rows in the Excel list, and then uses the drop_duplicates() method to deduplicate the Excel list, and finally saves the deduplicated Excel table as the source data after preliminary processing;

[0032] The second layer of the data processing module first uses the replace() method of Python's pandas library to remove all spaces and line breaks in the source data after preliminary processing, and then uses the find() method to obtain the index of the target keyword, and uses the obtained index to segment the large amount of natural language in the source data into multiple data item lists. Then, all data item lists are stored in a new Excel table to obtain the segmented source data. Finally, sensitive data columns such as names are removed from the segmented source data to obtain a desensitized data set, which is the data set after secondary processing.

[0033] The third layer of the data processing module first extracts all feature data items in the source data after secondary processing, and then uses different encoding methods to encode the feature data items of different categories to obtain the corresponding feature encoding results, and finally writes the feature encoding results back to the corresponding columns in the source data to obtain the preprocessed source data.

[0034] Preferably, the first layer of the random forest classification prediction model first extracts a feature data set composed of multiple feature data from the preprocessed source data, and then divides the feature data set into a training set, a validation set, and a test set according to a ratio, and finally separates the feature data and the target variable in the training set to obtain the training set feature data and the training set target variable; among which, for the three multi-label data of admission reason data, preliminary diagnosis data, and auxiliary examination data, additional matrix processing is required before the next step of data set division, specifically:

[0035] First, for different categories of multi-label data, the corresponding feature coding table is queried to obtain the coding list corresponding to each category of multi-label data.

[0036] Then, for each category of multi-label data, create a two-dimensional matrix with the number of rows equal to the total number of multi-label data of the category, each row corresponding to a multi-label data, and the number of columns of the matrix equal to the number of codes in the corresponding code list, and each column corresponding to a code;

[0037] Finally, for each category of multi-label data, the corresponding two-dimensional matrix is ​​filled, that is, the codes contained in it are traversed, and the serial number of the code is obtained in the corresponding code list, and then the element in the row where the multi-label data is located and the column where the serial number of its code is located in the corresponding two-dimensional matrix is ​​set to 1, and the other elements in the two-dimensional matrix are set to 0;

[0038] The second layer of the random forest classification prediction model first initializes the key parameters of the individual decision tree, then uses bootstrap sampling to extract samples from the training set data with replacement to construct multiple individual decision trees, and then in the splitting process of each node of each individual decision tree, a part of the features are randomly selected from all the features to calculate the Gini index, and the feature with the smallest Gini index is selected for splitting to determine the best splitting feature; then, for the obtained best splitting feature, a list of all corresponding candidate splitting points is generated, each candidate splitting point in the list is traversed, the corresponding Gini index is calculated, and the candidate splitting point with the smallest Gini index is selected as the best splitting point of the current node; thereafter, according to the selected best splitting feature and best splitting point, the sample of the current node is divided into multiple child nodes, and the aforementioned feature selection and splitting process is recursively performed on each child node until the stopping condition is met; when the splitting stop condition of an individual decision tree is met, the current node becomes a leaf node, and the prediction result of the sample in the node is stored in a binary two-dimensional matrix form as the prediction result, and the construction of the individual decision tree is completed.

[0039] The third layer of the random forest classification prediction model first integrates all the constructed individual decision trees into a random forest, and then obtains the coding list corresponding to the auxiliary inspection data according to the auxiliary inspection coding table. Then, for each code in the coding list, the prediction results of the code in all individual decision trees of the random forest are counted, and the prediction result with the most occurrences is selected as the final prediction result of the label corresponding to the code. Finally, the prediction results of all codes are summarized, and the predicted existing codes are used as the prediction results of the auxiliary inspection items.

[0040] Preferably, the auxiliary inspection item prediction model is obtained by training through the following steps:

[0041] (3-1) An original data set consisting of electronic medical record data of multiple inpatients is obtained, the original data set is preprocessed, and a data enhancement operation is performed on the preprocessed data set to obtain a data enhanced data set.

[0042] (3-2) Perform feature extraction and matrix processing on the data enhanced data set obtained in step (3-1) to obtain a matrixed data set.

[0043] (3-3) Divide the matrixed data set obtained in step (3-2) into a training set, a validation set, and a test set, and further separate the training set feature data and the training set target variables from the training set.

[0044] (3-4) Initialize the parameters of the individual decision tree to obtain an initialized individual decision tree.

[0045] (3-5) Perform bootstrap random sampling on the training set feature data and training set target variables obtained in step (3-3) to construct multiple individual decision trees in parallel.

[0046] (3-6) For each individual decision tree constructed in step (3-5), first select the most influential preliminary diagnostic feature from its current node, then randomly select one feature from all other features of the current node, calculate the Gini index corresponding to the two selected features, and select the feature with the smallest Gini index as the best splitting feature for the current node of the individual decision tree.

[0047] (3-7) For each individual decision tree constructed in step (3-5), a list of all corresponding candidate splitting points is generated according to the best splitting feature of the current node of the individual decision tree selected in step (3-6), each candidate splitting point in the list is traversed, the corresponding Gini index is calculated, and the candidate splitting point with the smallest Gini index is selected as the best splitting point of the current node of the individual decision tree.

[0048] (3-8) For each individual decision tree constructed in step (3-5), the sample of the current node of the individual decision tree is divided into a left node and a right node according to the optimal splitting feature of the current node of the individual decision tree selected in step (3-6) and the optimal splitting point of the current node of the individual decision tree selected in step (3-7).

[0049] (3-9) For each individual decision tree constructed in step (3-5), recursively repeat the above steps (3-6) to (3-8) for the individual decision tree until the preset end condition is reached, thereby obtaining the constructed individual decision tree.

[0050] (3-10) All individual decision trees constructed in step (3-9) are integrated into a random forest as a preliminarily trained random forest classification prediction model.

[0051] (3-11) Using the test set obtained in step (3-3), the random forest classification prediction model preliminarily trained in step (3-10) is verified to obtain the performance analysis results of the model. According to the performance analysis results, the parameters of the random forest classification prediction model are continuously adjusted, and cross-validation is performed to obtain the final trained random forest classification prediction model.

[0052] Preferably, step (3-2) specifically comprises first extracting gender data, age data, admission reason data, preliminary diagnosis data, auxiliary examination data and other feature item data from the data set after data enhancement obtained in step (3-1), and then, for each multi-label data (including admission reason data, preliminary diagnosis data and auxiliary examination data), creating a corresponding two-dimensional matrix for the multi-label data, thereby obtaining a matrixed data set;

[0053] Step (3-2) further includes: for each category of multi-label data, create a two-dimensional matrix, that is, create a two-dimensional matrix for the three categories of multi-label data: admission reason data, preliminary diagnosis data, and auxiliary examination data. The number of rows of the two-dimensional matrix is ​​equal to the total number of multi-label data of the category, and each row corresponds to a multi-label data. The number of columns of the matrix is ​​equal to the number of codes in the corresponding coding list, and each column corresponds to a code.

[0054] Step (3-3) is specifically as follows: first, the data set is randomly sampled in a ratio of 7:2:1 and divided into a training set, a validation set, and a test set; then, in the training set, the matrixed preliminary diagnosis data and admission reason data are further extracted to form the training set feature data, and the matrixed auxiliary examination item data are extracted to form the training set target variable.

[0055] Step (3-9) is specifically, for each individual decision tree, the process from step (3-6) to step (3-8) is continuously repeated, so that the individual decision tree is continuously split until the preset maximum depth is reached or the number of samples of the current node is less than the minimum number of sample splits. At this time, the current node becomes a leaf node, and the corresponding individual decision tree is completed.

[0056] Step (3-11) is specifically as follows: first, by calculating the feature importance score, drawing the confusion matrix, etc., the performance indicators of the initially trained random forest classification prediction model on the validation set are obtained; then, the model parameters in step (3-4) are continuously adjusted according to the performance indicators, and the model training process from (3-5) to (3-10) is repeated for cross-validation to optimize the model and obtain the final trained random forest classification prediction model.

[0057] According to another aspect of the present invention, a system for predicting patient auxiliary examination items based on supervised learning is provided, comprising:

[0058] The first module is used to obtain the electronic medical records of hospitalized patients and store them and the corresponding hospitalization numbers in the form of Excel tables as source data of hospitalized patients.

[0059] The second module is used to preprocess the source data of hospitalized patients obtained by the first module to obtain preprocessed source data.

[0060] The third module is used to input the source data preprocessed by the second module into the pre-trained auxiliary examination item prediction model to obtain the preliminary prediction results of the auxiliary examination items of hospitalized patients, and use the pre-established auxiliary examination coding table to reversely encode the preliminary prediction results to obtain the final prediction results of the auxiliary examination items of hospitalized patients.

[0061] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:

[0062] 1. The present invention adopts steps (3-5) to (3-6) to construct multiple independent decision trees, and each tree only needs to train part of the data, does not rely on a large amount of historical time series data, and can randomly select different samples and features for learning each time training. Therefore, it can solve the technical problem that the existing LSTM model cannot predict the auxiliary examination items of the patient's early disease due to complex training and dependence on a large amount of historical time series data;

[0063] 2. The present invention adopts steps (3-2) to (3-10), integrates multiple decision trees, and each tree can consider the complex relationship between features during training without making independence assumptions. Therefore, it can solve the technical problem that the existing Bayesian classifier ignores the correlation between features and has poor ability to process complex data, resulting in deviations in prediction results.

[0064] 3. The present invention adopts steps (3-5) to (3-10), randomly selects samples and features during the training process, and integrates the predictions of multiple decision trees, which can effectively avoid the interference of a single outlier on the model prediction, and does not need to calculate the distance between each test sample and all training samples. Therefore, it can solve the technical problem that the existing KNN method has high algorithm calculation complexity under large-scale data and is sensitive to outliers, resulting in poor model robustness;

[0065] 4. Since the present invention adopts steps (3-6) to (3-8), the training process is relatively simple and the splitting rule of each tree is based on intuitive and transparent indicators (such as Gini index, information gain, etc.), and there is no need to design a complex reward mechanism. Therefore, it can solve the technical problems of existing reinforcement learning methods, such as high requirements on data types and data set sizes, complex reward mechanism design, large resource consumption in the model training process, and lack of interpretability;

[0066] 5. Due to the adoption of step (1), the invention determines that the data source is a large medical institution, which ensures the data volume support and the authenticity and validity of the data, determines that the target group is inpatients, ensures the quality of electronic medical record data, and determines that the data format is unified in Excel form, which is more friendly to the data provider and can establish a unified data standard and interface;

[0067] 6. The present invention adopts the feature engineering technology implemented in step (2), selects multi-dimensional and multi-label feature data, and proposes a coding table method with strong generalization and fault tolerance in feature data encoding, which makes full use of the diagnosis data and symptom data in the electronic medical record. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 It is a flow chart of a method for predicting auxiliary examination items of patients based on supervised learning of the present invention;

[0069] Figure 2 It is a schematic diagram of the system framework of the present invention;

[0070] Figure 3 It is a schematic diagram of the model structure of the random forest classification prediction model based on supervised learning of the present invention;

[0071] Figure 4It is a schematic diagram of the training process of the auxiliary inspection item prediction model of the present invention;

[0072] Figure 5 It is a schematic diagram comparing the original data and the data after segmentation processing of the present invention;

[0073] Figure 6 It is a schematic diagram of data comparison before and after the data is feature encoded in the present invention;

[0074] Figure 7 It is a line graph of the accuracy of the auxiliary inspection item label prediction results of the present invention;

[0075] Figure 8 It is a schematic diagram of the confusion matrix of the present invention;

[0076] Fig. 9 It is a multi-model performance evaluation comparison bar chart of the present invention. DETAILED DESCRIPTION

[0077] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0078] The purpose of the present invention is to predict the auxiliary examination items of patients based on electronic medical records. Therefore, there are certain targeted requirements for medical record data during the research process, that is, the patients recorded in the electronic medical records have undergone auxiliary examinations. Through statistical analysis of the electronic medical record data of patients in the Department of Cardiology of a large tertiary hospital, it can be seen that in the four months of autumn and winter in 2023, the demand for auxiliary examinations for hospitalized patients reached tens of thousands of times per month, which is much higher than the demand for auxiliary examinations for routine outpatients. The demand for auxiliary examinations for hospitalized patients is more guaranteed and the demand is greater. Therefore, the research of the present invention based on the electronic medical record data of hospitalized patients has greater data friendliness and greater research value. Therefore, the present invention has certain target group targeting, that is, hospitalized patients.

[0079] Through the observation and analysis of the electronic medical record data of a large number of inpatients, it is found that most electronic medical records always follow the following basic diagnosis and treatment process: first record the patient's personal information, then inquire about symptoms and medical history, then the doctor gives a preliminary diagnosis, prescribes auxiliary examination items, and finally prescribes a diagnosis and treatment plan based on the examination results. Therefore, even if the medical record styles of different departments or different medical institutions are different, some common data forms can still be captured according to the basic diagnosis and treatment process. The present invention has studied the medical office system of the data provider in detail. A single outpatient process mainly includes seven processes: symptom inquiry, physical examination, specialist examination, preliminary diagnosis, auxiliary examination, differential diagnosis, and treatment plan. Correspondingly, the patient's electronic medical record will also record all detailed information along with the outpatient process. After investigating and investigating more large new medical institutions, it was found that no matter how different the diagnosis and treatment systems of different medical institutions are, and how different the electronic medical record styles are, the content record of the entire electronic medical record must at least include patient information, reasons for admission, preliminary diagnosis, auxiliary examinations, diagnosis and treatment plans, etc. This provides a greater universal reference value for the present invention when selecting feature data.

[0080] In terms of feature selection, an electronic medical record records the details of the patient's diagnosis and treatment process, including hospitalization number, gender, age, admission department, ward, bed number, admission time, admission reason, patient profile, current medical history, past history, auxiliary examination, preliminary diagnosis, diagnostic basis, differential diagnosis, and treatment plan, a total of 16 items. The purpose of this invention is to predict the auxiliary examinations that patients need to do based on medical record information, so it is necessary to consider features that have a greater impact on this aspect, such as medical history information, preliminary diagnosis, and admission reason, and filter out irrelevant items such as admission department, treatment plan, bed number, etc., and then remove some data items such as medical history information that may be missing or extremely difficult to process in different medical institutions based on text complexity and method universality requirements. Finally, the feature selection is determined to be: gender, age, admission reason, preliminary diagnosis, a total of four dimensions.

[0081] In terms of model selection and training, we chose to construct and train the prediction model based on random forest classification. When predicting auxiliary examination items, multi-dimensional and multi-label features are involved, including the patient's gender, age, reason for admission, initial diagnosis, etc. These features constitute a high-dimensional feature space. Random forest can effectively handle high-dimensional features because when constructing each node of the individual decision tree, it randomly selects a part of the features from many features for evaluation to find the optimal features and split points.

[0082] It is worth noting that considering the incomplete correspondence between diseases and symptoms (diseases of different categories may have the same symptoms, such as "dizziness" may be caused by cardiovascular disease, neurological disease, or cardiovascular disease), it is necessary to train different models according to different departments of medical institutions (cardiovascular medicine, urology, neurosurgery, traditional Chinese medicine, etc.). This manual mainly uses cardiovascular medicine as an example, and other departments can be handled similarly.

[0083] like Figure 1 As shown, the present invention provides a method for predicting patient auxiliary examination items based on supervised learning, comprising the following steps:

[0084] (1) Obtain the electronic medical records of inpatients and store them and the corresponding hospitalization number (which serves as a unique identifier) ​​in the form of an Excel spreadsheet as the source data of the inpatients.

[0085] Specifically, this step is to obtain the electronic medical records of hospitalized patients through the electronic medical record system of the medical institution.

[0086] The premise of selecting the electronic medical record in this step is that the patient needs auxiliary examination. Currently, the present invention is more targeted at inpatients, because inpatients have a greater demand for auxiliary examinations.

[0087] The advantage of this step (1) is that it is simple and direct to obtain data through the electronic medical record system of the medical institution. The medical institution stores a large amount of electronic medical record information of inpatients, most of which are detailed and of high quality, and new data are continuously available. This is a huge advantage for offline model training.

[0088] (2) Preprocessing the source data of the hospitalized patients obtained in step (1) to obtain preprocessed source data.

[0089] Specifically, this step is as follows: first, the obtained source data of hospitalized patients are cleaned and deduplicated (data with missing or repeated parts are deleted) to obtain the source data after preliminary processing; then, the source data after preliminary processing are successively subjected to keyword segmentation and data desensitization processing (i.e., large sections of natural language content in the data are divided into multiple columns of data items, and then sensitive data content, such as patient names, etc.) to obtain the source data after secondary processing; subsequently, the required feature item data (such as gender characteristics, age characteristics, characteristics of admission reasons, and feature item data of preliminary diagnosis characteristics) are extracted from the source data after secondary processing, and feature encoding is performed on the feature item data through the encoding method corresponding to the feature item data (i.e., the Chinese natural language data therein is converted into encoded data), and the source data finally obtained is the preprocessed source data.

[0090] The encoding method of the feature item data of the present invention is:

[0091] For the feature item data of gender characteristics, one-hot encoding is used to convert them into binary (male 1, female 0) feature encoding results to avoid the influence of the sequential relationship between genders on the training process of the auxiliary examination item prediction model.

[0092] For the feature item data of age characteristics, it is necessary to classify and process the inpatients according to their distribution in different age stages, and then perform unique hot encoding on each age stage to obtain the feature encoding result. Taking cardiovascular medicine as an example, the present invention marks the age under 50 as 0, 50-65 as 1, 66-80 as 2, and over 80 as 3.

[0093] As for the feature item data such as the characteristics of the reasons for admission and the preliminary diagnosis characteristics, since they are multi-label (meaning that one data item corresponds to multiple label data, such as a preliminary diagnosis data contains both "hypertension" and "heart disease" diagnosis results, that is, this data contains the two labels "hypertension" and "heart disease", this preliminary diagnosis data is multi-label), first extract all the labels of the feature item data to obtain a label list, then traverse the labels in the label list one by one, and obtain the corresponding code of each label in the pre-established feature coding table, and finally summarize the corresponding codes of all labels to obtain the feature coding result of the feature item data.

[0094] In other words, the feature coding table of the present invention is established according to the following steps:

[0095] First, all labels of each feature item data are extracted and summarized from a large amount of sample data to obtain the corresponding label list, and then the label list is deduplicated and normalized to obtain a standardized label list. Finally, the standardized label list is custom encoded to obtain the feature coding table corresponding to the feature item data.

[0096] The advantage of this step (2) is that the data provider only needs to provide the source data in Excel format, and does not need to perform much additional processing, so that the preprocessing operation of the data can be carried out in a highly standardized and systematic manner, and is more friendly to the data provider. At the same time, through standardized data cleaning, coding and feature extraction processes, it is ensured that all data items can be uniformly and standardizedly input into the model. This reduces manual intervention and errors, and makes the data processing process repeatable and scalable, which is convenient for subsequent updates and maintenance. In addition, this data preprocessing method can adapt to the data characteristics of different departments and has good versatility.

[0097] (3) The source data preprocessed in step (2) is input into a pre-trained auxiliary examination item prediction model to obtain preliminary prediction results of the auxiliary examination items of hospitalized patients (the prediction results are in coded form), and the preliminary prediction results are reversed (the coded form data is converted into text data) using a pre-established auxiliary examination coding table to obtain the final prediction results of the auxiliary examination items of hospitalized patients (which are in natural language text form).

[0098] Specifically, the auxiliary inspection code table of the present invention is established according to the following steps:

[0099] Firstly, all different labels of each auxiliary inspection data (the label of the auxiliary inspection data is the specific inspection item name) are extracted and summarized from a large amount of sample data to obtain a corresponding auxiliary inspection label list (the auxiliary inspection label list summarizes all different auxiliary inspection item names), and then the auxiliary inspection label list is deduplicated and normalized to obtain a standardized auxiliary inspection label list, and finally each label in the standardized auxiliary inspection label list is assigned a custom code starting with the letter "E" followed by three digits to obtain the auxiliary inspection code table corresponding to the auxiliary inspection data.

[0100] The advantage of this step (3) is that, through automated processing and standardized coding mechanisms, manual intervention is reduced, the risk of errors is reduced, and the efficiency and accuracy of data processing can be significantly improved. At the same time, unified coding table management improves the system's cross-department adaptability and scalability, facilitates the processing of multi-label information from different data sources, and supports subsequent expansion and maintenance.

[0101] The auxiliary inspection item prediction model of the present invention includes two parts: a data processing module and a random forest classification prediction model.

[0102] like Figure 2 As shown, the data processing module of the present invention has the source data of hospitalized patients as input and the source data after preprocessing as output, which includes a three-layer structure:

[0103] The first layer takes as input the source data of hospitalized patients, performs data cleaning and deduplication processing on the source data, and outputs the source data after preliminary processing.

[0104] Specifically, the first layer sorts the source data of hospitalized patients according to the hospitalization number to form an Excel list, and then uses the dropna() method of Python's pandas library to delete the missing data and blank rows in the Excel list. Then, the drop_duplicates() method is used to deduplicate the Excel list, and finally the deduplicated Excel table is saved as the source data after preliminary processing.

[0105] The second layer inputs the source data after preliminary processing output by the first layer. It performs content segmentation and data desensitization on the source data, and outputs the source data after secondary processing.

[0106] Specifically, the second layer first uses the replace() method of the pandas library of Python to remove all spaces and line breaks in the source data after preliminary processing, and then obtains the index of the target keyword (the present invention uses keywords such as "hospitalization number", "gender", "age", "preliminary diagnosis", "reason for admission", and "auxiliary examination") through the find() method, and uses the obtained index to segment the large amount of natural language in the source data into multiple data item lists. Then, all the data item lists are stored in a new Excel table to obtain the segmented source data. Finally, sensitive data columns such as name are removed from the segmented source data to obtain a desensitized data set, which is the data set after secondary processing.

[0107] like Figure 5 Shown is a comparison between the original electronic medical record data of hospitalized patients consisting of a large amount of natural language and the data after content segmentation and data desensitization processing.

[0108] The input of the third layer is the secondary processed source data output by the second layer. It encodes the feature data of the source data and outputs the preprocessed source data.

[0109] Specifically, the third layer first extracts all feature data items in the source data after secondary processing, and then uses different encoding methods to encode the feature data items of different categories to obtain corresponding feature encoding results, and finally writes the feature encoding results back to the corresponding columns in the source data to obtain the preprocessed source data.

[0110] In other words, the encoding method of the feature data of the present invention is carried out according to the following steps:

[0111] For the feature item data of gender characteristics, one-hot encoding is used to convert them into binary (male 1, female 0) feature encoding results to avoid the influence of the sequential relationship between genders on the training process of the auxiliary examination item prediction model.

[0112] For the feature item data of age characteristics, it is necessary to classify and process the inpatients according to their distribution in different age stages, and then perform unique hot encoding on each age stage to obtain the feature encoding result. Taking cardiovascular medicine as an example, the present invention marks the age under 50 as 0, 50-65 as 1, 66-80 as 2, and over 80 as 3.

[0113] For multi-label feature item data such as admission reason features and preliminary diagnosis features, since they are multi-label, all labels of the feature item data are first extracted to obtain a label list, and then the labels in the label list are traversed one by one and corresponded to the pre-established feature coding table to obtain the corresponding code of each label, and finally the corresponding codes of all labels are summarized to obtain the feature coding result. Figure 6 Shown is a comparison diagram of data examples before and after feature encoding according to the present invention.

[0114] Taking the feature item data of the preliminary diagnosis feature as an example, in the sample data after secondary processing, its feature item data is displayed in the form of a long natural language sentence separated by numbers and periods. First, the periods are keyword indexed and the long sentence content is decomposed to obtain multiple disease name labels, and then these disease name labels are summarized and counted to obtain the corresponding label list, and then each disease name label in the label list is assigned a code starting with the letter "D" followed by three digits to obtain the disease name coding table corresponding to the preliminary diagnosis feature. A similar processing method is also applicable to features such as reasons for admission and auxiliary examinations. The present invention implements three coding tables: a coding table for reasons for admission, a coding table for disease names, and a coding table for auxiliary examinations.

[0115] It is worth noting that after preliminary sorting, the disease name coding table contains a large number of labels, which may lead to data sparsity problems. Therefore, the disease name coding table data needs to be comprehensively sorted out to ensure that all possible disease names are covered, and further coding assimilation is required, that is, to unify the coding of various synonyms and abbreviations, and to unify the disease classification items with little difference in the inspection items. Taking the disease name coding table of cardiovascular medicine as an example, "coronary atherosclerotic heart disease" is sometimes also referred to as "coronary heart disease", so they need to be set to the same label or the same code; for another example, for "heart function level II" and "heart function level III", the examinations required are almost the same, that is, blood tests and electrocardiograms, then the same disease code can be set. The processing done by the present invention is carried out under professional and scientific medical guidance. After processing, the disease codes of the cardiovascular department are determined to be 78 items, covering more than 300 cardiovascular diseases with different names.

[0116] like Figure 3 As shown, the random forest classification prediction model of the present invention includes multiple individual decision trees, which are independent of each other and together constitute a powerful supervised learning model. The input of the random forest classification prediction model is the preprocessed source data output by the data processing module, and the output is the prediction result of the auxiliary inspection item (which is in coded form).

[0117] The random forest classification prediction model consists of the following three layers:

[0118] The input of the first layer is the preprocessed source data output by the data processing module, which performs feature extraction, data set division, data separation and other processing on the source data, and outputs data sets such as training set feature data, training set target variables, test set and validation set.

[0119] Specifically, the first layer first extracts a feature data set composed of multiple feature data (including gender data, age data, admission reason data, preliminary diagnosis data, auxiliary examination data, etc.) from the preprocessed source data (the data in the feature data set are all in coded form), and then divides the feature data set into training set, validation set and test set according to a certain proportion (such as the common 70%-80% as training set, and the rest as validation set and test set). Finally, the feature data and target variables in the training set are separated to obtain the training set feature data and the training set target variable.

[0120] It is worth noting that the feature data set of the present invention is multi-dimensional and contains multi-label data. For the three multi-label data of admission reason data, preliminary diagnosis data and auxiliary examination data, additional matrix processing is required before the next step of data set division, specifically:

[0121] First, for different categories of multi-label data (i.e., admission reason data, preliminary diagnosis data, and auxiliary examination data), the corresponding feature coding tables (i.e., admission reason coding table, disease name coding table, and auxiliary examination coding table) are queried to obtain the coding list corresponding to each category of multi-label data (including a serial number column and a corresponding coding column, and each code in the coding list corresponds to a label).

[0122] Then, for each category of multi-label data, create a two-dimensional matrix (i.e., a two-dimensional matrix needs to be created for each of the three categories of multi-label data: admission reason data, preliminary diagnosis data, and auxiliary examination data). The number of rows of the matrix is ​​equal to the total number of multi-label data of the category, and each row corresponds to a multi-label data. The number of columns of the matrix is ​​equal to the number of codes in the corresponding coding list, and each column corresponds to a code. For example, for 1,000 preliminary diagnosis data, if there are 50 codes for disease names in the corresponding coding list obtained according to the disease name coding table, then the size of the matrix is ​​1,000 × 50, and each row of the matrix corresponds to a preliminary diagnosis data, and each column corresponds to a code for a disease name.

[0123] Finally, for each category of multi-label data, fill the corresponding two-dimensional matrix;

[0124] Specifically, for each multi-label data of this category (which is in the form of coding), traverse the codes it contains, and obtain the serial number of the code in the corresponding code list, and then set the elements in the row of the multi-label data in the corresponding two-dimensional matrix and the column where the serial number of its code is located to 1, and set the other elements in the two-dimensional matrix to 0. For example, a preliminary diagnosis data is "pneumonia" and "hypertension". In the code list, the corresponding code serial number of "pneumonia" is 3, and the corresponding code serial number of "hypertension" is 7. Then, in the row of the two-dimensional matrix corresponding to this preliminary diagnosis data, the elements in the 3rd and 7th columns are set to 1, and the elements in the remaining columns are set to 0.

[0125] The input of the second layer is the training set feature data and training set target variables output by the first layer. Based on the classification and regression tree (CART) method, multiple individual decision trees are constructed according to the training set feature data and training set target variables (individual decision trees are the basic building blocks of the random forest classification prediction model. As they continue to grow, they will gradually form a tree structure) and output them. Each individual decision tree includes its own split structure and prediction results.

[0126] Specifically, the second layer first initializes the key parameters of the individual decision tree (including the maximum depth, the minimum number of sample splits, the minimum number of leaf node samples, etc.), which determine the growth process and final shape of the individual decision tree and can be continuously tuned during the model training process. Then, bootstrap is used to extract samples with replacement from the training set data (including the training set feature data and the training set target variable) (for example, for a training set with n samples and m features, n samples (possibly repeated) will be extracted each time) to construct multiple individual decision trees (the root node of the individual decision tree stores all the extracted samples). Then, in the splitting process of each node of each individual decision tree, a part of the features (gender, age, reason for admission, preliminary diagnosis, etc.) is randomly selected to calculate the Gini index from all the features (features such as gender, age, reason for admission, and preliminary diagnosis), and the feature with the smallest Gini index is selected for splitting to determine the best splitting feature; then, for the best splitting feature obtained, a list of all corresponding candidate splitting points (the splitting point is the core of the splitting process of the individual decision tree, which determines how to divide the samples of the current node into different sub-nodes according to the value of a certain feature, and its value is different values ​​or labels of the corresponding feature) is generated, each candidate splitting point in the list is traversed, the corresponding Gini index is calculated, and the candidate splitting point with the smallest Gini index is selected as the best splitting point for the current node; thereafter, according to the selected best splitting feature and best splitting point, the sample of the current node is divided into multiple sub-nodes, and the aforementioned feature selection and splitting process is recursively performed on each sub-node until the stopping condition is met (the stopping condition is any of the following three points: the node reaches the set maximum depth, the number of samples in the current node is less than the minimum number of sample splits, and the samples in the node all belong to the same category). When the split stop condition of an individual decision tree is met, the current node becomes a leaf node, and the prediction results of the samples in the node are stored in a binary two-dimensional matrix as the prediction results (each row of the matrix corresponds to a sample in the leaf node, and each column of the matrix corresponds to a label of the target variable, "1" represents the existence of the label prediction, and "0" represents the non-existence of the label prediction), and the construction of the individual decision tree is completed. Thus, multiple complete individual decision trees are obtained.

[0127] It is worth noting that for the above-mentioned process of randomly selecting features, since the data feature dimensions of the present invention vary greatly and the importance distribution is uneven, the preliminary diagnostic features with the greatest decision-making influence are preferentially selected, and then a feature is randomly selected from all other features (gender, age, reason for admission, etc.).

[0128] In addition, since the feature data items of the present invention have multiple labels, and the feature data of multiple labels are matrixed (each row of the matrixed feature data represents a data sample, and each column represents a feature label), the calculation of the Gini index of the feature data of the multi-label data in the node splitting process requires additional processing, as follows:

[0129] First, for each label in the feature data of the feature, calculate the Gini index. Assuming there are L labels in the feature data, for label L i (i-th label, corresponding to the i-th column in the matrixed feature data), calculate the sample proportion p0 with the label 0 and the sample proportion p1 with the label 1, and then use the following formula to get the Gini index Gini (L i ):

[0130]

[0131] Then, the Gini index L for each label in the feature data of this feature is i Perform weighted averaging to obtain the overall Gini index of the feature for the current node. node , which is the Gini index of this feature. The specific formula is as follows:

[0132]

[0133] Finally, during the splitting process of the individual decision tree, the optimal splitting feature is selected based on the Gini index of each feature (i.e., the feature with the smallest Gini index is selected).

[0134] The input of the third layer is the multiple individual decision trees constructed as the output of the second layer, which are integrated and the prediction results are combined to output the prediction results of the auxiliary inspection items (in coded form).

[0135] Specifically, the third layer first integrates all the constructed individual decision trees into a random forest, and then obtains the code list corresponding to the auxiliary inspection data according to the auxiliary inspection code table (including the sequence number column and the corresponding code column, each code in the list corresponds to an auxiliary inspection label, that is, the label of the target variable). Then, for each code in the code list, the prediction results of the code in all individual decision trees of the random forest are counted (the prediction result is in binary form, "1" represents existence and "0" represents non-existence), and the prediction result with the largest number of occurrences (that is, the one with more occurrences between "1" and "0") is selected as the final prediction result of the label corresponding to the code (that is, the final prediction result is the "majority vote" of the prediction results of all trees), and finally the prediction results of all codes are summarized, and the code predicted to exist (that is, the final prediction result is "1") is used as the prediction result of the auxiliary inspection item.

[0136] For example, the target variable has 3 labels (coded as A, B, and C), and there are 5 individual decision trees in the random forest. The prediction result of each individual decision tree for a sample is:

[0137] Individual decision tree 1: [1,0,1]

[0138] Individual decision tree 2: [0,1,1]

[0139] Individual decision tree 3: [1,0,0]

[0140] Individual decision tree 4: [1,1,1]

[0141] Individual decision tree 5: [0,0,1]

[0142] For label A, the prediction results of the five individual decision trees are 1, 0, 1, 1, 0, respectively, that is, 3 individual decision trees predict that label A exists, and 2 individual decision trees predict that label A does not exist, so the final prediction result of label A is 1 (label A exists). Similarly, the final prediction result of label B is 0 (label B does not exist), and the final prediction result of label C is 1 (label C exists). Therefore, the final prediction result is [1, 0, 1], and the output prediction result is [A, C].

[0143] It is worth noting that this paper also uses a series of numerical analysis and visualization methods to evaluate the model performance, which mainly include the following three performance evaluation and optimization methods:

[0144] The first item is the prediction accuracy evaluation. For the prediction of auxiliary inspection items, since the target output is multi-label and the labels of auxiliary inspections are all specific inspection items, the weights are not considered, so the macro average accuracy is used for evaluation. The steps to calculate the macro average accuracy are as follows:

[0145] a. For a specific sample data, its prediction result (several check codes) is compared with the actual check code of the sample. If there are identical check codes, the corresponding label prediction is considered accurate.

[0146] b. Count all predictions and get the actual prediction amount for each label.

[0147] c. Compare the actual prediction amount of each label with the amount of data of that label in the sample to obtain the prediction accuracy of the label.

[0148] d. Macro average accuracy = (accuracy of label 1 + accuracy of label 2 + ... + accuracy of label n) / n.

[0149] The second item is performance evaluation. The present invention uses a variety of evaluation indicators (confusion matrix, recall rate, AUC value, F value, etc.) to comprehensively evaluate the performance of the model. According to the evaluation results, the hyperparameters of the model can be adjusted, such as the number of individual decision trees, the maximum depth of the tree, the minimum number of sample splits, etc., to further optimize the performance of the model.

[0150] The third item is to score the feature importance. By evaluating the feature importance, we can find out which features have the greatest impact on the model's prediction results, so that we can reasonably adjust the feature weights and provide a basis for subsequent feature selection and model optimization.

[0151] In summary, the random forest classification prediction model adopted in the present invention can accurately predict the auxiliary examination items required by patients by effectively processing and supervising the electronic medical record data of hospitalized patients, providing valuable reference for medical decision-making.

[0152] like Figure 4 As shown, the auxiliary inspection item prediction model of the present invention is obtained by training through the following steps:

[0153] (3-1) Obtain an original data set consisting of electronic medical records of multiple inpatients (in this example, a total of 3,895 electronic medical records of inpatients in the Department of Cardiology, each of which is uniquely identified by the hospitalization number and stored in the form of an Excel table). Preprocess the original data set, and perform data enhancement on the preprocessed data set to obtain a data-enhanced data set.

[0154] It should be noted that the process of preprocessing the original data set in this step is exactly the same as the above step (2), and will not be repeated here.

[0155] In addition, although the supervised learning algorithm based on random forest used in the present invention has a certain ability to resist overfitting, its basic unit individual decision tree may still overfit the training data when the amount of data is insufficient. Therefore, when the amount of data is small, the individual decision tree may remember the details and noise of the training samples instead of learning the real data pattern, resulting in a decrease in the performance of the model on the test set or new data. Therefore, the present invention also uses data enhancement technology to generate more sample data.

[0156] In other words, the present invention has established a relatively reliable admission reason coding table. In view of the many-to-many relationship between the disease and the symptoms at the onset, some common symptoms of diseases in the department can be added to the admission reason data item of the data set to expand the data set. In this example, common symptoms of cardiovascular diseases include "chest pain", "dyspnea", "syncope", "palpitations", "fatigue", etc., so these common symptoms can be randomly added to the admission reason data to achieve the purpose of expanding the data set. After processing, the 3895 original data of this example obtained 3761 valid coded data, and after data enhancement, a total of 8000 valid data were obtained. Data from other departments can also be processed similarly, and this step can achieve effective data enhancement.

[0157] The advantage of this sub-step (3-10) is that the data set can be effectively expanded through data augmentation technology, thereby avoiding overfitting of the model to limited data during training. Especially when using a random forest-based algorithm, data augmentation helps improve the generalization ability of the model and avoids the model "remembering" the details and noise of the training data. This method is also applicable to other departments, so it can enhance the diversity and scale of data in different departments and ensure a wider range of data applications.

[0158] (3-2) Perform feature extraction and matrix processing on the data enhanced data set obtained in step (3-1) to obtain a matrixed data set.

[0159] Specifically, first, feature item data such as gender data, age data, admission reason data, preliminary diagnosis data, and auxiliary examination data are extracted from the data enhanced data set obtained in step (3-1), and then, for each multi-label data (including admission reason data, preliminary diagnosis data, and auxiliary examination data), a corresponding two-dimensional matrix is ​​created for the multi-label data to obtain a matrixed data set.

[0160] In other words, for each category of multi-label data, create a two-dimensional matrix (that is, a two-dimensional matrix needs to be created for the three types of multi-label data: admission reason data, preliminary diagnosis data, and auxiliary examination data). The number of rows of the matrix is ​​equal to the total number of multi-label data of the category, and each row corresponds to a multi-label data. The number of columns of the matrix is ​​equal to the number of codes in the corresponding coding list, and each column corresponds to a code.

[0161] Taking the preliminary diagnosis data as an example, assuming that the data set contains 8000 samples and 78 disease labels, we can create a two-dimensional matrix of 8000 × 78. For each sample, we traverse its disease label, and if the label exists, fill 1 in the corresponding position in the matrix, otherwise fill 0.

[0162] The advantage of this sub-step (3-2) is that, through matrix processing, multi-label data (such as preliminary diagnosis data) can be converted into a high-dimensional feature space, and each label is processed as an independent dimension, which improves the fine-grained expression of features. This not only provides more information for the model, but also simplifies the processing flow of multi-label data, avoiding tedious manual annotation and conversion work, thereby improving the efficiency of data processing and the prediction performance of the model.

[0163] (3-3) Divide the matrixed data set obtained in step (3-2) into a training set, a validation set, and a test set, and further separate the training set feature data and the training set target variables from the training set.

[0164] Specifically, first, the data set was randomly sampled in a ratio of 7:2:1 and divided into training set, validation set and test set; then, in the training set, the matrixed preliminary diagnosis data and admission reason data were further extracted to form the training set feature data, and the matrixed auxiliary examination item data were extracted to form the training set target variable.

[0165] In other words, assuming that the matrixed training set is data_matrix, where the first few columns are the training set feature data and the last few columns are the training set target variables, then:

[0166] Training set feature data: X_train = data_matrix[:,:-num_target_cols]

[0167] Training set target variable: y_train = data_matrix[:,-num_target_cols:]

[0168] Here, num_target_cols represents the number of columns of the target variable in the training set.

[0169] It is worth noting that the training set in the present invention is used to adjust parameter configurations such as trainable weights in the auxiliary inspection item prediction model, and the validation set is used to adjust hyperparameters such as the learning rate of the auxiliary inspection item prediction model. The test set does not participate in the training of the model and is used to statistically calculate the final prediction effect of the auxiliary inspection item prediction model.

[0170] The advantage of this sub-step (3-30) is that random sampling can effectively avoid data bias, ensure that each sample has a chance to be selected, and provide a wider data distribution, so that the training set can cover more sample features and enhance the robustness of the model.

[0171] (3-4) Initialize the parameters of the individual decision tree to obtain an initialized individual decision tree.

[0172] Specifically, the following parameters are set in this example:

[0173] n_estimators=200: indicates that the random forest contains 200 decision trees.

[0174] max_depth=10: The maximum depth of the decision tree is 10, which limits the growth depth of the tree.

[0175] min_samples_split=2: Stop further splitting when the number of samples in the current node is less than 2.

[0176] min_samples_leaf=1: A leaf node must contain at least 1 sample.

[0177] The advantage of this sub-step (3-4) is that, through such parameter settings, the complexity of the decision tree can be effectively controlled, preventing the tree from growing too deep and avoiding model overfitting. At the same time, ensuring that each leaf node contains at least one sample can ensure that each branch of the tree represents an actual class as much as possible.

[0178] (3-5) Perform bootstrap random sampling on the training set feature data and training set target variables obtained in step (3-3) to construct multiple individual decision trees in parallel.

[0179] When constructing each individual decision tree, 5,600 samples are randomly selected with replacement from the training set feature data.

[0180] The advantage of this sub-step (3-5) is that by randomly sampling the samples in the training set with replacement (i.e., bootstrap sampling), each individual decision tree sees a slightly different data set during training. This incomplete overlap of data increases the diversity of the model and helps reduce the overfitting phenomenon that may occur in a single individual decision tree.

[0181] (3-6) For each individual decision tree constructed in step (3-5), first select the most influential preliminary diagnostic feature from its current node, then randomly select one feature from all other features of the current node (gender, age, reason for admission, etc.), calculate the Gini index corresponding to the two selected features, and select the feature with the smallest Gini index as the best splitting feature for the current node of the individual decision tree.

[0182] The advantage of this sub-step (3-6) is that it can reduce redundancy and unnecessary computational overhead by prioritizing the most decision-making features. Focusing on the features that are most relevant to the target variable not only reduces training time, but also reduces noise during training, thereby improving model efficiency and accuracy. This process of feature selection helps improve the interpretability of the model and ensures that each individual decision tree is learned using the most relevant data.

[0183] (3-7) For each individual decision tree constructed in step (3-5), based on the best splitting feature of the current node of the individual decision tree selected in step (3-6), a list of all corresponding candidate splitting points (the candidate splitting points are all possible values ​​or labels of the best splitting feature) is generated, each candidate splitting point in the list is traversed, the corresponding Gini index is calculated, and the candidate splitting point with the smallest Gini index (i.e., the "purest" splitting point) is selected as the best splitting point for the current node of the individual decision tree.

[0184] Specifically, for the gender and age features in this example, the candidate split points are different values ​​of the feature. For multi-label features such as reasons for admission and preliminary diagnosis, the candidate split points are a certain label of the feature or a combination of different labels. For example, for the feature "age", the ages of all samples in this example are marked as "0", "1", "2", and "3", and the candidate split points are usually the "midpoints" of these marked values, such as "1" or "2". For the preliminary diagnosis feature, the candidate split points may be "hypertension" (whether the corresponding position in the corresponding matrix is ​​"1") or a combination of "hypertension" and "heart disease" (whether the corresponding positions in the corresponding matrix are all "1").

[0185] The advantage of this sub-step (3-7) is that, based on the Gini index calculation method, the impact of each feature on the label can be evaluated more carefully. By calculating the Gini index of each label, the contribution of different features to sample classification can be quantified more accurately, so as to make more reasonable decisions in the construction of individual decision trees.

[0186] (3-8) For each individual decision tree constructed in step (3-5), the sample of the current node of the individual decision tree is divided into a left node and a right node according to the optimal splitting feature of the current node of the individual decision tree selected in step (3-6) and the optimal splitting point of the current node of the individual decision tree selected in step (3-7).

[0187] For example, if the feature selected is age, and the split point is "30 years old", then the samples of the current node will be divided into samples younger than 30 years old and samples greater than or equal to 30 years old.

[0188] (3-9) For each individual decision tree constructed in step (3-5), recursively repeat the above steps (3-6) to (3-8) for the individual decision tree until the preset end condition is reached, thereby obtaining the constructed individual decision tree.

[0189] Specifically, for each individual decision tree, each of its child nodes will continuously repeat the process from step (3-6) to step (3-8), so that the individual decision tree will continue to split until it reaches the preset maximum depth or the number of samples at the current node is less than the minimum number of sample splits (in this example, the maximum depth is 10 and the minimum number of sample splits is 2). At this time, the current node will become a leaf node, and the corresponding individual decision tree will be completed.

[0190] The advantage of this sub-step (3-9) is that the recursive process allows the depth and structure of the tree to be dynamically adjusted according to the actual data, avoiding over-simplification or over-complication caused by fixed depth, and ensuring that individual decision trees can capture complex patterns while avoiding overfitting.

[0191] (3-10) All individual decision trees constructed in step (3-9) are integrated into a random forest as a preliminarily trained random forest classification prediction model.

[0192] Specifically, for a new sample, each individual decision tree in the random forest classification prediction model will give a prediction result. The final prediction result is determined by a voting mechanism, that is, the prediction category that appears most times in all individual decision trees is selected as the final prediction result.

[0193] The advantage of this sub-step (3-10) is that, through the voting mechanism of each individual decision tree, the random forest integrates the decisions of multiple trees, reducing the error caused by overfitting or bias of a single tree.

[0194] (3-11) Using the test set obtained in step (3-3), the random forest classification prediction model preliminarily trained in step (3-10) is verified to obtain the performance analysis results of the model. According to the performance analysis results, the parameters of the random forest classification prediction model are continuously adjusted, and cross-validation is performed to obtain the final trained random forest classification prediction model.

[0195] Specifically, first, the performance indicators (accuracy, F1 value, AUC value, etc.) of the initially trained random forest classification prediction model on the validation set are obtained by calculating the feature importance score and drawing the confusion matrix, and then the model parameters in step (3-4) are continuously adjusted according to the performance indicators, and the model training process from (3-5) to (3-10) is repeated for cross-validation to optimize the model and obtain the final trained random forest classification prediction model. For example, the number of individual decision trees can be adjusted, starting from 50, and gradually increased to 500 at intervals of 50, and the optimal number of decision trees can be determined according to the performance indicators.

[0196] The advantage of this sub-step (3-11) is that the multi-dimensional analysis and evaluation can fully understand the performance of the model. At the same time, by analyzing the feature importance scores, it is possible to identify which features are most critical in the decision-making process, which provides support for the interpretability of the model.

[0197] Experimental Results

[0198] In order to illustrate the effectiveness of the method of the present invention and its improvement on classification effect, the prediction results of auxiliary examination items in the cardiovascular internal medicine test set data are taken as an example to illustrate the effectiveness of the method.

[0199] Figure 7 The prediction results of each auxiliary inspection label are shown. As can be seen from the figure, there are 47 labels in total, and the accuracy of most label predictions fluctuates stably around the average value. In general, the macro average accuracy of the optimal model trained by the present invention is 84%.

[0200] Figure 8 The confusion matrix diagram drawn for the prediction structure of auxiliary examination items is a two-dimensional 47×47 table, which represents the classification prediction results of 47 different labels of auxiliary examinations. Each element Cij in the matrix represents the number of samples that actually belong to category i but are predicted to be category j. When i=j, it means that the actual label and the predicted label are completely matched, that is, the true positive. It can be seen that the color of the diagonal elements in the table is more prominent, that is, the number of predicted true positives is more prominent, which is a good performance of the prediction results. And based on the confusion matrix, indicators such as recall rate, F value, AUC value, etc. can be counted to evaluate the classification performance of the model.

[0201] In addition, in order to further illustrate the effectiveness of the method of the present invention, a detailed comparative analysis was conducted using the same data set with multiple models commonly used in the field of machine learning classification prediction (i.e., K-nearest neighbor algorithm KNN, support vector machine SVM, stochastic gradient descent SGD, multilayer perceptron MLP). The following table shows the detailed comparative evaluation numerical data, Fig. 9For comparison of the bar graph (RF is the present invention), it can be seen that compared with the other models, the auxiliary inspection item prediction model constructed by the present invention has the highest macro average accuracy, F-measure and area under the curve (AUC) are also the best, and the regression rate ranks second. Therefore, overall, the implementation effect of the present invention is better.

[0202] Model Macro average accuracy Regression rate F-number AUC value The present invention 0.836 0.810 0.857 0.910 K-Nearest Neighbor Algorithm 0.784 0.835 0.721 0.825 Support Vector Machine 0.711 0.771 0.733 0.785 Stochastic Gradient Descent 0.675 0.523 0.688 0.674

[0203] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for predicting patient auxiliary examination items based on supervised learning, characterized in that: The following steps are involved: (1) Obtain the electronic medical records of inpatients and store them and the corresponding hospitalization numbers in the form of an Excel spreadsheet as the source data of the inpatients. (2) Preprocessing the source data of the hospitalized patients obtained in step (1) to obtain preprocessed source data. (3) Inputting the source data preprocessed in step (2) into a pre-trained auxiliary examination item prediction model to obtain preliminary prediction results of the auxiliary examination items of hospitalized patients, and using a pre-established auxiliary examination coding table to reverse-encode the preliminary prediction results to obtain the final prediction results of the auxiliary examination items of hospitalized patients.

2. The method for predicting patient auxiliary examination items based on supervised learning according to claim 1, characterized in that: Step (2) specifically includes, first, cleaning and deduplication processing of the obtained source data of hospitalized patients to obtain the source data after preliminary processing; Then, the source data after the initial processing is subjected to keyword segmentation and data desensitization processing in order to obtain the source data after the secondary processing; Subsequently, the required feature item data is extracted from the secondary processed source data, and the feature item data is feature encoded using the encoding method corresponding to the feature item data, and the source data finally obtained is the preprocessed source data; The encoding method of feature item data is: For the feature item data of gender characteristics, one-hot encoding is used to convert it into a binary feature encoding result;. For the feature item data of age characteristics, it is necessary to classify them according to the distribution of hospitalized patients at different age stages, and then perform one-hot encoding on each age stage to obtain the feature encoding results; For feature item data including admission reason features and preliminary diagnosis features (which are multi-label), first extract all labels of the feature item data to obtain a label list, then traverse the labels in the label list one by one, and obtain the corresponding code of each label in the pre-established feature coding table, and finally summarize the corresponding codes of all labels to obtain the feature coding result of the feature item data.

3. The method for predicting patient auxiliary examination items based on supervised learning according to claim 1 or 2, characterized in that: The feature coding table is established according to the following steps: first, all labels of each feature item data are extracted and summarized from a large amount of sample data to obtain the corresponding label list, and then the label list is deduplicated and normalized to obtain a standardized label list. Finally, the standardized label list is custom-encoded to obtain the feature coding table corresponding to the feature item data. The auxiliary examination coding table is established according to the following steps: first, all different labels of each auxiliary examination data are extracted and summarized from a large amount of sample data to obtain the corresponding auxiliary examination label list (the auxiliary examination label list summarizes all different auxiliary examination item names), and then the auxiliary examination label list is deduplicated and normalized to obtain a standardized auxiliary examination label list, and finally, each label in the standardized auxiliary examination label list is assigned a custom code starting with the letter "E" followed by three digits to obtain the auxiliary examination coding table corresponding to the auxiliary examination data.

4. The method for predicting patient auxiliary examination items based on supervised learning according to any one of claims 1 to 3, characterized in that: The auxiliary inspection project prediction model includes a data processing module and a random forest classification prediction model; The input of the data processing module is the source data of hospitalized patients, and the output is the preprocessed source data, which includes a three-layer structure: The input of the first layer is the source data of hospitalized patients, which is cleaned and deduplicated, and the output is the source data after preliminary processing. The second layer inputs the source data after preliminary processing output by the first layer. It performs content segmentation and data desensitization on the source data, and outputs the source data after secondary processing. The input of the third layer is the secondary processed source data output by the second layer. It encodes the feature data of the source data and outputs the preprocessed source data.

5. The method for predicting patient auxiliary examination items based on supervised learning according to claim 4, characterized in that: The random forest classification prediction model includes multiple individual decision trees. The input of the random forest classification prediction model is the preprocessed source data output by the data processing module, and the output is the prediction result of the auxiliary inspection item. The random forest classification prediction model consists of the following three layers: The input of the first layer is the preprocessed source data output by the data processing module, which performs feature extraction, data set division, data separation and other processing on the source data, and outputs data sets such as training set feature data, training set target variables, test set and validation set. The input of the second layer is the training set feature data and training set target variables output by the first layer. Based on the classification and regression tree CART method, multiple individual decision trees are constructed and output according to the training set feature data and training set target variables. Each individual decision tree includes its own split structure and prediction results. The input of the third layer is the multiple individual decision trees constructed as the output of the second layer, which are integrated and the prediction results are combined to output the prediction results of the auxiliary inspection items.

6. The method for predicting patient auxiliary examination items based on supervised learning according to claim 5, characterized in that: The first layer of the data processing module sorts the source data of hospitalized patients according to their hospital numbers to form an Excel list, then uses the dropna() method of the python pandas library to delete missing data and blank rows in the Excel list, then uses the drop_duplicates() method to deduplicate the Excel list, and finally saves the deduplicated Excel table as the source data after preliminary processing; The second layer of the data processing module first uses the replace() method of Python's pandas library to remove all spaces and line breaks in the source data after preliminary processing, and then uses the find() method to obtain the index of the target keyword, and uses the obtained index to segment the large amount of natural language in the source data into multiple data item lists. Then, all data item lists are stored in a new Excel table to obtain the segmented source data. Finally, sensitive data columns such as names are removed from the segmented source data to obtain a desensitized data set, which is the data set after secondary processing. The third layer of the data processing module first extracts all feature data items in the source data after secondary processing, and then uses different encoding methods to encode the feature data items of different categories to obtain the corresponding feature encoding results, and finally writes the feature encoding results back to the corresponding columns in the source data to obtain the preprocessed source data.

7. The method for predicting patient auxiliary examination items based on supervised learning according to claim 6, characterized in that: The first layer of the random forest classification prediction model first extracts a feature data set composed of multiple feature data from the preprocessed source data, and then divides the feature data set into a training set, a validation set, and a test set according to the proportion. Finally, the feature data and the target variable in the training set are separated to obtain the training set feature data and the training set target variable. Among them, for the three multi-label data of admission reason data, preliminary diagnosis data, and auxiliary examination data, additional matrix processing is required before the next step of data set division, specifically: First, for different categories of multi-label data, the corresponding feature coding table is queried to obtain the coding list corresponding to each category of multi-label data. Then, for each category of multi-label data, create a two-dimensional matrix with the number of rows equal to the total number of multi-label data of the category, each row corresponding to a multi-label data, and the number of columns of the matrix equal to the number of codes in the corresponding code list, and each column corresponding to a code; Finally, for each category of multi-label data, the corresponding two-dimensional matrix is ​​filled, that is, the codes contained in it are traversed, and the serial number of the code is obtained in the corresponding code list, and then the element in the row where the multi-label data is located and the column where the serial number of its code is located in the corresponding two-dimensional matrix is ​​set to 1, and the other elements in the two-dimensional matrix are set to 0; The second layer of the random forest classification prediction model first initializes the key parameters of the individual decision tree, then uses bootstrap sampling to extract samples from the training set data with replacement to build multiple individual decision trees. Then, in the splitting process of each node of each individual decision tree, a part of the features is randomly selected from all the features to calculate the Gini index, and the feature with the smallest Gini index is selected for splitting to determine the best splitting feature; then, for the best splitting feature obtained, a list of all corresponding candidate splitting points is generated, each candidate splitting point in the list is traversed, the corresponding Gini index is calculated, and the candidate splitting point with the smallest Gini index is selected as the best splitting point for the current node; Afterwards, according to the selected optimal splitting feature and optimal splitting point, the sample of the current node is divided into multiple child nodes, and the aforementioned feature selection and splitting process is recursively performed on each child node until the stopping condition is met; when the splitting stopping condition of an individual decision tree is met, the current node becomes a leaf node, and the prediction results of the samples in the node are stored in a binary two-dimensional matrix form as the prediction result, and the construction of the individual decision tree is completed. The third layer of the random forest classification prediction model first integrates all the constructed individual decision trees into a random forest, and then obtains the coding list corresponding to the auxiliary inspection data according to the auxiliary inspection coding table. Then, for each code in the coding list, the prediction results of the code in all individual decision trees of the random forest are counted, and the prediction result with the most occurrences is selected as the final prediction result of the label corresponding to the code. Finally, the prediction results of all codes are summarized, and the predicted existing codes are used as the prediction results of the auxiliary inspection items.

8. The method for predicting patient auxiliary examination items based on supervised learning according to claim 7, characterized in that: The auxiliary inspection item prediction model is obtained through training in the following steps: (3-1) An original data set consisting of electronic medical record data of multiple inpatients is obtained, the original data set is preprocessed, and a data enhancement operation is performed on the preprocessed data set to obtain a data enhanced data set. (3-2) Perform feature extraction and matrix processing on the data enhanced data set obtained in step (3-1) to obtain a matrixed data set. (3-3) Divide the matrixed data set obtained in step (3-2) into a training set, a validation set, and a test set, and further separate the training set feature data and the training set target variables from the training set. (3-4) Initialize the parameters of the individual decision tree to obtain an initialized individual decision tree. (3-5) Perform bootstrap random sampling on the training set feature data and training set target variables obtained in step (3-3) to construct multiple individual decision trees in parallel. (3-6) For each individual decision tree constructed in step (3-5), first select the most influential preliminary diagnostic feature from its current node, then randomly select one feature from all other features of the current node, calculate the Gini index corresponding to the two selected features, and select the feature with the smallest Gini index as the best splitting feature for the current node of the individual decision tree. (3-7) For each individual decision tree constructed in step (3-5), a list of all corresponding candidate splitting points is generated according to the best splitting feature of the current node of the individual decision tree selected in step (3-6), each candidate splitting point in the list is traversed, the corresponding Gini index is calculated, and the candidate splitting point with the smallest Gini index is selected as the best splitting point of the current node of the individual decision tree. (3-8) For each individual decision tree constructed in step (3-5), the sample of the current node of the individual decision tree is divided into a left node and a right node according to the optimal splitting feature of the current node of the individual decision tree selected in step (3-6) and the optimal splitting point of the current node of the individual decision tree selected in step (3-7). (3-9) For each individual decision tree constructed in step (3-5), recursively repeat the above steps (3-6) to (3-8) for the individual decision tree until the preset end condition is reached, thereby obtaining the constructed individual decision tree. (3-10) All individual decision trees constructed in step (3-9) are integrated into a random forest as a preliminarily trained random forest classification prediction model. (3-11) Using the test set obtained in step (3-3), the random forest classification prediction model preliminarily trained in step (3-10) is verified to obtain the performance analysis results of the model. According to the performance analysis results, the parameters of the random forest classification prediction model are continuously adjusted, and cross-validation is performed to obtain the final trained random forest classification prediction model.

9. The method for predicting patient auxiliary examination items based on supervised learning according to claim 8, characterized in that: Step (3-2) is specifically as follows: first, extracting characteristic item data such as gender data, age data, admission reason data, preliminary diagnosis data, auxiliary examination data, etc. from the data set after data enhancement obtained in step (3-1); then, for each multi-label data (including admission reason data, preliminary diagnosis data, and auxiliary examination data), creating a corresponding two-dimensional matrix for the multi-label data, thereby obtaining a matrixed data set; Step (3-2) further includes: for each category of multi-label data, create a two-dimensional matrix, that is, create a two-dimensional matrix for the three categories of multi-label data: admission reason data, preliminary diagnosis data, and auxiliary examination data. The number of rows of the two-dimensional matrix is ​​equal to the total number of multi-label data of the category, and each row corresponds to a multi-label data. The number of columns of the matrix is ​​equal to the number of codes in the corresponding coding list, and each column corresponds to a code. Step (3-3) is specifically as follows: first, the data set is randomly sampled in a ratio of 7:2:1 and divided into a training set, a validation set, and a test set; then, in the training set, the matrixed preliminary diagnosis data and admission reason data are further extracted to form the training set feature data, and the matrixed auxiliary examination item data are extracted to form the training set target variable. Step (3-9) is specifically, for each individual decision tree, the process from step (3-6) to step (3-8) is continuously repeated, so that the individual decision tree is continuously split until the preset maximum depth is reached or the number of samples of the current node is less than the minimum number of sample splits. At this time, the current node becomes a leaf node, and the corresponding individual decision tree is completed. Step (3-11) is specifically as follows: first, by calculating the feature importance score, drawing the confusion matrix, etc., the performance indicators of the initially trained random forest classification prediction model on the validation set are obtained; then, the model parameters in step (3-4) are continuously adjusted according to the performance indicators, and the model training process from (3-5) to (3-10) is repeated for cross-validation to optimize the model and obtain the final trained random forest classification prediction model.

10. A system for predicting patient auxiliary examination items based on supervised learning, characterized in that: include: The first module is used to obtain the electronic medical records of hospitalized patients and store them and the corresponding hospitalization numbers in the form of Excel tables as source data of hospitalized patients. The second module is used to preprocess the source data of hospitalized patients obtained by the first module to obtain preprocessed source data. The third module is used to input the source data preprocessed by the second module into the pre-trained auxiliary examination item prediction model to obtain the preliminary prediction results of the auxiliary examination items of hospitalized patients, and use the pre-established auxiliary examination coding table to reversely encode the preliminary prediction results to obtain the final prediction results of the auxiliary examination items of hospitalized patients.

Citation Information

Patent Citations

  • Methylated biomarkers related to antipsychotic drug curative effect prediction

    CN113355406A

  • Diabetes ICU patient death risk prediction method based on artificial intelligence

    CN114664449A

  • Improved methods for identification of functional cell states

    US20240337647A1

  • Asthma diagnosis system based on decision tree and improved smote algorithms

    WO2022198761A1

Cited By

  • Parkinson's disease early recognition system and method based on multi-task learning

    CN120913884A