Automatic machine learning-based doctor seeing rationality analysis method and model

By constructing a medical consultation rationality analysis model based on automatic machine learning methods, the problems of single scenarios and small coverage in existing technologies are solved, efficient and accurate medical consultation rationality analysis is achieved, manual intervention is reduced, and analysis efficiency and accuracy are improved.

CN120708924APending Publication Date: 2025-09-26HUNAN AUTOMOTIVE ENG VOCATIONAL COLLEGE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510593260.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies focus on a single scenario in the analysis of the rationality of medical treatment, have a small coverage, and are unable to attribute the causes of specific violations, resulting in low efficiency and susceptibility to subjective factors.

Method used

An automatic machine learning-based method is used to construct a medical consultation rationality analysis model through automatic identification, preprocessing, feature engineering and model training. Algorithms such as random forest, LightGBM, AdaBoost, XGBoost and logistic regression are used, combined with recursive feature elimination and cross-validation to screen out the optimal features and models, and identify unreasonable medical consultation groups and violation characteristics.

Benefits of technology

It improves the efficiency and accuracy of medical rationality analysis, reduces the workload of manual review, enhances decision-making transparency, provides valuable evaluation results, reduces the influence of subjective factors, and can automatically process complex medical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708924A_ABST
    Figure CN120708924A_ABST
Patent Text Reader

Abstract

The invention discloses a doctor-seeing rationality analysis method and model based on automatic machine learning, and the method comprises the steps: building a doctor-seeing rationality analysis model through an automatic machine learning system, training the model through a training data set, enabling the model to automatically learn and recognize the rationality and irrationality in doctor-seeing data, and carrying out the recognition of the rationality and irrationality of the doctor-seeing data. And unreasonable types and behaviors are judged. Through automatic learning and optimization of a machine learning model, the method can cover a whole treatment process including links of diagnosis, treatment, medication and the like, can continuously adapt to new treatment data and changes, improves the generalization ability and robustness of the model, and solves the problems that in the prior art, a focusing scene is single, the coverage range is small, and the efficiency is high. And specific violation behaviors cannot be resolved, and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical consultation rationality analysis, and more specifically, relates to a medical consultation rationality analysis method and model based on automatic machine learning. Background Art

[0002] With the development of medical information technology, a vast amount of medical visit data is being collected and stored. Traditional methods for analyzing the rationality of medical visits often rely on manual screening, review, and empirical judgment from this vast amount of medical visit data. This is inefficient and susceptible to subjective factors. Therefore, how to effectively analyze and utilize this data to quickly determine the rationality of medical visits is crucial for improving medical efficiency and reducing medical costs. In particular, assessing the rationality of medical visits has become a pressing issue.

[0003] In order to detect violations in medical data with multi-source heterogeneous characteristics and complexity, CN 109934719 A discloses a method for detecting medical insurance violations, which includes the following steps: S1: obtaining medical insurance data and its corresponding medical data, and describing the medical data from multiple perspectives; S2: using a machine learning algorithm to extract features from the medical data and the medical insurance data to obtain feature vectors; S3: using a deep learning algorithm to train the feature vectors and the medical insurance data to generate a multidimensional feature model; S4: searching for medical data corresponding to the medical insurance data to be detected, inputting the medical insurance data to be detected and its corresponding medical data into the multidimensional feature model, and detecting the type of medical insurance behavior contained in the medical insurance data to be detected. Specifically, the patent uses machine learning algorithms to group patients according to disease diagnosis-related factors, such as age, gender, length of hospital stay, clinical diagnosis, symptoms, surgery, disease severity, comorbidities and complications, and prognosis. It analyzes medical data through the basic information and relationships of hospitals, doctors, and patients, and uses a variety of machine learning algorithms to standardize medical data. It also uses expert rules as feature inputs to obtain return result data of expert rules, and divides the return result data into multiple categories of data with different characteristics. Although the patent utilizes the learning behavior of machine learning algorithms to autonomously process large and complex medical data, it improves the accuracy, timeliness, and effectiveness of violations in the medical process. However, the patent focuses on a single scenario, has a small coverage, and cannot attribute the causes of specific violations. Summary of the Invention

[0004] In order to overcome the problems of existing analysis of the rationality of medical treatment in the medical process, such as a single focus scenario, a small coverage, and an inability to attribute specific violations, the present invention provides a method for analyzing the rationality of medical treatment based on automatic machine learning.

[0005] Another technical problem of the present invention is to provide a medical consultation rationality analysis model based on automatic machine learning.

[0006] The present invention is achieved through the following technical solutions:

[0007] A method for analyzing the rationality of medical consultation based on automatic machine learning, comprising the following steps:

[0008] S1. Collect data from the medical information form, settlement information form, and expense details form, identify the data characteristics and task types, and pre-process the data;

[0009] S2. Extracting features related to medical treatment rationality from the preprocessed data to obtain the data source;

[0010] S3. Use an automated machine learning system to build a model for analyzing the rationality of medical consultations and train the model using a training dataset.

[0011] S31. Identify unreasonable medical treatment groups from data sources;

[0012] S311. Using the data source as the first sample data, the first sample data is trained using random forest to obtain feature weight scores and sort them;

[0013] S312. Based on the sorting of feature weight scores, loop through each feature quantity of the first sample data, obtain the feature quantity with the highest score through recursive feature elimination and cross-validation, and select the selected features as the second sample data;

[0014] S313. Perform automatic learning and training on the second sample data to obtain an optimal model, and use the optimal model to identify the unreasonable medical treatment group;

[0015] S32. Based on automated machine learning, screen for violation characteristics and types of unreasonable medical treatment groups;

[0016] S4. Evaluate the visit data to be analyzed.

[0017] Furthermore, the preprocessing includes cleaning, removing duplicate, missing and abnormal data, and filling null value data.

[0018] Furthermore, the features related to the rationality of the medical treatment include extracting one or more of the medical treatment ID, gender, age, personnel type, industry, unit type, diagnosis primary code of the medical treatment ID, secondary code, hospitalization code, number of hospitalization days, frequency of medical treatment, medical institution, total cost, unified fund expenditure, total personal payment, and hospital grade.

[0019] Furthermore, 5-fold cross-validation was used for cross-validation, and the score was selected as the average of 5 validation scores.

[0020] Furthermore, Optuna is used to adjust parameters in automatic learning training.

[0021] A model for analyzing the rationality of medical consultations based on automated machine learning, including an automatic recognition module, a data preprocessing module, a feature engineering module, and a model training and evaluation module;

[0022] The automatic identification module includes a feature type identification unit and a task type identification unit, which can automatically identify the feature type and task type according to the data conditions of the data source;

[0023] Data preprocessing module, cleans data and removes duplicate, missing and abnormal data;

[0024] The feature engineering module constructs features for preprocessed data, calculates feature weights, and performs cross-validation to obtain the best feature selection;

[0025] The model training and evaluation module includes a model task selection unit, an algorithm selection unit, a model training unit, a hyperparameter optimization unit, a model fusion unit, and a model evaluation unit. The model task selection unit selects the corresponding model task according to the task type, the algorithm selection unit selects the corresponding algorithm according to the model task, the model training unit trains and cross-validates the base learners, the hyperparameter optimization unit automatically tunes the hyperparameters of each base learner, the model fusion unit fuses the base models, and the model evaluation unit evaluates and compares each base model and the fusion model, and outputs the optimal model.

[0026] Furthermore, feature types include discrete, integer, floating-point, text, and time; task types include binary classification tasks, multi-classification tasks, and regression tasks.

[0027] Furthermore, the algorithm selection unit includes random forest, LightGBM, AdaBoost, XGBoost and logistic regression classifiers for binary classification tasks and multi-classification tasks, and selects random forest, LightGBM, XGBoost classifiers for regression tasks.

[0028] Furthermore, the construction of discrete features includes first performing One-Hot encoding and then processing it using the OneHotEncoder in the sklean.preprocessing library; the construction of floating-point features includes first performing normalization and then processing it using the StandardScaler in the sklean.preprocessing library; the construction of integer features includes first performing binning and processing it using the KBinsDiscretizer in the sklean.preprocessing library; the text features are processed using the TfidfVectorizer in sklearn.feature_extraction.text; the time features extract the year, month, day, hour, week, day of the week, whether it is a weekday, whether it is the beginning of the month, and whether it is the end of the month as new feature columns.

[0029] Furthermore, the model fusion unit adopts the fusion strategy of Stacking and Voting to fuse the base models.

[0030] Compared with the prior art, the beneficial effects are:

[0031] The present invention adds systematic feature screening preprocessing to improve data quality, and performs corresponding feature processing according to the data release of the data, making data processing more flexible and without the need for human intervention. Then, through the automatic learning and optimization of the machine learning model, the weight score normalization process is introduced to make the feature importance comparable, and the feature subset is dynamically verified using recursive feature elimination combined with cross-validation, which can continuously adapt to new medical data and changes, improve the generalization ability and robustness of the model, and improve the efficiency and accuracy of the analysis of the rationality of medical treatment. The present invention can automatically complete tasks such as model selection, feature engineering and hyperparameter optimization, identify the irrational population, type and violation items of medical treatment, establish a visual relationship between the number of features and model performance, enhance decision-making transparency, reduce the workload of manual review and the influence of subjective factors, and the output evaluation results provide valuable references, which are convenient for personnel to better understand the medical situation and identify illegal medical records. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a process framework diagram for identifying unreasonable groups seeking medical treatment;

[0033] Figure 2 It is an indicator graph of the accuracy of the 10-fold cross validation of the three algorithms. DETAILED DESCRIPTION

[0034] The present invention will be further explained and illustrated below with reference to the examples, but the specific examples do not limit the present invention in any form. Unless otherwise specified, the methods and equipment used in the examples are conventional methods and equipment in the art, and the raw materials used are all conventional commercially available raw materials.

[0035] Example 1

[0036] This embodiment provides a medical consultation rationality analysis model based on automatic machine learning, including an automatic recognition module, a data preprocessing module, a feature engineering module, and a model training and evaluation module;

[0037] The automatic identification module, including the feature type identification unit and the task type identification unit, can automatically identify the feature type and task type according to the data conditions of the data source. Specifically:

[0038] Automatically identify the type of feature variables based on data distribution, including discrete, integer, floating-point, text, and time. Discrete variables are identified when the data type is object, bool, or category, and the field length is less than or equal to 25; integer variables are identified when the data type contains "int"; floating-point variables are identified when the data type contains "float"; text variables are identified when the data type is object, bool, or category, and the field length is greater than 25; and time variables are identified when the data type is object or category and can be processed using the to_datetime function in the pandas library.

[0039] Automatically identify the task type based on target variable data analysis, including binary classification, multi-classification, and regression. The task type identification criteria are: After removing duplicates from the target variable, count the number of categories. If the number of categories equals 2, it is identified as a binary classification task. If the directory variable is an integer and the number of categories after removing duplicates is less than 10% of the total number of rows, it is identified as a multi-classification task. Otherwise, it is identified as a regression problem.

[0040] The data preprocessing module cleans the data and removes duplicate, missing, and outlier data. For example, data with null values ​​for gender, age, and visit time are removed from medical records. Industry, unit type, and employee type are filled with 'A99999'; and null values ​​for hospitalization days are filled with 0. In expense details, outlier data with total charges equal to or less than 0 for the same medical insurance expense name is removed. In settlement data, outlier data with total charges equal to or less than 0 is deleted.

[0041] The feature engineering module extracts relevant features from the preprocessed data. For discrete feature columns, One-Hot encoding is performed on them using the OneHotEncoder in the sklean.preprocessing library. For floating-point features, they are normalized using the StandardScaler in the sklean.preprocessing library. For integer features, they are binned using the KBinsDiscretizer in the sklean.preprocessing library, where the number of bins is calculated using the Sturgess formula as num_bins = int(1+np.log2(n)), where int is rounded and np.log2 is a function in Numpy used to calculate the logarithm with base 2. For time features, the year, month, day, hour, week, day of the week, whether it is a working day, whether it is the beginning of the month, and whether it is the end of the month are extracted as new feature columns. For text features, TfidfVectorizer in sklearn.feature_extraction.text is used.

[0042] The weights of the features are then calculated and cross-validated to obtain the best feature selection.

[0043] The model training and evaluation module includes a model task selection unit, an algorithm selection unit, a model training unit, a hyperparameter optimization unit, a model fusion unit, and a model evaluation unit. Specifically:

[0044] The model task selection unit selects the corresponding model task according to the task type, including binary classification tasks, multi-classification tasks, and regression tasks. The algorithm selection unit selects the corresponding algorithm according to the model task. For binary classification tasks, five classifiers are selected as the first-layer base learners: RF (random forest), LightGBM (lightweight gradient boosting machine), AdaBoost (adaptive boosting), XGBoost (extreme gradient boosting tree), and LR (logistic regression); for multi-classification tasks, five classifiers are selected as the first-layer base learners: RF (random forest), LightGBM (lightweight gradient boosting machine), AdaBoost (adaptive boosting), XGBoost (extreme gradient boosting tree), and LR (logistic regression); for regression tasks, three regression algorithms are selected as the first-layer base learners: RF (random forest), LightGBM (lightweight gradient boosting machine), and XGBoost (extreme gradient boosting tree). The model training unit trains and cross-validates the base learners, the hyperparameter optimization unit automatically tunes the hyperparameters of each base learner, the model fusion unit fuses the base models based on the base learners using the Stacking and Voting fusion strategies, and the model evaluation unit evaluates and compares each base model and the fusion model and outputs the optimal model.

[0045] Example 2

[0046] This embodiment provides a method for analyzing the rationality of medical consultation based on automatic machine learning, characterized by the following steps:

[0047] S1. Collect data from the medical information form, settlement information form, and expense details form. The medical information form includes characteristics such as gender, age, unit type, personnel type, industry, and number of days of hospitalization; the expense details form includes characteristics such as the medical insurance catalog list of the medical ID, the corresponding expenses and quantities, and the medical insurance catalog type; the settlement information form includes characteristics such as the total cost of the medical treatment, the unified fund expenditure, the total personal payment, and the hospital grade.

[0048] S2. Preprocess the data. Specifically, first calculate the null value rate of each feature, filter out columns with a null value rate exceeding 80%, and remove them. Next, process missing values. For discrete features, use "-99999" to fill missing values; for integer features, use 0 to fill missing values; for floating-point features, use the average value of the feature column to fill missing values; for text features, use an empty string to fill missing values; for time features, use the previous row of the feature column that is not null to fill missing values. For example, eliminate data with null values ​​for gender, age, and visit time. Use 'A99999' to fill in the industry, unit type, and personnel type; and fill null values ​​for hospitalization days with 0. Eliminate abnormal data in the expense details table where the total cost for the same medical insurance expense name is equal to or less than 0. Delete abnormal data in the settlement data where the total cost is equal to or less than 0.

[0049] S3. By searching relevant business data and consulting professionals, extract features related to the rationality of medical visits from the preprocessed data to obtain the data source. Features related to the rationality of medical visits include basic information features such as the visit ID, gender, age, patient type, industry, and unit type from the medical information table, as well as relevant medical information such as the primary diagnosis code, secondary code, hospitalization code, length of stay, frequency of visits, and medical institution. Information related to the total cost, overall fund expenditure, total personal payment, and hospital grade of the visit ID are extracted from the settlement information table.

[0050] S4. Use an automated machine learning system to construct a model for analyzing the rationality of medical consultations and train the model using a training dataset.

[0051] S41. Identify unreasonable medical treatment groups from data sources;

[0052] S411. Using the data source as the first sample data, the first sample data is trained using a random forest to obtain feature weight scores, and the feature importance is made comparable based on the feature weight scores.

[0053] S412. For the first sample data, loop through each feature number, use recursive feature elimination (RFE) combined with 5-fold cross validation, gradually reduce the number of features through cyclic iteration, establish a relationship curve between the number of features and model performance (cross-validation score), and determine the optimal feature subset.

[0054] Specifically, in 5-fold cross-validation, the data is divided into five parts: four for training and one for validation. This process is repeated five times, with a different validation set selected each time. The average cross-validation score is the average of these five validation scores. As the number of features is cycled through, both RFE and 5-fold cross-validation are performed for each number, and the average score is calculated. The feature set with the highest average score is selected as the optimal feature combination.

[0055] S413. The determined optimal feature subset is used as the second sample data set. Automatic learning and training is performed on the second sample data. The sub-models of each sub-algorithm selected according to the task type are automatically tuned using the Optuna (Optimization Tuning) strategy to obtain the optimal model. The optimal model is used to identify the unreasonable medical treatment groups. The unreasonable medical treatment IDs are filtered out from the first sample data as a new data set to identify the specific medical insurance violation catalog for those medical treatment IDs.

[0056] S42. Similar to step S41, a model is trained based on automatic machine learning to predict whether medical treatment behavior is reasonable, and the importance of features that lead to unreasonable behavior is output. Then, feature importance analysis is performed to identify which medical insurance catalog entries have the greatest impact on the prediction results. Based on the results of the feature importance analysis, the specific medical treatment ID is associated with the medical insurance catalog entry that causes the behavior to be unreasonable, that is, the violation type.

[0057] S43. Similarly, based on automatic machine learning training, it is determined which specific medical insurance catalogue caused the unreasonableness.

[0058] S5. Evaluate the visit data to be analyzed.

[0059] Example 3

[0060] This embodiment uses a medical dataset for model training, the characteristics of which are shown in Table 1 below:

[0061] Table 1

[0062]

[0063]

[0064] The target features obtained are shown in Table 2:

[0065] Table 2

[0066] Medical ID Medical insurance catalog code Violating content 1111 002503040030000 Duplicate items 2222 2203010020000 Exceeding the maximum price 3333 001109000010200 Exceeded the scheduled times

[0067] Among them, the evaluation indicators of binary classification tasks include accuracy, precision, recall, F1-Score, area under the line (AUC), (Receiver operating characteristic curve, ROC) curve, PR curve and confusion matrix; the evaluation indicators of multi-classification tasks include accuracy, precision, recall, F1-Score, area under the line (AUC), (Receiver operating characteristic curve, ROC) curve, PR curve and confusion matrix; the evaluation indicators of regression tasks include MAE (Mean Absolute Error) mean absolute error, MSE (Mean Sequared Error) mean square error, MAPE (Mean Absolute Percentage Error) mean absolute percentage error, and R2 Score (coefficient of determination) regression score function. The calculation method of model evaluation indicators is as follows:

[0068] True Positive Rate (TPR or recall): The proportion of samples that are predicted to be positive among the true positive samples:

[0069]

[0070] Precision: Indicates the proportion of examples classified as positive that are actually positive:

[0071]

[0072] Accuracy: indicates the proportion of samples that are predicted to be normal to the total samples:

[0073]

[0074] TP represents the number of positive samples that are actually predicted to be positive; FN represents the number of negative samples that are actually predicted to be positive; FP represents the number of positive samples that are actually predicted to be negative; TN represents the number of negative samples that are actually predicted to be negative.

[0075] The F1 score is a commonly used metric for measuring the accuracy of a method, often used to judge the accuracy of an algorithm. Currently, in recognition and detection algorithms, precision and recall are often mentioned separately. The F-score considers both values ​​simultaneously, providing a balanced reflection of the accuracy of the algorithm.

[0076]

[0077] The comparison of model evaluation indicators calculated based on the above evaluation indicators is shown in Table 3:

[0078] Table 3

[0079]

[0080] Among them, SVC: Support Vector Classifier, a classification algorithm based on Support Vector Machine (SVM).

[0081] LR: Logistic Regression, a classic classification algorithm, especially suitable for binary classification problems.

[0082] The confusion matrix results of the test set prediction model are shown in Table 4:

[0083] Table 4

[0084]

[0085] The calculated evaluation index data is:

[0086]

[0087] The confusion matrix results of the validation set prediction model are shown in Table 5:

[0088] Table 5

[0089]

[0090] The calculated evaluation index data is:

[0091]

[0092] As can be seen from the above table, the method and model described in the present invention have a high accuracy rate in the diagnosis of the rationality of medical data. They can accurately identify illegal medical records, reduce the workload of manual review and the influence of subjective factors, and provide valuable reference based on the evaluation results, making it easier for personnel to quickly and better understand the medical situation.

[0093] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A method for analyzing the rationality of medical consultation based on automatic machine learning, characterized in that the steps include: S1. Collect data from the medical information form, settlement information form, and expense details form, identify the data characteristics and task types, and pre-process the data; S2. Extracting features related to medical treatment rationality from the preprocessed data to obtain the data source; S3. Use an automated machine learning system to build a model for analyzing the rationality of medical consultations and train the model using a training dataset. S31. Identify unreasonable medical treatment groups from data sources; S311. Using the data source as the first sample data, the first sample data is trained using random forest to obtain feature weight scores and sort them; S312. Based on the sorting of feature weight scores, loop through each feature quantity of the first sample data, obtain the feature quantity with the highest score through recursive feature elimination and cross-validation, and select the selected features as the second sample data; S313. Perform automatic learning and training on the second sample data to obtain an optimal model, and use the optimal model to identify the unreasonable medical treatment group; S32. Based on automated machine learning, screen for violation characteristics and types of unreasonable medical treatment groups; S4. Evaluate the visit data to be analyzed.

2. The method for analyzing the rationality of medical consultation based on automatic machine learning according to claim 1 is characterized in that: The preprocessing includes cleaning, removing duplicate, missing and abnormal data, and filling null value data.

3. The method for analyzing the rationality of medical consultation based on automatic machine learning according to claim 1 is characterized in that: The features related to the rationality of the medical treatment include extracting one or more of the medical treatment ID, gender, age, personnel type, industry, unit type, diagnosis primary code of the medical treatment ID, secondary code, hospitalization code, number of hospitalization days, frequency of medical treatment, medical institution, total cost, unified fund expenditure, total personal payment, and hospital grade.

4. The method for analyzing the rationality of medical consultation based on automatic machine learning according to claim 1 is characterized in that: The cross-validation used 5-fold cross-validation, and the score was the average of 5 validation scores.

5. The method for analyzing the rationality of medical consultation based on automatic machine learning according to claim 1 is characterized in that: Optuna is used to adjust parameters during automatic learning and training.

6. A model for analyzing the rationality of medical consultation based on automatic machine learning, characterized in that: Including automatic identification module, data preprocessing module, feature engineering module, model training and evaluation module; The automatic identification module includes a feature type identification unit and a task type identification unit, which can automatically identify the feature type and task type according to the data conditions of the data source; Data preprocessing module, cleans data and removes duplicate, missing and abnormal data; The feature engineering module constructs features for preprocessed data, calculates feature weights, and performs cross-validation to obtain the best feature selection; The model training and evaluation module includes a model task selection unit, an algorithm selection unit, a model training unit, a hyperparameter optimization unit, a model fusion unit, and a model evaluation unit. The model task selection unit selects the corresponding model task according to the task type, the algorithm selection unit selects the corresponding algorithm according to the model task, the model training unit trains and cross-validates the base learners, the hyperparameter optimization unit automatically tunes the hyperparameters of each base learner, the model fusion unit fuses the base models, and the model evaluation unit evaluates and compares each base model and the fusion model, and outputs the optimal model.

7. The medical consultation rationality analysis model based on automatic machine learning according to claim 6 is characterized in that: Feature types include discrete, integer, floating-point, text, and time; task types include binary classification, multi-classification, and regression.

8. The medical consultation rationality analysis model based on automatic machine learning according to claim 7 is characterized in that: The algorithm selection unit includes random forest, LightGBM, AdaBoost, XGBoost and logistic regression classifiers for binary classification tasks and multi-classification tasks, as well as random forest, LightGBM, and XGBoost classifiers for regression tasks.

9. The medical consultation rationality analysis model based on automatic machine learning according to claim 7 is characterized in that: The construction of discrete features includes first performing One-Hot encoding and then processing it using the OneHotEncoder in the sklean.preprocessing library; the construction of floating-point features includes first normalizing it and then processing it using the StandardScaler in the sklean.preprocessing library; the construction of integer features includes first binning it and processing it using the KBinsDiscretizer in the sklean.preprocessing library; the text features are processed using the TfidfVectorizer in sklearn.feature_extraction.text; the time features extract the year, month, day, hour, week, day of the week, whether it is a weekday, whether it is the beginning of the month, and whether it is the end of the month as new feature columns.

10. The medical consultation rationality analysis model based on automatic machine learning according to claim 6 is characterized in that: The model fusion unit uses the Stacking and Voting fusion strategies to fuse the base models.

Citation Information

Patent Citations

  • Detection method and detection device for medical insurance violation behavior and medical insurance fee control system

    CN109934719A