Medical insurance fraud behavior identification system and method

By combining unsupervised learning and supervised learning technologies to identify and quantify medical insurance fraud, the problems of insufficient data processing capabilities, insufficient model optimization, poor real-time performance and strong dependence on manual intervention in the existing technology are solved, and accurate identification and efficient management of medical behavior risks are achieved.

CN120047252APending Publication Date: 2025-05-27四川吉利学院
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411964517.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing medical insurance fraud detection technology has problems such as insufficient data processing capabilities, insufficient model optimization, poor real-time performance and strong dependence on manual intervention, resulting in inefficiency and high error rate.

Method used

A medical insurance fraud identification system is adopted to collect and integrate medical data, and use a combination of unsupervised learning and supervised learning to extract data features, cluster analysis and risk quantification, establish a risk index system, and improve the accuracy of risk warning through iterative optimization of the model.

Benefits of technology

It realizes the accurate identification and quantification of medical behavior risks, provides a comprehensive risk management solution, reduces the risk losses of medical insurance funds, improves the accuracy of risk warnings, and realizes highly intelligent and automated medical insurance risk control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047252A_ABST
    Figure CN120047252A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, particularly relates to a medical insurance fraud behavior identification system and method, realizes accurate identification and quantification of medical behavior risks, provides a comprehensive risk management solution which cannot be realized in the past, and improves the risk management efficiency by combining unsupervised learning and supervised learning technologies. According to the method, medical data can be monitored and analyzed in real time, potential fraud and abuse behaviors can be quickly identified, so that the risk loss of medical insurance funds is effectively reduced, the accuracy of risk early warning is improved, the workload of manual examination is gradually reduced by continuously optimizing training data and features, and the risk early warning efficiency is improved. Finally, highly intelligent and automatic medical insurance risk control is realized, and the innovative risk management mode significantly improves the operation efficiency and financial stability of medical insurance, and provides more scientific and reliable decision support for related parties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a medical insurance fraud behavior identification system and method. Background Art

[0002] With the continuous development of the medical insurance system, medical insurance fraud behaviors have gradually increased. The existing fraud detection technologies mainly rely on manual review and traditional rule engines. Although they can identify some fraud behaviors to a certain extent, due to the large amount and complexity of data, manual review often faces problems of low efficiency and high error rate. In addition, it is difficult for the rule settings of traditional rule engines to cover all possible fraud situations, resulting in many potential fraud behaviors not being discovered in time.

[0003] In recent years, the rapid development of artificial intelligence technology has provided new solutions for medical fraud detection. Especially the application of machine learning and deep learning technologies has made it possible to analyze and process massive medical data. However, there are still some deficiencies in the existing technical solutions:

[0004] 1. Insufficient data processing ability: Existing technologies often cannot effectively process a large amount of medical data from different sources and formats, resulting in low efficiency of data integration and feature extraction.

[0005] 2. Insufficient model optimization: Although some technical solutions use supervised learning and unsupervised learning models, they lack an effective iterative optimization mechanism, resulting in slow improvement of model performance and difficulty in adapting to constantly changing fraud behavior patterns.

[0006] 3. Poor real-time performance: Existing monitoring systems often cannot achieve real-time risk warning, resulting in the occurrence and losses of fraud behaviors not being controlled in time.

[0007] 4. Strong dependence on manual intervention: Many existing solutions rely on manual review and cannot achieve full automation, restricting the intelligent level of the system. Summary of the Invention

[0008] The purpose of the present invention is to provide a medical insurance fraud behavior identification system and method to solve the problems mentioned in the background art.

[0009] To achieve the above technical purpose, the technical solution adopted by the present invention is as follows:

[0010] A medical insurance fraud behavior identification system, the specific steps are as follows:

[0011] Step 1: Collect and integrate medical data and preprocess the data;

[0012] Step 2: Extract the original data features from the preprocessed data for use as indicators of potential risks;

[0013] Step 3: Use unsupervised learning and adopt the Gaussian mixture model for clustering to identify patterns of similar medical behaviors;

[0014] Step 4: Among the clusters generated by the unsupervised learning model, focus on identifying the clusters with fewer instances, and further analyze and annotate them to determine potential risks;

[0015] Step 5: After establishing the risk index system, use supervised learning to classify and quantify the identified risks, and train the machine learning model using the labeled data from the unsupervised clustering analysis;

[0016] Step 6: Evaluate the model performance;

[0017] Step 7: Based on the comparison between the model evaluation metrics and the results of manual review, determine the accuracy of the model and the review audit, further adjust and optimize the model, and repeatedly iterate this process to continuously improve the accuracy of the model's risk warning.

[0018] The data collected in Step 1 at least includes the patient's demographic information, medical history, medication records, and surgical procedures. The specific preprocessing method is as follows:

[0019] First, handle the missing values in the data and remove the duplicate data, then convert the data into the corresponding format. For the categorical variables, use one-hot encoding to process, separating each category into different label columns. For the quantitative variables with non-normal distribution, use logarithmic transformation. The calculation method is as follows:

[0020]

[0021] where, x i represents various variables in the data, represents the result of logarithmic transformation;

[0022] After completing the logarithmic transformation, perform data normalization processing. The calculation method is as follows:

[0023]

[0024] where, x i represents each original quantitative variable in the data, represents the result after normalization transformation, min(x) is the minimum value of the x variable set, and max(x) is the maximum value of the x variable set.

[0025] The specific method for feature extraction in Step 2 is as follows: After extracting the relevant attributes, use principal component analysis to extract the important information, and calculate the covariance matrix of the standardized features. The method is as follows:

[0026]

[0027] Among them, n represents the number of samples, that is, there are n samples in the data matrix, Z is the standardized form of the data matrix, and C is the covariance matrix of the standardized data matrix;

[0028] Calculate the explained variance for the covariance matrix after completing the analysis. The calculation method is as follows:

[0029] Cv i = λ i v i

[0030]

[0031] Among them, λ i is the i-th eigenvalue of the covariance matrix C, and v i is the corresponding unit eigenvector, and p represents the number of features in the original data.

[0032] The specific method of the third step is as follows: The Gaussian mixture model provides a description of the membership degree of each data point in each cluster. The method is as follows:

[0033]

[0034] Among them, x is the observed data, K is the number of Gaussian distributions (i.e., the number of clusters), π k is the mixing weight of the k-th Gaussian distribution, μ k is the mean vector of the k-th Gaussian distribution, and ∑ k is the covariance matrix of the k-th Gaussian distribution, and N(x|μ k , ∑ k ) represents the k-th Gaussian distribution that x follows;

[0035] Then, use the silhouette coefficient to evaluate the clustering quality. The method is as follows:

[0036]

[0037] Among them, a(i) is the average distance from sample i to other samples in the same cluster, b(i) is the average distance from sample i to the samples in the nearest neighbor cluster, and s(i) is the silhouette coefficient of sample i, and its value range is [-1, 1]:

[0038] When the silhouette coefficient is close to 1, it means that sample i is very close to its belonging cluster and far from other clusters;

[0039] When the silhouette coefficient is close to -1, it means that sample i is far from its belonging cluster and very close to other clusters;

[0040] When the silhouette coefficient is close to 0, it means that sample i is on the clustering boundary;

[0041] The optimal number of clusters K can be determined by maximizing the average silhouette coefficient.

[0042] The specific method of Step 4 is as follows: Combining domain knowledge and expert opinions, define risk thresholds and metrics. Behaviors that exceed the thresholds or deviate significantly from the norm are marked as potential risks and stored in the risk knowledge base for subsequent supervised learning and risk quantification.

[0043] The specific method of Step 5 is as follows: Train the machine learning model LightGBM using the labeled data from unsupervised clustering analysis. The method is as follows:

[0044] F m (x) = F m-1 (x) + γ m h m (x)

[0045] Among them, F m (x) represents the predicted value of the m-th tree, F m-1 (x) represents the cumulative predicted value of the first m - 1 trees, γ m represents the learning rate of the m-th tree, h m (x) represents the prediction function of the m-th tree.

[0046] The specific method of Step 6 is as follows:

[0047] Use recall rate, false positive rate, false negative rate, and the area enclosed by the ROC curve and the coordinate axes to evaluate the model performance. Among them, recall rate, false positive rate, and false negative rate are model evaluation metrics calculated based on the confusion matrix. The calculation methods are as follows:

[0048] Recall rate:

[0049] False positive rate:

[0050] False negative rate:

[0051] Among them, TP represents the number of positive instances correctly predicted as positive by the model, TN represents the number of negative instances correctly predicted as negative by the model, FP represents the number of negative instances wrongly predicted as positive by the model, and FN represents the number of positive instances wrongly predicted as negative by the model;

[0052] The calculation method of the area enclosed by the ROC curve and the coordinate axes is as follows:

[0053]

[0054] Among them, N + represents the number of positive samples, N - represents the number of negative samples, s iDenote the predicted score of the i-th positive sample as s j Denote the predicted score of the j-th negative sample as Π(s i > s j ) and its calculation rule is 1 when s i > s j and 0 otherwise;

[0055] The range of the AUC value is between [0, 1]. The larger the AUC, the better the classification performance of the model. When AUC = 0.5, it means the performance of the model is equivalent to random guessing.

[0056] The content adjusted in the seventh step at least includes optimizing machine annotation, index knowledge base, and business features.

[0057] A method for identifying medical insurance fraud behaviors is as follows:

[0058] First, when a patient sees a doctor or is hospitalized, a large amount of data generated by the patient is collected and stored. Then, based on these data, features are generated using a model formula according to business indicators, and corresponding risk exposure features are also generated. After that, an unsupervised learning model is trained for clustering and further analysis. Then, based on the unsupervised learning model, a supervised learning model is synchronously trained to classify and qualify risks. A comparison is made between the medical insurance supervision department and the classification results of the model, and the accuracy of the model and the review and audit is determined according to the comparison results, so as to further optimize the machine annotation of the model, and the model is trained and optimized during the repeated iteration process.

[0059] The present invention has at least the following advantages compared with the prior art:

[0060] The present invention realizes the accurate identification and quantification of medical behavior risks, provides a comprehensive risk management solution that has not been achieved in the past. By combining the technologies of unsupervised learning and supervised learning, it can monitor and analyze medical data in real time, quickly identify potential fraud and abuse behaviors, and thus effectively reduce the risk losses of medical insurance funds.

[0061] This method not only improves the accuracy of risk warning, but also gradually reduces the workload of manual review by continuously optimizing training data and features, and finally realizes highly intelligent and automated medical insurance risk control. This innovative risk management mode significantly improves the operation efficiency and financial stability of medical insurance, and provides more scientific and reliable decision-making support for relevant parties. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] The present invention can be further illustrated by the non-limiting embodiments given in the drawings.

[0063] Figure 1 It is a schematic diagram of the system flow structure of the present invention.

[0064] Figure 2 Schematic diagram of the development technology stack structure of the present invention. Detailed implementation manners

[0065] In order to enable those skilled in the art to better understand the present invention, the technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0066] As Figure 1-2 shown, a medical insurance fraud behavior identification system, the specific steps are as follows:

[0067] Step 1: Collect and integrate medical data and preprocess the data;

[0068] Step 2: Extract the original data features from the preprocessed data for use as indicators of potential risks;

[0069] Step 3: Use unsupervised learning and adopt the Gaussian mixture model for clustering to identify patterns of similar medical behaviors;

[0070] Step 4: Among the clusters generated by the unsupervised learning model, focus on identifying the clusters with fewer instances, and further analyze and annotate them to determine potential risks;

[0071] Step 5: After establishing the risk index system, use supervised learning to classify and quantify the identified risks, and train the machine learning model with the labeled data from the unsupervised clustering analysis;

[0072] Step 6: Evaluate the model performance;

[0073] Step 7: According to the comparison between the model evaluation indicators and the results of manual review, determine the accuracy of the model and the review and audit, further adjust and optimize the model, and repeatedly iterate this process to continuously improve the accuracy of the model's risk warning.

[0074] The data collected in Step 1 includes at least the patient's demographic information, medical history, medication records, and surgical procedures. The preprocessing method is specifically as follows:

[0075] First, handle the missing values in the data and remove the duplicate data, then convert the data into the corresponding format. For categorical variables, use one-hot encoding to separate each category into different label columns. For quantitative variables with non-normal distributions, use logarithmic transformation. The calculation method is as follows:

[0076]

[0077] where x i represents various variables in the data, represents the result of logarithmic transformation. Through the calculation in the formula, all x iAll will be converted into a form closer to normal distribution, thus eliminating the misleading influence of outliers on the model to a certain extent;

[0078] After completing the logarithmic transformation, the data is normalized and the calculation method is as follows:

[0079]

[0080] Among them, x i represents each original quantitative variable in the data, Represents the result after normalization transformation, min(x) is the minimum value of the x variable set, and max(x) is the maximum value of the x variable set.

[0081] Through the above method, we transformed the non-normally distributed quantitative variables using logarithmic transformation and then normalization. This method makes the data distribution closer to the normal distribution, reduces the misleading impact of outliers on the model, and allows different variables to be compared more accurately on the same scale, thereby ensuring the validity and reliability of the analysis.

[0082] The specific method of feature extraction in step 2 is: after extracting relevant attributes, use principal component analysis (PCA) to extract important information and calculate the covariance matrix of standardized features, the method is as follows:

[0083]

[0084] Where n represents the number of samples, that is, there are n samples in the data matrix, Z is the standardized form of the data matrix, C is the covariance matrix of the standardized data matrix, and the covariance matrix describes the correlation between the features. PCA is used to project the original high-dimensional data into a low-dimensional subspace while retaining the information of the original data as much as possible.

[0085] The explained variance is calculated for the covariance matrix after the analysis is completed. The calculation method is as follows:

[0086] Cv i =λ i v i

[0087]

[0088] Among them, λ i is the i-th eigenvalue of the covariance matrix C, v i is the corresponding unit eigenvector, p represents the number of features in the original data, and the formula uses the eigenvalues ​​to calculate the total variance and the variance contribution rate of each principal component, extracting the first N principal component features with an explanation variation ratio exceeding 0.99 to achieve greater feature transformation and dimensionality reduction, while maximally retaining feature interpretation information and reducing machine computing pressure.

[0089] The specific method of the third step is as follows: The Gaussian Mixture Models (GMM) provides a description of the membership degree of each data point in each cluster, and the method is as follows:

[0090]

[0091] where x is the observed data, K is the number of Gaussian distributions (i.e., the number of clusters), and π k is the mixing weight of the k-th Gaussian distribution, μ k is the mean vector of the k-th Gaussian distribution, ∑ k is the covariance matrix of the k-th Gaussian distribution, and N(x|μ k , ∑ k ) represents the k-th Gaussian distribution that x follows. This formula describes the probability density function of the observed data generated by the linear combination of K Gaussian distributions;

[0092] Then, the silhouette coefficient is used to evaluate the clustering quality, and the method is as follows:

[0093]

[0094] where a(i) is the average distance from sample i to other samples in the same cluster, b(i) is the average distance from sample i to samples in the nearest neighboring cluster, and s(i) is the silhouette coefficient of sample i, with a value range of [-1, 1]:

[0095] When the silhouette coefficient is close to 1, it means that sample i is very close to its belonging cluster and far from other clusters;

[0096] When the silhouette coefficient is close to -1, it means that sample i is far from its belonging cluster and very close to other clusters;

[0097] When the silhouette coefficient is close to 0, it means that sample i is on the clustering boundary;

[0098] By maximizing the average silhouette coefficient, the optimal number of clusters K can be determined.

[0099] The specific method of the fourth step is to define risk thresholds and metrics by combining domain knowledge and expert opinions. Behaviors that exceed the thresholds or deviate significantly from the norm are marked as potential risks and stored in the risk knowledge base for subsequent supervised learning and risk quantification. This method does not rely on predefined rules or labels and can efficiently detect and mark risks through clustering and outlier analysis.

[0100] The specific method of the fifth step is as follows: Use the labeled data from unsupervised clustering analysis to train the machine learning model LightGBM, and the method is as follows:

[0101] F m F(x) = F(x) + γh(x) m-1 + γh(x) m h m (x)

[0102] Among them, F(x) represents the predicted value of the m-th tree, F(x) represents the cumulative predicted value of the previous m - 1 trees, γ represents the learning rate of the m-th tree, and h(x) represents the prediction function of the m-th tree. By iteratively training new trees and correcting the prediction results of the previous tree at a certain learning rate, LightGBM can gradually improve the prediction performance of the model. Use grid search and cross-validation to test different combinations of hyperparameters to determine the optimal settings that can produce the best prediction accuracy. m (x) represents the predicted value of the m-th tree, F(x) represents the cumulative predicted value of the previous m - 1 trees, γ represents the learning rate of the m-th tree, and h(x) represents the prediction function of the m-th tree. By iteratively training new trees and correcting the prediction results of the previous tree at a certain learning rate, LightGBM can gradually improve the prediction performance of the model. Use grid search and cross-validation to test different combinations of hyperparameters to determine the optimal settings that can produce the best prediction accuracy. m-1 (x) represents the cumulative predicted value of the previous m - 1 trees, γ represents the learning rate of the m-th tree, and h(x) represents the prediction function of the m-th tree. By iteratively training new trees and correcting the prediction results of the previous tree at a certain learning rate, LightGBM can gradually improve the prediction performance of the model. Use grid search and cross-validation to test different combinations of hyperparameters to determine the optimal settings that can produce the best prediction accuracy. m represents the learning rate of the m-th tree, and h(x) represents the prediction function of the m-th tree. By iteratively training new trees and correcting the prediction results of the previous tree at a certain learning rate, LightGBM can gradually improve the prediction performance of the model. Use grid search and cross-validation to test different combinations of hyperparameters to determine the optimal settings that can produce the best prediction accuracy. m (x) represents the prediction function of the m-th tree. By iteratively training new trees and correcting the prediction results of the previous tree at a certain learning rate, LightGBM can gradually improve the prediction performance of the model. Use grid search and cross-validation to test different combinations of hyperparameters to determine the optimal settings that can produce the best prediction accuracy.

[0103] The specific method of the sixth step is as follows:

[0104] Use the True Positive Rate (TPR), False Positive Rate (FPR), False Negative Rate (FNR), and the area under the ROC curve (Area Under Curve, AUC) to evaluate the model performance. Among them, the recall rate, false positive rate, and false negative rate are model evaluation metrics calculated based on the confusion matrix. The calculation methods are as follows:

[0105] Recall rate:

[0106] False positive rate:

[0107] False negative rate:

[0108] Among them, TP (True Positive) represents the number of positive instances that the model correctly predicts as positive, TN (True Negative) represents the number of negative instances that the model correctly predicts as negative, FP (False Positive) represents the number of negative instances that the model wrongly predicts as positive, and FN (False Negative) represents the number of positive instances that the model wrongly predicts as negative. Based on these four metrics, the following three model evaluation metrics can be calculated: The recall rate measures the proportion of actual positive instances that the model correctly identifies, the FPR represents the proportion of actual negative instances that the model wrongly predicts as positive, and the FNR represents the proportion of actual positive instances that the model wrongly predicts as negative.

[0109] The calculation method for the area under the ROC curve is as follows:

[0110]

[0111] Among them, N + represents the number of positive samples, and N _ represents the number of negative samples, and s i represents the predicted score of the i-th positive sample, and s j represents the predicted score of the j-th negative sample, and Π(s i > s j ) is calculated as 1 when s i > s j , and 0 otherwise;

[0112] The range of the AUC value is between [0, 1]. The larger the AUC, the better the classification performance of the model. When AUC = 0.5, it means that the performance of the model is equivalent to random guessing;

[0113] AUC is a commonly used indicator to evaluate the performance of binary classification models, comprehensively considering the recall rate and precision of the model at different thresholds.

[0114] The content adjusted in step seven at least includes optimizing machine annotation, the indicator knowledge base, and business features; this technical method integrates unsupervised clustering and supervised learning, providing a comprehensive, data-driven health insurance fund risk management method. By leveraging the pattern recognition ability of unsupervised learning and the prediction ability of supervised models, the proposed framework can effectively identify, quantify, and handle various risks in the operation of medical funds.

[0115] Furthermore, the development of the present invention is implemented using a three-layer architecture, including a data layer, an AI logic layer, and an application layer, specifically as follows:

[0116] (1) Data layer: Based on big data development tools such as Apache Hive, Apache Spark, and Apache Flink, read, write, and manage medical insurance feature data. Use distributed data storage and offline and real-time processing technologies to ensure the stability and reliability of the data processing infrastructure.

[0117] (2) AI logic layer: Utilize machine learning algorithms in the AI logic layer, with the help of frameworks such as Python, Pandas, Scikit-learn, and Keras. For unsupervised learning models, algorithms such as K-means and Gaussian mixture models are used, and for supervised learning models, algorithms such as decision trees, random forests, and AdaBoost are used to build a fraud identification model to accurately identify potential fraud behaviors.

[0118] (3) Application layer: In the application layer, Vue.js is used as the main web application development tool. Vue.js has the ability to build user interfaces using HTML, CSS, and JavaScript, and uses a declarative component programming model to meet the needs of user interface development and business logic, and realize the presentation and dynamic response of business scenarios. In terms of data display, data visualization and operation are realized in the form of visual dashboards, allowing operators to intuitively understand and operate data.

[0119] There is also an embodiment of the present invention, which is specifically:

[0120] (1) During a patient's visit or hospitalization, a large amount of data will be generated, including the patient's basic information, the name, dosage, and frequency of medication, the type and level of surgery, etc. This data will be stored in the data warehouse.

[0121] (2) Based on the data warehouse, features are generated according to business indicators, such as the restricted mapping relationship between disease medication and surgery, the restricted mapping relationship between patient information and medication, etc. At the same time, risk exposure features are generated according to the threshold indicator knowledge base set by the medical insurance department. For indicators outside the knowledge base, risk exposure features based on average values ​​and outliers are generated. After completing ETL processing, feature vectors are generated.

[0122] (3) Train the unsupervised learning model, complete the clustering of all medical behaviors, conduct further qualitative analysis on each cluster, pay special attention to the clusters with fewer numbers, mark their illegal and irregular risks, optimize the feature labels, and create conditions for the training of the supervised learning model.

[0123] (4) Based on the optimization of feature labels by the unsupervised learning model, the supervised learning model is trained simultaneously to classify and characterize the risks of illegal and irregular activities. Doctors will receive risk warnings in real time and report them to the medical insurance regulatory department for review and audit.

[0124] (5) Compare the model classification results with the review and audit results of the medical insurance regulatory department, determine the accuracy of the model and the review and audit based on the positive and negative sample ratios and the confusion matrix, and then gradually adjust and optimize the machine labeling, optimize the indicator knowledge base and business characteristics, and improve the quality of training data.

[0125] (6) Repeat the above process to continuously improve the accuracy of the model risk warning, and combine it with the review and audit of the medical insurance regulatory department to continuously optimize the quality and characteristics of the training data, and repeat this process. Over time, the workload of manual review and audit will gradually decrease until a highly intelligent and automated medical insurance risk control is achieved.

[0126] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A medical insurance fraud identification system, characterized by: The specific steps are as follows: Step 1: Collect and integrate medical data and pre-process the data; Step 2: Extract raw data features from the preprocessed data to serve as an indicator of potential risks; Step 3: Using unsupervised learning, clustering was performed using a Gaussian mixture model to identify patterns of similar medical behaviors; Step 4: Among the clusters generated by the unsupervised learning model, focus on identifying clusters with fewer instances, and further analyze and annotate them to determine potential risks; Step 5: After establishing the risk indicator system, supervised learning is used to classify and quantify the identified risks, and the machine learning model is trained using labeled data from unsupervised cluster analysis; Step 6: Evaluate model performance; Step 7: Determine the accuracy of the model and audit by comparing the model evaluation indicators and manual review results, further adjust and optimize the model, iterate the process repeatedly, and continuously improve the accuracy of the model risk warning.

2. A medical insurance fraud identification system according to claim 1, characterized in that: The data collected in step 1 at least includes the patient's demographic information, medical history, medication records and surgical procedures, and the preprocessing method is specifically as follows: First, the data is processed for missing values ​​and duplicate data is removed, and then the data is converted to the corresponding format. Categorical variables are processed using one-hot encoding to separate each category into different label columns. For quantitative variables with non-normal distribution, logarithmic transformation is used. The calculation method is as follows: Among them, x i Represents various variables in the data, represents the result of logarithmic transformation; After completing the logarithmic transformation, the data is normalized and the calculation method is as follows: Among them, x i represents each original quantitative variable in the data, Represents the result after normalization transformation, min(x) is the minimum value of the x variable set, and max(x) is the maximum value of the x variable set.

3. A medical insurance fraud identification system according to claim 2, characterized in that: The specific method of feature extraction in step 2 is: after extracting relevant attributes, use principal component analysis to extract important information and calculate the covariance matrix of standardized features. The method is as follows: Where n represents the number of samples, that is, there are n samples in the data matrix, Z is the standardized form of the data matrix, and C is the covariance matrix of the standardized data matrix; The explained variance is calculated for the covariance matrix after the analysis is completed. The calculation method is as follows: Cv i =λ i v i Among them, λ i is the i-th eigenvalue of the covariance matrix C, v i is the corresponding unit eigenvector, and p represents the number of features in the original data.

4. A medical insurance fraud identification system according to claim 3, characterized in that: The specific method of step three is: the Gaussian mixture model provides a description of the degree of belonging of each data point in each cluster, and the method is as follows: Where x is the observed data, K is the number of Gaussian distributions (i.e. the number of clusters), and π k is the mixture weight of the kth Gaussian distribution, μ k is the mean vector of the kth Gaussian distribution, ∑ k is the covariance matrix of the kth Gaussian distribution, N(x|μ k ,∑ k ) indicates that it obeys the k-th Gaussian distribution in x; Then, the silhouette coefficient is used to assess the clustering quality as follows: Among them, a(i) is the average distance from sample i to other samples in the same cluster, b(i) is the average distance from sample i to the samples in the nearest neighbor cluster, and s(i) is the silhouette coefficient of sample i, which ranges from [-1,1]: When the silhouette coefficient is close to 1, it means that sample i is very close to the cluster to which it belongs and far away from other clusters; When the silhouette coefficient is close to -1, it means that sample i is far away from the cluster to which it belongs and is very close to other clusters; When the silhouette coefficient is close to 0, it means that sample i is located on the cluster boundary; By maximizing the average silhouette coefficient, the optimal number of clusters K can be determined.

5. A medical insurance fraud identification system according to claim 4, characterized in that: The specific method of step four is to define risk thresholds and indicators by combining domain knowledge and expert opinions. Behaviors that exceed the thresholds or deviate significantly from the norm are marked as potential risks and stored in the risk knowledge base for subsequent supervised learning and risk quantification.

6. A medical insurance fraud identification system according to claim 5, characterized in that: The specific method of step 5 is: use the labeled data from the unsupervised cluster analysis to train the machine learning model LightGBM, the method is as follows: F m (x)=F m-1 (x)+γ m h m (x) Among them, F m (x) represents the predicted value of the mth tree, F m-1 (x) represents the cumulative prediction value of the first m-1 trees, γ m represents the learning rate of the mth tree, h m (x) represents the prediction function of the mth tree.

7. A medical insurance fraud identification system according to claim 6, characterized in that: The specific method of step six is: The recall rate, false positive rate, false negative rate and the area under the ROC curve and the coordinate axis are used to evaluate the model performance. The recall rate, false positive rate and false negative rate are model evaluation indicators calculated based on the confusion matrix. The calculation method is as follows: Recall: False Positive Rate: False Negative Rate: Where TP represents the number of positive instances correctly predicted as positive by the model, TN represents the number of negative instances correctly predicted as negative by the model, FP represents the number of negative instances incorrectly predicted as positive by the model, and FN represents the number of positive instances incorrectly predicted as negative by the model; The calculation method of the area under the ROC curve and the coordinate axis is: Among them, N + Represents the number of positive samples, N - represents the number of negative samples, s i represents the prediction score of the i-th positive sample, s j represents the prediction score of the jth negative sample, ∏(s i >s j ) is calculated as follows: i >s j 1 when it is, otherwise 0; The AUC value range is between [0,1]. The larger the AUC is, the better the classification performance of the model is. When AUC = 0.5, it means that the performance of the model is equivalent to random guessing.

8. A medical insurance fraud identification system according to claim 7, characterized in that: The content adjusted in step seven at least includes optimizing machine annotation, indicator knowledge base and business characteristics.

9. A method for identifying medical insurance fraud, characterized by: The identification is performed using the identification system described in any one of claims 1 to 8, and the specific method is as follows: First, when patients visit the hospital or are hospitalized, a large amount of data is collected and stored. Then, based on this data, model formulas are used to generate features according to business indicators, and corresponding risk exposure features are generated. Thereafter, an unsupervised learning model is trained to perform clustering and further analysis. Next, based on the unsupervised learning model, a supervised learning model is trained simultaneously to classify and quantify risks. A comparison is made between the medical insurance regulatory department and the model classification results. The accuracy of the model and review audit is determined based on the comparison results, thereby further optimizing the model machine labeling and training and optimizing the model in a repeated iterative process.

Citation Information

Cited By

  • Fraud-related risk prevention and control method and system

    CN122155757A