Outpatient data anomaly identification method and device, electronic equipment and storage medium

By acquiring secondary diagnostic data from outpatient visits and utilizing a pre-trained secondary diagnostic impact prediction model and cost thresholds for outpatient visit groups, abnormal outpatient data can be automatically identified. This solves the problems of high subjectivity and low efficiency in the review of outpatient cost abuse in existing technologies, achieving higher identification accuracy and efficiency.

CN115472250BActive Publication Date: 2026-01-16PING AN HEALTH INSURANCE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211161892.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2026-01-16
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

In existing technologies, the review of outpatient expense abuse mainly relies on manual judgment, which has the problems of strong subjectivity, large workload and low efficiency. Moreover, the threshold setting for outpatient expense abuse judgment in existing claims risk control technologies is relatively crude and fails to effectively consider the combined impact of secondary diagnoses and other indicators, resulting in insufficient identification accuracy.

Method used

By acquiring secondary diagnostic data from outpatient visits, and utilizing a pre-trained secondary diagnostic impact prediction model, combined with the cost thresholds for multiple outpatient visit groups, the system automatically identifies whether outpatient data is abnormal. This includes constructing a secondary diagnostic impact prediction model, matching outpatient visit groups, and calculating cost thresholds.

Benefits of technology

It improved the accuracy and efficiency of outpatient data anomaly identification, reduced reliance on manual review, and increased the automation level of outpatient expense abuse review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115472250B_ABST
    Figure CN115472250B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an outpatient data anomaly identification method and device, electronic equipment and storage medium, comprising: obtaining outpatient treatment data to be identified, and obtaining corresponding secondary diagnosis data; inputting the secondary diagnosis data into a secondary diagnosis influence prediction model, and outputting a secondary diagnosis influence degree prediction value; obtaining a target outpatient treatment grouping of the outpatient treatment data according to at least one treatment index and the secondary diagnosis influence degree prediction value; determining whether the outpatient treatment data is abnormal data according to the cost data and the treatment cost threshold of the target outpatient treatment grouping; in the above manner, the secondary diagnosis influence degree is determined according to the outpatient treatment data, the secondary diagnosis influence degree is also taken as a treatment index, and then the target outpatient treatment grouping is matched according to each treatment index of the outpatient treatment data, and whether it is abnormal is judged according to the treatment cost threshold of the target outpatient treatment grouping, thereby improving the accuracy and efficiency of outpatient data anomaly identification.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of big data, and in particular to a clinic data anomaly identification method and device, an electronic device and a storage medium. BACKGROUND

[0002] With the help of the commercialization of medical insurance by the state and the continuous improvement of the public's insurance awareness, the number of customers purchasing health insurance is increasing day by day, and the number of claims is also growing. In the process of health insurance, the risk of claim settlement needs to be audited, which includes the audit of whether there is a fee abuse behavior in the clinic visit. At present, the audit of fee abuse in clinic visit mainly relies on the experience of auditors for manual judgment, which requires high ability of auditors. At the same time, it is necessary to consult case files, bills, image materials and other materials in the audit process, which takes a long time. Using manpower to audit cases one by one has the problems of strong subjectivity, large workload and low efficiency. The claim settlement risk control technology in the prior art generally sets a unified threshold for judging the fee abuse of clinic according to a single factor of disease, and pushes the abnormal cases exceeding the threshold to auditors for further audit. Although this method can reduce the audit manpower, the threshold for judging the fee abuse is relatively rough, and the comprehensive influence of secondary diagnosis and other indicators is not considered, which is not conducive to improving the accuracy of clinic data anomaly identification. SUMMARY

[0003] In view of the above problems, the embodiments of the present application provide a clinic data anomaly identification method and device, an electronic device and a storage medium to solve the above technical problems.

[0004] In a first aspect, the embodiments of the present application provide a clinic data anomaly identification method, comprising:

[0005] Obtaining clinic visit data to be identified, and obtaining corresponding secondary diagnosis data according to the clinic visit data, wherein the clinic visit data comprises at least one visit index and at least one fee data;

[0006] Inputting the secondary diagnosis data into a pre-trained secondary diagnosis influence prediction model, and outputting a secondary diagnosis influence degree prediction value corresponding to the secondary diagnosis data, wherein the secondary diagnosis influence prediction model is obtained by training historical secondary diagnosis data and historical secondary diagnosis influence degree calculation values;

[0007] Obtaining a target clinic visit group of the clinic visit data from a plurality of preset clinic visit groups according to the at least one visit index and the secondary diagnosis influence degree prediction value, wherein each clinic visit group comprises a plurality of standard visit indexes and at least one visit fee threshold;

[0008] Determine whether the outpatient visit data is abnormal data according to the cost data and the visit cost threshold of the target outpatient visit grouping.

[0009] Optionally, the acquiring corresponding secondary diagnosis data according to the outpatient visit data comprises:

[0010] Acquiring the number of secondary diagnosis diseases according to the outpatient visit data;

[0011] Acquiring the historical visit cost median of each secondary diagnosis disease in the outpatient visit data, and calculating the average value of the historical visit cost median;

[0012] Determining the secondary diagnosis disease with the highest disease level in the outpatient visit data;

[0013] Determining the number of secondary diagnosis diseases with disease level greater than a preset level in the outpatient visit data;

[0014] Acquiring the disease level difference between the secondary diagnosis disease with the highest disease level and the primary diagnosis disease in the outpatient visit data;

[0015] Constructing the secondary diagnosis data according to the number of secondary diagnosis diseases, the average value, the secondary diagnosis disease with the highest disease level, the number of secondary diagnosis diseases with disease level greater than a preset level, and the disease level difference.

[0016] Optionally, the training step of the secondary diagnosis influence prediction model comprises:

[0017] Acquiring at least one training sample, wherein the training sample comprises historical secondary diagnosis data, and the historical secondary diagnosis data is acquired according to historical outpatient visit data;

[0018] Acquiring a plurality of target historical outpatient visit data according to the primary diagnosis disease in the historical outpatient visit data, wherein the diagnosis disease in the target historical outpatient visit data is the primary diagnosis disease, and the target historical outpatient visit data does not contain secondary diagnosis disease;

[0019] Determining a basic visit cost according to the median of visit cost in the plurality of target historical outpatient visit data, and acquiring the historical secondary diagnosis influence degree calculation value according to the visit cost in the historical outpatient visit data and the basic visit cost;

[0020] Inputting the historical secondary diagnosis data into the secondary diagnosis influence prediction model, and outputting the historical secondary diagnosis influence degree prediction value;

[0021] A prediction error is calculated according to the historical secondary diagnosis influence degree prediction value and the historical secondary diagnosis influence degree calculation value, and parameters of the secondary diagnosis influence prediction model are adjusted according to the prediction error until the secondary diagnosis influence prediction model reaches a training convergence condition.

[0022] Optionally, the target outpatient visit grouping of the outpatient visit data is obtained from a plurality of preset outpatient visit groupings according to the at least one visit index and the secondary diagnosis influence degree prediction value, and the method comprises the following steps of:

[0023] A secondary diagnosis influence index of the outpatient visit data is determined according to the secondary diagnosis influence degree prediction value.

[0024] The at least one visit index and the secondary diagnosis influence index are matched with the standard visit index of the outpatient visit grouping, and a matched outpatient visit grouping is obtained.

[0025] The matched outpatient visit grouping with the largest number of standard visit indexes is taken as the target outpatient visit grouping.

[0026] Optionally, the step of calculating the visit expense threshold of the outpatient visit grouping comprises the following steps of:

[0027] A plurality of historical outpatient visit data matched with the outpatient visit grouping are obtained.

[0028] Target expense data matched with the type of the visit expense threshold in the historical outpatient visit data is obtained, and a corresponding target expense data set is constructed according to the target expense data.

[0029] The target expense data in the target expense data set is arranged in ascending order, a first preset quantile and a second preset quantile are obtained according to the arrangement result, and the first preset quantile is greater than the second preset quantile.

[0030] The visit expense threshold is calculated according to the first preset quantile, the second preset quantile and a preset coefficient, and the preset coefficient is greater than 1.

[0031] Optionally, after the visit expense threshold is calculated according to the first preset quantile, the second preset quantile and the preset coefficient, the method further comprises the following steps of:

[0032] A coefficient of variation of each target expense data in the target expense data set is obtained.

[0033] A median of each target expense data in the target expense data set is obtained, and a ratio of the visit expense threshold to the median is obtained.

[0034] According to the ratio, the sample quantity of the target fee data set, and the coefficient of variation, an accuracy evaluation value of the threshold of the visit fee is obtained.

[0035] Optionally, the determining whether the outpatient visit data is abnormal data according to the fee data and the threshold of the visit fee of the target outpatient visit group comprises:

[0036] If the fee data is greater than the threshold of the visit fee, it is determined that the outpatient visit data is abnormal data.

[0037] If the fee data is less than or equal to the threshold of the visit fee, it is determined that the outpatient visit data passes the audit.

[0038] In a second aspect, the embodiments of the present application further provide an outpatient data abnormality identification device, comprising:

[0039] a data acquisition module configured to acquire outpatient visit data to be identified, and acquire corresponding secondary diagnosis data according to the outpatient visit data, wherein the outpatient visit data comprises at least one visit index and at least one fee data;

[0040] a secondary diagnosis prediction module configured to input the secondary diagnosis data into a pre-trained secondary diagnosis influence prediction model, and output a secondary diagnosis influence degree prediction value corresponding to the secondary diagnosis data, wherein the secondary diagnosis influence prediction model is obtained by training historical secondary diagnosis data and historical secondary diagnosis influence degree calculation values;

[0041] a grouping matching module configured to acquire a target outpatient visit group of the outpatient visit data from a plurality of preset outpatient visit groups according to the at least one visit index and the secondary diagnosis influence degree prediction value, wherein each of the outpatient visit groups comprises a plurality of standard visit indexes and at least one threshold of visit fee;

[0042] an abnormality identification module configured to determine whether the outpatient visit data is abnormal data according to the fee data and the threshold of the visit fee of the target outpatient visit group.

[0043] In a third aspect, the embodiments of the present application further provide an electronic device, comprising a processor and a memory coupled to the processor, wherein the memory stores program instructions executable by the processor; and the processor executes the program instructions stored in the memory to implement the above-mentioned outpatient data abnormality identification method.

[0044] In a third aspect, the embodiments of the present application further provide a storage medium, wherein the storage medium stores program instructions, and the program instructions are executed by a processor to implement the above-mentioned outpatient data abnormality identification method.

[0045] The outpatient data anomaly identification method and device, the electronic device and the storage medium provided by the embodiments of the present application include: obtaining outpatient treatment data to be identified, obtaining corresponding secondary diagnosis data according to the outpatient treatment data, wherein the outpatient treatment data includes at least one treatment index and at least one cost data; inputting the secondary diagnosis data into a pre-trained secondary diagnosis influence prediction model, and outputting a secondary diagnosis influence degree prediction value corresponding to the secondary diagnosis data, wherein the secondary diagnosis influence prediction model is obtained by training according to historical secondary diagnosis data and historical secondary diagnosis influence degree calculation values; obtaining a target outpatient treatment group of the outpatient treatment data from a plurality of preset outpatient treatment groups according to the at least one treatment index and the secondary diagnosis influence degree prediction value, wherein each of the outpatient treatment groups includes a plurality of standard treatment indexes and at least one treatment cost threshold; determining whether the outpatient treatment data is abnormal data according to the cost data and the treatment cost threshold of the target outpatient treatment group; and through the above-mentioned manner, the secondary diagnosis influence degree is determined according to the outpatient treatment data, the secondary diagnosis influence degree is also used as a treatment index, and then the target outpatient treatment group is matched according to each treatment index of the outpatient treatment data, and it is determined whether it is abnormal according to the treatment cost threshold of the target outpatient treatment group, thereby improving the accuracy and efficiency of the outpatient data anomaly identification.

[0046] These aspects or other aspects of the present application will be more apparent in the following description of the embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0048] Figure 1 A flowchart of an outpatient data anomaly identification method provided by an embodiment of the present application is shown.

[0049] Figure 2 A structure diagram of an outpatient data anomaly identification device provided by an embodiment of the present application is shown.

[0050] Figure 3 A structure diagram of an electronic device provided by an embodiment of the present application is shown.

[0051] Figure 4 A structure diagram of a storage medium provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0052] The embodiments of the present application will be described in detail below with reference to the drawings, in which the same or similar components have the same reference numerals, and wherein:

[0053] In order to make the persons skilled in the art better understand the schemes of the present application, the technical schemes in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all the other embodiments obtained by the persons skilled in the art without creative labor fall within the protection scope of the present application.

[0054] In the embodiments of the present application, at least one refers to one or more; multiple refers to two or more than two. In the description of the present application, the terms "first", "second", "third" and the like are only used for distinguishing the purposes of description, and cannot be understood as indicating or implying relative importance, nor indicating or implying sequence.

[0055] In the description of the present application, the reference "one embodiment" or "some embodiments" and the like means that the specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Therefore, in the description of the present application, the terms "include", "contain", "have" and their variants all mean "include but not limited to", unless otherwise specifically emphasized.

[0056] It should be noted that in the embodiments of the present application, the association relationship of the associated objects described by "and / or" can represent three kinds of relationships, for example, A and / or B can represent three cases of A alone, A and B together, and B alone.

[0057] It should be noted that in the embodiments of the present application, "connection" can be understood as electrical connection, and the connection between two electrical elements can be direct or indirect connection between two electrical elements. For example, A and B are connected, which can be direct connection between A and B, or indirect connection between A and B through one or more other electrical elements.

[0058] An embodiment of the present application provides a clinic data anomaly identification method. An execution subject of the clinic data anomaly identification method includes, but is not limited to, at least one of an electronic device, such as a server, a terminal, or the like, which can be configured to execute the clinic data anomaly identification method provided by the embodiment of the present application. In other words, the clinic data anomaly identification method can be executed by software or hardware installed in a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster, and the like.

[0059] Referring to FIG. 1, Figure 1 FIG. 1 is a flowchart of a clinic data anomaly identification method provided by an embodiment of the present application. It should be noted that the method of the present application is not limited to the order of the flowchart shown in FIG. 1, as long as the same results are achieved. Figure 1 In this embodiment, the clinic data anomaly identification method includes the following steps:

[0060] S10, obtaining clinic data to be identified, and obtaining corresponding secondary diagnosis data according to the clinic data, wherein the clinic data includes at least one clinic index and at least one cost data;

[0061] In this embodiment, the clinic data includes at least one clinic index and at least one cost data.

[0062] As an implementation manner, the clinic index included in the clinic data can include, but is not limited to, the following: a primary diagnosis disease, at least one secondary diagnosis disease, a hospital property, a hospital level, a development degree of a clinic area, and an age of a clinic person. The primary diagnosis disease is a clinic disease of a clinic person, and the secondary diagnosis disease is another incidental disease found in addition to the primary diagnosis disease and related to the primary diagnosis disease, for example, a clinic person A goes to a clinic due to general edema and fever, the primary diagnosis disease is kidney abscess, and it is found that the blood sugar of the clinic person A is high, which will affect the treatment of kidney abscess if the blood sugar is not controlled well. Therefore, the treatment of diabetes is related to the treatment of kidney abscess, and the secondary diagnosis disease is diabetes. For another example, a clinic person B goes to a clinic due to lung pain, the primary diagnosis disease is pneumonia, and it is found that the clinic person B has a cold, which will affect the body state of the clinic person B and further affect the recovery speed of pneumonia. Therefore, the treatment of the cold is related to the treatment of pneumonia, and the secondary diagnosis disease is the cold.

[0063] In some embodiments, the primary diagnosis disease can be a disease ICD (International Classification of Diseases) top three code, and the secondary diagnosis disease can also be a disease ICD top three code.

[0064] In some embodiments, the hospital nature can be: high-end hospital, private hospital or public hospital, wherein the high-end hospital can refer to expensive hospital designated by insurance company, international department, special department, etc.; the hospital level can be first-class, second-class, third-class, ungraded, foreign capital, private, other, wherein foreign capital and private generally refer to private hospital; based on the developed degree of the area where the hospital is located, the treatment area is labeled with four levels reflecting different levels of development, for example, the developed degree of the treatment area can be first-line, second-line, third-line, or fourth-line and below; at the same time, the treatment person will also be labeled with age range tags of [0, 5], (5, 15], (15, 25], (25, 45], (45, 60], (60, 70], (>70) according to age, for example, the age of the treatment person can be [0, 5], (5, 15], (15, 25], (25, 45], (45, 60], (60, 70], or (>70).

[0065] As an embodiment, the cost data included in the outpatient treatment data can include but is not limited to: total treatment cost, examination cost, drug cost, treatment cost, and other costs. The total treatment cost is the sum of the examination cost, the drug cost, the treatment cost, and the other costs. In this embodiment, the total treatment cost and each single cost can be respectively identified for abnormalities, and it is respectively judged whether the total treatment cost, the examination cost, the drug cost, the treatment cost, and the other costs are abnormal.

[0066] As an embodiment, the main diagnostic disease, the hospital nature, the hospital level, the developed degree of the treatment area, and the age of the treatment person are directly used as a treatment index respectively; the secondary diagnostic disease cannot be directly used as a treatment index, and needs to be predicted through the subsequent step S20, and then the treatment index related to the secondary diagnostic disease is determined according to the prediction result.

[0067] As an embodiment, in step S10, the corresponding secondary diagnostic data is obtained according to the outpatient treatment data, including the following steps:

[0068] S11, the number of secondary diagnostic diseases is obtained according to the outpatient treatment data;

[0069] S12, the median of the historical treatment cost of each secondary diagnostic disease in the outpatient treatment data is obtained, and the average of the historical treatment cost medians is calculated;

[0070] The median of the historical visit cost of each secondary diagnosis disease can be obtained in the following manner: obtaining a plurality of target historical outpatient visit data of the current secondary diagnosis disease as the primary diagnosis disease and not containing the secondary diagnosis disease; constructing a target historical visit cost data set according to the cost data in the obtained target historical outpatient visit data; sorting all cost data in the target historical visit cost data set in ascending order, and determining the historical visit cost median of the current herbal diagnosis disease according to the sorting result, wherein the cost data can be one or more of the total visit cost, the examination cost, the drug cost, the treatment cost or other costs.

[0071] When the number of secondary diagnosis diseases is multiple, the arithmetic mean of the historical visit cost medians of the multiple secondary diagnosis diseases is calculated as the average value; when the number of secondary diagnosis diseases is one, the historical visit cost median of the only secondary diagnosis disease is the average value.

[0072] S13, determining the secondary diagnosis disease with the highest disease level in the outpatient visit data;

[0073] As an implementation manner, the diseases can be divided into 20 levels according to the severity of different diseases, wherein the higher the disease level, the higher the severity of the disease, and the higher the visit cost of the disease. The disease level of each secondary diagnosis disease in the outpatient visit data is determined respectively, and the secondary diagnosis disease with the highest disease level is determined.

[0074] S14, determining the number of secondary diagnosis diseases with a disease level greater than a preset level in the outpatient visit data;

[0075] Wherein, the disease level of each secondary diagnosis disease in the outpatient visit data is determined respectively, and the secondary diagnosis disease with a disease level greater than a preset level is determined. The preset level can be set as needed, for example, the preset level can be 10 levels. The more the number of secondary diagnosis diseases greater than the preset level, the higher the visit cost.

[0076] S15, obtaining the disease level difference between the secondary diagnosis disease with the highest disease level and the primary diagnosis disease in the outpatient visit data;

[0077] Wherein, the greater the disease level difference between the secondary diagnosis disease with the highest disease level and the primary diagnosis disease, the greater the influence of the secondary diagnosis disease on the visit cost.

[0078] S16, constructing the secondary diagnosis data according to the number of secondary diagnosis diseases, the average value, the secondary diagnosis disease with the highest disease level, the number of secondary diagnosis diseases with a disease level greater than a preset level, and the disease level difference;

[0079] As an implementation manner, each of the number of the secondary diagnosis diseases, the average value, the secondary diagnosis disease with the highest disease grade, the number of the secondary diagnosis diseases with the disease grade greater than a preset grade, and the disease grade difference can be one-hot coded respectively to obtain corresponding one-hot coded data, and then the one-hot coded data is input into an Embedding layer to output Embedding coded data represented by Embedding, and then each Embedding coded data is constructed into secondary diagnosis data as an input of a secondary diagnosis influence prediction model in a subsequent step.

[0080] In some embodiments, the secondary diagnosis data can further include hospital property, hospital grade, development level of a treatment area, and age of a patient.

[0081] S20, inputting the secondary diagnosis data into a pre-trained secondary diagnosis influence prediction model to output a secondary diagnosis influence degree prediction value corresponding to the secondary diagnosis data, wherein the secondary diagnosis influence prediction model is trained according to historical secondary diagnosis data and historical secondary diagnosis influence degree calculation values.

[0082] The secondary diagnosis influence degree prediction value is used to represent a ratio of a sum of treatment fees of the primary diagnosis disease and the secondary diagnosis disease to a treatment fee of the primary diagnosis disease alone, and the higher the secondary diagnosis influence degree prediction value is, the greater the influence of the secondary diagnosis disease on the treatment fee is.

[0083] As an implementation manner, the training step of the secondary diagnosis influence prediction model includes:

[0084] S21, obtaining at least one training sample, wherein the training sample includes historical secondary diagnosis data, and the historical secondary diagnosis data is obtained according to historical outpatient treatment data.

[0085] The historical outpatient treatment data is similar to the outpatient treatment data in step S10, and can include but is not limited to the following: a primary diagnosis disease, at least one secondary diagnosis disease, hospital property, hospital grade, development level of a treatment area, and age of a patient. The historical outpatient treatment data includes both the primary diagnosis disease and the secondary diagnosis disease. The historical secondary diagnosis data is similar to the secondary diagnosis data described above, and the structure and obtaining manner of the historical secondary diagnosis data are described in steps S11 to S16 about the secondary diagnosis data, which will not be described here.

[0086] S22, obtaining a plurality of target historical outpatient visit data according to the main diagnosis disease in the historical outpatient visit data, wherein the diagnosis disease in the target historical outpatient visit data is the main diagnosis disease and the target historical outpatient visit data does not include the secondary diagnosis disease;

[0087] The main diagnosis disease in the target historical outpatient visit data is the same as the main diagnosis disease in the historical outpatient visit data, and the target historical outpatient visit data only includes the main diagnosis disease and does not include any secondary diagnosis disease.

[0088] As an implementation form, in the process of processing the expense data in the historical outpatient visit data, when the time span of the historical data is large, in order to increase the reliability of the historical expense data, the various expense data is multiplied by the corresponding medical inflation coefficient according to the visit time to eliminate the medical inflation, wherein the medical inflation coefficient is obtained according to the smooth trend analysis of the outpatient expense under the conditions of the same disease, the same region, the same hospital property, the same hospital level and the change of time.

[0089] S23, determining a basic visit expense according to the median of the visit expense in the plurality of target historical outpatient visit data, and obtaining the historical secondary diagnosis influence degree calculation value according to the visit expense in the historical outpatient visit data and the basic visit expense;

[0090] The visit expense in the target historical outpatient visit data is sorted in ascending order, and the median of the visit expense is obtained according to the sorting result, which is used as the basic visit expense. For one historical outpatient visit data, the historical secondary diagnosis influence degree calculation value is the value obtained by dividing the visit expense by the basic visit expense. The visit expense in the historical outpatient visit data includes the expense of the main diagnosis disease and the expense of the secondary diagnosis disease.

[0091] As an optional implementation form, the visit expense can be the total visit expense, the examination expense, the drug expense, the treatment expense or other expenses.

[0092] S24, inputting the historical secondary diagnosis data into the secondary diagnosis influence prediction model to output the historical secondary diagnosis influence degree prediction value;

[0093] S25, calculating a prediction error according to the historical secondary diagnosis influence degree prediction value and the historical secondary diagnosis influence degree calculation value, adjusting the parameters of the secondary diagnosis influence prediction model according to the prediction error until the secondary diagnosis influence prediction model reaches a training convergence condition;

[0094] As an implementation form, the secondary diagnosis influence prediction model can be an XGBOOST regression model, which belongs to the category of Gradient Boosting Decision Tree (GBDT) model. The basic idea of GBDT is to let a new base model fit the bias of the previous stage base model, so as to continuously reduce the bias of the additive model. The base model is a CART (Classification and Regression Tree) model, which can only control the problem of overfitting by controlling the number of iterations and the depth of the tree. Compared with the classic GBDT, XGBOOST has made some improvements, so that the effect and performance are obviously improved. The objective function of the XGBOOST regression model is:

[0095]

[0096] wherein, is a loss function used to measure the difference between the predicted value and the true value y i (the calculated value obtained in step S23 as the true value), x i is the i-th sample in the training sample, f t is a function in the function space F, and Ω(f i ) is a penalty term of model complexity. The XGBOOST regression model used in the embodiment adds a regular term Ω(f i ) reflecting the complexity in the objective function. Then, the above objective function is Taylor expanded to obtain the expansion of the objective function: XGBOOST adds a regular term in the objective function to control the complexity of the model, which can effectively prevent overfitting. At the same time, by introducing the second-order derivative through the second-order Taylor expansion of the objective function, not only the accuracy is increased, but also the objective function can be customized, so that XGBOOST is suitable for both classification and regression problems.

[0097] S30, obtaining the target outpatient service data grouping of the outpatient service data from a plurality of preset outpatient service data groupings according to the at least one medical visit indicator and the secondary diagnosis influence degree prediction value, wherein each of the outpatient service data groupings comprises a plurality of standard medical visit indicators and at least one medical visit cost threshold;

[0098] As an implementation form, step S30 further comprises the following steps:

[0099] S31, determining a secondary diagnosis influence indicator of the outpatient service data according to the secondary diagnosis influence degree prediction value;

[0100] The secondary diagnosis influence indicator can be a low secondary diagnosis influence degree, a medium secondary diagnosis influence degree, or a high secondary diagnosis influence degree. In some embodiments, the trained secondary diagnosis influence prediction model is applied to predict the influence degree of new secondary diagnosis data on the overall cost in production, and the influence degree prediction value is divided into three levels of low, medium, and high, such as (<1.3), [1.3, 1.8), (>=1.8), and the like. The obtained secondary diagnosis influence degree level label is used as one of the important indicators for grouping the outpatient visits. Specifically, if the secondary diagnosis influence degree prediction value is less than 1.3, it is determined that the secondary diagnosis influence degree is low; if the secondary diagnosis influence degree prediction value is greater than or equal to 1.3 and less than 1.8, it is determined that the secondary diagnosis influence degree is medium; and if the secondary diagnosis influence degree prediction value is greater than or equal to 1.8, it is determined that the secondary diagnosis influence degree is high. Furthermore, the secondary diagnosis influence degree is used as an indicator for the visits.

[0101] S32, matching the at least one visit indicator and the secondary diagnosis influence indicator with the standard visit indicator of the outpatient visit grouping, to obtain a matched outpatient visit grouping;

[0102] The standard visit indicators in the outpatient visit grouping include the main diagnostic disease, the influence degree of the secondary diagnostic disease, the hospital nature, the hospital level, the developed degree of the visit area, and the age of the visit person. The types of the standard visit indicators in the outpatient visit grouping can be consistent with the types of the visit indicators in the outpatient visit data. The types of the visit indicators in the outpatient visit data can be more abundant than the types of the standard visit indicators in the outpatient visit grouping. For example, the outpatient visit grouping 1: the main diagnostic disease (disease ICD top three digits XX1), the low influence degree of the secondary disease, the hospital nature (public hospital), the hospital level (three levels), the developed degree of the visit area (first line), and the age of the visit person (greater than 25 years old and less than or equal to 45 years old, (25, 45]). The outpatient visit grouping 2: the main diagnostic disease (disease ICD top three digits XX5), the high influence degree of the secondary disease, the hospital nature (public hospital), the hospital level (three levels), the developed degree of the visit area (first line), and the age of the visit person (greater than 25 years old and less than or equal to 45 years old, (25, 45]). The outpatient visit grouping 3: the main diagnostic disease (disease ICD top three digits XX1), the low influence degree of the secondary disease, the hospital nature (public hospital), the hospital level (three levels), and the developed degree of the visit area (first line). When the outpatient visit data is matched with each outpatient visit grouping, the visit indicators in the outpatient visit data are matched with the standard visit indicators in the outpatient visit grouping. If the standard visit indicators in the outpatient visit grouping can be matched with corresponding visit indicators in the outpatient visit data, the outpatient visit grouping is the matched outpatient visit grouping of the outpatient visit data. For example, the outpatient visit data T: the main diagnostic disease (disease ICD top three digits XX1), the low influence degree of the secondary disease, the hospital nature (public hospital), the hospital level (three levels), the developed degree of the visit area (first line), and the age of the visit person (greater than 25 years old and less than or equal to 45 years old, (25, 45]). The matched outpatient visit grouping of the outpatient visit data T is the outpatient visit grouping 1 and the outpatient visit grouping 3.

[0103] S33, the matched outpatient visit grouping with the most standard visit indicators is taken as the target outpatient visit grouping;

[0104] For two or more matched outpatient visit groupings, the more standard visit indicators, the finer the granularity of the outpatient visit grouping, which is conducive to improving the identification accuracy. For example, the matched outpatient visit groupings of the outpatient visit data T are the outpatient visit grouping 1 and the outpatient visit grouping 3. The standard visit indicators of the outpatient visit grouping 1 are more than those of the outpatient visit grouping 3. Therefore, the target outpatient visit grouping of the outpatient visit data T is the outpatient visit grouping 1.

[0105] As an implementation form, the calculation step of the visit expense threshold of the outpatient visit grouping includes:

[0106] S41, Obtain multiple historical outpatient visit data that match the outpatient visit group;

[0107] The matching method between outpatient visit groups and historical outpatient visit data is the same as the matching method between outpatient visit groups and outpatient visit data recorded in step S32, as detailed above.

[0108] S42, obtain target cost data in the historical outpatient visit data that matches the type of the visit cost threshold, and construct a corresponding target cost dataset based on the target cost data;

[0109] For example, if the threshold for medical expenses is the total cost of outpatient visits, then the target cost data is the total cost of outpatient visits in the historical outpatient data.

[0110] S43, arrange the target cost data in the target cost dataset in ascending order, and obtain a first preset quantile and a second preset quantile based on the arrangement result, wherein the first preset quantile is greater than the second preset quantile;

[0111] As one implementation method, the first preset quantile can be the third-quarter quantile in the permutation result, and the second preset quantile can be the quarter quantile in the permutation result.

[0112] S44, calculate the medical expense threshold based on the first preset quantile, the second preset quantile, and the preset coefficient, wherein the preset coefficient is greater than 1;

[0113] As one implementation method, the box plot extreme value method is used to calculate the medical expense threshold. Specifically, the medical expense threshold can be obtained according to the following formula:

[0114] threshold=q 3 -α*(q 3 -q 1 )

[0115] Where threshold is the threshold for medical expenses, q 3 The first preset quantile (e.g., the third-quarter quantile in the permutation result), the second preset quantile (e.g., the quarter quantile in the permutation result), and α are preset coefficients.

[0116] In some implementations, q 3 It is the third-quarter digit of the data set after sorting the data in ascending order, q 1 The value is a quarter quantile, and the preset coefficient α can be 1.5, which can be adjusted according to the actual situation.

[0117] In some embodiments, to improve the accuracy of outpatient visit cost abuse judgment in actual production, the method sets α to 3 through continuous verification. The maximum threshold of the total cost and each cost in each outpatient visit group is calculated by the extreme value method of the box plot, which is used as the threshold for judging the cost abuse of each outpatient visit group.

[0118] In some embodiments, after step S44, the following steps are further included:

[0119] S51, obtaining the coefficient of variation of each target cost data in the target cost data set;

[0120] As an embodiment, when the outpatient visit groups are grouped, the smaller the coefficient of variation of the data in the group, the smaller the discrete degree of the data in the group, and the more reliable the grouping. Therefore, the present embodiment can exclude the outpatient visit groups whose coefficient of variation is greater than a preset coefficient of variation threshold. For example, the preset coefficient of variation threshold can be 0.7.

[0121] S52, obtaining the median of each target cost data in the target cost data set, and obtaining the ratio of the visit cost threshold to the median;

[0122] Wherein, under the premise of meeting the sample size and discrete degree, the greater the ratio of the maximum threshold to the median of the outpatient visit group, the more valuable the judgment of cost abuse of the group.

[0123] S53, obtaining the accuracy evaluation value of the visit cost threshold according to the ratio, the sample size of the target cost data set, and the coefficient of variation;

[0124] As an embodiment, the accuracy of the visit cost threshold of the outpatient visit group can be evaluated according to the following formula:

[0125]

[0126] Wherein, Score is the accuracy evaluation value of the visit cost threshold, max med is the ratio of the visit cost threshold to the median, count is the sample size of the target cost data set, and CV is the coefficient of variation obtained in step S51.

[0127] In some embodiments, the accuracy evaluation values of the visit cost thresholds of all outpatient visit groups can be sorted in descending order. To ensure the accuracy of outpatient cost abuse judgment in actual production, only the top 50% of the outpatient visit groups are extracted to construct the threshold library for judging the cost abuse of outpatient visits.

[0128] In some embodiments, the obtaining step of the plurality of outpatient visit groups comprises:

[0129] S61, obtain a plurality of historical outpatient visit data, the historical outpatient visit data comprising a first preset number of types of potential visit indicators and a second preset number of types of cost data;

[0130] The visit indicator types of the potential visit indicators can be main diagnosis diseases, secondary diagnosis disease influence degrees, hospital natures, hospital grades, visit area development degrees, and visit person ages.

[0131] S62, obtain, by using a control variable method, an influence degree value of each of the potential visit indicators on the visit cost, wherein the visit cost is a sum of the cost data.

[0132] S63, sort the visit indicator types in descending order of the influence degree values.

[0133] The visit indicator types with large influence degrees on the cost data are arranged in front, and the visit indicator types with small influence degrees on the cost data are arranged in back. For example, the sorting result is main diagnosis diseases, secondary diagnosis disease influence degrees, hospital natures, hospital grades, visit area development degrees, and visit person ages.

[0134] S64, select the visit indicator types from front to back according to the sorting result, and sequentially construct outpatient visit groups of different granularities.

[0135] Specifically, when constructing the outpatient visit groups, the main diagnosis diseases, the secondary diagnosis disease influence degrees, and the hospital natures in the first three positions can be used to construct the groups, for example, an outpatient visit group A1: disease X1, high secondary diagnosis disease influence degree, and public hospital nature; an outpatient visit group A2: disease X1, low secondary diagnosis disease influence degree, and public hospital nature; and an outpatient visit group A3: disease X1, medium secondary diagnosis disease influence degree, and private hospital nature.

[0136] On the basis of the groups of the three indicator types, the fourth indicator type, the hospital grade, is added, for example, on the basis of the outpatient visit group A1, the following groups can be obtained: an outpatient visit group B1: disease X1, high secondary diagnosis disease influence degree, public hospital nature, and first-grade hospital; an outpatient visit group B2: disease X1, high secondary diagnosis disease influence degree, public hospital nature, and second-grade hospital; an outpatient visit group B3: disease X1, high secondary diagnosis disease influence degree, public hospital nature, and third-grade hospital; and an outpatient visit group B4: disease X1, high secondary diagnosis disease influence degree, public hospital nature, and non-rated hospital.

[0137] On the basis of the grouping of the four index types, the fifth index type of the developed degree of the visiting area is added, and the index types are sequentially added until all the index types are constructed.

[0138] S40, determining whether the outpatient visit data is abnormal data according to the fee data and the visit fee threshold of the target outpatient visit grouping;

[0139] If the fee data is greater than the visit fee threshold, it is determined that the outpatient visit data is abnormal data; if the fee data is less than or equal to the visit fee threshold, it is determined that the outpatient visit data passes the audit.

[0140] As shown in Figure 2 An embodiment of the present application provides a kind of outpatient data anomaly identification device, and the outpatient data anomaly identification device 20 includes: data acquisition module 21, secondary diagnosis prediction module 22, grouping matching module 23 and anomaly identification module 24, wherein data acquisition module 21, for obtaining the outpatient visit data to be identified, obtains corresponding secondary diagnosis data according to the outpatient visit data, wherein the outpatient visit data includes at least one visit index and at least one fee data;Secondary diagnosis prediction module 22 is used to input the secondary diagnosis data into pre-trained secondary diagnosis influence prediction model, and output the secondary diagnosis influence degree prediction value corresponding to the secondary diagnosis data, wherein the secondary diagnosis influence prediction model is obtained by training according to historical secondary diagnosis data and historical secondary diagnosis influence degree calculation value;Grouping matching module 23 is used to obtain the target outpatient visit grouping of the outpatient visit data from the plurality of preset outpatient visit groupings according to the at least one visit index and the secondary diagnosis influence degree prediction value, wherein each of the outpatient visit groupings includes a plurality of standard visit indexes and at least one visit fee threshold;Anomaly identification module 24 is used to determine whether the outpatient visit data is abnormal data according to the fee data and the visit fee threshold of the target outpatient visit grouping.

[0141] In some embodiments, the data acquisition module 21 is further configured to: acquire the number of secondary diagnosis diseases according to the outpatient visit data; acquire the historical visit expense median of each of the secondary diagnosis diseases in the outpatient visit data, calculate the average of the historical visit expense median; determine the secondary diagnosis disease with the highest disease level in the outpatient visit data; determine the number of secondary diagnosis diseases with disease level greater than a preset level in the outpatient visit data; acquire the disease level difference between the secondary diagnosis disease with the highest disease level and the primary diagnosis disease in the outpatient visit data; and construct the secondary diagnosis data according to the number of secondary diagnosis diseases, the average, the secondary diagnosis disease with the highest disease level, the number of secondary diagnosis diseases with disease level greater than the preset level, and the disease level difference.

[0142] In some embodiments, the secondary diagnosis prediction module 22 is further configured to: acquire at least one training sample, the training sample including historical secondary diagnosis data, the historical secondary diagnosis data being acquired according to historical outpatient visit data; acquire a plurality of target historical outpatient visit data according to the primary diagnosis disease in the historical outpatient visit data, wherein the diagnosis disease in the target historical outpatient visit data is the primary diagnosis disease and the target historical outpatient visit data does not include secondary diagnosis disease; determine a basic visit expense according to the median of the visit expense in the plurality of target historical outpatient visit data, acquire the historical secondary diagnosis influence degree calculation value according to the visit expense in the historical outpatient visit data and the basic visit expense; input the historical secondary diagnosis data into the secondary diagnosis influence prediction model, and output the historical secondary diagnosis influence degree prediction value; calculate a prediction error according to the historical secondary diagnosis influence degree prediction value and the historical secondary diagnosis influence degree calculation value, adjust the parameters of the secondary diagnosis influence prediction model according to the prediction error, until the secondary diagnosis influence prediction model reaches a training convergence condition.

[0143] In some embodiments, the grouping matching module 23 is further configured to: determine the secondary diagnosis influence indicator of the outpatient visit data according to the secondary diagnosis influence degree prediction value; match the at least one visit indicator and the secondary diagnosis influence indicator with the standard visit indicator of the outpatient visit grouping, to obtain a matched outpatient visit grouping; and take the matched outpatient visit grouping with the largest number of standard visit indicators as the target outpatient visit grouping.

[0144] In some embodiments, the grouping matching module 23 is further configured to: obtain a plurality of historical outpatient visit data matched with the outpatient visit grouping; obtain target cost data in the historical outpatient visit data matched with the type of the visit cost threshold, construct a corresponding target cost data set according to the target cost data; arrange the target cost data in the target cost data set in ascending order, obtain a first preset quantile and a second preset quantile according to the arrangement result, the first preset quantile being greater than the second preset quantile; and calculate the visit cost threshold according to the first preset quantile, the second preset quantile, and a preset coefficient, the preset coefficient being greater than 1.

[0145] In some embodiments, the grouping matching module 23 is further configured to: obtain a coefficient of variation of each target cost data in the target cost data set; obtain a median of each target cost data in the target cost data set, and obtain a ratio of the visit cost threshold to the median; and obtain an accuracy evaluation value of the visit cost threshold according to the ratio, a sample quantity of the target cost data set, and the coefficient of variation.

[0146] In some embodiments, the abnormality identification module 24 is further configured to: if the cost data is greater than the visit cost threshold, determine that the outpatient visit data is abnormal data; and if the cost data is less than or equal to the visit cost threshold, determine that the outpatient visit data passes the audit.

[0147] Figure 3 FIG. 1 is a structural schematic diagram of an electronic device according to an embodiment of the present application. As shown in FIG. 1, the electronic device 30 includes a processor 31 and a memory 32 coupled to the processor 31. Figure 3 The memory 32 stores program instructions for implementing the outpatient data abnormality identification method of any of the above embodiments.

[0148] The memory 32 stores program instructions for implementing the outpatient data abnormality identification method of any of the above embodiments.

[0149] The processor 31 is configured to execute the program instructions stored in the memory 32 to perform the outpatient data abnormality identification.

[0150] The processor 31 can also be referred to as a CPU (Central Processing Unit). The processor 31 can be an integrated circuit chip having a processing capability of signals. The processor 31 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application-Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0151] Reference is made to Figure 4 , Figure 4 The structure diagram of the storage medium of an embodiment of the present application is shown. The storage medium 40 of the embodiment of the present application stores program instructions 41 capable of implementing all the methods described above, wherein the program instructions 41 can be stored in the storage medium in the form of a software product, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media capable of storing program codes, or a terminal device such as a computer, a server, a mobile phone, or a tablet.

[0152] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the above-described device embodiments are merely schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.

[0153] In addition, each function unit in the various embodiments of the present application can be integrated into a processing unit, or each unit can exist alone physically, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware, or can be implemented in the form of a software function unit. The above is merely a preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent flow transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied to other related technical fields, is also included in the patent protection scope of the present application.

[0154] The above is merely a preferred embodiment of the present application, and does not limit the present application in any form. Although the present application has been disclosed as above with the preferred embodiment, it is not intended to limit the present application. Any person skilled in the art can make some minor changes or modifications to the equivalent embodiments of the equivalent changes within the scope of the technical solutions of the present application, but as long as it does not deviate from the technical solutions of the present application, any simplification, modification, equivalent change and modification of the above embodiments according to the technical essence of the present application are still within the scope of the technical solutions of the present application.

Claims

1. A method for identifying abnormal outpatient data, characterized in that, The method comprises the following steps: obtaining outpatient data to be identified, and obtaining corresponding secondary diagnosis data according to the outpatient data, wherein the outpatient data comprises at least one medical index and at least one expense data; inputting the secondary diagnosis data into a pre-trained secondary diagnosis influence prediction model to output a secondary diagnosis influence degree prediction value corresponding to the secondary diagnosis data, wherein the secondary diagnosis influence prediction model is trained according to historical secondary diagnosis data and historical secondary diagnosis influence degree calculation values; obtaining a target outpatient grouping of the outpatient data from a plurality of pre-set outpatient groupings according to the at least one medical index and the secondary diagnosis influence degree prediction value, wherein each of the outpatient groupings comprises a plurality of standard medical indexes and at least one medical expense threshold; determining whether the outpatient data is abnormal data according to the expense data and the medical expense threshold of the target outpatient grouping; wherein the training step of the secondary diagnosis influence prediction model comprises: obtaining at least one training sample, wherein the training sample comprises historical secondary diagnosis data obtained according to historical outpatient data; obtaining a plurality of target historical outpatient data according to a primary diagnosis disease in the historical outpatient data, wherein the diagnosis disease in the target historical outpatient data is the primary diagnosis disease and the target historical outpatient data does not contain a secondary diagnosis disease; determining a basic medical expense according to the median of the medical expenses in the plurality of target historical outpatient data, and obtaining the historical secondary diagnosis influence degree calculation value according to the medical expenses in the historical outpatient data and the basic medical expense; inputting the historical secondary diagnosis data into the secondary diagnosis influence prediction model to output the historical secondary diagnosis influence degree prediction value; calculating a prediction error according to the historical secondary diagnosis influence degree prediction value and the historical secondary diagnosis influence degree calculation value, adjusting the parameters of the secondary diagnosis influence prediction model according to the prediction error, and stopping until the secondary diagnosis influence prediction model reaches a training convergence condition. 2.The outpatient data anomaly identification method of claim 1, wherein, The method of obtaining the corresponding secondary diagnosis data according to the outpatient data comprises: obtaining the number of secondary diagnosis diseases according to the outpatient data; obtaining the median of historical medical expenses of each of the secondary diagnosis diseases in the outpatient data, and calculating the average value of the median of the historical medical expenses; determining the secondary diagnosis disease with the highest disease level in the outpatient data; determining the number of secondary diagnosis diseases with a disease level greater than a preset level in the outpatient data; obtaining the disease level difference between the secondary diagnosis disease with the highest disease level and the primary diagnosis disease in the outpatient data; constructing the secondary diagnosis data according to the number of secondary diagnosis diseases, the average value, the secondary diagnosis disease with the highest disease level, the number of secondary diagnosis diseases with a disease level greater than a preset level, and the disease level difference. 3.The outpatient data anomaly identification method of claim 1, wherein, The target outpatient visit group of the outpatient visit data is obtained from a plurality of preset outpatient visit groups according to the at least one visit index and the secondary diagnosis influence degree prediction value, and the target outpatient visit group comprises: determining a secondary diagnosis influence index of the outpatient visit data according to the secondary diagnosis influence degree prediction value; matching the at least one visit index and the secondary diagnosis influence index with the standard visit index of the outpatient visit group to obtain a matched outpatient visit group; taking the matched outpatient visit group with the largest number of standard visit indexes as the target outpatient visit group. 4.The outpatient data anomaly identification method of claim 1, wherein, The calculation step of the visit expense threshold of the outpatient visit group comprises: obtaining a plurality of historical outpatient visit data matched with the outpatient visit group; obtaining target expense data matched with the type of the visit expense threshold in the historical outpatient visit data, and constructing a corresponding target expense data set according to the target expense data; arranging the target expense data in the target expense data set in ascending order, and obtaining a first preset quantile and a second preset quantile according to the arrangement result, wherein the first preset quantile is greater than the second preset quantile; calculating the visit expense threshold according to the first preset quantile, the second preset quantile and a preset coefficient, wherein the preset coefficient is greater than 1. 5.The outpatient data anomaly identification method of claim 4, wherein, After the visit expense threshold is calculated according to the first preset quantile, the second preset quantile and a preset coefficient, the method further comprises: obtaining a coefficient of variation of each target expense data in the target expense data set; obtaining a median of each target expense data in the target expense data set, and obtaining a ratio of the visit expense threshold to the median; obtaining an accuracy evaluation value of the visit expense threshold according to the ratio, a sample number of the target expense data set and the coefficient of variation. 6.The outpatient data anomaly identification method of claim 1, wherein, The method for determining whether the outpatient visit data is abnormal data according to the expense data and the visit expense threshold of the target outpatient visit group comprises: if the expense data is greater than the visit expense threshold, determining that the outpatient visit data is abnormal data; if the expense data is less than or equal to the visit expense threshold, determining that the outpatient visit data passes the review.

7. An outpatient data anomaly identification device characterized by comprising: The method comprises: a data acquisition module configured to acquire outpatient visit data to be identified, and obtain corresponding secondary diagnosis data according to the outpatient visit data, wherein the outpatient visit data comprises at least one visit index and at least one expense data; The secondary diagnosis prediction module is configured to input the secondary diagnosis data into a pre-trained secondary diagnosis influence prediction model, and output a secondary diagnosis influence degree prediction value corresponding to the secondary diagnosis data, wherein the secondary diagnosis influence prediction model is trained according to historical secondary diagnosis data and historical secondary diagnosis influence degree calculation values; wherein the training of the secondary diagnosis influence prediction model comprises: obtaining at least one training sample, wherein the training sample comprises historical secondary diagnosis data, and the historical secondary diagnosis data is obtained according to historical outpatient diagnosis data; obtaining a plurality of target historical outpatient diagnosis data according to a primary diagnosis disease in the historical outpatient diagnosis data, wherein the diagnosis disease in the target historical outpatient diagnosis data is the primary diagnosis disease, and the target historical outpatient diagnosis data does not contain a secondary diagnosis disease; determining a basic treatment cost according to a median value of treatment costs in the plurality of target historical outpatient diagnosis data, and obtaining the historical secondary diagnosis influence degree calculation value according to a treatment cost in the historical outpatient diagnosis data and the basic treatment cost; inputting the historical secondary diagnosis data into the secondary diagnosis influence prediction model, and outputting a historical secondary diagnosis influence degree prediction value; calculating a prediction error according to the historical secondary diagnosis influence degree prediction value and the historical secondary diagnosis influence degree calculation value, adjusting parameters of the secondary diagnosis influence prediction model according to the prediction error, and stopping until the secondary diagnosis influence prediction model reaches a training convergence condition; The grouping matching module is configured to obtain a target outpatient diagnosis grouping of the outpatient diagnosis data from a plurality of preset outpatient diagnosis groupings according to the at least one treatment index and the secondary diagnosis influence degree prediction value, wherein each of the plurality of outpatient diagnosis groupings comprises a plurality of standard treatment indexes and at least one treatment cost threshold value; The abnormality identification module is configured to determine whether the outpatient diagnosis data is abnormal data according to the cost data and the treatment cost threshold value of the target outpatient diagnosis grouping.

8. An electronic device, comprising: The storage medium stores program instructions, and the program instructions are executed by the processor to implement the outpatient data abnormality identification method.

9. A storage medium, characterized by The storage medium stores program instructions, and the program instructions are executed by the processor to implement the outpatient data abnormality identification method.

Citation Information

Patent Citations

  • A modeling method for evaluating expenses for chronic diseases

    CN106407686A