Abnormal centralized reimbursement identification analysis method and system based on dynamic window
By using a dynamic window-based method for identifying and analyzing abnormal centralized reimbursements, and employing sliding windows and density clustering algorithms, fraudulent activities at the end of the annual medical insurance settlement period can be identified. This solves the problem of insufficient dynamic perception capabilities in traditional methods and enables precise supervision and early warning of medical insurance funds.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DAREWAY SOFTWARE
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-17
AI Technical Summary
In the existing management of medical insurance funds, traditional supervision and identification methods are difficult to effectively capture the time concentration and cost anomalies of concentrated reimbursement fraud at the end of the medical insurance year-end settlement. They lack dynamic perception capabilities, have insufficient adaptability, and have low efficiency in post-event verification, making it difficult to provide timely warnings and interventions.
An abnormal centralized reimbursement identification and analysis method based on dynamic windows is adopted. By acquiring the medical treatment and settlement data of insured persons, a medical treatment time series is generated. By using sliding window calculation and density clustering algorithm, abnormal data groups are identified, and medical treatment abnormality index and reimbursement abnormality index are calculated to accurately identify suspicious insured persons and medical institutions.
It enables accurate identification of fraudulent reimbursement activities at the end of the medical insurance settlement period, identifies suspected insured persons and medical institutions, assists medical insurance institutions in strengthening fund supervision, and improves the timeliness and accuracy of supervision.
Smart Images

Figure CN121883178A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical insurance anti-fraud technology, and in particular relates to a method and system for identifying and analyzing abnormal centralized reimbursement based on dynamic windows. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] With the continuous expansion of basic medical insurance coverage and the sustained improvement of protection levels, the safety and sustainable operation of the medical insurance fund face increasingly severe challenges. In the management practice of the medical insurance fund, in order to control the risk of fund expenditure and ensure the fairness of the system, an annual maximum payment limit is usually set, commonly known as the "ceiling". Medical expenses exceeding this limit are generally no longer paid by the pooled fund.
[0004] However, driven by illicit profits, a typical fraudulent pattern has emerged in the medical insurance field: some designated medical institutions or insured individuals, taking advantage of the end of the annual settlement cycle, artificially create high medical expense records in a short period by organizing concentrated card swiping and rushing to prescribe drugs or treatments, attempting to quickly deplete or approach their annual reimbursement limit, thereby achieving the illegal purpose of misappropriating medical insurance funds. Such behavior typically exhibits significant temporal concentration and abnormal costs: its occurrence is highly concentrated at the end of the medical insurance year, which is seriously inconsistent with the normal distribution of medical treatment time; at the same time, the settlement costs and reimbursement amounts generated in a single instance or within a short period often far exceed the patient's actual reasonable medical needs. This "peak-rush" or "concentrated reimbursement" fraud not only puts unreasonable large expenditure pressure on the medical insurance fund in a short period, accelerating fund depletion, but may also trigger the risk of fund exhaustion in some areas, seriously eroding the safety pool of the medical insurance fund, and ultimately harming the fair protection rights of the vast majority of law-abiding insured individuals.
[0005] Currently, traditional regulatory and identification methods for this type of fraudulent behavior with temporal clustering characteristics are mostly based on static rules or post-event manual verification, such as setting thresholds for single reimbursement amounts or monitoring the frequent prescribing of specific drugs. These methods have obvious limitations: first, they lack the dynamic perception of "concentration" over time, making it difficult to effectively capture abnormal patterns of behavior in the short term; second, the rules are relatively fixed, lacking adaptability and flexibility in the face of constantly evolving fraudulent methods; and third, post-event verification is inefficient, making timely warnings and intervention difficult. Summary of the Invention
[0006] To overcome the shortcomings of the prior art, this invention provides a method and system for identifying and analyzing abnormal centralized reimbursement based on dynamic windows. This method can accurately identify fraudulent centralized reimbursement activities at the end of the medical insurance settlement period, identify suspected insured persons, medical institutions and related clues, and assist medical insurance institutions in strengthening fund supervision.
[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: The first aspect of this invention provides a method for identifying and analyzing abnormal centralized expense reimbursements based on dynamic windows; An abnormal centralized expense reimbursement identification and analysis method based on dynamic windows includes: Obtain all medical treatment settlement data of insured individuals within a settlement year, and based on the medical treatment settlement data, create a medical treatment time series for each insured individual; A sliding window calculation method is used to calculate the relevant values of all patient data objects in each window; Analyze the objects in each time window, divide different data into different data clusters, and classify the objects in the free time window that do not belong to any data cluster into the suspected data group; evaluate the data in the suspected data group, and list the data with scores exceeding a certain threshold as abnormal data; Based on the aforementioned abnormal data, the medical treatment abnormality index and reimbursement abnormality index of each insured person, as well as the abnormal transaction frequency index and abnormal reimbursement ratio index of each designated medical institution, are calculated respectively within the abnormal monitoring period. Based on the calculated index values, suspected insured individuals and suspected designated medical institutions are identified, and clue data containing relevant abnormal medical records is generated.
[0008] As a further technical solution, the generation of the medical visit time series includes: Obtain data related to insured individuals and their medical treatment, including insured individual information, medical insurance settlement information, and information on designated medical institutions; The acquired data undergoes preliminary processing, forming a medical visit data object for each visit of the insured person. This object includes the settlement time, settlement cost, reimbursement cost, and settlement medical institution. A time series of medical visit information is then generated based on the visit time of the medical visit data object.
[0009] As a further technical solution, the analysis of objects within each time window, dividing different data into different data clusters, includes: Treat all insured individuals’ time window objects as a point, and map all time window objects to a two-dimensional interval according to the two dimensions of total calculated amount and total reimbursement amount of medical treatment data objects within the window, to form the original dataset. By calculating the Euclidean distance between objects in different time windows, different data are divided into different data clusters according to the distance between objects and the number of close objects. Objects in free time windows that do not belong to any data cluster are classified into the suspect data group.
[0010] As a further technical solution, the step of classifying detached time window objects that do not belong to any data cluster into a suspected data group includes: Set distance range thresholds and intimacy thresholds; Choose any time window object, calculate its Euclidean distance to neighboring time window objects, mark time window objects whose Euclidean distance is less than the distance range threshold as close objects, and mark time window objects as core objects if the number of close objects of the time window object exceeds the closeness threshold. Iterate through all time window objects and mark all core objects. Select any core object to mark it as a data cluster. Add the same data cluster mark to close objects within the distance range threshold. If a close object within the distance range threshold is also a core object, iterate through the marking of that close object until all close objects within the distance range threshold of the core object in the data cluster are included in the data cluster. The size of the data clusters is analyzed. If the number of data objects in a data cluster exceeds a certain threshold, the data in that data cluster is removed from the original dataset. Data is then selected again from the remaining original dataset for data cluster labeling and data removal, until the remaining data no longer meets the requirements. The remaining time window objects are then classified as suspect data.
[0011] As a further technical solution, the evaluation of the suspected data group, classifying data with scores exceeding a certain threshold as abnormal data, includes: Calculate the average cost across all time windows in the original dataset. Evaluate the suspicious data group based on the average cost within each time window and the reimbursement limit. Data with evaluation scores exceeding a certain threshold are classified as abnormal. The method for calculating the abnormality score for the abnormal data group is as follows:
[0012] in, It is abnormal data group data Abnormal index, , These are weighting coefficients. Settle the fees on an average basis across all time windows. This is the annual reimbursement limit.
[0013] As a further technical solution, based on abnormal data, the abnormal medical treatment index and abnormal reimbursement index of each insured person during the abnormal monitoring period are calculated, and suspicious insured persons and clue data are identified, including: The abnormal window data of the interval is classified according to the insured persons. The abnormal time window objects of the insured persons are parsed to obtain the medical treatment data objects corresponding to the abnormal time windows. Duplicate medical treatment time objects are removed to generate a set of abnormal medical treatment data objects of the interval. Calculate the abnormal medical treatment index and abnormal reimbursement index during the abnormal monitoring period to obtain the abnormal value of the insured person; Insured individuals whose abnormal values exceed a certain threshold are marked as suspected cases of centralized reimbursement, and the set of abnormal medical visit times corresponding to these insured individuals is marked as abnormal clue data for centralized reimbursement.
[0014] As a further technical solution, based on abnormal data, the abnormal transaction frequency index and abnormal reimbursement ratio index of each designated medical institution during the abnormal monitoring period are calculated, and suspected designated medical institutions and clue data are identified, including: Get all the outpatient data objects in all time windows corresponding to the abnormal group data, remove duplicate outpatient data objects that appear in multiple windows, classify the outpatient data objects according to the settlement medical institutions, and generate several abnormal medical institution outpatient data sets. Obtain all outpatient data objects of medical institutions within the abnormal monitoring period and generate a medical institution outpatient data set; The abnormality index of settlement medical institutions is calculated from two dimensions: abnormal drug purchase frequency and abnormal reimbursement ratio. Medical institutions whose abnormal index exceeds a certain threshold are marked as suspected centralized reimbursement institutions, and the abnormal objects corresponding to these institutions are used as clue data for suspected centralized reimbursement institutions.
[0015] A second aspect of the present invention provides an abnormal centralized reimbursement identification and analysis system based on dynamic windows.
[0016] An anomaly centralized expense reimbursement identification and analysis system based on dynamic windows includes: The medical visit sequence generation module is configured to: obtain all medical visit settlement data of insured persons within a settlement year, and generate a medical visit time sequence for each insured person based on the medical visit settlement data; The sliding window management module is configured to use a sliding window calculation method to calculate the relevant values of all patient data objects in each window; The abnormal window data screening module is configured to: analyze the objects in each time window, divide different data into different data clusters, classify the free time window objects that do not belong to any data cluster into the suspected data group; evaluate the data in the suspected data group, and list those with scores exceeding a certain threshold as abnormal data; The suspected clue management module is configured to: calculate the medical treatment abnormality index and reimbursement abnormality index of each insured person and the abnormal transaction frequency index and abnormal reimbursement ratio index of each designated medical institution during the abnormal monitoring period based on the abnormal data; identify suspected insured persons and suspected designated medical institutions according to the calculated index values, and generate clue data containing relevant abnormal medical records.
[0017] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of the method for identifying and analyzing abnormal centralized reimbursement based on dynamic windows as described in the first aspect of the present invention.
[0018] A fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the abnormal centralized reimbursement identification and analysis method based on dynamic windows as described in the first aspect of the present invention.
[0019] The above one or more technical solutions have the following beneficial effects: This invention constructs time-series data for each insured individual's medical expenses based on multi-dimensional characteristics such as the time of occurrence, total actual expenses, and reimbursement amount. Then, using a sliding time window algorithm, the time series is divided into fixed-size time windows. Within each time window, indicators such as total medical settlement expenses and reimbursement amounts are calculated. Anomaly-based window data is filtered out using density clustering. Based on this, abnormal concentrated reimbursement window data is further filtered using the insured individual's annual indicator information, providing clues to detect fraudulent medical treatment behavior. Furthermore, medical institutions with a large amount of suspicious data are subject to comprehensive monitoring to trace the source and accurately identify institutions suspected of concentrated reimbursement.
[0020] This invention forms a medical data object from each medical visit of an insured person, including settlement time, settlement cost, reimbursement cost, and settlement medical institution. It generates a time series of medical information according to the medical time of the medical data object, forming a correlation between medical behavior, medical time, and medical cost. The free medical data is arranged according to the time distribution. Combined with the time window algorithm, it is beneficial to fit the short-term concentrated reimbursement behavior of insured persons.
[0021] In the process of identifying suspicious data, this invention first generates abnormal data groups by filtering abnormal time windows through density clustering; second, it evaluates the medical treatment situation of insured individuals by combining the data in the abnormal data groups with their daily medical treatment behavior and the capping data to obtain abnormal behaviors; and third, it analyzes the abnormal medical treatment behavior in the abnormal data groups and links it with medical institutions, and assesses whether medical institutions are inducing insured individuals to rush to the top reimbursement limit by combining abnormal drug purchase frequency and abnormal reimbursement ratio, thus achieving accurate identification of centralized reimbursement behavior by comprehensively considering multiple aspects.
[0022] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0023] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0024] Figure 1 This is a flowchart of the method in the first embodiment.
[0025] Figure 2 This is a system structure diagram of the second embodiment. Detailed Implementation
[0026] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0027] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0028] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0029] Example 1 This embodiment discloses an abnormal centralized reimbursement identification and analysis method based on dynamic windows. It constructs a time series of medical data of insured persons, divides the time series through dynamic windows, analyzes and identifies abnormal data groups in the time window data, and further evaluates the abnormal data groups to obtain suspected insured persons, suspected institutions and corresponding clue data. It effectively identifies abnormal reimbursement behavior that exceeds the reimbursement limit, screens out suspected persons and medical institutions, and assists medical security institutions in supervising and combating medical insurance fraud.
[0030] Specifically, such as Figure 1As shown, a method for identifying and analyzing abnormal centralized expense reimbursements based on dynamic windows includes: Step S1: Obtain all medical treatment settlement data of the insured person within a settlement year, and based on the medical treatment settlement data, generate a medical treatment time series for each insured person.
[0031] Step S11: Obtain all medical treatment settlement data of the insured person within a settlement year, including insured person information, medical insurance settlement information, designated medical institution information, etc.
[0032] The information of the insured person should include the insured person's ID number and name.
[0033] The medical insurance settlement information includes settlement information generated through various methods such as online reimbursement and manual reimbursement at the center. Each piece of settlement information should include at least the medical institution ID, settlement time, insured person's ID number, settlement amount, and reimbursement amount.
[0034] The information on designated medical institutions includes the medical institution ID and the name of the designated medical institution.
[0035] Step S12: Perform preliminary processing on the medical settlement data. Each medical visit data is formed into a medical data object. Generate a time sequence of medical information according to the medical time of the medical data object.
[0036] The initial data processing includes deduplication of duplicate data and deletion of data with missing important data fields.
[0037] Each medical visit of an insured person is recorded as a medical data object, which includes the settlement time, settlement cost, reimbursement cost, and settlement medical institution.
[0038] A unique medical visit time series is created for each insured person within the same settlement year. The medical visit data objects are associated with the time points corresponding to the medical visit time series of the insured person according to the settlement time, and a series is generated for each insured person.
[0039] .
[0040] in, Representing insured persons Time series, This refers to the medical visit data of the insured person during the j-th visit within the current year.
[0041] Step S2: Using the sliding window calculation method, calculate the relevant values of all medical data objects in each window.
[0042] Step S21: Set a fixed sliding window (e.g., one month or 7 days), and a fixed step size, smaller than the sliding window size, to ensure the continuity of window data. Divide the time series of each insured person into a fixed number of windows using the sliding window, and calculate the relevant values of the medical data objects in each window, including the total calculated amount and total reimbursement amount for the time window objects. See below:
[0043]
[0044] in, It is a sequence of time windows for insured person i. The object representing the k-th time window of insured person i. This indicates the k-th window marker for insured person i. For insured person i, calculate the total amount for all medical records within the k-th window. This represents the total reimbursement amount for all medical records of insured person i within the k-th window.
[0045] Step S3: Analyze the objects in each time window, divide the different data into different data clusters, and classify the objects in the free time window that do not belong to any data cluster into the suspected data group; evaluate the data in the suspected data group, and list the data with scores exceeding a certain threshold as abnormal data.
[0046] Step S31: Treat all insured individuals' time window objects as a point, and map all time window objects to a two-dimensional interval according to the two dimensions of the total calculated amount and the total reimbursement amount of the medical data objects within the window, forming the original dataset.
[0047] Step S32: By calculating the Euclidean distance between objects in different time windows in the original dataset, the data is divided into different data clusters according to the distance between objects and the number of close objects. Objects in free time windows that do not belong to any data cluster are assigned to the suspect data group. Specifically: Step S321: Set a distance range threshold and an intimacy threshold. The distance range threshold is used to measure the distance between objects in different time windows, and the intimacy threshold is used to measure the number of close objects adjacent to the object in the time window.
[0048] Step S322: Select any time window object in the original dataset, calculate its Euclidean distance to neighboring time window objects, and mark time window objects whose Euclidean distance is less than a distance threshold as close objects. If the number of close objects of a time window object exceeds the closeness threshold, then mark the time window object as a core object. Use this method to traverse all time window objects and mark all core objects.
[0049] Step S323: Select any core object and mark it as a data cluster. Add the same data cluster mark to intimate objects within the distance range threshold. If an intimate object within the distance range threshold is also a core object, further iterate the search for that intimate object until all intimate objects within the distance range threshold of the core object in the data cluster are included in the data cluster.
[0050] Step S324 analyzes the size of the data cluster. If the number of data objects in the data cluster exceeds a certain threshold, the data of the data cluster is removed from the original dataset.
[0051] Step S325: Repeat steps S321-S324 until the remaining data no longer meets the requirements, and classify the remaining time window objects as suspect data.
[0052] Step S33: Calculate the average cost of all time windows in the original dataset. Based on the average cost of the time windows and the reimbursement limit, further evaluate the data of the suspected data group and classify those with evaluation scores exceeding a certain threshold as abnormal.
[0053] The method for calculating anomaly scores for abnormal data groups is as follows:
[0054] in, It is abnormal data group data Abnormal index, , These are weighting coefficients. Settle the fees on an average basis across all time windows. This is the annual reimbursement limit.
[0055] Step S4: Based on the abnormal data, calculate the medical treatment abnormality index and reimbursement abnormality index for each insured person during the abnormal monitoring period, as well as the abnormal transaction frequency index and abnormal reimbursement ratio index for each designated medical institution.
[0056] Step S41: Define a period of time at the end of the annual medical insurance settlement cycle as the abnormal monitoring period, and obtain the abnormal time window object data in the abnormal data group that is located in the abnormal monitoring period to form the interval abnormal window data.
[0057] Step S42: Generate a set of abnormal medical visit time objects from the abnormal window data of the interval, calculate the abnormal medical visit index and the abnormal reimbursement index within the abnormal monitoring period to generate abnormal values of insured persons, and obtain data on suspected centralized reimbursement suspects and clues based on the abnormal values of insured persons.
[0058] Step S421: Classify the abnormal window data of the interval according to the insured persons, parse the abnormal time window object to which the insured persons belong, obtain the medical treatment data object corresponding to the abnormal time window, remove duplicate medical treatment time objects, and generate a set of abnormal medical treatment data objects of the interval.
[0059] .
[0060] in Let i be the set of abnormal medical visit times within a given time interval. This refers to the m-th interval of abnormal medical visits by insured person i.
[0061] Step S422: Calculate the abnormal medical treatment index and abnormal reimbursement index during the abnormal monitoring period to obtain the abnormal value of the insured person;
[0062]
[0063]
[0064] in, It is a medical visit abnormality index. It is the reimbursement anomaly index. For the insured person's abnormal index, It is the sum of the total settlement costs for insured individuals with abnormal medical visit times. It is the sum of the total settlement costs for all medical visits of the insured person i. It is the sum of the total reimbursement expenses for insured individuals with abnormal medical visit times. This is the annual reimbursement limit. , It is a weighting coefficient, and .
[0065] Step S423: Mark insured persons whose abnormal values exceed a certain threshold as suspected cases of centralized reimbursement, and mark the set of abnormal medical visit times corresponding to the insured persons as abnormal clue data for centralized reimbursement.
[0066] Step S43: Obtain the medical visit data objects from the interval abnormal window data, obtain all medical visit data objects within the abnormal monitoring period, calculate the abnormal index of the settlement medical institution from two dimensions: abnormal drug purchase frequency and abnormal reimbursement ratio, and mark the abnormal index exceeding a certain threshold as a suspected medical institution.
[0067] Step S431: Obtain all medical visit data objects in all time windows corresponding to the abnormal window data of the interval, remove duplicate medical visit data objects that appear in multiple windows, classify the medical visit data objects according to the settlement medical institutions, and generate several abnormal medical institution medical visit data sets.
[0068]
[0069] in This is the vth abnormal patient to visit the medical institution.
[0070] Step S432: Obtain all outpatient data objects of medical institutions within the abnormal monitoring period and generate a medical institution outpatient data set.
[0071]
[0072] in It is the z-th patient visit data object of medical institution u within the abnormal monitoring period.
[0073] Step S433: Calculate the abnormality index of the settlement medical institution from two dimensions: abnormal drug purchase frequency and abnormal reimbursement ratio. The calculation formula is as follows:
[0074] in, It is an abnormal index of medical institutions. It refers to the number of patients with abnormal medical records at medical institutions. This refers to the number of patients treated at the medical institution during the abnormal monitoring period. This is the total reimbursement amount for patients with abnormal medical visits at medical institutions. This is the sum of reimbursement amounts for patients treated at medical institutions during the abnormal monitoring period. , It is the weighting coefficient.
[0075] Step S434: Medical institutions whose abnormal index exceeds a certain threshold are marked as suspected centralized reimbursement institutions, and the abnormal objects corresponding to the institutions are used as suspected centralized reimbursement institution clue data.
[0076] Step S5: Based on the calculated index value, identify the suspected insured persons and suspected designated medical institutions, and generate clue data containing relevant abnormal medical records.
[0077] Based on the calculated abnormal medical treatment index and abnormal reimbursement index, insured individuals exceeding preset thresholds are marked as "suspected insured persons," and their detailed medical records within all corresponding abnormal time windows are extracted as clue data. Simultaneously, based on the abnormal transaction frequency index and abnormal reimbursement ratio index of medical institutions, institutions exceeding thresholds are marked as "suspected designated medical institutions," and all related abnormal medical records are linked to them to form a structured clue report.
[0078] Example 2 This embodiment discloses an abnormal centralized reimbursement identification and analysis system based on dynamic windows; like Figure 2 As shown, an abnormal centralized reimbursement identification and analysis system based on dynamic windows includes: a medical appointment sequence generation module 1, a sliding window management module 2, an abnormal window data screening module 3, and a suspicious clue management module 4.
[0079] The medical visit sequence generation module 1 is used to acquire data on insured persons and related medical treatment information. After preliminary processing of the data, it generates a medical visit time series for each insured person based on the medical visit information. It includes a data preprocessing module 101 and a medical visit time series generation module 102.
[0080] The data preprocessing module 101 acquires insured persons and treatment-related data for preliminary data processing, including deduplication of duplicate data and deletion of data with missing important data fields.
[0081] The medical visit time series generation module 102 forms a medical visit data object from each medical visit of the insured person, and generates a medical visit information time series according to the medical visit time of the medical visit data object.
[0082] The sliding window management module 2 divides the medical visit time series into multiple windows and calculates the total amount and total reimbursement amount of all medical visit data objects in each window. It includes a sliding window setting module 201 and a window data generation module 202.
[0083] The sliding window setting module 201 sets a sliding window of a fixed size and a fixed step size, the step size being smaller than the size of the sliding window.
[0084] The window data generation module 202 divides the time series of each insured person into a fixed number of windows through a sliding window, and calculates the relevant values of the medical data objects in each window, including the total amount of the time window object and the total reimbursement amount.
[0085] The abnormal window data screening module 3 uses the distance between time window objects to divide the data into different data clusters, generating suspected data group data and abnormal data group data, including original dataset generation 301, suspected data group generation 302, and abnormal data group data generation 303.
[0086] The original data generation module 301 focuses all the time window objects of all insured persons on a single point, and maps all the time window objects to a two-dimensional interval according to the two dimensions of the total calculated amount and the total reimbursement amount of the medical treatment data objects within the window, thus forming the original dataset.
[0087] The suspected data group generation module 302 calculates the Euclidean distance between objects in different time windows, divides different data into different data clusters according to the distance between objects and the number of close objects, and assigns free time window objects that do not belong to any data cluster to the suspected data group.
[0088] The abnormal data group generation module 303 calculates the average cost of all time windows in the original dataset, and further evaluates the suspected data group data based on the average cost of the time window and the overall reimbursement limit, listing the data with evaluation scores exceeding a certain threshold as abnormal data group data.
[0089] The suspected clue management module 4 is used to generate suspected insured persons and clue data, suspected medical institutions and clue data, including the suspected insured persons and clue generation module 401 and the suspected insured institutions and clue generation module 402.
[0090] The suspected insured person and clue generation module 401 classifies the abnormal data group data according to the insured persons, calculates the abnormal medical treatment index and abnormal reimbursement index of the insured persons within the abnormal monitoring period, obtains the abnormal value of the insured persons, and obtains the suspected insured persons and clue data. The suspected insured institution and clue generation module 402 classifies the abnormal data group data according to medical institutions, calculates the abnormal index of the settlement medical institution from two dimensions: abnormal drug purchase frequency and abnormal reimbursement ratio, and obtains the suspected medical institution and clue data. Example 3 The purpose of this embodiment is to provide a computer-readable storage medium.
[0091] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the dynamic window-based abnormal centralized reimbursement identification and analysis method described in Example 1.
[0092] Example 4 The purpose of this embodiment is to provide an electronic device.
[0093] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the abnormal centralized reimbursement identification and analysis method based on dynamic windows as described in Embodiment 1.
[0094] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0095] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.
[0096] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for identifying and analyzing abnormal centralized expense reimbursements based on dynamic windows, characterized in that, include: Obtain all medical treatment settlement data of insured individuals within a settlement year, and based on the medical treatment settlement data, create a medical treatment time series for each insured individual; A sliding window calculation method is used to calculate the relevant values of all patient data objects in each window; Analyze the objects in each time window, divide different data into different data clusters, and classify the free time window objects that do not belong to any data cluster into the suspected data group; The suspected data group is evaluated, and data with scores exceeding a certain threshold are listed as abnormal data. Based on the aforementioned abnormal data, the medical treatment abnormality index and reimbursement abnormality index of each insured person, as well as the abnormal transaction frequency index and abnormal reimbursement ratio index of each designated medical institution, are calculated respectively within the abnormal monitoring period. Based on the calculated index values, suspected insured individuals and suspected designated medical institutions are identified, and clue data containing relevant abnormal medical records is generated.
2. The abnormal centralized reimbursement identification and analysis method based on dynamic windows as described in claim 1, characterized in that, The generation of the medical visit time series includes: Obtain data related to insured individuals and their medical treatment, including insured individual information, medical insurance settlement information, and information on designated medical institutions; The acquired data undergoes preliminary processing, forming a medical visit data object for each visit of the insured person. This object includes the settlement time, settlement cost, reimbursement cost, and settlement medical institution. A time series of medical visit information is then generated based on the visit time of the medical visit data object.
3. The abnormal centralized reimbursement identification and analysis method based on dynamic windows as described in claim 1, characterized in that, The analysis of objects within each time window, dividing different data into different data clusters, includes: Treat all insured individuals’ time window objects as a point, and map all time window objects to a two-dimensional interval according to the two dimensions of total calculated amount and total reimbursement amount of medical treatment data objects within the window, to form the original dataset. By calculating the Euclidean distance between objects in different time windows, different data are divided into different data clusters according to the distance between objects and the number of close objects. Objects in free time windows that do not belong to any data cluster are classified into the suspect data group.
4. The abnormal centralized reimbursement identification and analysis method based on dynamic windows as described in claim 3, characterized in that, The step of assigning detached time window objects that do not belong to any data cluster to the suspected data group includes: Set distance range thresholds and intimacy thresholds; Choose any time window object, calculate its Euclidean distance to neighboring time window objects, mark time window objects whose Euclidean distance is less than the distance range threshold as close objects, and mark time window objects as core objects if the number of close objects of the time window object exceeds the closeness threshold. Iterate through all time window objects and mark all core objects. Select any core object to mark it as a data cluster. Add the same data cluster mark to close objects within the distance range threshold. If a close object within the distance range threshold is also a core object, iterate through the marking of that close object until all close objects within the distance range threshold of the core object in the data cluster are included in the data cluster. The size of the data clusters is analyzed. If the number of data objects in a data cluster exceeds a certain threshold, the data in that data cluster is removed from the original dataset. Data is then selected again from the remaining original dataset for data cluster labeling and data removal, until the remaining data no longer meets the requirements. The remaining time window objects are then classified as suspect data.
5. The abnormal centralized reimbursement identification and analysis method based on dynamic windows as described in claim 1, characterized in that, The evaluation of the suspected data group, which classifies data with scores exceeding a certain threshold as abnormal data, includes: Calculate the average cost across all time windows in the original dataset. Evaluate the suspicious data group based on the average cost within each time window and the reimbursement limit. Data with evaluation scores exceeding a certain threshold are classified as abnormal. The method for calculating the abnormality score for the abnormal data group is as follows: in, It is abnormal data group data Abnormal index, , These are weighting coefficients. Settle the fees on an average basis across all time windows. This is the annual reimbursement limit.
6. The abnormal centralized reimbursement identification and analysis method based on dynamic windows as described in claim 1, characterized in that, Based on abnormal data, the medical visit abnormality index and reimbursement abnormality index of each insured person were calculated during the abnormal monitoring period, and suspicious insured persons and clue data were identified, including: The abnormal window data of the interval is classified according to the insured persons. The abnormal time window objects of the insured persons are parsed to obtain the medical treatment data objects corresponding to the abnormal time windows. Duplicate medical treatment time objects are removed to generate a set of abnormal medical treatment data objects of the interval. Calculate the abnormal medical treatment index and abnormal reimbursement index during the abnormal monitoring period to obtain the abnormal value of the insured person; Insured individuals whose abnormal values exceed a certain threshold are marked as suspected cases of centralized reimbursement, and the set of abnormal medical visit times corresponding to these insured individuals is marked as abnormal clue data for centralized reimbursement.
7. The abnormal centralized reimbursement identification and analysis method based on dynamic windows as described in claim 1, characterized in that, Based on the abnormal data, the abnormal transaction frequency index and abnormal reimbursement ratio index of each designated medical institution during the abnormal monitoring period were calculated, and suspected designated medical institutions and clue data were identified, including: Get all the outpatient data objects in all time windows corresponding to the abnormal window data of the interval, remove duplicate outpatient data objects that appear in multiple windows, classify the outpatient data objects according to the settlement medical institutions, and generate several sets of outpatient data of abnormal medical institutions. Obtain all outpatient data objects of medical institutions within the abnormal monitoring period and generate a medical institution outpatient data set; The abnormality index of settlement medical institutions is calculated from two dimensions: abnormal drug purchase frequency and abnormal reimbursement ratio. Medical institutions whose abnormal index exceeds a certain threshold are marked as suspected centralized reimbursement institutions, and the abnormal objects corresponding to these institutions are used as clue data for suspected centralized reimbursement institutions.
8. A dynamic window-based abnormal centralized expense reimbursement identification and analysis system, characterized in that, include: The medical visit sequence generation module is configured to: obtain all medical visit settlement data of insured persons within a settlement year, and generate a medical visit time sequence for each insured person based on the medical visit settlement data; The sliding window management module is configured to use a sliding window calculation method to calculate the relevant values of all patient data objects in each window; The abnormal window data screening module is configured to: analyze the objects in each time window, divide different data into different data clusters, and classify the free time window objects that do not belong to any data cluster into the suspected data group. The suspected data group is evaluated, and data with scores exceeding a certain threshold are listed as abnormal data. The suspected clue management module is configured to: calculate the medical treatment abnormality index and reimbursement abnormality index of each insured person and the abnormal transaction frequency index and abnormal reimbursement ratio index of each designated medical institution during the abnormal monitoring period based on the abnormal data; identify suspected insured persons and suspected designated medical institutions according to the calculated index values, and generate clue data containing relevant abnormal medical records.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the abnormal centralized reimbursement identification and analysis method based on dynamic windows as described in any one of claims 1-7.
10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the abnormal centralized reimbursement identification and analysis method based on dynamic windows as described in any one of claims 1-7.