Medical insurance fund supervision-oriented medium and high risk clue generation method and system

By preprocessing and cleaning medical insurance business data, combining multi-dimensional rule supervision and similarity scoring, and using the density-hierarchical clustering algorithm for adaptive risk assessment, the problems of low data processing efficiency and insufficient adaptability in the existing medical insurance fund supervision model are solved, and the accurate generation of medium and high-risk clues is achieved.

CN120655436APending Publication Date: 2025-09-16SUN YAT SEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510923035.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

The existing medical insurance fund supervision model is inefficient in data cleaning, feature construction and anomaly identification. It lacks in-depth exploration of group collaborative behavior, lacks mutual verification between individual anomalies and group anomalies, and the clue generation model relies on expert experience to set evaluation rules, lacking adaptability and versatility.

Method used

A strategy combining multi-dimensional rule supervision and similarity scoring is adopted. By preprocessing and cleaning the medical insurance business data, the individual anomaly index and the same-frequency card similarity scoring method are used to mine abnormal cards, and the density-hierarchical clustering algorithm is combined for adaptive risk assessment to generate medium and high-risk clues.

Benefits of technology

It has achieved accurate mining of high-risk clues in medical insurance fund supervision, overcome the problem of difficulty in identifying group collaborative criminal behavior in existing technologies, and improved the adaptability and accuracy of risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655436A_ABST
    Figure CN120655436A_ABST
Patent Text Reader

Abstract

The invention discloses a medical insurance fund supervision-oriented medium-and-high-risk clue generation method and a medical insurance fund supervision-oriented medium-and-high-risk clue generation system. The method comprises the following steps: preprocessing medical insurance business data, and outputting cleaned medical insurance settlement data; determining an individual anomaly index based on a preset supervision rule threshold by using the cleaned medical insurance settlement data, and outputting an individual anomaly card set; mining abnormal cards by using the individual abnormal card set in combination with a multi-dimensional rule supervision and same-frequency card similarity scoring method to obtain a final abnormal card set; and performing risk clustering division on the final abnormal card set by using an adaptive risk assessment method based on a density-hierarchical clustering algorithm to obtain classification thresholds of different risk levels, and extracting medium-high risk groups as medium-high risk clue assessment results by using the classification thresholds. According to the scheme, the dependence on expert experience is effectively reduced while accurate identification of the group abnormal medical insurance card is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of big data information processing, and more specifically, relates to a method and system for generating medium- and high-risk clues for medical insurance fund supervision. Background Art

[0002] As the scale of medical insurance funds continues to expand, their supervision faces unprecedented challenges. Traditional supervisory methods, such as large-scale human resources and manual review, are unable to adapt to today's complex and ever-changing regulatory needs. To address these challenges, it is necessary to apply big data and artificial intelligence technologies to build efficient medical insurance legal supervision models. These technologies can rapidly analyze massive amounts of data, improving the efficiency and accuracy of legal supervision.

[0003] Existing legal supervision models still face numerous shortcomings in their application to the healthcare sector. For one thing, existing methods are inefficient in data cleaning, feature construction, and anomaly identification for complex, multi-source healthcare business data, making it difficult to quickly extract high-quality legal supervision leads. Furthermore, existing lead generation models primarily focus on detecting individual anomalous behaviors and lack the ability to deeply explore coordinated group behavior, making it difficult to identify group anomalous leads, such as drug trafficking rings. Furthermore, the lack of cross-verification and dynamic feedback mechanisms between individual and group anomalies leads to isolated leads and high false positive rates. Furthermore, existing risk assessment models fail to fully utilize relevant information after generating leads, generally relying on expert experience to set assessment rules and thresholds. This leads to a high degree of subjectivity, making it difficult to adapt to the risk identification needs of diverse regions, healthcare policies, or regulatory systems. Their lack of versatility and adaptability severely hinders the effectiveness of legal supervision.

[0004] The invention patent with the prior art publication number CN118378201A proposes a method and device for detecting abnormal behavior in a medical insurance group. The method includes: obtaining medical insurance settlement data and extracting key features of medical insurance settlement after preprocessing; constructing three-dimensional dynamic graphs based on the key features, namely a dynamic graph representing the number of drug purchases, a dynamic graph representing the diversity of drug types, and a dynamic graph representing the amount of drug purchases; fusing the three-dimensional dynamic graphs to obtain a fusion graph, and performing a community discovery search for the medical insurance group on the fusion graph, and constructing a community graph based on the searched community set; using an autoencoder to encode and decode the community graph, and calculating the reconstruction error of the community graph as an anomaly indicator to screen abnormal communities in the community graph and realize the detection of abnormal behavior in the medical insurance group. This scheme lacks rule pre-screening, has a high risk of false alarms, and the threshold still relies on experience and lacks objective self-adaptation. The autoencoder of this scheme needs to manually set the reconstruction error threshold, relies on experience, and lacks objective self-adaptation. Summary of the Invention

[0005] The present invention aims to overcome the problems existing in the existing supervision methods based on medical insurance business data, including the insufficient efficiency of mining complex and diverse business data, the lack of effective methods to mine group anomalies, the insufficient utilization of the generated clue-related information, and the high reliance of risk assessment on thresholds preset by experts. The present invention provides a method and system for generating medium- and high-risk clues for medical insurance fund supervision.

[0006] The primary purpose of the present invention is to solve the above technical problems, and the technical solutions of the present invention are as follows: A first aspect of the present invention provides a method for generating medium- and high-risk clues for medical insurance fund supervision, comprising the following steps: Pre-process medical insurance business data and output cleaned medical insurance settlement data; Using the cleaned medical insurance settlement data, based on a preset supervision rule threshold, an individual anomaly index is determined, and an individual anomaly card set is output; The individual abnormal card set is combined with multi-dimensional rule supervision and the same-frequency card similarity scoring method to mine abnormal cards and obtain the final abnormal card set; The final abnormal card set is divided into risk clusters using an adaptive risk assessment method based on a density-hierarchical clustering algorithm to obtain classification thresholds of different risk levels, and the classification thresholds are used to extract medium and high risk groups as medium and high risk clue assessment results.

[0007] Furthermore, pre-processing of the medical insurance business data includes performing outlier cleaning and sensitive field desensitization on the medical insurance business data, and performing outlier cleaning on the medical insurance business data includes: When a field is empty, delete it if it is relevant only to a few patients; When the field value is 0, it is retained if it is related to the amount, otherwise it is deleted; When a record is empty, delete it or fill it with the mean or mode; Only one identical record is retained; Restore the correct position of misplaced records; Records with a small number of different fields will be retained as long as there is no conflict; For outliers, delete records whose values ​​exceed n times the standard deviation of the field mean; In conflict detection, verify whether numerical data is within a reasonable range, including age and amount; Sensitive fields are desensitized to ensure that the processed data cannot be traced back to the patient or institution's personal identity information, including: Map fields such as medical insurance card number and institution code to random strings and retain the mapping table; The patient's name is retained only by the last name, and the ID number is retained only by the first 6 and 7-10 digits to extract the region and age; Delete the home address, unit name, and unit address fields; The medical institution field retains the region and hospital category information.

[0008] Among the records that have undergone the cleaning and desensitization processing, records of medical-related fields are retained, and the cleaned medical insurance settlement data is output.

[0009] Furthermore, the cleaned medical insurance settlement data is used to determine the individual abnormality index based on the preset supervision rule threshold, and the individual abnormality card set is output, including the following steps: Set corresponding supervision rule thresholds under the three supervision dimensions of medical frequency, medical expenses, and medical behavior. Calculate each rule of the medical insurance business data based on the supervision rule threshold, and calculate the violation score of each rule; The number of times a patient hits each supervision rule is counted, and weighted aggregation is performed based on the weight of each rule to calculate the individual abnormality index of each patient; Compare the individual abnormality index with the set threshold, filter out individual medical insurance cards whose abnormality index exceeds the threshold, and output the individual abnormal card set.

[0010] Furthermore, the individual abnormal card set is combined with multi-dimensional rule supervision and the same-frequency card similarity scoring method to mine abnormal cards, and the final abnormal card set is obtained, including the following steps: Using the individual abnormal card set, combined with the preset time threshold and medical institution information, we can determine the same-frequency card set that has settlement behavior at the same medical institution within the preset time window. The expression for determining the same-frequency card set is as follows:

[0011] Among them, the adjacent settlement time is determined by the preset time threshold control, Represents a collection of medical institutions. Represents the settlement time set , represents a medical institution, i and j represent any two different settlement record numbers, 、 Respectively represent the settlement time of settlement record i and settlement record j, 、 represent the medical institutions of settlement record i and settlement record j respectively; Eliminate normal cards whose settlement amount and drug expense ratio do not meet the preset threshold in the same-frequency card set, and output a set of abnormal same-frequency cards to be identified; The same-frequency card set to be identified as abnormal is scored and screened using the same-frequency card similarity scoring method to obtain the final abnormal card set.

[0012] Furthermore, the abnormal same-frequency card set to be identified is scored and screened using the same-frequency card similarity scoring method to obtain the final abnormal card set, including the following steps: The similarity of medical information of the identified abnormal same-frequency candidate cards within n years is calculated from the time dimension and content dimension respectively. The comprehensive similarity score is calculated by constructing a similarity function and combining it with the weighted coefficient. , The expression is as follows:

[0013] Among them, A and B represent two different medical insurance cards. is the score of the i-th content similarity dimension, is the corresponding weight, N is the total number of dimensions, and t represents the settlement time interval. Indicates the i-th content dimension within the time; The comprehensive similarity score is compared with the abnormality judgment threshold. For the set of suspected abnormal same-frequency candidate cards to be identified that exceeds the preset abnormality judgment threshold, if the comprehensive similarity score between any candidate card and a card in the individual abnormal card set is higher than the preset abnormality judgment threshold, the set of suspected abnormal same-frequency candidate cards to be identified is included in the abnormal same-frequency card set, and the abnormal same-frequency card set is output as the final abnormal card set.

[0014] Furthermore, the final abnormal card set is divided into risk clusters using an adaptive risk assessment method based on a density-hierarchical clustering algorithm to obtain classification thresholds of different risk levels, and the classification thresholds are used to extract medium and high risk groups as medium and high risk clue assessment results, including the following steps: Objectively weight the abnormal card set based on the entropy weight method to calculate the abnormal index; Using the TOPSIS method to determine the relative closeness of the abnormal index, and output the risk aggregation index; The risk aggregation indicators are clustered using the FLASC density-hierarchical clustering algorithm to obtain classification thresholds for different risk levels, and the classification thresholds are used to extract medium and high risk groups as medium and high risk clue assessment results.

[0015] Furthermore, the final abnormal card set is objectively weighted based on the entropy weight method to calculate the abnormal index, which includes the following steps: The information entropy is calculated for the indicator values ​​corresponding to each supervision rule. The expression is as follows:

[0016] Among them, m represents the total number of medical insurance cards, represents the value in the jth column of the i-th medical insurance card, represents the entropy value of the j-th supervision rule; Based on the information entropy value of each indicator, the entropy weight method is used to calculate the weight of each supervision rule. The expression is as follows:

[0017] in, is the indicator aggregation weight of the jth supervision rule, and n represents the total number of supervision rules; Perform weighted summation of the indicator values ​​and corresponding weights under each supervision rule dimension, and output the abnormality index corresponding to medical insurance card i , the expression is as follows: .

[0018] Furthermore, the TOPSIS method is used to determine the relative closeness of the abnormality index and output a risk aggregation index, including the following steps: According to the maximum and minimum values ​​of each supervision rule indicator, the ideal solution and negative ideal solution are constructed, which represent the theoretical optimal state and the worst state respectively; Calculate the weighted Euclidean distance between the anomaly score of each medical insurance card and the ideal solution , the expression is as follows:

[0019] Where n is the total number of indicators, is the weight of the j-th indicator, is the distance between the value of the i-th medical insurance card on the j-th indicator and the ideal solution; Calculate the weighted Euclidean distance between the anomaly score of each medical insurance card and the negative ideal solution , the expression is as follows:

[0020] in, is the distance between the value of the i-th medical insurance card on the j-th indicator and the negative ideal solution; Combine the positive and negative ideal solution distances to calculate the relative closeness of each medical insurance card , as a risk aggregation indicator, the expression is as follows: .

[0021] Furthermore, the risk aggregation indicators are clustered using the FLASC density-hierarchical clustering algorithm to obtain classification thresholds for different risk levels, and medium- and high-risk groups are extracted using the classification thresholds as medium- and high-risk clue assessment results, including the following steps: Based on the risk aggregation index of the medical insurance card, a high-density data point set is constructed, and high-density areas are identified and divided into initial clustering units; For each cluster unit, the eccentricity of each data point relative to its cluster centroid is calculated to describe the topological position of the point in the cluster structure. Points with low eccentricity are usually located in the core area, and points with high eccentricity are more likely to be at the edge of the cluster or at the end of the branch. The expression is as follows:

[0022] Among them, i is the data point number, j is the cluster number, is the data point, The cluster to which it belongs The center of mass, represents the distance metric function; Two types of approximate graphs are constructed based on eccentricity and distance between points, including: Full approximate graph: retains the connection relationship between all points whose distance is no greater than the maximum edge length in the minimum spanning tree of the cluster, which is used to describe the complete topological structure; Core approximation graph: only retains the connection relationship between points whose distance does not exceed the maximum value of their respective core distances, which is used to extract the cluster core substructure; Based on the structure diagram, a hierarchical aggregation strategy is adopted to merge adjacent data clusters step by step to build a complete risk level hierarchical structure, and adaptively divide the classification thresholds of low-risk clues and medium- and high-risk clues; Using the classification threshold, medium and high risk groups are extracted as medium and high risk clue assessment results.

[0023] The second aspect of the present invention provides a medium- and high-risk clue generation system for medical insurance fund supervision, including a memory and a processor. The memory includes a medium- and high-risk clue generation method program for medical insurance fund supervision. When the medium- and high-risk clue generation method program for medical insurance fund supervision is executed by the processor, it implements the steps of a medium- and high-risk clue generation method for medical insurance fund supervision.

[0024] Compared with the prior art, the beneficial effects of the technical solution of the present invention are: This paper proposes a novel method for detecting abnormal medical insurance cards in groups. This method combines multi-dimensional rule supervision with similarity scoring to accurately identify and distinguish abnormal medical insurance cards, overcoming the difficulty in identifying and distinguishing coordinated group behavior in existing technologies. By introducing a risk assessment method based on density-hierarchical clustering, the abnormal card collection is adaptively clustered into risk categories, generating classification thresholds corresponding to different risk levels. This effectively addresses the existing technical issues of relying heavily on expert experience and insufficient utilization of clue-related information. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to make the purpose and technical solution of the present invention clearer, the present invention provides the following drawings and descriptions: Figure 1 A flow chart of a method provided by an embodiment of the present invention; Figure 2 This is a flowchart for evaluating medium- and high-risk clues provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0026] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present application and the features therein can be combined with each other.

[0027] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.

[0028] Example 1: The present invention provides a method for generating medium and high risk clues for medical insurance fund supervision, such as Figure 1 The following is a flowchart of a method for generating medium- and high-risk leads for medical insurance fund supervision. The specific steps are as follows: S1: Preprocess the medical insurance business data and output the cleaned medical insurance settlement data.

[0029] The preprocessing of medical insurance business data includes performing outlier cleaning and sensitive field desensitization on the medical insurance business data. The outlier cleaning of medical insurance business data is shown in Table 1, including: When a field is empty, delete it if it is relevant only to a few patients; When the field value is 0, it is retained if it is related to the amount, otherwise it is deleted; When a record is empty, delete it or fill it with the mean or mode; Only one identical record is retained; Records with a small number of different fields will be retained as long as there is no conflict; For outliers, delete records whose values ​​exceed n (in this example, n is 5) times the standard deviation of the field mean; In conflict detection, verify whether numerical data is within a reasonable range, including age and amount; Table 1

[0030] Restoring the correct location of misplaced records includes: For field misalignment, where some field values ​​appear in the wrong column, structural consistency is verified by checking delimiters (such as the number of delimiters in CSV format), validating the field format using regular expression rules, and repositioning and correcting fields based on known field format patterns (such as date and ID code structures). In the case of an entire row being misplaced, that is, a row of data is inserted at the wrong position, continuity detection of record identifiers (such as primary key ID or business index) is used to identify the abnormal row position; In the case of partial data misalignment, that is, a record is incorrectly split to the end of the previous line or the beginning of the next line, by analyzing features such as abnormal number of fields and record length deviation, combined with sliding window technology to match and identify potential split locations, cross-row reconstruction can be achieved.

[0031] For severely misaligned data whose structure cannot be effectively restored, they are directly cleared to ensure the integrity and accuracy of the structure after data cleaning.

[0032] Sensitive fields are desensitized to ensure that the processed data cannot be traced back to the patient or institution's personal identity information, as shown in Table 2, including: Map fields such as medical insurance card number and institution code to random strings and keep the mapping table.

[0033] Only the last name of the patient is retained, and only the first 6 and 7-10 digits of the ID number are retained to extract the region and age.

[0034] Delete the Home Address, Work Name, and Work Address fields.

[0035] The medical institution field retains the region and hospital category information.

[0036] Table 2

[0037] Among the records after the cleaning and desensitization processing, the records of medical-related fields (including drug items and drug payments) will be retained, and the cleaned medical insurance settlement data will be output.

[0038] S2: Using the cleaned medical insurance settlement data, based on a preset supervision rule threshold, determine the individual anomaly index and output an individual anomaly card set.

[0039] The specific process is: Corresponding supervision rule thresholds are set under the three supervision dimensions of medical frequency, medical expenses, and medical behavior, as shown in Table 3. The supervision rule thresholds include but are not limited to: a) Under the frequency dimension, set rules such as "more than 15 outpatient visits per month", "more than 4 outpatient visits per day for 3 or more days per month", and "100 or more outpatient visits per year".

[0040] b) Under the medical expenses dimension, set rules such as "monthly outpatient expenses cumulatively exceeding 5,000 yuan", "annual outpatient and emergency expenses cumulatively exceeding 25,000 yuan", and "annual medical insurance total settlement expenses exceeding 30,000 yuan with drug costs accounting for more than 80% and laboratory test costs accounting for less than 10%".

[0041] c) Under the dimension of medical treatment behavior, set rules such as "settling the same drug at multiple medical institutions within one week" and "settling more than 10 types of drugs within one week".

[0042] Table 3

[0043] Based on the supervision rule threshold, the medical insurance business data (medical records, expense data) is calculated rule by rule, and the violation score of each rule is counted.

[0044] The number of times a patient hits each supervision rule is counted, and weighted aggregation is performed based on the weight of each rule to calculate the individual abnormality index of each patient.

[0045] The individual abnormality index is compared with a set threshold (in this embodiment, the threshold is 0.75), individual medical insurance cards whose abnormality index exceeds the threshold are screened out, and an individual abnormal card set is output.

[0046] S3: Use the individual abnormal card set combined with multi-dimensional rule supervision and the same-frequency card similarity scoring method to mine abnormal cards and obtain the final abnormal card set.

[0047] The specific process is: Using the individual abnormal card set, combined with the preset time threshold and medical institution information, we can determine the same-frequency card set that has settlement behavior at the same medical institution within the preset time window. The expression for determining the same-frequency card set is as follows:

[0048] Among them, the adjacent settlement time is determined by the preset time threshold Control (in this embodiment, the time threshold is set to the settlement time difference within ten minutes), Represents a collection of medical institutions. Represents the settlement time set , represents a medical institution, i and j represent any two different settlement record numbers, 、 Respectively represent the settlement time of settlement record i and settlement record j, 、 Represent the medical institutions of settlement record i and settlement record j respectively.

[0049] Normal cards in the same-frequency card set whose settlement amount and drug expense ratio do not meet the preset thresholds (in this embodiment, the threshold values ​​are settlement amount greater than RMB 30,000 and drug expense ratio greater than 0.8) are excluded, and a set of abnormal same-frequency cards to be identified is output.

[0050] The abnormal same-frequency card set to be identified is scored and screened using the same-frequency card similarity scoring method to obtain the final abnormal card set, including the following steps: For the set of abnormal same-frequency candidate cards to be identified, the similarity of medical information within n years (in this embodiment, n is 1) is calculated from the time dimension (such as settlement time interval) and content dimension (such as drug type, amount, frequency of use, etc.) between the same-frequency candidate cards. By constructing a similarity function and combining it with the weighting coefficient, the comprehensive similarity score is calculated. , The expression is as follows:

[0051] Among them, A and B represent two different medical insurance cards. is the score of the i-th content similarity dimension, is the corresponding weight, N is the total number of dimensions (dimensions include the type, amount, and frequency of use of medical insurance items), t represents the settlement time interval, Represents the i-th content dimension (drug type, amount, frequency of use) within that time.

[0052] The calculated comprehensive similarity score is compared with the anomaly judgment threshold (in this embodiment, the anomaly judgment threshold is 0.6). For the same-frequency candidate card set that exceeds the anomaly judgment threshold, if the comprehensive score between an individual abnormal card in the individual abnormal card set and a candidate card in the same-frequency candidate card set is higher than the anomaly judgment threshold, then the same-frequency candidate card set is included in the abnormal same-frequency card set.

[0053] Through the above iterative strategy, the set of same-frequency cards with similar input abnormal behavior patterns is continuously expanded to achieve complete coverage of group abnormal cards and obtain the final abnormal card set.

[0054] S4: Use the adaptive risk assessment method based on density-hierarchical clustering algorithm to perform risk clustering on the final abnormal card set to obtain classification thresholds of different risk levels, and use the classification thresholds to extract medium and high risk groups as medium and high risk clue assessment results, such as Figure 2 shown.

[0055] The specific process is: The abnormal card set is objectively weighted based on the entropy weight method to calculate the abnormal index, including the following steps: The information entropy is calculated for the indicator values ​​corresponding to each supervision rule to quantify the amount of information and uncertainty of each indicator. The more evenly the distribution of an indicator value is, the higher its entropy value is, indicating that the discrimination of the indicator is weaker, and vice versa. The expression is as follows:

[0056] Among them, m represents the total number of medical insurance cards, represents the value in the jth column of the i-th medical insurance card, represents the entropy value of the j-th supervision rule; Based on the information entropy value of each indicator, the entropy weight method is used to calculate the weight of each supervision rule. The expression is as follows:

[0057] in, is the indicator aggregation weight of the jth supervision rule, and n represents the total number of supervision rules; Perform weighted summation of the indicator values ​​and corresponding weights under each supervision rule dimension, and output the abnormality index corresponding to medical insurance card i , the expression is as follows: .

[0058] The TOPSIS method is used to determine the relative closeness of the abnormal index and output the risk aggregation index, including the following steps: According to the maximum and minimum values ​​of each supervision rule indicator, the ideal solution and negative ideal solution are constructed, which represent the theoretical optimal state and the worst state respectively; Calculate the weighted Euclidean distance between the anomaly score of each medical insurance card and the ideal solution , the expression is as follows:

[0059] Where n is the total number of indicators, is the weight of the j-th indicator, is the distance between the value of the i-th medical insurance card on the j-th indicator and the ideal solution; Calculate the weighted Euclidean distance between the anomaly score of each medical insurance card and the negative ideal solution , the expression is as follows:

[0060] in, is the distance between the value of the i-th medical insurance card on the j-th indicator and the negative ideal solution; Combine the positive and negative ideal solution distances to calculate the relative closeness of each medical insurance card , as a risk aggregation indicator, the expression is as follows: .

[0061] The risk aggregation indicators are clustered using the FLASC density-hierarchical clustering algorithm to obtain classification thresholds for different risk levels. The classification thresholds are used to extract medium and high risk groups as medium and high risk clue assessment results, including the following steps: Based on the risk aggregation index of the medical insurance card, a high-density data point set is constructed, and high-density areas are identified and divided into initial clustering units; For each cluster unit, the eccentricity of each data point relative to its cluster centroid is calculated to describe the topological position of the point in the cluster structure. Points with low eccentricity are usually located in the core area, and points with high eccentricity are more likely to be at the edge of the cluster or at the end of the branch. The expression is as follows:

[0062] Among them, i is the data point number, j is the cluster number, is the data point, The cluster to which it belongs The center of mass, represents the distance metric function; Two types of approximate graphs are constructed based on eccentricity and distance between points, including: Full approximate graph: retains the connection relationship between all points whose distance is no greater than the maximum edge length in the minimum spanning tree of the cluster, which is used to describe the complete topological structure; Core approximation graph: only retains the connection relationship between points whose distance does not exceed the maximum value of their respective core distances, which is used to extract the cluster core substructure; Based on the structure diagram, a hierarchical aggregation strategy is adopted to merge adjacent data clusters step by step to build a complete risk level hierarchical structure, and adaptively divide the classification thresholds of low-risk clues and medium- and high-risk clues; Using the classification threshold, medium and high risk groups are extracted as medium and high risk clue assessment results.

[0063] This paper proposes a novel method for detecting abnormal medical insurance cards in groups. This method combines multi-dimensional rule supervision with similarity scoring to accurately identify and distinguish abnormal medical insurance cards, overcoming the difficulty in identifying and distinguishing coordinated group behavior in existing technologies. By introducing a risk assessment method based on density-hierarchical clustering, the abnormal card collection is adaptively clustered into risk categories, generating classification thresholds corresponding to different risk levels. This effectively addresses the existing technical issues of relying heavily on expert experience and insufficient utilization of clue-related information.

[0064] Example 2: This embodiment provides a medium- and high-risk clue generation system for medical insurance fund supervision, including a memory and a processor. The memory includes a medium- and high-risk clue generation method program for medical insurance fund supervision. When the medium- and high-risk clue generation method program for medical insurance fund supervision is executed by the processor, the steps of a medium- and high-risk clue generation method for medical insurance fund supervision as described in Example 1 are implemented.

[0065] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A method for generating medium- and high-risk clues for medical insurance fund supervision, characterized by: The following steps are involved: Pre-process medical insurance business data and output cleaned medical insurance settlement data; Using the cleaned medical insurance settlement data, based on a preset supervision rule threshold, an individual anomaly index is determined, and an individual anomaly card set is output; The individual abnormal card set is combined with multi-dimensional rule supervision and the same-frequency card similarity scoring method to mine abnormal cards and obtain the final abnormal card set; The final abnormal card set is divided into risk clusters using an adaptive risk assessment method based on a density-hierarchical clustering algorithm to obtain classification thresholds of different risk levels, and the classification thresholds are used to extract medium and high risk groups as medium and high risk clue assessment results.

2. A method for generating medium- and high-risk clues for medical insurance fund supervision according to claim 1, characterized in that: Preprocessing of medical insurance business data includes performing outlier cleaning and sensitive field desensitization on medical insurance business data. Performing outlier cleaning on medical insurance business data includes: When a field is empty, delete it if it is relevant only to a few patients; When the field value is 0, it is retained if it is related to the amount, otherwise it is deleted; When a record is empty, delete it or fill it with the mean or mode; Only one identical record is retained; Restore the correct position of misplaced records; Records with a small number of different fields will be retained as long as there is no conflict; For outliers, delete records whose values ​​exceed n times the standard deviation of the field mean; In conflict detection, verify whether numerical data is within a reasonable range, including age and amount; Sensitive fields are desensitized to ensure that the processed data cannot be traced back to the patient or institution's personal identity information, including: Map fields such as medical insurance card number and institution code to random strings and retain the mapping table; The patient's name is retained only by the last name, and the ID number is retained only by the first 6 and 7-10 digits to extract the region and age; Delete the home address, unit name, and unit address fields; The medical institution field retains the region and hospital category information. Among the records that have undergone the cleaning and desensitization processing, records of medical-related fields are retained, and the cleaned medical insurance settlement data is output.

3. The method for generating medium- and high-risk clues for medical insurance fund supervision according to claim 1 is characterized in that: Using the cleaned medical insurance settlement data, based on the preset supervision rule threshold, the individual anomaly index is determined and the individual anomaly card set is output, including the following steps: Set corresponding supervision rule thresholds under the three supervision dimensions of medical frequency, medical expenses, and medical behavior. Calculate each rule of the medical insurance business data based on the supervision rule threshold, and calculate the violation score of each rule; The number of times a patient hits each supervision rule is counted, and weighted aggregation is performed based on the weight of each rule to calculate the individual abnormality index of each patient; Compare the individual abnormality index with the set threshold, filter out individual medical insurance cards whose abnormality index exceeds the threshold, and output the individual abnormal card set.

4. The method for generating medium- and high-risk clues for medical insurance fund supervision according to claim 1 is characterized in that: The individual abnormal card set is combined with multi-dimensional rule supervision and the same-frequency card similarity scoring method to mine abnormal cards and obtain the final abnormal card set. The following steps are involved: Using the individual abnormal card set, combined with the preset time threshold and medical institution information, we can determine the same-frequency card set that has settlement behavior at the same medical institution within the preset time window. The expression for determining the same-frequency card set is as follows: Among them, the adjacent settlement time is determined by the preset time threshold control, Represents a collection of medical institutions. Represents the settlement time set , represents a medical institution, i and j represent any two different settlement record numbers, 、 Respectively represent the settlement time of settlement record i and settlement record j, 、 represent the medical institutions of settlement record i and settlement record j respectively; Eliminate normal cards whose settlement amount and drug expense ratio do not meet the preset threshold in the same-frequency card set, and output a set of suspected abnormal same-frequency cards to be identified; The same-frequency card set to be identified as suspected abnormal is scored and screened using the same-frequency card similarity scoring method to obtain the final abnormal card set.

5. The method for generating medium- and high-risk clues for medical insurance fund supervision according to claim 4 is characterized in that: The same-frequency card set to be identified as suspected abnormal is scored and screened using the same-frequency card similarity scoring method to obtain the final abnormal card set, including the following steps: For the set of suspected abnormal same-frequency candidate cards, the similarity of medical information within n years between the same-frequency candidate cards is calculated from the time dimension and content dimension respectively. The comprehensive similarity score is calculated by constructing a similarity function and combining it with the weighted coefficient. , The expression is as follows: Among them, A and B represent two different medical insurance cards. is the score of the i-th content similarity dimension, is the corresponding weight, N is the total number of dimensions, and t represents the settlement time interval. Indicates the i-th content dimension within the time; The comprehensive similarity score is compared with the abnormality judgment threshold. For the set of suspected abnormal same-frequency candidate cards to be identified that exceeds the preset abnormality judgment threshold, if the comprehensive similarity score between any candidate card and a card in the individual abnormal card set is higher than the preset abnormality judgment threshold, the set of suspected abnormal same-frequency candidate cards to be identified is included in the abnormal same-frequency card set, and the abnormal same-frequency card set is output as the final abnormal card set.

6. The method for generating medium- and high-risk clues for medical insurance fund supervision according to claim 1 is characterized in that: The final abnormal card set is divided into risk clusters using an adaptive risk assessment method based on a density-hierarchical clustering algorithm to obtain classification thresholds for different risk levels. The classification thresholds are used to extract medium and high risk groups as medium and high risk clue assessment results, including the following steps: Objectively weight the abnormal card set based on the entropy weight method to calculate the abnormal index; Using the TOPSIS method to determine the relative closeness of the abnormal index, and output the risk aggregation index; The risk aggregation indicators are clustered using the FLASC density-hierarchical clustering algorithm to obtain classification thresholds for different risk levels, and the classification thresholds are used to extract medium and high risk groups as medium and high risk clue assessment results.

7. The method for generating medium- and high-risk clues for medical insurance fund supervision according to claim 6 is characterized in that: The final abnormal card set is objectively weighted based on the entropy weight method to calculate the abnormal index, including the following steps: The information entropy is calculated for the indicator values ​​corresponding to each supervision rule. The expression is as follows: Among them, m represents the total number of medical insurance cards, represents the value in the jth column of the i-th medical insurance card, represents the entropy value of the j-th supervision rule; Based on the information entropy value of each indicator, the entropy weight method is used to calculate the weight of each supervision rule. The expression is as follows: in, is the indicator aggregation weight of the jth supervision rule, and n represents the total number of supervision rules; Perform weighted summation of the indicator values ​​and corresponding weights under each supervision rule dimension, and output the abnormality index corresponding to medical insurance card i , the expression is as follows: 。 8. The method for generating medium- and high-risk clues for medical insurance fund supervision according to claim 6 is characterized in that: The TOPSIS method is used to determine the relative closeness of the abnormal index and output the risk aggregation index, including the following steps: According to the maximum and minimum values ​​of each supervision rule indicator, the ideal solution and negative ideal solution are constructed, which represent the theoretical optimal state and the worst state respectively; Calculate the weighted Euclidean distance between the anomaly score of each medical insurance card and the ideal solution , the expression is as follows: Where n is the total number of indicators, is the weight of the j-th indicator, is the distance between the value of the i-th medical insurance card on the j-th indicator and the ideal solution; Calculate the weighted Euclidean distance between the anomaly score of each medical insurance card and the negative ideal solution , the expression is as follows: in, is the distance between the value of the i-th medical insurance card on the j-th indicator and the negative ideal solution; Combine the positive and negative ideal solution distances to calculate the relative closeness of each medical insurance card , as a risk aggregation indicator, the expression is as follows: 。 9. The method for generating medium- and high-risk clues for medical insurance fund supervision according to claim 6 is characterized in that: The risk aggregation indicators are clustered using the FLASC density-hierarchical clustering algorithm to obtain classification thresholds for different risk levels. The classification thresholds are used to extract medium and high risk groups as medium and high risk clue assessment results, including the following steps: Based on the risk aggregation index of the medical insurance card, a high-density data point set is constructed, and high-density areas are identified and divided into initial clustering units; For each cluster unit, the eccentricity of each data point relative to its cluster centroid is calculated to describe the topological position of the point in the cluster structure. Points with low eccentricity are usually located in the core area, and points with high eccentricity are more likely to be at the edge of the cluster or at the end of the branch. The expression is as follows: Among them, i is the data point number, j is the cluster number, is the data point, The cluster to which it belongs The center of mass, represents the distance metric function; Two types of approximate graphs are constructed based on eccentricity and distance between points, including: Full approximate graph: retains the connection relationship between all points whose distance is no greater than the maximum edge length in the minimum spanning tree of the cluster, which is used to describe the complete topological structure; Core approximation graph: only retains the connection relationship between points whose distance does not exceed the maximum value of their respective core distances, which is used to extract the cluster core substructure; Based on the structure diagram, a hierarchical aggregation strategy is adopted to merge adjacent data clusters step by step to build a complete risk level hierarchical structure, and adaptively divide the classification thresholds of low-risk clues and medium- and high-risk clues; Using the classification threshold, medium and high risk groups are extracted as medium and high risk clue assessment results.

10. A medium- and high-risk clue generation system for medical insurance fund supervision, characterized by: The system includes: a memory and a processor, wherein the memory includes a method program for generating medium- and high-risk clues for medical insurance fund supervision. When the method program for generating medium- and high-risk clues for medical insurance fund supervision is executed by the processor, the steps of a method for generating medium- and high-risk clues for medical insurance fund supervision as described in any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Medical insurance group abnormal behavior detection method and device

    CN118378201A