Methods, devices, equipment, and media for identifying medical insurance fraud based on neighborhood similarity

CN115829760BActive Publication Date: 2026-09-01XIAMEN UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211488104.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-25
Publication Date
2026-09-01
Estimated Expiration
2042-11-25

AI Technical Summary

Technical Problem

这种情况传统的医保欺诈检测方法无法充分利用用户之间的交互关系,导致难以正确检测出欺诈行为

Benefits of technology

本发明实施例的医保欺诈识别方法通过异构图将患者的行为转化为计算机可以识别并处理的数据,通过采样得到不同行为模式的数据,通过注意力机制聚合邻居节点的信息和元路径的信息,减少了噪声节点和低相关元路径干扰,能够让取得更能表达患者行为的最终嵌入表示,从而大大提高后续判断的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115829760B_ABST
    Figure CN115829760B_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, device, and medium for identifying medical insurance fraud based on neighborhood similarity, relating to the field of medical big data technology. The medical insurance fraud identification method includes: S1, constructing a medical heterogeneous graph based on medical data; S2, sampling the meta-paths of various behavioral patterns to obtain a heterogeneous sub-graph; S3, encoding the heterogeneous sub-graph to obtain an initial neighborhood set; S4, calculating and filtering the similarity of each neighborhood based on the initial neighborhood set to obtain a final neighborhood set; S5, fusing the final neighborhood sets using a first attention mechanism to obtain the embedding representation of each patient node under each behavioral pattern; S6, determining the importance of various behavioral patterns based on the embedding representations; S7, fusing the embedding representations based on importance using a second attention mechanism to obtain the final embedding representation of each patient node; and S8, classifying the final embedding representations to determine whether each patient node is a patient involved in medical insurance fraud.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical big data technology, and more specifically, to a method, apparatus, device, and medium for identifying medical insurance fraud based on neighborhood similarity. Background Technology

[0002] The widespread availability of medical insurance has provided healthcare security for the public. However, while offering convenience, it has also created new avenues for fraud. Medical insurance fraud takes many forms, such as issuing false invoices or fake receipts to insured individuals, illegally using medical insurance credentials, obtaining medications and supplies through deception, and reselling them for illegal profit. Medical insurance fraud severely harms the interests of policyholders, making it an urgent task to uncover potential fraudsters within the complex medical insurance data.

[0003] Traditional methods for identifying healthcare fraud include rule-based methods, supervised learning methods, and unsupervised learning methods. Rule-based methods require domain experts to analyze past fraudulent activities to construct possible fraud patterns and establish corresponding rules to screen for suspicious behavior. Supervised learning methods treat fraud as a binary classification problem, training a fraud classifier to distinguish fraudulent behavior. Unsupervised learning methods, such as outlier detection, utilize various statistical, distance, and density metrics to describe the degree of alienation between data samples and other samples, thereby identifying outliers with high alienation.

[0004] Among these methods, rule-based approaches are labor-intensive, inefficient, and not always accurate in detecting fraud. Supervised learning methods require a large number of labels to achieve good results, thus incurring significant time and cost for data annotation, making the workload extremely heavy. Unsupervised learning methods are unsuitable for biased datasets (e.g., medical insurance datasets). Furthermore, traditional methods often focus only on feature attributes, neglecting other attributes in medical insurance datasets, leading to consistently low detection accuracy.

[0005] In the process of medical insurance fraud, fraudulent users may exhibit unusual characteristics and behaviors. For example, a fraudulent patient might simultaneously obtain large quantities of the same medications from multiple hospitals, or obtain numerous prescriptions for medications unrelated to a particular hospital department. Traditional medical insurance fraud detection methods cannot fully leverage user interactions in such cases, making it difficult to accurately detect fraudulent behavior.

[0006] In view of this, the applicant hereby submits this application after studying the existing technology. Summary of the Invention

[0007] The present invention provides a method, apparatus, device and medium for identifying medical insurance fraud based on neighborhood similarity, in order to improve at least one of the above-mentioned technical problems.

[0008] First aspect

[0009] This invention provides a method for identifying medical insurance fraud based on neighborhood similarity, which includes steps S1 to S8.

[0010] S1. Acquire medical data and construct a medical heterogeneous graph based on the medical data. The medical heterogeneous graph includes patient nodes.

[0011] S2. Obtain the meta-paths of various behavior patterns of patient nodes, and sample the medical heterogeneous graph based on the meta-paths to obtain heterogeneous subgraphs of various behavior patterns.

[0012] S3. Based on the heterogeneous subgraphs of various behavioral patterns, obtain the initial neighborhood set of each patient node under various behavioral patterns through the relational rotation encoder.

[0013] S4. Calculate the similarity of each neighborhood based on the initial neighborhood set, and filter them using an adaptive filtering threshold to obtain the final neighborhood set for each patient node under various behavioral patterns.

[0014] S5. By using the first attention mechanism, the final neighborhood sets of each patient node under various behavioral patterns are fused to obtain the embedding representation of each patient node under each behavioral pattern.

[0015] S6. Based on the embedded representation of each patient node in each behavioral pattern, obtain the importance of various behavioral patterns.

[0016] S7. Based on the importance of various behavioral patterns, the embedding representations of each patient node under various behavioral patterns are fused through the second attention mechanism to obtain the final embedding representation of each patient node.

[0017] S8. Classify the final embedded representation of each patient node to determine whether each patient node is a patient who has committed medical insurance fraud.

[0018] The second aspect This invention provides a medical insurance fraud identification device based on neighborhood similarity, comprising: The heterogeneous graph construction module is used to acquire medical data and construct a medical heterogeneous graph based on the medical data. This medical heterogeneous graph includes patient nodes.

[0019] The sampling module is used to obtain the meta-paths of various behavioral patterns of patient nodes, and to sample the medical heterogeneous graph based on the meta-paths to obtain heterogeneous subgraphs of various behavioral patterns.

[0020] The initial neighborhood acquisition module is used to obtain the initial neighborhood set of each patient node under various behavioral patterns by means of a relational rotation encoder, based on the heterogeneous subgraphs of various behavioral patterns.

[0021] The final neighborhood acquisition module is used to calculate the similarity of each neighborhood based on the initial neighborhood set, and to filter them using an adaptive filtering threshold to obtain the final neighborhood set of each patient node under various behavioral patterns.

[0022] The first fusion module is used to fuse the final neighborhood sets of each patient node under various behavioral patterns through the first attention mechanism, and obtain the embedding representation of each patient node under each behavioral pattern.

[0023] The importance acquisition module is used to acquire the importance of various behavioral patterns based on the embedded representation of each patient node in each behavioral pattern.

[0024] The second fusion module is used to fuse the embedded representations of various behavioral patterns of each patient node through a second attention mechanism based on the importance of various behavioral patterns, so as to obtain the final embedded representation of each patient node.

[0025] The judgment module is used to classify the final embedded representation of each patient node in order to determine whether each patient node is a patient who has committed medical insurance fraud.

[0026] Third aspect This invention provides a medical insurance fraud identification device based on neighborhood similarity, comprising a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement the medical insurance fraud identification method based on neighborhood similarity as described in any paragraph of the first aspect.

[0027] Fourth aspect This invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device containing the computer-readable storage medium to perform the medical insurance fraud identification method based on neighborhood similarity as described in any paragraph of the first aspect.

[0028] By adopting the above technical solution, the present invention can achieve the following technical effects: The medical insurance fraud identification method of this invention transforms patient behavior into data that can be recognized and processed by a computer through heterogeneous graphs. It obtains data of different behavior patterns by sampling and aggregates information of neighboring nodes and meta-paths through an attention mechanism, reducing interference from noisy nodes and low-relevance meta-paths. This enables the acquisition of a final embedded representation that better expresses patient behavior, thereby greatly improving the accuracy of subsequent judgments. Attached Figure Description

[0029] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a flowchart illustrating a medical insurance fraud identification method based on neighborhood similarity.

[0031] Figure 2 This is a logic diagram of a medical insurance fraud identification method based on neighborhood similarity.

[0032] Figure 3 It is a meta-path diagram of heterogeneous graphs and behavioral patterns.

[0033] Figure 4 This is a schematic diagram of a medical insurance fraud identification device based on neighborhood similarity. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0035] Example 1 Please see Figures 1 to 3 The first embodiment of the present invention provides a method for identifying medical insurance fraud based on neighborhood similarity, which can be executed by a medical insurance fraud identification device based on neighborhood similarity (hereinafter referred to as: identification device). In particular, it is executed by one or more processors in the identification device to implement steps S1 to S8.

[0036] S1. Acquire medical data and construct a medical heterogeneous graph based on the medical data. The medical heterogeneous graph includes patient nodes.

[0037] Specifically, by using heterogeneous graph modeling to model real-world medical insurance scenarios, the problem of medical insurance fraud detection is modeled as a patient node classification problem within a heterogeneous graph. This provides a theoretical foundation for subsequent steps in solving the medical insurance fraud detection problem and has significant practical implications.

[0038] It is understood that the identification device may be an electronic device with computing power, such as a portable laptop computer, desktop computer, server, smartphone or tablet computer.

[0039] Based on the above embodiments, in an optional embodiment of the present invention, step S1 specifically includes steps S11 to S12.

[0040] S11. Obtain medical data and extract medical records based on the medical data.

[0041] S12. Based on medical records, construct a medical heterogeneous graph using patients, hospital departments, dates, and medications as entities. Hospitals and departments are treated as a single entity, while departments with the same name in different hospitals are treated as different entities. Date entities are refined to the day level.

[0042] Specifically, the medical insurance dataset contains millions of transaction records from a large number of users. To better understand patient behavior, this embodiment of the invention constructs a medical insurance heterogeneous graph. All medical records of the selected patients are extracted, and four entities are constructed from them: patient, hospital department, date, and medication. To further refine the spatial representation, hospitals and departments are treated as a single entity, meaning that even departments with the same name in different hospitals are treated as different entities. The date entity is refined to the day.

[0043] S2. Obtain the meta-paths of various behavior patterns of patient nodes, and sample the medical heterogeneous graph based on the meta-paths to obtain heterogeneous subgraphs of various behavior patterns.

[0044] Specifically, different behavioral patterns correspond to different meta-paths in the heterogeneous graph. Sampling the heterogeneous graph based on the meta-paths can yield patient groups with different behavioral characteristics, thereby obtaining the behavioral features in the medical heterogeneous graph.

[0045] like Figure 3 As shown in the above embodiments, in an optional embodiment of the present invention, step S1 specifically includes steps S21 to S23.

[0046] S21. Obtain the meta-paths for the three behavioral patterns of the patient node. The meta-paths for the three behavioral patterns include "patient-hospital department-patient", "patient-drug-patient", and "patient-date-patient".

[0047] S22. Sample the medical heterogeneous graph based on the meta-paths of the three behavioral patterns to obtain the initial sub-graphs of the three behavioral patterns.

[0048] Specifically, such as Figure 3 As shown, for ease of explanation, we only display three entities in the heterogeneous diagram: patient (P), hospital department (H), and drug (M). (See attached diagram.) Figure 1 As shown in 'a', patients P1, P2, and P3 all received treatment at Hospital H1. The semantic meta-path in PHP (patient-hospital department-patient) is shown below. Figure 1As shown in b in the diagram. Using semantic meta-path sampling of neighbors can be viewed as a heterogeneous graph starting from the patient node, traversing between nodes of different types according to the order of the meta-path, and finally returning to the patient node. For example: starting from fraudster P2, passing through hospital H1, and finally returning to fraudster P3. The semantic information of multiple semantic meta-paths in PHP can be understood as patients seeing doctors in the same hospital department.

[0049] In this embodiment, three types of meta-paths are used for sampling. Multiple meta-paths are used simultaneously to decompose the graph into three subgraph structures at different levels. Besides PHP, other structures include PDP (Patient-Date-Patient) and PMP (Patient-Medication-Patient), representing patients who visited the doctor on the same day and patients using the same medication, respectively. In other embodiments, the heterogeneous graph can contain more types of nodes and meta-paths; this invention does not specify these in detail.

[0050] S23. Project all node features from the initial subgraph onto the same feature space to obtain a heterogeneous subgraph with three behavioral modes. The projection model is as follows: .

[0051] In the formula, It is the feature representation of the patient node v after projection. It is the parameter weight matrix of the patient node. It is the feature representation of the patient node v before projection.

[0052] Specifically, nodes and edges in heterogeneous graphs have different types, and node attributes of different types have feature vectors of different dimensions. Even if nodes happen to have the same dimension, they may belong to different feature spaces. Therefore, in this embodiment, the features of heterogeneous nodes are projected into the same feature space.

[0053] S3. Based on the heterogeneous subgraphs of various behavioral patterns, obtain the initial neighborhood set of each patient node under various behavioral patterns through the relational rotation encoder.

[0054] Specifically, structural and semantic information embedded in the target node, metapath-based neighboring nodes, and the context between them is learned by encoding metapath instances (i.e., paths between the target patient node and neighboring patient nodes in a heterogeneous subgraph).

[0055] Based on the above embodiments, in an optional embodiment of the present invention, step S3 specifically includes steps S31 to S32.

[0056] S31. Based on the heterogeneous subgraphs of various behavior patterns, obtain the set of meta-path instances for each patient node under various behavior patterns.

[0057] S32. Using a relational rotation encoder, each metapath instance in the metapath instance set is encoded into a vector representation to obtain the neighborhood of the patient node, thus acquiring the initial neighborhood set for each patient node under various behavioral patterns. The relational rotation encoder is: .

[0058] In the formula, The vector representation of the metapath instances from the target patient node v to the neighboring patient node u under behavior pattern M. For encoding functions, Feature representation of the target patient node v after projection. Feature representation of neighboring patient node u after projection. intermediate node Feature representation after projection Let be the set of intermediate nodes between the target patient node v and the neighboring patient node u under the behavior pattern M.

[0059] Specifically, a relational rotation encoder is used to transform the meta-path instance of each patient node in the subgraph into a vector. The relational rotation encoder, proposed by RotatE for knowledge graph embedding, is a meta-path instance encoder based on relational rotation in a complex space. As the encoding function, the relational rotary encoder can be specifically represented as: .

[0060] .

[0061] .

[0062] .

[0063] In the formula, Let M be the vector representation of the metapath instance from the target patient node v to the neighboring patient node u (i.e., the neighborhood of the target patient node v), and the metapath instance. ), , , The intermediate vector of the target patient node V, The number of nodes in the metapath instance, For the first Vector representation after projection of each node For the first The intermediate vector of each node, For matrix dot product of the same dimension, For the first The node and the first The relationship between nodes.

[0064] After encoding metapath instances into vector representations, for a target patient node v, a metapath instance based on target node v is considered as a neighborhood of target node v. .

[0065] S4. Calculate the similarity of each neighborhood based on the initial neighborhood set, and filter them using an adaptive filtering threshold to obtain the final neighborhood set for each patient node under various behavioral patterns.

[0066] This invention calculates the neighborhood similarity of a target patient node based on a neighborhood similarity metric. A single-layer MLP is used as the node predictor, and the prediction scores of the target node and its neighbors are used for the similarity metric.

[0067] Based on the above embodiments, in an optional embodiment of the present invention, step S4 specifically includes steps S41 to S42.

[0068] S41. Based on the initial neighborhood set, calculate the similarity of each neighborhood of the patient node using a neighborhood similarity metric. The neighborhood similarity metric model is as follows: .

[0069] In the formula, The neighborhood of patient node v Similarity For activation function, For single-layer perceptron, For the neighborhood The vector representation of .

[0070] S42. Based on the similarity of each patient node's neighborhood, an adaptive filtering threshold is used to select neighborhoods, obtaining the final neighborhood set for each patient node under various behavioral patterns. The adaptive filtering threshold... for: .

[0071] .

[0072] In the formula, For behavior pattern r, the first Average similarity score over a period of time For behavior pattern r, the first Average similarity score over a period of time For the number of patient nodes, The neighborhood of patient node v under behavior pattern r In the Similarity within each cycle.

[0073] Specifically, a reinforcement learning-based similarity-aware neighborhood selector performs adaptive filtering to automatically select the optimal number of similar neighbors, thus avoiding the high cost of data annotation. In this embodiment, sampling is used in conjunction with an adaptive filtering threshold to select similar neighbors under each relation, and a reinforcement learning (RL) algorithm is used during GNN training to identify the optimal threshold.

[0074] Specifically, during the training phase, for the target patient node v in the current batch under the metapath, a set of similarity scores is first calculated using a neighborhood similarity measurement model. Then, its neighborhoods are sorted in descending order according to the similarity scores, retaining the neighborhoods with the highest similarity in the current batch and discarding the rest. Other neighborhoods discarded in the current batch will not participate in the aggregation process.

[0075] To optimize the computational efficiency of neighbor (neighborhood) selection, this embodiment of the invention uses a reinforcement learning (RL) framework to find the optimal threshold. Given an initial threshold ,Will Defined as a neighborhood selector to increase or decrease selection. A fixed small value Optimal The goal is to find the most similar neighborhood of the target node under relation r. The average similarity score of period e under relation r is as follows: .

[0076] Then, a reward mechanism is designed based on the average similarity score difference between two consecutive batches. The reward for period e is defined as follows: .

[0077] Note that the reward is positive when the average distance of the newly selected neighborhood in period e is less than that in the previous period; otherwise, the reward is negative.

[0078] The embodiments of the present invention do not require a greedy search strategy and use immediate rewards to update actions.

[0079] S5. By using the first attention mechanism, the final neighborhood sets of each patient node under various behavioral patterns are fused to obtain the embedding representation of each patient node under each behavioral pattern.

[0080] Specifically, after selecting the best neighborhood, local aggregation is used, and an attention mechanism is employed to perform the aggregation based on the target node. Metapath instance (i.e., neighborhood set) is weighted.

[0081] Based on the above embodiments, in an optional embodiment of the present invention, the first attention mechanism is: .

[0082] .

[0083] In the formula, Let M be the embedding representation of patient node v under behavioral pattern M, and T be the number of independent attention mechanisms. For activation function, For neighboring patients nodes, Let M be the set of neighboring patient nodes of the target patient node v. The weights of neighboring patient node u relative to target patient node v under behavior pattern M. The vector representation of the metapath instances from the target patient node v to the neighboring patient node u under behavior pattern M. It is the parameterized attention vector of behavior pattern M. Feature representation of the target patient node v after projection. Let V be a vector representation of the metapath instances from the target patient node v to the neighboring patient node k under behavioral pattern M.

[0084] Specifically, the learning process can be stabilized through a multi-head attention mechanism. In this embodiment, T independent attention mechanisms are executed, and their outputs are then concatenated to reduce the high variance caused by heterogeneous graphs.

[0085] S6. Based on the embedded representation of each patient node in each behavioral pattern, obtain the importance of various behavioral patterns.

[0086] Specifically, after aggregating the information of nodes within each behavioral pattern (i.e., the neighborhood set) at the local aggregation layer, a global aggregation layer is used to combine the embedded representations of different behavioral patterns of the target patient node (i.e., the semantic information of different meta-paths). Different behavioral patterns have varying importance in the medical heterogeneous graph. Therefore, this embodiment of the invention first calculates the importance of each behavioral pattern, and then uses an attention mechanism to aggregate different behavioral patterns based on their importance.

[0087] Based on the above embodiments, in an optional embodiment of the present invention, the calculation model for the importance of various behavioral patterns is as follows: .

[0088] .

[0089] In the formula, For the first The importance of individual behavioral patterns For behavioral patterns weights, It is the number of behavioral patterns, For behavioral patterns weights, Parameterized attention vector for patient nodes, For transpose, For the set of patient nodes, For behavioral patterns Embedded representation of patient node v and These are learnable parameters.

[0090] S7. Based on the importance of various behavioral patterns, the embedding representations of each patient node under various behavioral patterns are fused through the second attention mechanism to obtain the final embedding representation of each patient node.

[0091] Based on the above embodiments, in an optional embodiment of the present invention, the model of the second attention mechanism is as follows: .

[0092] In the formula, It is the final embedded representation of the target patient node v. A collection of behavioral patterns The importance of behavioral pattern M Let V be the embedding representation of patient node v under behavior pattern M.

[0093] Specifically, when each behavioral pattern is calculated Importance We can then use this attention coefficient to perform a weighted summation of the embedding vectors under different behavioral patterns of the target node v to obtain the final embedding vector.

[0094] Finally, an additional linear transformation with a nonlinear function is used to project the node embeddings into a vector space with the desired output dimension. The additional linear transformation is as follows: .

[0095] In the formula, The output feature vector of the target patient node and It's just a difference in dimensions. It is an activation function. It is a weight matrix.

[0096] S8. Classify the final embedded representation of each patient node to determine whether each patient node is a patient who has committed medical insurance fraud.

[0097] In this embodiment, the final embedding is classified using a multilayer perceptron. In other embodiments, other existing classification models can be used to classify the final embedding to determine whether a patient node is a patient who has committed medical insurance fraud.

[0098] Traditional medical insurance fraud detection methods often focus only on feature attributes, neglecting the rich behavioral attributes involved in the medical insurance process. This invention constructs a heterogeneous graph based on real medical insurance data, and uses the interaction relationships between entities in the graph to represent these behavioral attributes. Then, a graph neural network is used to classify nodes, solving the problem of determining whether a patient is a fraudster. Graph neural networks are a form of semi-supervised learning, therefore requiring only a small number of outlier samples, making them well-suited for medical insurance data with only a very limited number of fraud records.

[0099] The medical insurance fraud identification method of this invention transforms patient behavior into data that can be recognized and processed by a computer through heterogeneous graphs. It obtains data of different behavior patterns by sampling and aggregates information of neighboring nodes and meta-paths through an attention mechanism, reducing interference from noisy nodes and low-relevance meta-paths. This enables the acquisition of a final embedded representation that better expresses patient behavior, thereby greatly improving the accuracy of subsequent judgments.

[0100] Example 2 This invention provides a medical insurance fraud identification device based on neighborhood similarity, comprising: Heterogeneous graph construction module 1 is used to acquire medical data and construct a medical heterogeneous graph based on the medical data. The medical heterogeneous graph includes patient nodes.

[0101] Sampling module 2 is used to obtain the meta-paths of various behavioral patterns of patient nodes, and to sample the medical heterogeneous graph based on the meta-paths to obtain heterogeneous subgraphs of various behavioral patterns.

[0102] The initial neighborhood acquisition module 3 is used to acquire the initial neighborhood set of each patient node under various behavioral patterns by means of a relational rotation encoder, based on the heterogeneous subgraphs of various behavioral patterns.

[0103] The final neighborhood acquisition module 4 is used to calculate the similarity of each neighborhood based on the initial neighborhood set, and to filter them through an adaptive filtering threshold to obtain the final neighborhood set of each patient node under various behavioral patterns.

[0104] The first fusion module 5 is used to fuse the final neighborhood sets of each patient node under various behavioral patterns through the first attention mechanism, and obtain the embedding representation of each patient node under each behavioral pattern.

[0105] Importance acquisition module 6 is used to acquire the importance of various behavioral patterns based on the embedded representation of each patient node in each behavioral pattern.

[0106] The second fusion module 7 is used to fuse the embedded representations of various behavioral patterns of each patient node through a second attention mechanism based on the importance of various behavioral patterns, so as to obtain the final embedded representation of each patient node.

[0107] The judgment module 8 is used to classify the final embedded representation of each patient node in order to determine whether each patient node is a patient who has committed medical insurance fraud.

[0108] Based on the above embodiments, in an optional embodiment of the present invention, the heterogeneous graph construction module 1 specifically includes: The medical record extraction unit is used to acquire medical data and extract medical records based on the medical data.

[0109] The heterogeneous graph construction unit is used to construct a medical heterogeneous graph based on medical records, using patients, hospital departments, dates, and medications as entities. Hospitals and departments are treated as a single entity, while departments with the same name in different hospitals are treated as distinct entities. Date entities are refined to the day level.

[0110] Based on the above embodiments, in an optional embodiment of the present invention, the step sampling module 2 specifically includes: The meta-path acquisition unit is used to acquire the meta-paths of three behavioral patterns of the patient node. The meta-paths of the three behavioral patterns include "patient-hospital department-patient", "patient-drug-patient", and "patient-date-patient".

[0111] The sampling unit is used to sample the medical heterogeneous graph based on the meta-paths of the three behavioral modes to obtain the initial subgraphs of the three behavioral modes.

[0112] The projection unit projects all node features from the initial subgraph onto the same feature space, obtaining a heterogeneous subgraph with three behavioral modes. The projection model is as follows: In the formula, It is the feature representation of the patient node v after projection. It is the parameter weight matrix of the patient node. It is the feature representation of the patient node v before projection.

[0113] Based on the above embodiments, in an optional embodiment of the present invention, the initial neighborhood acquisition module 3 specifically includes: The metapath instance set acquisition unit is used to acquire the metapath instance set of each patient node under various behavioral patterns based on the heterogeneous subgraphs of various behavioral patterns.

[0114] The initial neighborhood set acquisition unit is used to encode the meta-path instances in the meta-path instance set into vector representations using a relational rotation encoder, thereby obtaining the neighborhood of each patient node and acquiring the initial neighborhood set for each patient node under various behavioral patterns. The relational rotation encoder is as follows: . In the formula, The vector representation of the metapath instances from the target patient node v to the neighboring patient node u under behavior pattern M. For encoding functions, Feature representation of the target patient node v after projection. Feature representation of neighboring patient node u after projection. intermediate node Feature representation after projection Let be the set of intermediate nodes between the target patient node v and the neighboring patient node u under the behavior pattern M.

[0115] Based on the above embodiments, in an optional embodiment of the present invention, the final neighborhood acquisition module 4 specifically includes: The similarity calculation unit is used to calculate the similarity of each neighborhood of a patient node based on the initial neighborhood set, using a neighborhood similarity metric. The neighborhood similarity metric model is as follows: . In the formula, The neighborhood of patient node v Similarity For activation function, For single-layer perceptron, For the neighborhood The vector representation of .

[0116] The neighborhood selection unit is used to select neighborhoods based on the similarity of each patient node's neighborhoods using an adaptive filtering threshold, thereby obtaining the final neighborhood set for each patient node under various behavioral patterns. The adaptive filtering threshold... for: . . In the formula, For behavior pattern r, the first Average similarity score over a period of time For behavior pattern r, the first Average similarity score over a period of time For the number of patient nodes, The neighborhood of patient node v under behavior pattern r In the Similarity within each cycle.

[0117] Based on the above embodiments, in an optional embodiment of the present invention, the first attention mechanism is: . . In the formula, Let M be the embedding representation of patient node v under behavioral pattern M, and T be the number of independent attention mechanisms. For activation function, For neighboring patients nodes, Let M be the set of neighboring patient nodes of the target patient node v. The weights of neighboring patient node u relative to target patient node v under behavior pattern M. The vector representation of the metapath instances from the target patient node v to the neighboring patient node u under behavior pattern M. It is the parameterized attention vector of behavior pattern M. Feature representation of the target patient node v after projection. Let V be a vector representation of the metapath instances from the target patient node v to the neighboring patient node k under behavioral pattern M.

[0118] Based on the above embodiments, in an optional embodiment of the present invention, the calculation model for the importance of various behavioral patterns is as follows: .

[0119] .

[0120] In the formula, For the first The importance of individual behavioral patterns For behavioral patterns weights, It is the number of behavioral patterns, For behavioral patterns weights, Parameterized attention vector for patient nodes, For transpose, For the set of patient nodes, For behavioral patterns Embedded representation of patient node v and These are learnable parameters.

[0121] Based on the above embodiments, in an optional embodiment of the present invention, the model of the second attention mechanism is as follows: .

[0122] In the formula, It is the final embedded representation of the target patient node v. A collection of behavioral patterns The importance of behavioral pattern M Let V be the embedding representation of patient node v under behavior pattern M.

[0123] Example 3 This invention provides a medical insurance fraud identification device based on neighborhood similarity, comprising a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement the medical insurance fraud identification method based on neighborhood similarity as described in any paragraph of Embodiment 1.

[0124] Example 4 This invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program, which, when executed, controls the device containing the computer-readable storage medium to perform the medical insurance fraud identification method based on neighborhood similarity as described in any paragraph of Embodiment 1.

[0125] In the several embodiments provided in this invention, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0126] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0127] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, electronic device, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. In the absence of further restrictions, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0128] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0129] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0130] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0131] The use of "first" and "second" in the embodiments is merely to distinguish similar objects and does not represent a specific ordering of objects. It is understood that "first" and "second" can be interchanged in a specific order or sequence where permitted. It should be understood that the objects distinguished by "first" and "second" can be interchanged where appropriate so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0132] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for identifying medical insurance fraud based on neighborhood similarity, characterized in that, Include: Acquire medical data and construct a medical heterogeneous graph based on the medical data; wherein the medical heterogeneous graph includes patient nodes; Obtain the meta-paths of various behavioral patterns of patient nodes, and sample the medical heterogeneous graph based on the meta-paths to obtain heterogeneous subgraphs of various behavioral patterns. Based on the heterogeneous subgraphs of the various behavioral patterns, a set of meta-path instances for each patient node under each behavioral pattern is obtained; the meta-path instances in the set of meta-path instances are encoded into vector representations using a relational rotation encoder to obtain the neighborhood of the patient node, thus obtaining the initial neighborhood set for each patient node under each behavioral pattern; wherein, the relational rotation encoder is: ; In the formula, The vector representation of the metapath instances from the target patient node v to the neighboring patient node u under behavior pattern M. For encoding functions, Feature representation of the target patient node v after projection. Feature representation of neighboring patient node u after projection. intermediate node Feature representation after projection Let M be the set of intermediate nodes between the target patient node v and the neighboring patient nodes u under the behavior pattern M; Based on the initial neighborhood set, the similarity of each neighborhood of a patient node is calculated using a neighborhood similarity metric. Based on the similarity of each neighborhood of the patient node, neighborhoods are selected using an adaptive filtering threshold to obtain the final neighborhood set for each patient node under various behavioral patterns. The neighborhood similarity metric model is as follows: ; Adaptive filtering threshold for: ; ; In the formula, The neighborhood of patient node v Similarity For activation function, For single-layer perceptron, For the neighborhood Vector representation of; For behavior pattern r, the first Average similarity score over a period of time For behavior pattern r, the first Average similarity score over a period of time For the number of patient nodes, The neighborhood of patient node v under behavior pattern r In the Similarity in each cycle; By fusing the final neighborhood sets of each patient node under various behavioral patterns through the first attention mechanism, the embedded representations of each patient node under each behavioral pattern are obtained. Based on the embedding representations of each behavioral pattern in each patient node, the importance of each behavioral pattern is obtained; wherein, the importance of each behavioral pattern is: ; ; In the formula, For the first The importance of individual behavioral patterns For behavioral patterns weights, It is the number of behavioral patterns, For behavioral patterns weights, Parameterized attention vector for patient nodes, For transpose, For the set of patient nodes, For behavioral patterns Embedded representation of patient node v and These are learnable parameters; Based on the importance of the various behavioral patterns, the embedding representations of each patient node under various behavioral patterns are fused through a second attention mechanism to obtain the final embedding representation of each patient node; wherein, the second attention mechanism is: ; In the formula, It is the final embedded representation of the target patient node v. A collection of behavioral patterns The importance of behavioral pattern M For the embedding representation of patient node v under behavior pattern M; The final embedded representations of each patient node are classified to determine whether each patient node is a patient who has committed medical insurance fraud.

2. The medical insurance fraud identification method based on neighborhood similarity according to claim 1, characterized in that, The first attention mechanism is: ; ; In the formula, Let M be the embedding representation of patient node v under behavioral pattern M, and T be the number of independent attention mechanisms. For activation function, For neighboring patients nodes, Let M be the set of neighboring patient nodes of the target patient node v. The weights of neighboring patient node u relative to target patient node v under behavior pattern M. The vector representation of the metapath instances from the target patient node v to the neighboring patient node u under behavior pattern M. It is the parameterized attention vector of behavior pattern M. Feature representation of the target patient node v after projection. Let V be a vector representation of the metapath instances from the target patient node v to the neighboring patient node k under behavioral pattern M.

3. The medical insurance fraud identification method based on neighborhood similarity according to any one of claims 1 to 2, characterized in that, Acquire medical data and construct a medical heterogeneous graph based on the medical data; wherein, the medical heterogeneous graph includes patient nodes, specifically including: Acquire medical data and extract medical records based on the medical data; Based on the medical records, a medical heterogeneous graph is constructed with the patient, hospital department, date, and medication as entities; where the hospital and department are a whole, and departments with the same name in different hospitals are different entities; the date entity is refined to the day.

4. The medical insurance fraud identification method based on neighborhood similarity according to any one of claims 1 to 2, characterized in that, Obtain the meta-paths of various behavioral patterns of patient nodes, and sample the medical heterogeneous graph based on the meta-paths to obtain heterogeneous subgraphs of various behavioral patterns, specifically including: Obtain the meta-paths of three behavioral patterns of the patient node; wherein, the meta-paths of the three behavioral patterns include patient-hospital department-patient, patient-medication-patient, and patient-date-patient; The medical heterogeneous graph is sampled based on the meta-paths of the three behavioral patterns to obtain the initial sub-graphs of the three behavioral patterns. Projecting all node features from the initial subgraph onto the same feature space yields a heterogeneous subgraph with three behavioral modes; wherein the projection model is... In the formula, It is the feature representation of the patient node v after projection. It is the parameter weight matrix of the patient node. It is the feature representation of the patient node v before projection.

5. A medical insurance fraud identification device based on neighborhood similarity, characterized in that, Used to perform the medical insurance fraud identification method based on neighborhood similarity as described in any one of claims 1 to 4; The medical insurance fraud detection device includes: A heterogeneous graph construction module is used to acquire medical data and construct a medical heterogeneous graph based on the medical data; wherein, the medical heterogeneous graph includes patient nodes; The sampling module is used to obtain the meta-paths of various behavioral patterns of patient nodes, and to sample the medical heterogeneous graph according to the meta-paths to obtain heterogeneous subgraphs of various behavioral patterns. The initial neighborhood acquisition module is used to acquire the initial neighborhood set of each patient node under various behavioral patterns by means of a relational rotation encoder, based on the heterogeneous subgraph of the various behavioral patterns. The final neighborhood acquisition module is used to calculate the similarity of each neighborhood according to the initial neighborhood set, and to filter them through an adaptive filtering threshold to obtain the final neighborhood set of each patient node under various behavioral patterns. The first fusion module is used to fuse the final neighborhood sets of each patient node under various behavioral patterns through the first attention mechanism to obtain the embedded representation of each patient node under each behavioral pattern. The importance acquisition module is used to acquire the importance of various behavioral patterns based on the embedded representation of each behavioral pattern of each patient node. The second fusion module is used to fuse the embedding representations of each patient node under various behavioral patterns according to the importance of the various behavioral patterns through a second attention mechanism to obtain the final embedding representation of each patient node. The judgment module is used to classify the final embedded representation of each patient node in order to determine whether each patient node is a patient who has committed medical insurance fraud.

6. A medical insurance fraud detection device based on neighborhood similarity, characterized in that, It includes a processor, a memory, and a computer program stored in the memory; the computer program can be executed by the processor to implement the medical insurance fraud identification method based on neighborhood similarity as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the medical insurance fraud identification method based on neighborhood similarity as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Medical insurance fraud detection algorithm and system based on multilayer attention mechanism graph neural network

    CN114463141A