Medical insurance anomaly detection method and device for guiding truncation affinity maximization based on anomaly density, equipment and medium
By employing an anomalous density-guided truncation affinity maximization method, and utilizing graph neural networks and dynamic threshold pruning techniques, the problem of abnormal pattern recognition in medical insurance data was solved. This approach achieves efficient and accurate medical insurance anomaly detection, reduces dependence on labeled data, and improves the model's adaptability and robustness.
Patent Information
- Application Number
- CN202510748496.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-10-31
AI Technical Summary
Existing methods for detecting anomalies in medical insurance data are insufficient in handling the complexity of medical insurance data, the concealment of anomaly patterns, and the imbalance of categories. They are difficult to accurately identify anomaly patterns in medical insurance data and are highly dependent on labeled data.
We employ an anomalous density-guided truncation affinity maximization method. By acquiring the heterogeneous graph of medical insurance data, we calculate the anomalous connection density, perform graph pruning by combining global and local thresholds, train the model using a graph neural network, and calculate the anomalous score to determine whether the data is anomalous.
It achieves efficient and accurate detection of medical insurance anomalies, reduces dependence on labeled data, improves the adaptability and robustness of the model, and can adapt to medical insurance data of different scales and distributions.
Smart Images

Figure CN120875897A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data analytics, and more specifically, to a method, apparatus, equipment, and medium for detecting medical insurance anomalies based on anomaly density-guided truncation affinity maximization. Background Technology
[0002] Existing methods for detecting anomalies in medical insurance mainly include supervised learning, unsupervised learning, data reconstruction-based methods, and self-supervised learning. Supervised learning methods rely on a large amount of labeled data to train classification models, such as logistic regression and random forests. However, labeled medical insurance data is scarce in reality, severely limiting their practical application. Unsupervised learning methods, while suitable for data lacking labels, such as algorithms based on local outliers and isolated forests, struggle to capture complex and hidden anomaly patterns in medical insurance data, resulting in limited detection effectiveness. Data reconstruction-based methods, such as autoencoders, suffer from overfitting risks and struggle to capture high-order relationships between nodes, leading to inaccurate anomaly identification. Self-supervised learning methods learn node representations through surrogate tasks, reducing reliance on labeled data. However, during training, they are prone to blurring the distinction between normal and abnormal behavior, resulting in a high false positive rate and sensitivity to data noise and uneven distribution.
[0003] Furthermore, traditional pruning strategies based on fixed thresholds or probabilities struggle to accurately remove duplicate edges, impacting the model's accuracy in processing high-dimensional heterogeneous graph structures. Existing technologies have significant limitations in handling the complexity of medical insurance data, the concealment of anomaly patterns, and class imbalance. Therefore, there is an urgent need for an efficient, accurate, and adaptable method for detecting medical insurance anomalies to overcome these shortcomings. Summary of the Invention
[0004] The present invention provides a method, apparatus, device and medium for detecting medical insurance anomalies based on abnormal density-guided truncation affinity maximization, in order to improve at least one of the above-mentioned technical problems.
[0005] In a first aspect, the present invention provides a medical insurance anomaly detection method based on abnormal density-guided truncation affinity maximization, which includes steps S1 to S5.
[0006] S1. Obtain medical insurance data and integrate it into a heterogeneous graph. Then, extract meta-paths and combine them into a meta-graph. Finally, calculate the abnormal connection density of each meta-graph to filter the meta-graphs and obtain the semantic graph of medical insurance data.
[0007] S2. Based on the medical insurance data semantic graph, calculate the global threshold and the local threshold and merge them into a fusion threshold. Then, perform a pruning operation on the medical insurance data semantic graph based on the fusion threshold to obtain the sequence truncated adjacency matrix.
[0008] S3. Truncate the adjacency matrix according to the sequence, and train the function with the objective function of maximizing neighborhood affinity. Layered graph neural network.
[0009] S4. Calculate the anomaly score based on the graph neural network.
[0010] S5. Based on the abnormal score, determine whether the medical insurance data is abnormal data.
[0011] As a further aspect of the present invention, step S1 includes steps S11 to S15.
[0012] S11. Collect medical insurance data and integrate it into a heterogeneous graph. The medical insurance data includes entity data and relational data. The entity data includes patients, hospitals, drugs, and departments. The relational data refers to the relationships between entities. Nodes in the heterogeneous graph represent entities, and edges represent the relationships between entities.
[0013] S12. Based on the heterogeneous graph and semantic knowledge of medical insurance business, define multiple meta-paths.
[0014] S13. By common sampling, the first preset number of metapaths are sequentially fused to obtain multiple metagraphs.
[0015] S14. Calculate the outlier connectivity density of each metagraph. In the formula, For abnormal connection density, Represents the set of all nodes. A set of abnormal nodes Represents a node To the node edge information, This indicates that there is an edge. It means there is no boundary.
[0016] S15. After calculating the abnormal connection density of all metagraphs, select the metagraph with the highest density to obtain the medical insurance data semantic graph.
[0017] As a further aspect of the present invention, step S2 includes steps S21 to S27.
[0018] S21. Calculate the distance between each node in the medical insurance data semantic graph. In the formula, Represents a node To the node distance, For the number of dimensions, For the dimension index, For nodes In the Values in each dimension For nodes In the Values in each dimension.
[0019] S22. Calculate the optimal number of clusters based on the semantic graph of the medical insurance data. In the formula, For the optimal number of clusters, The point set function of the maximum value of the independent variable To dynamically determine the maximum number of candidate clusters using dual boundary constraints, When the number of clusters is Clustering quality evaluation function at time This is the Calinski-Harabasz index.
[0020] S23. Obtain the global threshold based on the optimal number of clusters. In the formula, For global threshold, For the cluster center number, For the first Cluster centers, This represents the number of samples belonging to this cluster.
[0021] S24. Calculate the dynamic quantiles based on the semantic graph of the medical insurance data. In the formula, For dynamic quantiles, For the baseline quantile, For adjustment coefficient, For node degree, This represents the total number of nodes.
[0022] S25. Obtain the local threshold based on the dynamic quantile. In the formula, For local threshold, For quantile functions, Represents a node , This represents the set of direct neighbors.
[0023] S26. Based on the global threshold and the local threshold, a fusion threshold is obtained by fusing them together. In the formula, For nodes fusion threshold, This is a hyperparameter.
[0024] S27. Based on the fusion threshold, perform a pruning operation to obtain the truncated adjacency matrix. The pruning conditions are: In the formula, Let's define the semantic adjacency matrix of medical insurance data. The edges in For nodes The fusion threshold.
[0025] As a further aspect of the present invention, step S2 also includes: setting Pruning operations are performed at each truncation depth to extract sequence truncation adjacency matrices at different depths.
[0026] As a further aspect of the present invention, step S2 specifically involves: processing the original medical insurance data semantic graph. implement The operation "Based on the medical insurance data semantic graph, calculate the global threshold and local threshold and fuse them into a fusion threshold, then perform pruning operations on the medical insurance data semantic graph according to the fusion threshold to obtain the sequence truncated adjacency matrix" is set each time. Cut off depth, get A set of truncated adjacency matrices of sequences.
[0027] As a further aspect of the present invention, the graph neural network model is as follows: . In the formula, Indicates the first Layer node representation matrix, For activation function, For the degree matrix of the semantic graph of medical insurance data, For medical insurance data semantic adjacency matrix, Indicates the first Layer node representation matrix, It is the first Layer weight parameters.
[0028] As a further aspect of the present invention, the objective function is: . In the formula, Indicates minimizing loss, The learnable parameters of the GNN model, Represents anomaly scoring function, For medical insurance data semantic adjacency matrix, For node features, For regularization hyperparameters, The difference operation of sets Represents the set of all nodes Remove nodes The set of neighboring nodes The remaining set of nodes For nodes The representation vector, For nodes The representation vector.
[0029] As a further aspect of the present invention, calculating anomaly scores based on the graph neural network specifically includes: When no settings are set Repeat and At a cutoff depth, the outlier score is: .
[0030] When setting Repeat and At a cutoff depth, the outlier score is: .
[0031] In the formula, For abnormal scores, Represents a node Anomaly scoring function The learnable parameters of the GNN model, For the first The th repetition Learnable parameters of a GNN model with a cutoff depth For medical insurance data semantic adjacency matrix, These are node features.
[0032] As a further aspect of the present invention, the anomaly scoring function is: .
[0033] In the formula, For nodes The set of neighboring nodes For nodes The representation vector, For nodes The representation vector.
[0034] Secondly, the present invention provides a medical insurance anomaly detection device based on anomaly density-guided truncation affinity maximization, which includes a semantic graph module, a pruning module, a training module, a score calculation module, and an anomaly judgment module.
[0035] The semantic graph module is used to acquire medical insurance data and integrate it into a heterogeneous graph. Then, it extracts meta-paths and combines them into a meta-graph. Finally, it calculates the abnormal connection density of each meta-graph to filter the meta-graphs and obtain the semantic graph of medical insurance data.
[0036] The pruning module is used to calculate a global threshold and a local threshold based on the medical insurance data semantic graph and merge them into a fusion threshold. Then, it performs a pruning operation on the medical insurance data semantic graph based on the fusion threshold to obtain a sequence truncated adjacency matrix.
[0037] The training module is used to truncate the adjacency matrix according to the sequence and train a function that maximizes neighborhood affinity. Layered graph neural network.
[0038] The scoring calculation module is used to calculate anomaly scores based on the graph neural network.
[0039] The anomaly detection module is used to determine whether the medical insurance data is abnormal based on the anomaly score.
[0040] Thirdly, the present invention provides a medical insurance anomaly detection device based on anomaly density-guided truncation affinity maximization, comprising a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement a medical insurance anomaly detection method based on anomaly density-guided truncation affinity maximization as described in any paragraph of the first aspect.
[0041] Fourthly, the present invention provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform a medical insurance anomaly detection method based on anomaly density-guided truncation affinity maximization as described in any paragraph of the first aspect.
[0042] By adopting the above technical solution, the present invention can achieve the following technical effects: The medical insurance anomaly detection method based on anomaly density-guided truncation affinity maximization in this invention not only enables efficient anomaly detection and reduces dependence on labeled data, but also has the beneficial effect of strong adaptability. Attached Figure Description
[0043] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the specific embodiments of the present invention will be briefly introduced below. It should be understood that the following drawings only show some specific embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.
[0044] Figure 1 This is a flowchart illustrating a medical insurance anomaly detection method based on abnormal density-guided truncation affinity maximization.
[0045] Figure 2 This is a logic diagram of a medical insurance anomaly detection method based on abnormal density-guided truncation affinity maximization.
[0046] Figure 3 This is an example diagram illustrating the heterogeneous graph, metapath, and metagraph of a medical insurance dataset.
[0047] Figure 4 This is a flowchart of medical insurance data preprocessing. Detailed Implementation
[0048] The technical solutions of the present invention will now be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention.
[0049] Example 1, please refer to Figures 1 to 4 The first embodiment of this invention provides a Health Insurance Anomaly Detection via AnomalyDensity-Guided Truncated Affinity Maximization (HAD-TAM) method. The ultimate goal of the HAD-TAM model is to obtain an anomaly score to assess whether abnormal health insurance behavior exists.
[0050] like Figure 2 As shown, the HAD-TAM model framework consists of three modules: Module 1 is the graph processing module, which performs data preprocessing on medical insurance data; Module 2 is Dynamic Normal Structure Keeping Graph Truncation (DNSGT) for pruning; and Module 3 is the Neighborhood Affinity Maximization Network (NAMNet), which maximizes neighborhood affinity through training. DNSGT, Dynamic Normal Structure Keeping Graph Truncation, is a module in the HAD-TAM model used for graph structure pruning. NAMNet, Neighborhood Affinity Maximization Network, is another module in the HAD-TAM model used to learn node representations by optimizing the objective function. Neighborhood affinity is a metric that measures the similarity between a node and its neighbors, used to distinguish between normal and abnormal nodes.
[0051] A brief introduction to HAD-TAM: First, the graph processing module obtains a semantic graph of medical insurance data with high anomaly density. Next, the DNSGT module performs dynamic pruning based on local and global thresholds, further truncating anomalous edges to obtain a cleaner, normal graph structure. Finally, the NAMNet module maximizes neighborhood affinity, enabling the model to learn more normal medical insurance patterns and enhancing its anomaly detection performance.
[0052] The specific operating mechanism of the HAD-TAM model architecture is as follows: First, the graph processing module dynamically generates all possible combinations of meta-paths and calculates the abnormal connection density between different combinations of meta-paths for logical rule filtering, automatically selecting the medical insurance data semantic graph with the highest abnormal connection density as the input graph data. Figure 2 The graph processing section on the left is a simplified flowchart of the data processing section. The specific flowchart is as follows: Figure 3 As shown.
[0053] Secondly, the DNSGT module employs a global-local fusion threshold strategy to iteratively prune non-similar edges in the medical insurance data semantic graph. Preferably, this is performed independently on the medical insurance data semantic graph. Each DNSGT operation sets... Each cutoff depth captures structural features of the graph with different sparsity, generating... The set of truncated adjacency matrices of group sequences, each group containing Sequence truncation adjacency matrices with different truncation depths.
[0054] NAMNet is then trained on each adjacency matrix, ultimately forming a matrix containing... An ensemble HAD-TAM model of NAMNets. Each NAMNet independently trains a GNN on a sequence-truncated adjacency matrix.
[0055] Finally, NAMNet is trained using the graph processing module and the neighborhood affinity maximization objective of the graph structure generated by DNSGT, ultimately obtaining the anomaly score for each node. Preferably, NAMNet optimizes the objective function. By learning node representations, the purity of the graph structure is gradually improved and the distinguishability between normal and abnormal nodes is enhanced.
[0056] An embodiment of the present invention provides a medical insurance anomaly detection method based on anomaly density-guided truncation affinity maximization, which can be executed by a medical insurance anomaly detection device based on anomaly density-guided truncation affinity maximization (hereinafter referred to as: medical insurance anomaly detection device). Specifically, it is executed by one or more processors in the medical insurance anomaly detection device to implement steps S1 to S5.
[0057] S1. Obtain medical insurance data and integrate it into a heterogeneous graph. Then, extract meta-paths and combine them into a meta-graph. Finally, calculate the abnormal connection density of each meta-graph to filter the meta-graphs and obtain the semantic graph of medical insurance data.
[0058] Figure 3 This is a flowchart of the medical insurance data preprocessing process of the present invention. Figure 3 Part a of the diagram shows the heterogeneous medical insurance data obtained by processing medical insurance data. Figure 3 Part b in the diagram illustrates the sampling of basic meta-paths and the combination of multiple semantic meta-paths. Figure 3Part c in the diagram shows the calculation of the anomaly density scores for all combinatorial metapaths. Figure 3 The d part in the diagram shows the semantic graph of medical insurance data obtained by selecting meta-path combinations with high anomaly density scores.
[0059] Based on the above embodiments, in an optional embodiment of the present invention, step S1 specifically includes steps S11 to S15.
[0060] S11. Collect medical insurance data and integrate it into a heterogeneous graph. The medical insurance data includes entity data and relational data. The entity data includes patients, hospitals, drugs, and departments. The relational data refers to the relationships between entities. Nodes in the heterogeneous graph represent entities, and edges represent the relationships between entities.
[0061] Specifically, medical insurance data includes entity information such as patients, hospitals, drugs, and departments, as well as relationship data between them, such as patient medical records and drug usage records. Integrating medical insurance data into a heterogeneous graph, where nodes represent different entities and edges represent relationships between entities, results in a heterogeneous graph structure like this: Figure 3 part a in Figure 4 As shown in part a of the diagram.
[0062] S12. Based on the heterogeneous graph and semantic knowledge of medical insurance business, define multiple meta-paths.
[0063] Specifically, based on the semantics of medical insurance business, potential abnormal patterns are captured. For example, frequent visits to different departments or unusual drug combinations may indicate abnormalities.
[0064] Metapath such as: Patient - Appointment Time - Patient ( Patient-Department-Patient ( ), Patient-Medication-Patient ( These meta-paths capture higher-order semantic relationships between different entity types. Meta-paths include... Figure 3 part b in Figure 4 As shown in part b of the diagram. Then, the corresponding meta-path information is extracted from the heterogeneous graph based on the meta-path.
[0065] Metapath The definition of is: In the formula, Representing the patient entity, Indicates the actual time of medical visit, Indicates the patient and appointment time The relationship between them Indicates the time of medical visit and patients The relationship between them.
[0066] S13. By common sampling, a first preset number of meta-paths are sequentially merged to obtain a meta-graph. In this embodiment, the first preset number is 3, and any 3 meta-paths are combined into 1 meta-graph.
[0067] The metagraph construction method of this invention, by sampling and combining metapath information with various semantics, can better represent the complex semantic relationships in the medical insurance data network.
[0068] Figure 3 Part b describes the method for combining multiple semantic meta-paths. First, it involves combining two meta-paths... and By performing joint sampling, a first metagraph containing the semantic information of these two metapaths is obtained. It incorporates the patient's Department ,drug and time Semantic information such as metapaths. Then, this study will... and the first element graph Perform another common path sampling to obtain a second-dimensional graph containing three types of meta-path semantic information. ,like Figure 4 As shown in part b of the document.
[0069] The entire process of combining multiple semantic meta-paths is as follows: In the formula, For the first element graph, Represents the common path sampling function, For the meta-pathway patient-department-patient, For the meta-pathway patient-medicine-patient, For second element graph, For the meta-path: Patient - Visiting Time - Patient.
[0070] S14. Calculate the abnormal connection density of each metagraph.
[0071] Due to the massive volume of medical insurance data, the proportion of anomalies in real medical insurance data is relatively small. Therefore, the model needs to process data with a certain proportion of anomalies to ensure sufficient data to distinguish between normal and abnormal patterns. To increase the proportion of potentially anomaly-containing data, this invention designs anomaly connection density as a screening indicator for medical insurance data. Anomaly connection density is an indicator used to measure the proportion of anomaly connections in the metagraph, assisting in screening graph structures related to abnormal behavior, such as... Figure 4 As shown in part c.
[0072] The calculation model for abnormal connection density is as follows: .
[0073] In the formula, For abnormal connection density, Represents the set of all nodes. A set of abnormal nodes Represents a node To the node The edge information. This indicates that there is an edge. It means there is no boundary.
[0074] Abnormal nodes are identified by the abnormal labels used when loading data. These labels are only used for the final evaluation and do not participate in the model training process. The anomaly density score for each metagraph is obtained by calculating the ratio of the total number of anomalous connections to the total number of edges in the graph. A higher anomalous connection density in a metagraph indicates a higher proportion of anomalous patterns.
[0075] S15. After calculating the abnormal connection density of all metagraphs, select the metagraph with the highest density to obtain the medical insurance data semantic graph.
[0076] Specifically, selecting the meta-graph with the highest density helps improve anomaly detection performance. Assuming the second meta-graph is the highest density meta-graph, the second meta-graph... The semantic adjacency matrix is Node characteristics are Let the semantic adjacency matrix of medical insurance data be defined. Then the second element graph The corresponding semantic adjacency matrix and node features Combined into a medical insurance data semantic graph ,like Figure 3 As shown in part d in the diagram.
[0077] This invention accurately captures potential anomaly patterns by calculating the density of abnormal connections and selecting the metagraph with the highest density. This mechanism avoids the bias of manually pre-setting metapaths and improves the model's ability to identify complex anomaly patterns.
[0078] S2. Based on the medical insurance data semantic graph, calculate the global threshold and the local threshold and merge them into a fusion threshold. Then, perform a pruning operation on the medical insurance data semantic graph based on the fusion threshold to obtain the sequence truncated adjacency matrix.
[0079] In medical insurance anomaly detection scenarios, non-identical edges can severely interfere with the optimization process of maximizing neighborhood affinity. To address this, this invention proposes the DNSGT mechanism, which combines global and local thresholds to achieve more refined graph pruning. The DNSGT mechanism comprises two core modules: a global threshold based on clustering analysis and a local threshold based on quantile statistics. The global threshold is a unified standard set based on the overall data distribution, suitable for universal anomaly detection. The local threshold is dynamically adjusted based on individual node characteristics to adapt to the feature differences in local data. The global threshold provides a benchmark, while the local threshold achieves more refined adaptation, jointly optimizing the detection effect.
[0080] Based on the above embodiments, in an optional embodiment of the present invention, step S2 specifically includes steps S21 to S27.
[0081] S21. Calculate the distance between each node in the medical insurance data semantic graph.
[0082] .
[0083] In the formula, Represents a node To the node distance, For the number of dimensions, For the dimension index, For nodes In the Values in each dimension For nodes In the Values in each dimension.
[0084] S22. Calculate the optimal number of clusters based on the semantic graph of the medical insurance data.
[0085] .
[0086] In the formula, For the optimal number of clusters, The point set function of the maximum value of the independent variable To dynamically determine the maximum number of candidate clusters using dual boundary constraints, When the number of clusters is Clustering quality evaluation function at time This is the Calinski-Harabasz index.
[0087] The optimal number of clusters is determined by the Calinski-Harabasz exponent ( ). Dynamic selection. The maximum number of candidate clusters is dynamically determined through double boundary constraints. , Take 5 and the number of valid edges The smaller of the two values is used, while ensuring that the lower limit of the number of candidates is 2. This avoids overfitting to sparse data and limits the computational complexity of dense data to the number of effective edges. constraint.
[0088] S23. Obtain the global threshold based on the optimal number of clusters.
[0089] .
[0090] In the formula, For global threshold, For the cluster center number, For the first Cluster centers, This represents the number of samples belonging to this cluster.
[0091] The core function of a global threshold is to provide an adaptive baseline reference value for edge pruning in graph structures. However, relying solely on a global threshold for graph pruning has limitations: if the global threshold is too strict for the local distribution of some nodes, it may mistakenly delete homopair edges. If the global threshold is too lenient for some nodes, it may leave non-homopair edges.
[0092] Therefore, this invention proposes an adaptive pruning strategy that combines global and local thresholds. The local threshold is introduced to address the limitations of the global threshold. By incorporating the node's own neighborhood statistics, the local threshold provides a more refined pruning standard for each node, minimizing the risk of false deletions or retention of non-same edges due to probability-based or rigid pruning strategies.
[0093] S24. Calculate the dynamic quantiles based on the semantic graph of the medical insurance data.
[0094] .
[0095] In the formula, For dynamic quantiles, For the baseline quantile, For adjustment coefficient, For node degree, This represents the total number of nodes.
[0096] The baseline quantile is a hyperparameter, which was set to 0.9 during the experiment. An adjustment coefficient controls the strength of the influence of node degree on the dynamic quantile.
[0097] Since the medical insurance graph data was preprocessed in the early stages of the experiment, a relatively normal graph structure was selected through anomaly density calculation. According to the neighborhood affinity theory, nodes with tighter connections in the graph structure are more likely to be normal nodes, while those with sparse connections are more likely to be anomalies. Dynamic quantiles ensure that when the node degree is large, the dynamic quantile value is larger, thus retaining more homopairs; when the node degree is small, non-homopairs are strictly pruned.
[0098] S25. Obtain the local threshold based on the dynamic quantile.
[0099] .
[0100] In the formula, For local threshold, For quantile functions, Represents a node , This represents the set of direct neighbors.
[0101] Specifically, by combining nodes and nodes Euclidean distance between and dynamic quantiles This generates adaptive local thresholds. By dynamically sensing the distribution characteristics of local neighborhoods, more precise topology pruning is achieved.
[0102] S26. Based on the global threshold and the local threshold, a fusion threshold is obtained by fusing them together.
[0103] .
[0104] In the formula, For nodes fusion threshold, This is a hyperparameter.
[0105] Its function is to balance the ratio of global threshold to local threshold, and the resulting fusion threshold is used for conditional pruning.
[0106] S27. Perform a pruning operation based on the fusion threshold to obtain the sequence truncated adjacency matrix.
[0107] The pruning conditions are: .
[0108] In the formula, Let's define the semantic adjacency matrix of medical insurance data. The edges in For nodes The fusion threshold.
[0109] Based on the above embodiments, in an optional embodiment of the present invention, the DNSGT module can be configured with different cutoff depths (i.e., the maximum number of candidate clusters). This is used to extract sequence truncation adjacency matrices at different depths. In this embodiment, each time step S2 is executed, the following settings are configured: Perform the operation at different cutoff depths to obtain... Sequence truncation adjacency matrices with different truncation depths.
[0110] Preferably, in this embodiment, the semantic graph of the original medical insurance data is... Independent execution In step S2, each time a setting is made Different cutoff depths were obtained, ultimately yielding... A set of truncated adjacency matrices. By executing... Step S2 enables the construction of an ensemble model to improve the robustness of the model.
[0111] This embodiment designs a more precise adaptive pruning mechanism to overcome the limitations of existing pruning strategies based on fixed thresholds or probabilistic probabilistics. This improves the accuracy of removing non-identical edges in high-dimensional medical insurance data and reduces the risk of erroneous deletion or retention of abnormal edges. Specifically, the global threshold is dynamically determined using K-means clustering and the Calinski-Harabasz exponent, while the local threshold is dynamically adjusted based on the node's own neighborhood statistics. This pruning strategy, which integrates global and local thresholds, effectively removes interference from non-identical edges, improving the model's processing efficiency and accuracy for high-dimensional medical insurance data.
[0112] S3. Truncate the adjacency matrix according to the sequence, and train the function with the objective function of maximizing neighborhood affinity. Layered graph neural networks learn a mapping function. .
[0113] In this embodiment, the DNSGT module uses a global-local fusion thresholding strategy to iteratively prune non-same edges in the semantic graph of the medical insurance data to obtain a truncated adjacency matrix. Then, a GNN is trained based on this truncated adjacency matrix to capture the structural features of the graph.
[0114] In an optional embodiment, the DNSGT module extracts sequence truncated adjacency matrices at different depths by setting different truncation depths. Then, a GNN is trained independently for each sequence truncated adjacency matrix to capture the structural features of the graph at different sparsities.
[0115] More preferably, the original image Independent execution Each DNSGT operation sets the following parameters: A cutoff depth, thus generating Set of truncated adjacency matrices Each group contains The set of sequence truncated adjacency matrices consisting of sequence truncated adjacency matrices with different truncation depths. .pass The operation of building an integrated model can effectively improve the detection robustness of the model.
[0116] NAMNet is then trained on the adjacency matrix of each sequence truncation, ultimately forming a matrix containing... An ensemble HAD-TAM model of NAMNet. Note that during NAMNet training, the neighborhood affinity calculation in the global threshold formula is still based on the original medical insurance data semantic adjacency matrix. This design is based on two considerations: preserving the fundamental role of the original graph structure in anomaly detection; and using the graph truncation operation solely to eliminate biases during graph convolution.
[0117] This invention sets up a neighborhood affinity maximization network (NAMNet) to learn a mapping function using a graph neural network (GNN). This enhances the neighborhood affinity between nodes and their neighbors, while suppressing the neighborhood affinity of nodes connected to abnormal or irrelevant nodes.
[0118] NAMNet utilizes A layered graph neural network maps nodes in a graph to a new representation space. This process can be generally represented as: .
[0119] In the formula, Indicates the first Layer node representation matrix, For graph neural networks, For medical insurance data semantic adjacency matrix, Indicates the first Layer node representation matrix, It is the first Layer weight parameters.
[0120] In this embodiment, the graph neural network uses a graph convolutional network (GCN) operation. When no different truncation depths are set, there is only one sequence truncation adjacency matrix, which is used to replace the original medical insurance data semantic adjacency matrix in the graph convolution operation.
[0121] use The specific representation of a layered graph neural network mapping nodes in a graph to a new representation space is as follows: .
[0122] In the formula, Indicates the first Layer node representation matrix, For activation function, For the degree matrix of the semantic graph of medical insurance data, For medical insurance data semantic adjacency matrix, Indicates the first Layer node representation matrix, It is the first Layer weight parameters.
[0123] Represents a node In the Layer representation. In this embodiment, the graph neural network is configured with... Layer, then This represents the node of the last GCN layer. Mapping function. The order mapping of graph convolution in the formula for calculating the optimal number of clusters.
[0124] This invention uses an objective function that maximizes neighborhood affinity to optimize the mapping function. .
[0125] The objective function is: .
[0126] In the formula, Indicates minimizing loss, The learnable parameters of the GNN model, Represents anomaly scoring function, For medical insurance data semantic adjacency matrix, For node features, For regularization hyperparameters, The difference operation of sets Represents the set of all nodes Remove nodes The set of neighboring nodes The remaining set of nodes (i.e., nodes) (set of non-neighbor nodes) For nodes The representation vector, For nodes The representation vector.
[0127] This objective function is a key optimization objective function used to train the GNN model, enabling the model to learn node representations by maximizing neighborhood affinity. This objective function not only improves the model's ability to identify normal nodes but also enhances its ability to perform regularization. This prevents the model from becoming overly smooth. By introducing a regularization term, this embodiment of the invention ensures that the representation of each node is not only similar to its neighbors but also maintains a certain degree of distinguishability from non-neighbor nodes. Furthermore, by pre-processing the adjacency matrix through the DNSGT component, the optimization process of NAMNet can be effectively avoided from being affected by edges connecting normal and abnormal nodes.
[0128] Based on the above embodiments, in an optional embodiment of the present invention, a progressive graph truncation strategy is used to train NAMNet hierarchically, thereby effectively detecting anomalous nodes at different graph truncation scales. Specifically, according to... GNN models with different truncation depths are trained using sequence truncation adjacency matrices of different truncation depths, and their parameter sets are denoted as follows: Each NAMNet employs a truncated adjacency matrix in graph convolution operations. Replace the original adjacency matrix This process completed the training of the base model based on HAD-TAM, which was developed by... It consists of several GNNs.
[0129] In this embodiment, by optimizing the objective function to maximize the neighborhood affinity between normal nodes and their neighbors, the distinguishability between normal and abnormal nodes is significantly improved. This module is implemented using a graph neural network (GCN), which can learn the representation vectors of nodes, thereby improving the model's detection performance and enabling it to more accurately identify abnormal medical insurance behaviors, thus enhancing the model's detection performance on imbalanced datasets.
[0130] S4. Calculate the anomaly score based on the embedded representation.
[0131] When no settings are set Repeat and At a cutoff depth, the outlier score is: .
[0132] When setting Repeat and At a cutoff depth, the outlier score is: .
[0133] In the formula, For abnormal scores, Represents a node Anomaly scoring function The learnable parameters of the GNN model, For the first The th repetition Learnable parameters of a GNN model with a cutoff depth For medical insurance data semantic adjacency matrix, For node features, During the inference phase, anomaly scores based on neighborhood affinity are obtained for each node from each NAMNet, and finally aggregated from all... The overall anomaly score is calculated using the neighborhood affinity scores of each NAMNet. This reflects the strength of the neighborhood affinity of nodes in the representation space learned at different graph truncation scales. The weaker the neighborhood affinity exhibited in the representation space learned at different graph truncation scales, the more it means... More likely, it's an abnormal node with an abnormal score. The larger it is.
[0134] Because the original medical insurance data may contain many irrelevant attributes, and some abnormal nodes may be connected to similar abnormal nodes, the medical insurance anomaly detection method based on anomaly density-guided truncation affinity maximization in this embodiment of the invention uses NAMNet to enhance the neighborhood affinity of relevant nodes and reduce the neighborhood affinity of abnormal or irrelevant nodes. Non-similar edges are pruned using the DNSGT module to eliminate these interferences. The optimization process makes the neighborhood affinity of normal nodes significantly higher than that of abnormal nodes, giving the model more robust anomaly detection performance.
[0135] The HAD-TAM-based anomaly score that maximizes neighborhood affinity is: .
[0136] In the formula, For nodes The set of neighboring nodes For nodes The representation vector, For nodes The representation vector.
[0137] For nodes The anomaly scoring function represents the node. The degree of abnormality is obtained by taking the negative of the definition of neighborhood affinity. This represents the learnable parameters of the GNN model, used to learn node representations. .
[0138] In this embodiment, by independently performing multiple DNSGT operations on the original graph and generating multiple sets of truncated adjacency matrices, HAD-TAM forms an ensemble model. This ensemble model can comprehensively consider the graph structure features at different truncation depths, improving the model's robustness and detection performance.
[0139] S5. Based on the anomaly score, determine whether the medical insurance data is abnormal. The anomaly score can be used to determine whether the medical insurance data is abnormal by setting a threshold. The threshold can be the average of the anomaly scores of normal medical insurance data, or it can be calculated using other methods. This invention does not specifically limit the method for obtaining the anomaly score threshold.
[0140] The medical insurance anomaly detection method based on anomaly density-guided truncation affinity maximization in this invention not only enables efficient anomaly detection and reduces dependence on labeled data, but also has the beneficial effect of strong adaptability.
[0141] Highly efficient anomaly detection: Through an anomaly density-guided meta-path selection mechanism, HAD-TAM accurately identifies complex semantic features related to medical insurance anomalies, reducing the bias of manually preset meta-paths. The Dynamic Normal Graph Structure Preservation Truncation (DNSGT) module, combined with an adaptive pruning strategy using global and local thresholds, effectively eliminates interference from non-identical edges, improving the model's efficiency and accuracy in processing high-dimensional medical insurance data. The Neighborhood Affinity Maximization Network (NAMNet) optimizes the objective function to maximize the neighborhood affinity between normal nodes and their neighbors, significantly improving the distinguishability between normal and anomaly nodes.
[0142] By reducing dependence on labeled data, HAD-TAM, as an unsupervised learning method, does not rely on a large amount of labeled data for training, thus reducing its dependence on labeled data and improving the feasibility and generalization ability of the model in practical applications.
[0143] With its strong adaptability, HAD-TAM can adapt to medical insurance data of different sizes and distributions through dynamic pruning and neighborhood affinity maximization mechanisms, demonstrating broad applicability and robustness.
[0144] It is understood that the medical insurance anomaly detection device can be an electronic device with computing power, such as a portable laptop computer, desktop computer, server, smartphone, or tablet computer.
[0145] Example 2: The present invention provides a medical insurance anomaly detection device based on anomaly density-guided truncation affinity maximization, which includes a semantic graph module, a pruning module, a training module, a score calculation module, and an anomaly judgment module.
[0146] The semantic graph module is used to acquire medical insurance data and integrate it into a heterogeneous graph. Then, it extracts meta-paths and combines them into a meta-graph. Finally, it calculates the abnormal connection density of each meta-graph to filter the meta-graphs and obtain the semantic graph of medical insurance data.
[0147] The pruning module is used to calculate a global threshold and a local threshold based on the medical insurance data semantic graph and merge them into a fusion threshold. Then, it performs a pruning operation on the medical insurance data semantic graph based on the fusion threshold to obtain a sequence truncated adjacency matrix.
[0148] The training module is used to truncate the adjacency matrix according to the sequence and train a function that maximizes neighborhood affinity. Layered graph neural network.
[0149] The scoring calculation module is used to calculate anomaly scores based on the graph neural network.
[0150] The anomaly detection module is used to determine whether the medical insurance data is abnormal based on the anomaly score.
[0151] Example 3: This invention provides a medical insurance anomaly detection device based on anomaly density-guided truncation affinity maximization, comprising a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement a medical insurance anomaly detection method based on anomaly density-guided truncation affinity maximization as described in any paragraph of Example 1.
[0152] Example 4: The present invention provides a computer-readable storage medium, the computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute a medical insurance anomaly detection method based on anomaly density-guided truncation affinity maximization as described in any paragraph of Example 1.
[0153] Obviously, the embodiments described above are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0154] In the several embodiments provided in this invention, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0155] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0156] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, electronic device, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks. It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0157] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0158] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0159] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0160] The terms "first" and "second" used in the embodiments are merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first" and "second" can be interchanged in a specific order or sequence where permitted. It should be understood that the objects distinguished by "first" and "second" can be interchanged where appropriate so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0161] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for detecting medical insurance anomalies based on anomaly density-guided truncation affinity maximization, characterized in that, Include: The medical insurance data is acquired and integrated into a heterogeneous graph. Then, meta-paths are extracted and combined into a meta-graph. Finally, the abnormal connection density of each meta-graph is calculated to filter the meta-graphs and obtain the semantic graph of the medical insurance data. Based on the medical insurance data semantic graph, a global threshold and a local threshold are calculated and fused into a fusion threshold. Then, a pruning operation is performed on the medical insurance data semantic graph based on the fusion threshold to obtain a sequence truncated adjacency matrix. The adjacency matrix is truncated based on the sequence, and a target function that maximizes neighborhood affinity is trained and configured. Layered graph neural networks; Anomaly scores are calculated based on the graph neural network. Based on the aforementioned anomaly score, determine whether the medical insurance data is abnormal.
2. The medical insurance anomaly detection method based on abnormal density-guided truncation affinity maximization as described in claim 1, characterized in that, The process involves acquiring medical insurance data and integrating it into a heterogeneous graph. Metapaths are then extracted and combined into a metagraph. Finally, the abnormal connection density of each metagraph is calculated to filter the metagraphs and obtain a semantic graph of the medical insurance data. This process specifically includes: Collect medical insurance data and integrate it into a heterogeneous graph; wherein, medical insurance data includes entity data and relational data; the entity data includes patients, hospitals, drugs and departments; the relational data is the data between entities; the nodes in the heterogeneous graph are entities, and the edges are the relationships between entities; Based on the heterogeneous graph, and semantic knowledge of medical insurance business, multiple meta-paths are defined; By jointly sampling, a first preset number of meta-paths are sequentially fused to obtain multiple meta-graphs; Calculate the outlier connectivity density of each metagraph separately; where, In the formula, For abnormal connection density, Represents the set of all nodes. A set of abnormal nodes Represents a node To the node edge information, Indicates that there is an edge. It means there is no boundary; After calculating the abnormal connection density of all metagraphs, the metagraph with the highest density is selected to obtain the semantic graph of medical insurance data.
3. The medical insurance anomaly detection method based on abnormal density-guided truncation affinity maximization as described in claim 1, characterized in that, Based on the medical insurance data semantic graph, a global threshold and a local threshold are calculated and fused into a fused threshold. Then, a pruning operation is performed on the medical insurance data semantic graph based on the fused threshold to obtain a sequence truncated adjacency matrix, specifically including: Calculate the distances between each node in the semantic graph of the medical insurance data; In the formula, Represents a node To the node distance, For the number of dimensions, For the dimension index, For nodes In the Values in each dimension For nodes In the Values in each dimension; Calculate the optimal number of clusters based on the semantic graph of the medical insurance data; In the formula, For the optimal number of clusters, The point set function of the maximum value of the independent variable To dynamically determine the maximum number of candidate clusters using dual boundary constraints, When the number of clusters is Clustering quality evaluation function at time The Calinski-Harabasz index; Based on the optimal number of clusters, obtain the global threshold; In the formula, For global threshold, For the cluster center number, For the first Cluster centers, This represents the number of samples belonging to this cluster; Calculate the dynamic quantiles based on the semantic graph of the medical insurance data; In the formula, For dynamic quantiles, For the baseline quantile, For adjustment coefficient, For node degree, This represents the total number of nodes; Based on the dynamic quantile, obtain the local threshold; In the formula, For local threshold, For quantile functions, Represents a node , Represents the set of direct neighbors; A fusion threshold is obtained by fusing the global threshold and the local threshold. In the formula, For nodes fusion threshold, For hyperparameters; Based on the fusion threshold, a pruning operation is performed to obtain the truncated adjacency matrix of the sequence; the pruning condition is: In the formula, Let's define the semantic adjacency matrix of medical insurance data. The edges in For nodes The fusion threshold.
4. The medical insurance anomaly detection method based on abnormal density-guided truncation affinity maximization as described in claim 3, characterized in that, By setting Pruning operations are performed at each truncation depth to extract sequence truncation adjacency matrices at different depths; Based on the medical insurance data semantic graph, a global threshold and a local threshold are calculated and fused into a fused threshold. Then, a pruning operation is performed on the medical insurance data semantic graph based on the fused threshold to obtain a sequence truncated adjacency matrix, specifically: Semantic graph of raw medical insurance data implement The operation "Based on the medical insurance data semantic graph, calculate the global threshold and local threshold and fuse them into a fused threshold, then perform pruning operations on the medical insurance data semantic graph according to the fused threshold to obtain the sequence truncated adjacency matrix" is performed, setting each time. Cut off depth, get A set of truncated adjacency matrices of sequences.
5. The medical insurance anomaly detection method based on abnormal density-guided truncation affinity maximization as described in claim 1, characterized in that, The graph neural network model is: In the formula, Indicates the first Layer node representation matrix, For activation function, For the degree matrix of the semantic graph of medical insurance data, For medical insurance data semantic adjacency matrix, Indicates the first Layer node representation matrix, It is the first Layer weight parameters; The objective function is: In the formula, Indicates minimizing loss, The learnable parameters of the GNN model, Represents anomaly scoring function, For medical insurance data semantic adjacency matrix, For node features, For regularization hyperparameters, The difference operation of sets Represents the set of all nodes Remove nodes The set of neighboring nodes The remaining set of nodes For nodes The representation vector, For nodes The representation vector.
6. A method for detecting medical insurance anomalies based on anomaly density-guided truncation affinity maximization according to any one of claims 1 to 5, characterized in that, Based on the graph neural network, anomaly scores are calculated, specifically including: When no settings are set Repeat and At a cutoff depth, the outlier score is: ; When setting Repeat and At a cutoff depth, the outlier score is: ; In the formula, For abnormal scores, Represents a node Anomaly scoring function The learnable parameters of the GNN model, For the first The th repetition Learnable parameters of a GNN model with a cutoff depth For medical insurance data semantic adjacency matrix, These are node features.
7. The medical insurance anomaly detection method based on abnormal density-guided truncation affinity maximization as described in claim 6, characterized in that, The anomaly scoring function is: ; In the formula, For nodes The set of neighboring nodes For nodes The representation vector, For nodes The representation vector.
8. A medical insurance anomaly detection device based on abnormal density-guided truncation affinity maximization, characterized in that, Include: The semantic graph module is used to acquire medical insurance data and integrate it into a heterogeneous graph. Then, it extracts meta-paths and combines them into a meta-graph. Finally, it calculates the abnormal connection density of each meta-graph to filter the meta-graphs and obtain the semantic graph of medical insurance data. The pruning module is used to calculate a global threshold and a local threshold based on the medical insurance data semantic graph and fuse them into a fusion threshold. Then, it performs a pruning operation on the medical insurance data semantic graph based on the fusion threshold to obtain a sequence truncated adjacency matrix. The training module is used to truncate the adjacency matrix according to the sequence and train a function that maximizes neighborhood affinity. Layered graph neural networks; The scoring calculation module is used to calculate anomaly scores based on the graph neural network. The anomaly detection module is used to determine whether the medical insurance data is abnormal based on the anomaly score.
9. A medical insurance anomaly detection device based on abnormal density-guided truncation affinity maximization, characterized in that, It includes a processor, a memory, and a computer program stored in the memory; the computer program can be executed by the processor to implement a medical insurance anomaly detection method based on anomaly density-guided truncation affinity maximization as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform a medical insurance anomaly detection method based on anomaly density-guided truncation affinity maximization as described in any one of claims 1 to 7.