Medical insurance fraud detection method and device, computer equipment and storage medium
By performing feature encoding and timing feature extraction on heterogeneous medical data, combined with dynamic attention mechanism and conformal calibration, the comprehensiveness and accuracy of medical insurance fraud detection are solved, especially fraud detection for medical providers, achieving more efficient fraud identification.
Patent Information
- Application Number
- CN202510771761.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
现有技术中,医疗保险欺诈检测主要聚焦于参保人的欺诈行为,检测准确性较低,且对医疗提供方的欺诈检测研究较少,导致检测不够全面。
By obtaining heterogeneous medical data for standardized encoding and feature fusion, long-term and short-term memory networks are used for timing feature extraction, dynamic attention mechanism is used for feature aggregation, and feature refining and conformal calibration are combined with full-connected networks and classifiers to generate a set of fraud predictions that meet the preset signal level.
It improves the comprehensiveness and accuracy of medical insurance fraud detection, can capture the timing information of claims, and enhances the detection ability of medical providers to fraud.
Smart Images

Figure CN120298010A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a medical insurance fraud detection method, device, computer device, and storage medium. Background Technique
[0002] With the rapid development of the medical insurance industry in recent years, the amount of funds has expanded on a large scale. How to effectively manage these funds has become a major challenge. At the same time, the increasing number of insured persons, more diversified diagnosis and treatment models, and medical projects have also increased business relationships, making the operation process of medical insurance more complex. Therefore, it is necessary to focus on the behavior of "fraud" and try to identify medical insurance fraud through advanced technologies.
[0003] To identify medical insurance fraud, it is first necessary to clarify the different categories of fraud. Around the entities participating in medical insurance, there are mainly insured persons, medical providers, and third-party payment institutions. Classification based on the participating parties can be mainly divided into fraud by a single entity and fraud by multiple parties in combination. However, in fact, the fraud of insured persons is usually small-scale and sudden, and has a relatively limited impact on the medical insurance system. In contrast, the fraud of medical providers is often more concealed and large-scale. However, most current studies mainly focus on capturing the fraud behavior of insured persons, including identity theft, fictitious bills, drug abuse, etc., and there are relatively few studies on the fraud detection of medical providers, resulting in incomplete medical insurance fraud detection and low detection accuracy. Summary of the Invention
[0004] The purpose of the embodiments of this application is to propose a medical insurance fraud detection method, device, computer device, and storage medium to improve the accuracy of medical insurance fraud detection.
[0005] To solve the above technical problems, the embodiments of this application provide a medical insurance fraud detection method, including: Obtain heterogeneous medical data, perform standardized coding and feature fusion on the heterogeneous medical data, and generate a high-dimensional fusion feature sequence; Extract temporal features from the high-dimensional fusion feature sequence according to the claim time interval through a long short-term memory network, and generate a temporal feature sequence; Adopt a dynamic attention mechanism to aggregate features of the temporal feature sequence, and generate a fraud representation vector with a fixed dimension; Perform feature refinement processing on the fraud representation vector through a fully connected network to generate a high-density feature vector; Generate a prediction result based on the high-density feature vector through a classifier, and perform conformal calibration on the prediction result to generate a fraud prediction set that meets the preset confidence level.
[0006] To solve the above technical problems, an embodiment of the present application provides a medical insurance fraud detection device, including: A feature fusion module, configured to obtain heterogeneous medical data, perform standardized encoding and feature fusion on the heterogeneous medical data, and generate a high-dimensional fusion feature sequence; A temporal feature extraction module, configured to perform temporal feature extraction on the high-dimensional fusion feature sequence according to the claim time interval through a long short-term memory network, and generate a temporal feature sequence; A dynamic attention aggregation module, configured to perform feature aggregation on the temporal feature sequence by adopting a dynamic attention mechanism, and generate a fraud representation vector with a fixed dimension; A feature refinement module, configured to perform feature refinement processing on the fraud representation vector through a fully connected network, and generate a high-density feature vector; A conformal calibration module, configured to generate a prediction result based on the high-density feature vector through a classifier, and perform conformal calibration on the prediction result to generate a fraud prediction set that meets the preset confidence level.
[0007] To solve the above technical problems, a technical solution adopted by the present invention is: to provide a computer device, including one or more processors; a memory, configured to store one or more programs, so that the one or more processors implement the medical insurance fraud detection method described in any one of the above.
[0008] To solve the above technical problems, a technical solution adopted by the present invention is: a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the medical insurance fraud detection method described in any one of the above is implemented.
[0009] An embodiment of the present invention provides a medical insurance fraud detection method, device, computer device, and storage medium. The method includes: obtaining heterogeneous medical data, performing standardized coding and feature fusion on the heterogeneous medical data to generate a high-dimensional fusion feature sequence; performing temporal feature extraction on the high-dimensional fusion feature sequence according to the claim time interval through a long short-term memory network to generate a temporal feature sequence; adopting a dynamic attention mechanism to perform feature aggregation on the temporal feature sequence to generate a fraud representation vector with a fixed dimension; performing feature refinement processing on the fraud representation vector through a fully connected network to generate a high-density feature vector; generating a prediction result based on the high-density feature vector through a classifier, and performing conformal calibration on the prediction result to generate a fraud prediction set that meets the preset confidence level. By encoding features of heterogeneous medical data, extracting features according to the claim time interval through a long short-term memory network, and performing prediction calibration in a conformal calibration manner, the embodiment of the present invention enables the present application to detect medical fraud for medical providers, improve the comprehensiveness of detection, and capture the temporal information of claims, which is beneficial to improving the accuracy of medical fraud detection. Description of the Drawings
[0010] To more clearly illustrate the solutions in this application, the following will briefly introduce the drawings required for the description of the embodiments of this application. Obviously, the following-described drawings are some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 It is a flowchart for implementing the medical insurance fraud detection method provided by the embodiment of this application; Figure 2 It is a flowchart for implementing the first sub-process in the medical insurance fraud detection method provided by the embodiment of this application; Figure 3 It is a flowchart for implementing the second sub-process in the medical insurance fraud detection method provided by the embodiment of this application; Figure 4 It is a flowchart for implementing the third sub-process in the medical insurance fraud detection method provided by the embodiment of this application; Figure 5 It is a flowchart for implementing the fourth sub-process in the medical insurance fraud detection method provided by the embodiment of this application; Figure 6 It is a flowchart for implementing the fifth sub-process in the medical insurance fraud detection method provided by the embodiment of this application; Figure 7 It is a schematic diagram of the medical insurance fraud detection device provided by the embodiment of this application; Figure 8It is a schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.
[0013] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0014] To enable those skilled in the technical field to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0015] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0016] It should be noted that the medical insurance fraud detection method provided by the embodiments of this application is generally executed by a server. Correspondingly, the medical insurance fraud detection device is generally configured in the server.
[0017] Please refer to Figure 1 , Figure 1 which shows a specific implementation manner of the medical insurance fraud detection method.
[0018] It should be noted that if there are substantially the same results, the method of the present invention is not limited to Figure 1 the process sequence shown, and the method includes the following steps: S1: Obtain heterogeneous medical data, perform standardized encoding and feature fusion on the heterogeneous medical data, and generate a high-dimensional fusion feature sequence.
[0019] Specifically, in the embodiments of the present application, an encoding-based feature fusion layer is adopted to perform standardized encoding and feature fusion on heterogeneous medical data, generating a high-dimensional fusion feature sequence to achieve the standardized expression of heterogeneous medical data. Among them, heterogeneous medical data refers to a collection of data of various types, formats, or structures in the medical field. These data come from different medical systems, devices, or departments and have diverse forms and characteristics. For example, heterogeneous medical data can be tabular data (such as patient ID, age, diagnosis code, drug code, etc.) in electronic health records (EHRs), laboratory test results (numerical data such as blood glucose values, blood cell counts, etc.), medical reports in XML or JSON format, medical images, and electronic medical records, etc.
[0020] Please refer to Figure 2 , Figure 2 which shows a specific implementation manner of step S1, described in detail as follows: S11: Obtain the heterogeneous medical data; S12: Map the discrete categorical data in the heterogeneous medical data to continuous vectors of a fixed dimension through an embedding layer, obtaining a first vector; S13: Extract semantic embeddings from the text data in the heterogeneous medical data to generate a second vector; S14: Perform standardization processing on the connected numerical data in the heterogeneous medical data to generate a third vector; S15: Perform feature concatenation on the first vector, the second vector, and the third vector, and perform non-linear fusion on the concatenated features to generate the high-dimensional fusion feature sequence.
[0021] Specifically, obtain heterogeneous medical data, which is divided into different types of data, including discrete categorical features (such as diagnosis codes, drug codes), text features (such as case descriptions), and numerical features (such as costs, time intervals). For discrete categorical data (such as diagnosis codes, drug codes), use an embedding layer (Embedding Layer), and in the embodiments of the present application, map it to continuous vectors of a fixed dimension (such as 128 dimensions) to generate a first vector. For text features (such as medical record descriptions), use pre-trained word vectors (such as Word2Vec) or a BERT model to extract semantic embeddings to generate a second vector. Numerical feature standardization: Perform standardization (Z-Score) processing on continuous numerical features (such as costs, time intervals) to generate a third vector. Concatenate the first vector, the second vector, and the third vector along the feature dimension to form a unified high-dimensional feature vector, and finally perform non-linear fusion on the concatenated features through a multi-layer fully connected network (such as two layers of 512 dimensions, with the activation function being ReLU) to output a standardized feature matrix, obtaining a high-dimensional fusion feature sequence.
[0022] S2: Use a long short-term memory network to extract temporal features from the high-dimensional fusion feature sequence according to the claim time interval, generating a temporal feature sequence.
[0023] Specifically, existing research in the field of medical insurance fraud has paid little attention to the sequential features of claims. Even if there is, the features contained in the sequence are basically treated as a covariate. However, such a method cannot directly embed the time characteristics of the claim sequence into the model for learning, resulting in the inability to extract the features in the sequence well. When facing sequence data with dynamic time dependence, traditional machine learning methods can usually only capture static features. For long sequences, long short-term memory networks usually have better processing effects. Since the embodiments of the present application are directed to long sequence data, and most of the sequence lengths are above 500, the embodiments of the present application use long short-term memory networks for feature extraction. In the embodiments of the present application, the structure of the long short-term memory network is improved so that it can capture and selectively forget the time intervals in the long sequence.
[0024] Please refer to Figure 3 , Figure 3 which shows a specific implementation manner of step S2, described in detail as follows: S21: Use the high-dimensional fusion feature sequence and the claim time interval as the input of the long short-term memory network; S22: Calculate the forget gate based on the high-dimensional fusion feature sequence and the claim time interval through a first preset formula; S23: Dynamically adjust the amount of information in the candidate memory unit through the claim time interval, calculate the output hidden state at each time step, and splice the output hidden states to generate a temporal feature sequence.
[0025] In the embodiments of the present application, the claim time interval (such as the number of days between two claims) is introduced into the input high-dimensional fusion feature sequence as the input of the long short-term memory network. The first preset formula (1) for calculating the forget gate by introducing the claim time interval is: (1); where is the forget gate, 、 、 are randomly initialized weights, is a bias randomly initialized, is the high-dimensional fusion feature sequence, is the claim time interval, 、 is the sequence 、 The claim moment.
[0026] Calculate the input gate through formula (2), the memory unit through formula (3) and formula (4), the output gate through formula (5), the hidden state through formula (6), and the final output through formula (7), where formulas (2)-(7) are as follows:
[0027] Wherein, are the input gate, candidate memory unit, memory unit, output gate, hidden unit, and final output, are randomly initialized weights, is the bias generated by random initialization. In the embodiments of the present application, the improved forget gate can capture information at the timestamp level. For long sequences, the forget gate can determine whether the intervals between timestamps can be remembered during training. The role of the input gate is to determine whether to ignore the input data. The role of the output gate is to determine whether to use the hidden state. The hidden state and memory unit enter the next long short-term memory network for calculation at the next time step, which is the output collected at each time step. In the embodiments of the present application, the amount of information of the candidate memory unit is dynamically adjusted through the claim time interval, the output hidden state at each time step is calculated, and the output hidden states are concatenated to generate a temporal feature sequence. The improved long short-term memory network can not only perform feature processing on long sequences but also capture the interval features of sequence timestamps, which is beneficial to improving the accuracy of feature extraction and further improving the accuracy of fraud detection.
[0028] S3: Adopt a dynamic attention mechanism to aggregate the features of the temporal feature sequence to generate a fraud representation vector with a fixed dimension.
[0029] Specifically, each sample consists of several claim records, that is, the shape of each sample input into the model is different, which is impossible for the model to process. Therefore, it is necessary to aggregate the data with a single sample as the dimension. In the embodiments of the present application, a dynamic attention mechanism is adopted to aggregate the features of the temporal feature sequence to generate a fraud representation vector with a fixed dimension, so as to realize generating a weight distribution of a variable-length sequence and compressing it into a fixed-dimension feature vector.
[0030] Please refer to Figure 4 , Figure 4 shows a specific implementation manner of step S3, which is described in detail as follows: S31: Map the aggregation target to a query vector through a fully connected layer, and construct key-value pairs for the temporal feature sequence; S32: Calculate the score of each value based on the query vector and the key-value pair using a scaled dot product function to obtain an initial score, and normalize the initial score to obtain a normalized score; S33: Perform a weighted sum of the normalized score and the input value to aggregate the features of the time series feature sequence and generate the fraud representation vector of a fixed dimension.
[0031] Specifically, three concepts are introduced in the attention mechanism: the casual cue is called a query (query, q), each input is a key-value pair of a value (value, v) and a non-casual cue (key, k). That is, in the embodiments of the present application, the aggregation target (such as the provider ID) is mapped to a query vector through a fully connected layer, and the time series feature sequence is constructed into a key-value pair. By using a scaled dot product function to calculate the score of each value based on the query vector and the key-value pair, so as to achieve a targeted query of each query vector for the value, an initial score is obtained, and the initial score is normalized to obtain a normalized score. Finally, a weighted sum of the normalized score and the input value is performed to aggregate the features of the time series feature sequence and generate the fraud representation vector of a fixed dimension. Among them, the formulas corresponding to the scaled dot product function, normalization processing, and weighted sum are formula (8), formula (9), and formula (10) respectively, as follows:
[0032] Among them, is a introduced scaling factor. Since the attention mechanism pays attention to the importance of the sequence, the embodiments of the present application provide interpretability for the neural network to a certain extent. During the calculation process, the attention weights of each time step of the sequence are obtained, and the attention weights are used for subsequent claim positioning. The attention weights can reflect the degree of importance of this position, which is an important manifestation of the interpretability of the attention mechanism.
[0033] S4: Perform feature refinement processing on the fraud representation vector through a fully connected network to generate a high-density feature vector.
[0034] Specifically, perform feature refinement processing on the fraud representation vector through a fully connected network to achieve mining the fraud pattern representation at the provider level and generate a high-density feature vector.
[0035] Please refer to Figure 5 , Figure 5 shows a specific implementation manner of step S4, which is described in detail as follows: S41: Perform a multi-layer non-linear transformation on the fraud representation vector through the fully connected network to obtain a transformed feature vector; S42: Compress the changed feature vector to generate an initial high-density feature vector; S43: If there are static features, refine the static features and concatenate the refined static features with the dynamic features in the initial high-density feature vector to generate the high-density feature vector.
[0036] Specifically, perform multi-layer non-linear transformation on the fraud representation vector through a fully-connected network. In a specific embodiment, the first fully-connected layer (512 dimensions, activation function is ReLU), followed by batch normalization (BatchNorm) and Dropout (probability 0.3); the second fully-connected layer (256 dimensions, activation function is ReLU), followed by batch normalization; the final fully-connected layer (128 dimensions, activation function is Tanh) compresses the feature dimension and outputs a high-density representation. If there are other static features (such as provider history), concatenate them with the refined dynamic features to form the final feature vector and obtain the high-density feature vector.
[0037] S5: Generate a prediction result based on the high-density feature vector through a classifier, and perform conformal calibration on the prediction result to generate a fraud prediction set that meets the preset confidence level.
[0038] Specifically, the embodiment of the present application uses a multi-strategy conformal prediction layer to output a confidence calibration block to form a closed-loop optimization system through cascaded feature transfer, providing a fraud determination set for suspicious claim location and review priority determination while ensuring detection accuracy.
[0039] Please refer to Figure 6 , Figure 6 which shows a specific implementation of step S5, described in detail as follows: S51: Divide the label set corresponding to the high-density feature vector to generate multiple independent calibration samples; S52: Classify each independent calibration sample based on the high-density feature vector through the classifier to generate the prediction result corresponding to each independent calibration sample; S53: Calculate the non-conformity score based on the prediction result according to the preset conformal prediction method to generate a quantile threshold, and perform prediction screening according to the quantile threshold and the preset confidence level to generate the fraud prediction set. Wherein, the preset conformal prediction method is any one of the coverage priority method and the prediction size priority method.
[0040] Embodiments of this application are expected to perform analysis in the post - stage of classifier prediction. Conformal prediction is a confidence predictor that can calculate the confidence of prediction results based on machine learning algorithms. This framework outputs the P - value of each result to measure the degree of consistency between the prediction set and the true set. The size of the prediction set and the class coverage are two metrics for evaluating the prediction set. Therefore, there are two design ideas for conformal prediction: "prediction size first" and "coverage first", that is, the preset conformal prediction method in the embodiments of this application is either the coverage - priority method or the prediction - size - priority method.
[0041] The coverage - priority method focuses on covering as many true results as possible at a given confidence level. Its goal is to provide a higher coverage rate, even if the confidence interval may be large. After a slight correction to the set size, Equation (11) calculates the minimum... on the calibration dataset that just meets the confidence level of 1 - α. 。 It can help determine whether each prediction in the test set can enter the prediction set. Specifically, represents the set of labels corresponding to the feature vector whose sum of probabilities is , n represents the sample size of the dataset, and α is the confidence level. Equation (12) determines whether each predicted label in the test set can enter the prediction set. Among them, Equation (11) and Equation (12) are respectively: ; 。
[0042] The prediction - size - priority method is to minimize the size of the prediction interval as much as possible at a given confidence level. Its goal is to narrow the prediction interval as much as possible to improve the certainty of the prediction. As shown in Equation (13), where is the sum of the probabilities of the labels with probabilities greater than that of y, represents the probability of label y, and u is a randomized term. The last term is a regularization term. Equation (14) generates the prediction set on the test set. Among them, a matching upper bound is added to the set - valued function to limit the size of the prediction set to a certain extent, so that the set will not expand without restraint to meet the confidence level of 1 - α. Among them, Equation (13) and Equation (14) are respectively: 。
[0043] In a specific embodiment, after step S5, it further includes: combining attention weights and long short-term memory network hidden state analysis to output an interpretable fraud localization report. Specifically, it extracts the weights at each time step in the dynamic attention mechanism, identifies high-risk claim records with weights higher than a preset threshold; combines the hidden state of the long short-term memory network to reverse-track claims with abnormal time intervals, generates a visual analysis of the time series anomaly pattern, and generates an interpretable fraud localization report.
[0044] In the embodiment of the present application, heterogeneous medical data is obtained, and the heterogeneous medical data is subjected to standardized coding and feature fusion to generate a high-dimensional fusion feature sequence; the long short-term memory network is used to extract time series features from the high-dimensional fusion feature sequence according to the claim time interval to generate a time series feature sequence; a dynamic attention mechanism is used to aggregate the features of the time series feature sequence to generate a fraud representation vector with a fixed dimension; a fully connected network is used to perform feature refinement processing on the fraud representation vector to generate a high-density feature vector; a classifier generates a prediction result based on the high-density feature vector, and performs conformal calibration on the prediction result to generate a fraud prediction set that meets the preset confidence level. The embodiment of the present invention encodes features for heterogeneous medical data, extracts features according to the claim time interval through the long short-term memory network, and at the same time uses the method of conformal calibration for prediction calibration, so that the present application can detect medical fraud for medical providers, improve the comprehensiveness of detection, and can capture the time series information of claims, which is beneficial to improving the accuracy of medical fraud detection.
[0045] In the embodiment of the present application, for capturing time series information, an improved recurrent neural network is used to capture the time series information of claims. The attention mechanism can aggregate the sequence, enabling the neural network to process it. At the same time, the weight analysis of the time series-based attention mechanism can focus on suspicious claims in the time series. The embodiment of the present application provides confidence guarantee for the prediction result. By introducing conformal prediction, this framework can obtain a prediction set for each sample, realizing providing priorities for subsequent manual review and making up for the probability deviation caused by unbalanced data.
[0046] Please refer to Figure 7 , as an implementation of the above Figure 1 shown method, the present application provides an embodiment of a medical insurance fraud detection device. This device embodiment corresponds to the Figure 1 shown method embodiment, and this device can be specifically applied to various electronic devices.
[0047] As Figure 7 shown, the medical insurance fraud detection device of this embodiment includes: a feature fusion module 61, a time series feature extraction module 62, a dynamic attention aggregation module 63, a feature refinement module 64, and a conformal calibration module 65, where: A feature fusion module 61, configured to obtain heterogeneous medical data, perform standardized coding and feature fusion on the heterogeneous medical data, and generate a high-dimensional fusion feature sequence; A temporal feature extraction module 62, configured to perform temporal feature extraction on the high-dimensional fusion feature sequence according to a claim time interval through a long short-term memory network, and generate a temporal feature sequence; A dynamic attention aggregation module 63, configured to perform feature aggregation on the temporal feature sequence by adopting a dynamic attention mechanism, and generate a fraud representation vector with a fixed dimension; A feature refinement module 64, configured to perform feature refinement processing on the fraud representation vector through a fully connected network, and generate a high-density feature vector; A conformal calibration module 65, configured to generate a prediction result based on the high-density feature vector through a classifier, and perform conformal calibration on the prediction result to generate a fraud prediction set that meets a preset confidence level.
[0048] Further, the temporal feature extraction module 62 includes: An input confirmation unit, configured to use the high-dimensional fusion feature sequence and the claim time interval as inputs of the long short-term memory network; A calculation unit, configured to calculate a forgetting gate based on the high-dimensional fusion feature sequence and the claim time interval through a first preset formula; A temporal feature sequence generation unit, configured to dynamically adjust the amount of information of a candidate memory unit according to the claim time interval, calculate an output hidden state at each time step, and splice the output hidden states to generate a temporal feature sequence.
[0049] Further, the first preset formula is: ; Wherein, is the forgetting gate, 、 、 is a randomly initialized weight, is a bias randomly initialized, is the high-dimensional fusion feature sequence, is the claim time interval, 、 is the sequence 、 's claim time.
[0050] Further, the dynamic attention aggregation module 63 includes: A key-value pair construction unit, configured to map an aggregation target to a query vector through a fully-connected layer, and construct key-value pairs from the time series feature sequence; A normalization processing unit, configured to calculate the score of each value based on the query vector and the key-value pairs by using a scaled dot product function to obtain an initial score, and perform normalization processing on the initial score to obtain a normalized score; A feature aggregation unit, configured to perform weighted summation on the normalized score and the input value to perform feature aggregation on the time series feature sequence, and generate the fraud representation vector of a fixed dimension.
[0051] Further, the conformal calibration module 65 includes: A sample partitioning unit, configured to partition the label set corresponding to the high-density feature vector to generate a plurality of independent calibration samples; A sample classification unit, configured to classify each of the independent calibration samples based on the high-density feature vector through the classifier to generate the prediction result corresponding to each of the independent calibration samples; A prediction screening unit, configured to calculate a non-conformity score based on the prediction result according to a preset conformal prediction method to generate a quantile threshold, and perform prediction screening according to the quantile threshold and the preset confidence level to generate the fraud prediction set, where the preset conformal prediction method is any one of a coverage priority method and a prediction size priority method.
[0052] Further, the feature fusion module 61 includes: A data acquisition unit, configured to acquire the heterogeneous medical data; A first vector generation unit, configured to map the discrete categorical data in the heterogeneous medical data to a continuous vector of a fixed dimension through an embedding layer to obtain a first vector; A second vector generation unit, configured to extract semantic embeddings from the text data in the heterogeneous medical data to generate a second vector; A third vector generation unit, configured to perform standardization processing on the connected numerical data in the heterogeneous medical data to generate a third vector; A non-linear fusion unit, configured to perform feature splicing on the first vector, the second vector, and the third vector, and perform non-linear fusion on the spliced features to generate the high-dimensional fusion feature sequence.
[0053] Further, the feature refinement module 64 includes: A non-linear transformation unit, configured to perform multi-layer non-linear transformation on the fraud representation vector through the fully-connected network to obtain a transformed feature vector; A feature compression unit, configured to compress the changed feature vector to generate an initial high-density feature vector; A feature splicing unit, configured to, if there are static features, refine the static features and splice the refined static features with the dynamic features in the initial high-density feature vector to generate the high-density feature vector.
[0054] To solve the above technical problems, an embodiment of the present application further provides a computer device. For details, please refer to Figure 8 , Figure 8 which is the basic structural block diagram of the computer device in this embodiment.
[0055] The computer device 7 includes a memory 71, a processor 72, and a network interface 73 that are communicatively connected to each other through a system bus. It should be noted that Figure 8 only a computer device 7 with three components, namely, a memory 71, a processor 72, and a network interface 73, is shown in
[0056] However, it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0057] The memory 71 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 71 may be an internal storage unit of the computer device 7, such as the hard disk or memory of the computer device 7. In other embodiments, the memory 71 may also be an external storage device of the computer device 7, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 7. Of course, the memory 71 may also include both the internal storage unit and the external storage device of the computer device 7. In this embodiment, the memory 71 is generally used to store the operating system and various application software installed on the computer device 7, such as the program code of the medical insurance fraud detection method. In addition, the memory 71 may also be used to temporarily store various data that have been output or will be output.
[0058] In some embodiments, the processor 72 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 72 is generally used to control the overall operation of the computer device 7. In this embodiment, the processor 72 is used to run the program code stored in the memory 71 or process data, such as running the program code of the above-mentioned medical insurance fraud detection method to implement various embodiments of the medical insurance fraud detection method.
[0059] The network interface 73 may include a wireless network interface or a wired network interface, and the network interface 73 is generally used to establish a communication connection between the computer device 7 and other electronic devices.
[0060] This application also provides another implementation manner, that is, to provide a computer-readable storage medium storing a computer program, and the computer program can be executed by at least one processor to enable at least one processor to execute the steps of a medical insurance fraud detection method as described above.
[0061] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of various embodiments of the present application.
[0062] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The drawings give preferred embodiments of the present application, but do not limit the scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is similarly within the protection scope of the present application.
Claims
1. A medical insurance fraud detection method, characterized in that, Including: Obtain heterogeneous medical data, perform standardized coding and feature fusion on the heterogeneous medical data, and generate a high-dimensional fusion feature sequence; Perform temporal feature extraction on the high-dimensional fusion feature sequence according to the claim time interval through a long short-term memory network to generate a temporal feature sequence; Adopt a dynamic attention mechanism to aggregate features of the temporal feature sequence to generate a fraud representation vector with a fixed dimension; Perform feature refinement processing on the fraud representation vector through a fully connected network to generate a high-density feature vector; Generate a prediction result based on the high-density feature vector through a classifier, and perform conformal calibration on the prediction result to generate a fraud prediction set that meets the preset confidence level.
2. The medical insurance fraud detection method according to claim 1, characterized in that The performing temporal feature extraction on the high-dimensional fusion feature sequence according to the claim time interval through a long short-term memory network to generate a temporal feature sequence includes: Use the high-dimensional fusion feature sequence and the claim time interval as the input of the long short-term memory network; Calculate a forget gate based on the high-dimensional fusion feature sequence and the claim time interval through a first preset formula; Dynamically adjust the amount of information of the candidate memory unit according to the claim time interval, calculate the output hidden state at each time step, and splice the output hidden states to generate a temporal feature sequence.
3. The medical insurance fraud detection method according to claim 2, characterized in that, The first preset formula is: ; Among them, is the forgetting gate, 、 、 are randomly initialized weights, is the bias generated by random initialization, is the high-dimensional fusion feature sequence, is the claim time interval, 、 is the sequence 、 of claim moments.
4. The medical insurance fraud detection method according to claim 1, characterized in that The adopting a dynamic attention mechanism to aggregate features of the temporal feature sequence to generate a fraud representation vector with a fixed dimension includes: Map the aggregation target to a query vector through a fully connected layer, and construct key-value pairs for the temporal feature sequence; Use a scaled dot product function to calculate the score of each value based on the query vector and the key-value pairs to obtain an initial score, and perform normalization processing on the initial score to obtain a normalized score; Perform weighted summation on the normalized score and the input value to aggregate features of the temporal feature sequence to generate the fraud representation vector with a fixed dimension.
5. The medical insurance fraud detection method according to claim 1, characterized in that The generating a prediction result based on the high-density feature vector through a classifier, and performing conformal calibration on the prediction result to generate a fraud prediction set that meets the preset confidence level includes: Divide the label set corresponding to the high-density feature vector to generate multiple independent calibration samples; Classify each independent calibration sample based on the high-density feature vector through the classifier to generate the prediction result corresponding to each independent calibration sample; Calculate a non-conformity score based on the prediction result according to a preset conformal prediction method to generate a quantile threshold, and perform prediction screening according to the quantile threshold and the preset confidence level to generate the fraud prediction set, where the preset conformal prediction method is any one of a coverage priority method and a prediction size priority method.
6. The medical insurance fraud detection method according to claim 1, wherein, The obtaining heterogeneous medical data, performing standardized coding and feature fusion on the heterogeneous medical data, and generating a high-dimensional fusion feature sequence includes: Obtain the heterogeneous medical data; Map the discrete categorical data in the heterogeneous medical data to a continuous vector with a fixed dimension through an embedding layer to obtain a first vector; Extract semantic embeddings from the text data in the heterogeneous medical data to generate a second vector; Perform normalization processing on the connection-type numerical data in the heterogeneous medical data to generate a third vector; Perform feature concatenation on the first vector, the second vector, and the third vector, and perform non-linear fusion on the concatenated features to generate the high-dimensional fusion feature sequence.
7. The medical insurance fraud detection method according to any one of claims 1 to 6, characterized in that The feature refinement processing of the fraud representation vector through the fully connected network to generate a high-density feature vector includes: Perform multi-layer non-linear transformation on the fraud representation vector through the fully connected network to obtain a transformed feature vector; Perform feature compression on the transformed feature vector to generate an initial high-density feature vector; If there are static features, perform feature refinement on the static features, and concatenate the refined static features with the dynamic features in the initial high-density feature vector to generate the high-density feature vector.
8. A medical insurance fraud detection device, characterized in that, Include: A feature fusion module for obtaining heterogeneous medical data and performing standardized encoding and feature fusion on the heterogeneous medical data to generate a high-dimensional fusion feature sequence; A time series feature extraction module for extracting time series features from the high-dimensional fusion feature sequence according to the claim time interval through a long short-term memory network to generate a time series feature sequence; A dynamic attention aggregation module for aggregating features of the time series feature sequence by using a dynamic attention mechanism to generate a fraud representation vector with a fixed dimension; A feature refinement module for performing feature refinement processing on the fraud representation vector through a fully connected network to generate a high-density feature vector; A conformal calibration module for generating a prediction result based on the high-density feature vector through a classifier and performing conformal calibration on the prediction result to generate a fraud prediction set that meets the preset confidence level.
9. A computer device, characterized in that, It includes a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it implements the medical insurance fraud detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the medical insurance fraud detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Medical insurance group abnormal behavior detection method and device
CN118378201A
Healthcare fraud detection using language modeling and co-morbidity analysis
US20140149129A1
Organized healthcare fraud detection
US20140172439A1
Cited By
Medical data monitoring and early warning method, device, equipment and medium
CN120492988A
Medical data monitoring and early warning method, device, equipment and medium
CN120492988B
DIP and DRG medical insurance risk prediction method and device, equipment and medium
CN121504637A
Discrimination and interpretability analysis fused medical insurance fraud behavior detection method and system
CN122066526A