Medical insurance fraud identification method and system

By performing spatiotemporal embedding and feature extraction on medical insurance data, and optimizing the medical insurance fraud identification model using multi-layer convolutional neural networks and self-attention mechanisms, the problems of behavioral pattern complexity and data imbalance in medical insurance fraud identification are solved, thereby improving the accuracy of medical insurance fraud identification.

CN120598694BActive Publication Date: 2025-10-28CENT SOUTH UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511098010.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-10-28
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

Existing methods for identifying medical insurance fraud fail to effectively preserve the complex patterns of claimant behavior and struggle to balance feature diversity and rationality when dealing with imbalances in medical insurance data categories, resulting in insufficient identification accuracy.

Method used

By receiving claim requests from insured individuals, obtaining historical medical data, extracting claim features, and performing spatiotemporal embedding based on time and location, the model utilizes multi-layer convolutional neural networks and self-attention mechanisms to extract claim behavior features, and combines multi-component joint loss functions to optimize the medical insurance fraud identification model.

Benefits of technology

It improves the accuracy of medical insurance fraud identification, can capture complex patterns and deep semantic information of claims behavior, and accurately establishes the correlation between claims behavior and medical insurance fraud behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598694B_ABST
    Figure CN120598694B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for identifying medical insurance fraud. The method includes receiving or responding to a claim request from an insured person and obtaining the insured person's historical medical records; extracting claim features from the historical medical records indicating the insured person's claim behavior; spatiotemporally embedding the claim features according to the time and location of the claim behavior to obtain a spatiotemporal embedding matrix; inputting the spatiotemporal embedding matrix into a pre-trained medical insurance fraud identification model, which then outputs the medical insurance fraud identification result for the insured person. This invention can improve the accuracy of medical insurance fraud identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical insurance fraud identification technology, specifically relating to a medical insurance fraud identification method and system. Background Technology

[0002] Medical insurance fraud identification is the first and crucial step in combating medical insurance fraud. Accurate fraud identification methods can improve the intelligence and efficiency of medical insurance supervision, thereby enhancing the effectiveness of medical insurance fund utilization. Current research has the following main shortcomings:

[0003] (1) Due to the complex nature of medical insurance data, existing methods fail to effectively preserve the complex patterns of claimant behavior when extracting claim features. Current methods typically use a single claim record as the smallest sample unit, or extract claim feature vectors without temporal information based on all claimant data, which ignores the changes in claimant behavior over time. In fact, according to Self-Determination Theory, when an individual's fraudulent behavior is driven by both intrinsic and extrinsic motivations, they often use continuous behavioral patterns to spread risk and increase profits. Therefore, preserving temporal attribute information based on all individual claim records helps the model capture the changing patterns of claimant behavior, thereby improving fraud detection capabilities.

[0004] (2) Regarding the imbalance of categories in medical insurance data, existing methods cannot adequately balance feature diversity and rationality. To balance the number of positive and negative samples, some methods use undersampling to reduce the number of majority class samples, but this may lead to the loss of information from the majority class samples. Most methods increase the number of minority class samples through oversampling or data generation, which can avoid information loss, but the synthesized new data is difficult to accurately capture the complex patterns of fraudulent behavior, often introducing too much noise, resulting in unrepresentative data and thus limiting the performance of the model in fraud detection.

[0005] The aforementioned shortcomings limit the performance of existing medical insurance fraud detection methods and reduce the accuracy of medical insurance fraud detection. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a method and system for identifying medical insurance fraud, so as to improve the accuracy of medical insurance fraud identification.

[0007] In a first aspect, the present invention provides a method for identifying medical insurance fraud, the method comprising the following steps:

[0008] Receive or respond to claims from insured individuals and obtain their historical medical records;

[0009] Extract claim characteristics of insured individuals from historical medical records;

[0010] Based on the time and location of the claim, the claim features are spatiotemporally embedded to obtain a spatiotemporal embedding matrix;

[0011] The spatiotemporal embedding matrix is ​​input into a pre-trained medical insurance fraud detection model, which outputs the medical insurance fraud detection results for insured individuals. The medical insurance fraud detection model includes a feature extraction module for extracting medical insurance fraud features of claim behavior, and a medical insurance fraud detection module for identifying medical insurance fraud behavior corresponding to claim behavior based on medical insurance fraud features. The medical insurance fraud detection model is optimized using a multi-component joint loss function during training, which includes binary cross-entropy loss, non-uniform center loss, and cosine distance loss.

[0012] Optional features include claim behavior features, disease diagnosis features, cost features, insured person identity features, and drug reimbursement features.

[0013] Claim behavior characteristics include frequency of visits, number of days of visits, and number of hospitals visited;

[0014] Disease diagnostic characteristics include the number of types of diseases visited and the repetition rate of diseases visited;

[0015] Cost characteristics include the amount of medical expenses;

[0016] The insured person's identity characteristics are used to indicate whether the insured person is entitled to preferential treatment;

[0017] The characteristics of drug reimbursement include drug type, drug price, and drug quantity.

[0018] Optionally, the claim features are spatiotemporally embedded based on the time and location of the claim, resulting in a spatiotemporal embedding matrix, including:

[0019] A linear layer is used to perform dimensional transformation on the claim features to obtain the dimensional transformation matrix;

[0020] Location information matrix is ​​obtained using location encoding;

[0021] The time information is encoded using torch's Embedding layer to obtain a time information matrix;

[0022] The location information matrix and the time information matrix are embedded into the dimension transformation matrix to obtain the spatiotemporal embedding matrix.

[0023] Optionally, the characteristics of medical insurance fraud in extracting claims include:

[0024] Local features of the spatiotemporal embedding matrix are extracted using a multi-layer convolutional neural network; these local features are used to characterize the correlation between claims behavior in adjacent time periods.

[0025] The self-attention mechanism is used to extract long-distance dependencies between claims behaviors in non-adjacent time periods from local features;

[0026] By integrating local features and long-range dependencies using a forward propagation layer, medical insurance fraud features can be obtained.

[0027] Optionally, the expression for the spatiotemporal embedding matrix is:

[0028]

[0029] in, Represents the spatiotemporal embedding matrix The Middle Feature vectors for each time period , , This represents the dimension after the claim feature transformation. Indicates the total number of time periods. This represents the batch normalization function. The dimension transformation matrix represents the first... Feature vectors for each time period , Indicates the first The time information vector corresponding to the feature vector of each time period. , Indicates the Location information vector of feature vectors for each time period , Indicates a linear layer. Indicates the extracted first Claim feature vectors for each time period, , This represents the total number of claim features extracted within a single time period. The first element representing the location information vector for a certain time period One dimension, , Indicates the The position vector of the i-th time period Values ​​of each dimension express The number of dimensions, here " / / " represents the integer division operator, and "%" represents the remainder division operator.

[0030] Optionally, local features of the spatiotemporal embedding matrix can be extracted using a multi-layer convolutional neural network, including:

[0031] Through calculation formula

[0032]

[0033]

[0034] Obtain a representation of the correlation between claims in adjacent time periods. ;in, This represents the output of the first convolutional layer.

[0035] Through calculation formula

[0036]

[0037] The local features are obtained .

[0038] Optionally, the expression for long-distance dependencies is:

[0039]

[0040]

[0041]

[0042]

[0043]

[0044]

[0045] in, Indicates long-distance dependency. This represents the feature matrix after merging the output matrices of all attention heads along the feature dimension. Indicates the The output matrix of each attention head, , Indicates the number of attention heads. Represents the query matrix. Represents the key matrix. Represents a value matrix, , , Depend on Obtained through linear transformation, Both represent weight matrices. express Feature dimensions, express Feature dimensions after linear transformation.

[0046] Optionally, the expression for the characteristics of medical insurance fraud is: .

[0047] Optionally, the expression for optimizing the multi-component joint loss function is:

[0048]

[0049]

[0050]

[0051]

[0052] in, This represents the multi-component joint loss function. This represents the binary classification cross-entropy loss. Indicates the cosine distance loss. Indicates non-uniform central loss. Represents the total sample size. Indicates the One sample, , Indicates the The true label of each sample, labeled as either a fraudulent sample or a normal sample. Indicates the Fraud probability prediction for each sample , These represent the center loss weights for normal samples and fraudulent samples, respectively. , These represent the number of normal samples and the number of fraudulent samples, respectively. , Indicates the The feature vector of each sample , Let represent the center vectors of normal samples and fraudulent samples, respectively, which are learnable parameters.

[0053] Secondly, the present invention provides a medical insurance fraud identification system, comprising:

[0054] The data acquisition module is used to receive or respond to claims requests from insured individuals and to acquire their historical medical records.

[0055] The first functional module is used to extract the claim characteristics of insured persons who have made claims from historical medical records.

[0056] The second functional module is used to perform spatiotemporal embedding of claim features based on the time and location of the claim behavior to obtain a spatiotemporal embedding matrix.

[0057] The data output module is used to input the spatiotemporal embedding matrix into the pre-trained medical insurance fraud identification model, and the medical insurance fraud identification model outputs the medical insurance fraud identification results of the insured persons. The medical insurance fraud identification model includes a feature extraction module for extracting medical insurance fraud features of claim behavior, and a medical insurance fraud identification module for identifying medical insurance fraud behavior corresponding to the claim behavior based on the medical insurance fraud features. The medical insurance fraud identification model is optimized using a multi-component joint loss function during training. The multi-component joint loss function includes binary cross-entropy loss, non-uniform center loss, and cosine distance loss.

[0058] The beneficial effects of this invention are:

[0059] The medical insurance fraud identification method provided by this invention performs spatiotemporal embedding of claim features based on the time and location of the claim behavior to obtain a spatiotemporal embedding matrix. This facilitates the subsequent parallel extraction of semantic information using a self-attention mechanism, thereby improving the efficiency of model training and inference and enhancing the accuracy of medical insurance fraud identification. The constructed medical insurance fraud identification model has a feature extraction module used to extract medical insurance fraud features of claim behavior, which can capture the complex patterns of claim behavior. Its medical insurance fraud identification module is used to identify the medical insurance fraud behavior corresponding to the claim behavior based on the medical insurance fraud features, which can capture the deep semantic information of claim behavior, accurately establish the correlation between claim behavior and medical insurance fraud behavior, and improve the accuracy of medical insurance fraud identification. Attached Figure Description

[0060] Figure 1 This is a flowchart of a medical insurance fraud identification method in one embodiment of this application;

[0061] Figure 2 This is a model structure diagram of a medical insurance fraud detection model in one embodiment of this application; wherein, Figure 2 (a) represents the overall structure of the medical insurance fraud detection model. Figure 2 (b) represents the self-attention mechanism. Figure 2 (c) represents a multi-layer convolutional neural network. Figure 2 (d) represents a fully connected layer. Figure 2 (e) indicates the forward propagation layer;

[0062] Figure 3 This is a schematic diagram of the medical insurance fraud identification system in one embodiment of this application. Detailed Implementation

[0063] To address the issue of poor accuracy in traditional medical insurance fraud identification methods, this invention discloses a medical insurance fraud identification method and system. This method performs spatiotemporal embedding of claim features based on the time and location of the claim, obtaining a spatiotemporal embedding matrix. This facilitates the subsequent parallel extraction of semantic information using a self-attention mechanism, thereby improving the efficiency of model training and inference, and ultimately enhancing the accuracy of medical insurance fraud identification. The constructed medical insurance fraud identification model has a feature extraction module used to extract medical insurance fraud features from claim behaviors, capturing complex patterns in claim behaviors. Its medical insurance fraud identification module is used to identify the corresponding medical insurance fraud behaviors based on these features, capturing deep semantic information from claim behaviors and accurately establishing the correlation between claim behaviors and medical insurance fraud behaviors, thus improving the accuracy of medical insurance fraud identification.

[0064] The medical insurance fraud identification method provided by this invention will be described in detail below.

[0065] like Figure 1 As shown, this method for identifying medical insurance fraud includes the following steps:

[0066] Step 11: Receive or respond to the insured person's claim request and obtain the insured person's historical medical data.

[0067] In this embodiment of the invention, the aforementioned historical medical records can be obtained by the medical insurance regulatory department from the patient database maintained by the medical institution.

[0068] In one feasible implementation, historical medical data includes patient identity information, consultation time, hospital visited, diagnosis results, consultation costs, and medication information prescribed by the doctor.

[0069] After obtaining the insured person's historical medical records, the historical medical records will be normalized.

[0070] Step 12: Extract the claim characteristics of insured individuals who made claims from historical medical records.

[0071] In embodiments of the present invention, the claim characteristics include claim behavior characteristics, disease diagnosis characteristics, cost characteristics, insured person identity characteristics, and drug reimbursement characteristics.

[0072] Among the characteristics of claim behavior are frequency of medical visits, number of days of medical visits, and number of hospitals visited.

[0073] Disease diagnostic characteristics include the number of types of diseases visited and the repetition rate of diseases visited.

[0074] Cost characteristics include the amount of medical expenses incurred.

[0075] The insured person's identity characteristics are used to indicate whether the insured person is entitled to preferential treatment. For example: whether they are disabled or military personnel.

[0076] The characteristics of drug reimbursement include drug type, drug price, and drug quantity.

[0077] Step 13: Spatiotemporally embed the claim features according to the time and location of the claim behavior to obtain the spatiotemporal embedding matrix.

[0078] Specifically, this includes steps 13.1 to 13.4.

[0079] Step 13.1: Use a linear layer to perform dimensional transformation on the claim features to obtain the dimensional transformation matrix.

[0080] In one feasible implementation, through calculation formula The dimensional transformation matrix is ​​obtained. .in, Indicates a linear layer. Indicates the extracted first Claim feature vectors for each time period, This indicates the total number of claim features extracted within a single time period.

[0081] It should be noted that the significance of dimensionality transformation lies in enabling feature interaction and fusion. For example, suppose the input feature vector is... After dimensionality transformation through a linear layer (outputting 3 nodes), we can obtain... .

[0082] Step 13.2: Use location encoding to obtain the location information matrix.

[0083] In one feasible implementation, the location information matrix is ​​obtained by calculating the following formula. .

[0084]

[0085] in, The first element representing the location information vector for a certain time period One dimension, , Indicates the The position vector of the i-th time period Values ​​of each dimension , express The number of dimensions, here " / / " represents the integer division operator, and "%" represents the remainder division operator. This indicates the total number of time periods.

[0086] Step 13.3: Encode the time information using torch's Embedding layer to obtain a time information matrix.

[0087] In one feasible implementation, through calculation formula , obtained the Time information vector corresponding to the feature vector of each time period .

[0088] Step 13.4: Embed the location information matrix and time information matrix into the dimension transformation matrix to obtain the spatiotemporal embedding matrix.

[0089] Specifically, through calculation formula The spatiotemporal embedding matrix is ​​obtained. . This represents the batch normalization function. , , The dimensional transformation matrix represents the first dimension. Feature vectors for each time period .

[0090] Each element in the spatiotemporal embedding matrix corresponds to a feature vector formed by fusing claim features, claim location, and claim time over a period of time.

[0091] Step 14: Input the spatiotemporal embedding matrix into the pre-trained medical insurance fraud identification model, and the medical insurance fraud identification model outputs the medical insurance fraud identification results of the insured persons.

[0092] The medical insurance fraud detection model includes a feature extraction module for extracting medical insurance fraud characteristics of claim behavior, and a medical insurance fraud detection module for identifying medical insurance fraud behavior corresponding to the claim behavior based on the medical insurance fraud characteristics. In one feasible implementation, the model structure of the medical insurance fraud detection model is as follows: Figure 2 As shown, Figure 2 (a) represents the overall structure of the medical insurance fraud identification model.

[0093] In an embodiment of the present invention, the feature extraction module and the medical insurance fraud identification module are connected in sequence.

[0094] The feature extraction module and the medical insurance fraud detection module are explained below.

[0095] For the feature extraction module, extract medical insurance fraud features of claim behavior, including: steps 14.1.1 to 14.1.3.

[0096] Step 14.1.1: Extract local features of the spatiotemporal embedding matrix using a multi-layer convolutional neural network.

[0097] In this embodiment of the invention, local features are used to characterize the correlation between claims in adjacent time periods.

[0098] One feasible implementation utilizes a multi-layer convolutional neural network (such as...) Figure 2 (c) The process of extracting local features of the spatiotemporal embedding matrix includes steps 14.1.1.1 to 14.1.1.2.

[0099] Step 14.1.1.1: Using the embedding matrix as input, a two-layer convolutional network is used to capture the correlation between behavioral features in adjacent time periods in the claim data.

[0100] Specifically, the correlation between claims in adjacent time periods is obtained by calculating the following formula. .

[0101]

[0102]

[0103] in, This represents the output of the first convolutional layer.

[0104] Step 14.1.1.2: Use linear layers and residual layers to integrate local information, solve the gradient vanishing and overfitting problems in deep networks, and obtain a feature matrix that integrates behavioral information from nearby time periods.

[0105] Specifically, local features are obtained by calculating the following formula. .

[0106] .

[0107] Step 14.1.2: Use the self-attention mechanism to obtain the long-distance dependency between claim behaviors in non-adjacent time periods from local features.

[0108] One feasible implementation utilizes a self-attention mechanism (such as...) Figure 2 (b) The process of obtaining long-distance dependencies between claims behaviors in non-adjacent time periods from local features includes steps 14.1.2.1 to 14.1.2.3.

[0109] Step 14.1.2.1, for Perform linear transformations to obtain the query matrices respectively. Key matrix Sum matrix And divide each matrix into feature dimensions. One point of attention.

[0110] The formal expression is as follows:

[0111]

[0112]

[0113]

[0114] in, Represents the query matrix. Represents the key matrix. Represents a value matrix, , , Depend on Obtained through linear transformation, Both represent weight matrices. express The feature dimensions.

[0115] Step 14.1.2.2: For each attention head, calculate the attention score matrix separately, and further adaptively and dynamically fuse the value vectors of all time periods into the feature vector of the target time period in order to capture long-distance dependencies.

[0116] Specifically, the first number is obtained by calculating the following formula. The output matrix of each attention head.

[0117]

[0118] in, Indicates the The output matrix of each attention head, , Indicates the number of attention heads. express Feature dimensions after linear transformation.

[0119] Step 14.1.2.3: The output matrices of all attention heads are merged along the feature dimension, and then a linear layer is used to integrate all the information. A residual layer is used to prevent the model from overfitting, resulting in the output matrix of the multi-head attention mechanism layer, i.e., the long-distance dependency.

[0120] Specifically, through the merging formula The feature matrix after merging multiple attention points is obtained, and then further calculated using the formula... To obtain long-distance dependencies .

[0121] Step 14.1.3, utilize the forward propagation layer (e.g., Figure 2 (e) integrates local features and long-distance dependencies to obtain medical insurance fraud features.

[0122] In an embodiment of the present invention, the expression for the characteristics of medical insurance fraud is:

[0123] .

[0124] In embodiments of the present invention, a fully connected layer (such as...) is utilized. Figure 2 (d) Establish a mapping relationship between claim behavior characteristics and medical insurance fraud, so that the model can output the medical insurance fraud probability of each claimant.

[0125] Specifically, through calculation formula Obtain the probability of medical insurance fraud .

[0126] The training process of the medical insurance fraud detection model in this embodiment of the invention is described below, specifically including:

[0127] Using pre-acquired samples of identified medical insurance fraud, a medical insurance fraud detection model is trained. During training, a multi-component joint loss function is used for optimization. Training terminates when the multi-component joint loss value of the medical insurance fraud detection model is less than a preset loss threshold, resulting in a pre-trained medical insurance fraud detection model. For example, the Adam optimizer can be used to optimize the medical insurance fraud detection model, with the learning rate set to decay with increasing training epochs.

[0128] In this embodiment of the invention, the multi-component joint loss function includes binary cross-entropy loss, non-uniform center loss, and cosine distance loss.

[0129] It's important to note that a single task can only ensure that the model captures deep information within that specific task. This information-capturing ability may be biased, leading to an overemphasis on certain information while neglecting other information helpful to the task. Therefore, in deep learning research, researchers often utilize multi-task collaborative training methods, directly manifested in the use of joint loss functions. This allows the model to focus on a wider range of features. For example, in natural language processing, BERT encoders are trained using a combination of cloze test and sentence matching tasks. In this embodiment of the invention, to improve the model's performance in learning the classification boundaries of positive and negative samples, a center loss and a cosine loss are introduced on top of the classification loss (cross-entropy) to constrain the model's learning of the feature space of positive and negative samples, thereby making the classification boundaries of positive and negative samples more explicit.

[0130] In one feasible implementation, the expression for optimizing the multi-component joint loss function is as follows:

[0131]

[0132]

[0133]

[0134]

[0135] in, This represents the joint loss function of multiple components. Represents the cross-entropy loss in binary classification. Indicates the cosine distance loss. Indicates non-uniform central loss. Represents the total sample size. Indicates the One sample, , Indicates the The true label of each sample, labeled as either a fraudulent sample or a normal sample. Indicates the Fraud probability prediction for each sample , These represent the center loss weights for normal samples and fraudulent samples, respectively. , These represent the number of normal samples and the number of fraudulent samples, respectively. , Indicates the The feature vector of each sample , Let represent the center vectors of normal samples and fraudulent samples, respectively, which are learnable parameters.

[0136] The medical insurance fraud identification method provided by this invention has the following advantages: It performs spatiotemporal embedding of claim features based on the time and location of the claim behavior to obtain a spatiotemporal embedding matrix, which facilitates the subsequent parallel extraction of semantic information using a self-attention mechanism, thereby improving the efficiency of model training and inference. This is beneficial to improving the accuracy of medical insurance fraud identification. The constructed medical insurance fraud identification model has a feature extraction module used to extract medical insurance fraud features from claim behavior, capable of capturing complex patterns of claim behavior. Its medical insurance fraud identification module is used to identify the corresponding medical insurance fraud behavior based on the medical insurance fraud features, capable of capturing deep semantic information of claim behavior, accurately establishing the correlation between claim behavior and medical insurance fraud behavior, and improving the accuracy of medical insurance fraud identification.

[0137] The medical insurance fraud identification system provided by this invention will be described below.

[0138] like Figure 3 As shown, the medical insurance fraud detection system 300 includes:

[0139] The data acquisition module 301 is used to receive or respond to the claim requests of insured persons and acquire the historical medical data of insured persons;

[0140] The first functional module 302 is used to extract the claim characteristics of insured persons who have made claims from historical medical data;

[0141] The second functional module 303 is used to perform spatiotemporal embedding of claim features based on the time and location of the claim behavior to obtain a spatiotemporal embedding matrix.

[0142] The data output module 304 is used to input the spatiotemporal embedding matrix into the pre-trained medical insurance fraud identification model, and the medical insurance fraud identification model outputs the medical insurance fraud identification results of the insured persons. The medical insurance fraud identification model includes a feature extraction module for extracting medical insurance fraud features of claim behavior, and a medical insurance fraud identification module for identifying medical insurance fraud behavior corresponding to the claim behavior based on the medical insurance fraud features. The medical insurance fraud identification model is optimized by a multi-component joint loss function during training. The multi-component joint loss function includes binary cross-entropy loss, non-uniform center loss and cosine distance loss.

[0143] It should be noted that the information interaction and execution process between the above modules are based on the same concept as the method embodiments of this application. Their specific functions and technical effects can be found in the method embodiments section, and will not be repeated here. Those skilled in the art will understand that, for ease of description and brevity, the division of the above functional units and modules is only used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0144] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of protection of this application is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of one or more embodiments of this application as described above, which are not provided in detail for the sake of brevity.

[0145] One or more embodiments in this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments in this application should be included within the protection scope of this application.

Claims

1. A method for identifying medical insurance fraud, characterized in that, include: Receive or respond to claims from insured individuals and obtain their historical medical records. Extract the claim characteristics of the insured person's claim behavior from the historical medical data; Based on the time and location of the claim, the claim features are spatiotemporally embedded to obtain a spatiotemporal embedding matrix; The spatiotemporal embedding matrix is ​​input into a pre-trained medical insurance fraud detection model, which outputs the medical insurance fraud detection result for the insured person. The medical insurance fraud detection model includes a feature extraction module for extracting medical insurance fraud features of claim behavior, and a medical insurance fraud detection module for identifying medical insurance fraud behavior corresponding to the claim behavior based on the medical insurance fraud features. The medical insurance fraud detection model is optimized using a multi-component joint loss function during training, which includes binary cross-entropy loss, non-uniform center loss, and cosine distance loss. The extraction of medical insurance fraud features of claim behavior includes: extracting local features of the spatiotemporal embedding matrix using a multi-layer convolutional neural network; the local features are used to characterize the correlation between claim behaviors in adjacent time periods; using a self-attention mechanism to obtain the long-distance dependency between claim behaviors in non-adjacent time periods from the local features; and using a forward propagation layer to integrate the local features and the long-distance dependency to obtain the medical insurance fraud features. The result of the medical insurance fraud identification is the probability of medical insurance fraud; the medical insurance fraud identification module uses a fully connected layer to establish a mapping relationship between claim behavior characteristics and medical insurance fraud. The multi-component joint loss function, based on the binary cross-entropy loss, introduces non-uniform center loss and cosine distance loss to constrain the medical insurance fraud identification model's learning of the feature space of normal samples and fraud samples, thereby making the classification boundaries of normal samples and fraud samples more obvious.

2. The medical insurance fraud identification method according to claim 1, characterized in that, The claim characteristics include claim behavior characteristics, disease diagnosis characteristics, cost characteristics, insured person identity characteristics, and drug reimbursement characteristics. The characteristics of the claim behavior include the frequency of medical visits, the number of days of medical visits, and the number of hospitals visited; The disease diagnostic characteristics include the number of types of diseases treated and the repetition rate of diseases treated. The cost characteristics include the amount of medical expenses; The insured person's identity characteristics are used to indicate whether the insured person is entitled to preferential treatment; The characteristics of drug reimbursement include drug type, drug price, and drug quantity.

3. The medical insurance fraud identification method according to claim 2, characterized in that, The step of spatiotemporally embedding the claim features based on the time and location of the claim to obtain a spatiotemporal embedding matrix includes: The claim features are dimensionally transformed using a linear layer to obtain a dimensional transformation matrix; Location information matrix is ​​obtained using location encoding; The time information is encoded using torch's Embedding layer to obtain a time information matrix; The location information matrix and the time information matrix are embedded into the dimensional transformation matrix to obtain the spatiotemporal embedding matrix.

4. The medical insurance fraud identification method according to claim 3, characterized in that, The expression for the spatiotemporal embedding matrix is: in, Represents the spatiotemporal embedding matrix The Middle Feature vectors for each time period , , This represents the dimension after the claim feature transformation. Indicates the total number of time periods. This represents the batch normalization function. The dimension transformation matrix represents the first... Feature vectors for each time period , Indicates the first The time information vector corresponding to the feature vector of each time period. , Indicates the Location information vector of feature vectors for each time period , Indicates a linear layer. Indicates the extracted first Claim feature vectors for each time period, , This represents the total number of claim features extracted within a single time period. The first element representing the location information vector for a certain time period One dimension, , Indicates the The position vector of the i-th time period Values ​​of each dimension express The number of dimensions, here " / / " represents the integer division operator, and "%" represents the remainder division operator.

5. The medical insurance fraud identification method according to claim 4, characterized in that, The extraction of local features from the spatiotemporal embedding matrix using a multi-layer convolutional neural network includes: Through calculation formula Obtain a representation of the correlation between claims in adjacent time periods. ;in, This represents the output of the first convolutional layer. Through calculation formula The local features are obtained .

6. The medical insurance fraud identification method according to claim 5, characterized in that, The expression for the long-distance dependency is: in, Indicates long-distance dependency. This represents the feature matrix after merging the output matrices of all attention heads along the feature dimension. Indicates the The output matrix of each attention head, , Indicates the number of attention heads. Represents the query matrix. Represents the key matrix. Represents a value matrix, , , Depend on Obtained through linear transformation, Both represent weight matrices. express Feature dimensions, express Feature dimensions after linear transformation.

7. The medical insurance fraud identification method according to claim 6, characterized in that, The expression for the characteristics of medical insurance fraud is: .

8. The medical insurance fraud identification method according to claim 1, characterized in that, The expression for optimizing the multi-component joint loss function is as follows: in, This represents the multi-component joint loss function. This represents the binary classification cross-entropy loss. Indicates the cosine distance loss. Indicates non-uniform central loss. Represents the total sample size. Indicates the One sample, , Indicates the The true label of each sample, labeled as either a fraudulent sample or a normal sample. Indicates the Fraud probability prediction for each sample , These represent the center loss weights for normal samples and fraudulent samples, respectively. , These represent the number of normal samples and the number of fraudulent samples, respectively. , Indicates the The feature vector of each sample , Let represent the center vectors of normal samples and fraudulent samples, respectively, which are learnable parameters.

9. A medical insurance fraud detection system, characterized in that, include: The data acquisition module is used to receive or respond to the claim requests of insured persons and acquire the historical medical data of the insured persons. The first functional module is used to extract the claim characteristics of the insured person's claim behavior from the historical medical data; The second functional module is used to perform spatiotemporal embedding of the claim features based on the time and location of the claim behavior to obtain a spatiotemporal embedding matrix. The data output module is used to input the spatiotemporal embedding matrix into a pre-trained medical insurance fraud detection model, and the medical insurance fraud detection model outputs the medical insurance fraud detection results of the insured person. The medical insurance fraud detection model includes a feature extraction module for extracting medical insurance fraud features of claim behavior, and a medical insurance fraud detection module for identifying medical insurance fraud behavior corresponding to the claim behavior based on the medical insurance fraud features. The medical insurance fraud detection model is optimized using a multi-component joint loss function during training. The multi-component joint loss function includes binary cross-entropy loss, non-uniform center loss, and cosine distance loss. The extraction of medical insurance fraud features of claim behavior includes: extracting local features of the spatiotemporal embedding matrix using a multi-layer convolutional neural network; the local features are used to characterize the correlation between claim behaviors in adjacent time periods; using a self-attention mechanism to obtain the long-distance dependency relationship between claim behaviors in non-adjacent time periods from the local features; and using a forward propagation layer to integrate the local features and the long-distance dependency relationship to obtain the medical insurance fraud features. The result of the medical insurance fraud identification is the probability of medical insurance fraud; the medical insurance fraud identification module uses a fully connected layer to establish a mapping relationship between claim behavior characteristics and medical insurance fraud. The multi-component joint loss function, based on the binary cross-entropy loss, introduces non-uniform center loss and cosine distance loss to constrain the medical insurance fraud identification model's learning of the feature space of normal samples and fraud samples, thereby making the classification boundaries of normal samples and fraud samples more obvious.

Citation Information

Patent Citations

  • Method for identifying patient with insurance fraud behavior based on improved generative adversarial network

    CN113628057A