Medical insurance fraud detection method and device based on particle swarm optimization generative adversarial network
Through particle swarm optimization generation adversarial network generation and screening synthetic sample data, the training sample set is constructed, which solves the problem of imbalance in medical insurance data categories and improves the accuracy of medical insurance fraud detection.
Patent Information
- Application Number
- CN202510743933.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The accuracy of medical insurance fraud detection in the prior art is mainly due to the extreme category imbalance of medical insurance data, which has significantly reduced the sensitivity of traditional supervision models to a few categories, resulting in high missed detection rates and low generalization performance.
The method of generating adversarial networks based on particle swarm optimization is adopted, and synthetic sample data is generated by generating adversarial networks, and synthetic sample data is screened using particle swarm optimization algorithm to build a training sample set, improve category balance, and train a medical insurance fraud detection model.
The accuracy of medical insurance fraud detection is improved, and the performance of the model is enhanced by augmenting a few types of samples and selecting feature, and the problem of category imbalance is solved.
Smart Images

Figure CN120277418A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a medical insurance fraud detection method and device based on a particle swarm optimization generative adversarial network. Background Art
[0002] In the prior art, there are ways to use traditional supervised models (such as logistic regression and random forest) for medical insurance fraud detection. However, due to the extremely low proportion of real samples of medical insurance fraud, there is an extreme class imbalance problem in medical insurance data. When training traditional supervised models, since the loss function is dominated by the majority class, the sensitivity to the minority class is significantly reduced, resulting in a high missed detection rate and low generalization performance. The accuracy of using supervised models to detect medical insurance fraud in the prior art is low. Summary of the Invention
[0003] The present invention provides a medical insurance fraud detection method and device based on a particle swarm optimization generative adversarial network, which is used to solve the defect of low accuracy in medical insurance fraud detection in the prior art and achieve the effect of improving the accuracy of medical insurance fraud detection.
[0004] The present invention provides a medical insurance fraud detection method based on a particle swarm optimization generative adversarial network, including: Inputting random noise into the generator of the trained generative adversarial network to obtain synthetic sample data output by the generator. The synthetic sample data includes synthetic sample data to be detected and fraud labels corresponding to the synthetic sample data to be detected. The generative adversarial network further includes a discriminator. The training data of the generative adversarial network includes real sample data. The real sample data includes real sample data to be detected and fraud labels corresponding to the real sample data to be detected; Based on the first particle swarm optimization algorithm, screening the synthetic sample data to obtain a plurality of screened synthetic sample data. In the first particle swarm optimization algorithm, each particle corresponds to a sample set. The sample set includes a plurality of the synthetic sample data. The fitness value of each particle is determined based on the global diversity of the sample set corresponding to the particle and the local density of each synthetic sample data in the sample set; Combining the plurality of screened synthetic sample data and the plurality of real sample data to obtain a training sample set, and training a classification model based on the training sample set to obtain a medical insurance fraud detection model; Inputting the data to be detected into the medical insurance fraud detection model to obtain a fraud detection result output by the medical insurance fraud detection model.
[0005] A medical insurance fraud detection method based on a particle swarm optimization generative adversarial network provided by the present invention. Before inputting random noise into the generator of the trained generative adversarial network, it includes: Obtain real medical insurance data, where the real medical insurance data includes various types of features; Use the second particle swarm optimization algorithm to determine the target feature type among the various types of features. Among them, the particles in the second particle swarm optimization algorithm correspond to a feature type set, and the feature type set includes at least one feature type. The fitness value of the particles in the second particle swarm optimization algorithm reflects the discrimination ability and independence of the feature types in the feature type set corresponding to the particles; Construct the real sample data based on the target feature type.
[0006] A medical insurance fraud detection method based on a particle swarm optimization generative adversarial network provided by the present invention. The fitness value of the particles in the second particle swarm optimization algorithm is determined based on the first formula; The first formula is: ; Among them, represents the fitness value of the particle corresponding to the feature type set X, , , are weight coefficients, represents the mutual information between X and the fraud label set Y, , n is the total number of samples, represents the value vector of the feature type of the i-th sample in the feature type set X, is the fraud label of the i-th sample, represents the average value vector of the feature types of all samples in the feature type set X, represents the mean value of the fraud labels of all samples, represents the number of feature types in the feature type set X, and are the standard deviation vectors of X and Y respectively.
[0007] A medical insurance fraud detection method based on a particle swarm optimization generative adversarial network provided by the present invention. The local density calculation formula in the fitness value of the particles in the first particle swarm optimization algorithm is: ; ; Among them, is the local density of the synthetic sample data s, is the estimated probability density of the synthetic sample data s, k is the total number of the synthetic sample data, represents the Euclidean distance between the synthetic sample data s and the synthetic sample data ; is the standard deviation of the real sample data; d is the feature dimension in the synthetic sample data, is a normalization constant.
[0008] According to a medical insurance fraud detection method based on particle swarm optimization generative adversarial network provided by the present invention, the global diversity calculation formula in the fitness value of particles in the first particle swarm optimization algorithm is: ; wherein, represents the global diversity of the sample set corresponding to the particle, represents the 1-Wasserstein distance between the distribution of the sample set and the distribution of the real sample data.
[0009] According to a medical insurance fraud detection method based on particle swarm optimization generative adversarial network provided by the present invention, the number of the synthetic sample data is N times the difference between the number of samples of the majority class and the minority class in the real sample data, and N is greater than or equal to 2.
[0010] The present invention also provides a medical insurance fraud detection device based on particle swarm optimization generative adversarial network, including: A sample synthesis module, configured to input random noise into a generator of a trained generative adversarial network to obtain synthetic sample data output by the generator, where the synthetic sample data includes synthetic sample data to be detected and fraud labels corresponding to the synthetic sample data to be detected, the generative adversarial network further includes a discriminator, and the training data of the generative adversarial network includes real sample data, and the real sample data includes real sample data to be detected and fraud labels corresponding to the real sample data to be detected; A sample screening module, configured to screen the synthetic sample data based on the first particle swarm optimization algorithm to obtain a plurality of screened synthetic sample data. In the first particle swarm optimization algorithm, each particle corresponds to a sample set, the sample set includes a plurality of the synthetic sample data, and the fitness value of each particle is determined based on the global diversity of the sample set corresponding to the particle and the local density of each synthetic sample data in the sample set; A training module, configured to combine a plurality of the screened synthetic sample data and a plurality of the real sample data to obtain a training sample set, and train a classification model based on the training sample set to obtain a medical insurance fraud detection model; A detection module for inputting data to be detected into the medical insurance fraud detection model to obtain a fraud detection result output by the medical insurance fraud detection model.
[0011] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the medical insurance fraud detection method based on particle swarm optimization generative adversarial network as described in any one of the above is implemented.
[0012] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the medical insurance fraud detection method based on particle swarm optimization generative adversarial network as described in any one of the above is implemented.
[0013] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the medical insurance fraud detection method based on particle swarm optimization generative adversarial network as described in any one of the above is implemented.
[0014] The medical insurance fraud detection method and device based on particle swarm optimization generative adversarial network provided by the present invention. The medical insurance fraud detection method based on particle swarm optimization generative adversarial network includes: inputting random noise into the generator of the trained generative adversarial network to obtain synthetic sample data output by the generator. The synthetic sample data includes synthetic sample data to be detected and fraud labels corresponding to the synthetic sample data to be detected. The generative adversarial network also includes a discriminator. The training data of the generative adversarial network includes real sample data. The real sample data includes real sample data to be detected and fraud labels corresponding to the real sample data to be detected; screening the synthetic sample data based on the first particle swarm optimization algorithm to obtain multiple screened synthetic sample data. In the first particle swarm optimization algorithm, each particle corresponds to a sample set, and the sample set includes multiple synthetic sample data. The fitness value of each particle is determined based on the global diversity of the sample set corresponding to the particle and the local density of each synthetic sample data in the sample set; combining the multiple screened synthetic sample data and multiple real sample data to obtain a training sample set, training a classification model based on the training sample set to obtain a medical insurance fraud detection model, and inputting the data to be detected into the medical insurance fraud detection model to obtain a fraud detection result output by the medical insurance fraud detection model.
[0015] In this way, by training the generative adversarial network using real sample data, the generator in the generative adversarial network can generate synthetic sample data. Then, through the particle swarm optimization algorithm, the synthetic sample data generated by the generator is screened to obtain screened synthetic sample data with better global diversity and local density. Together with the real sample data, they form a training sample set to train the classification model. This can expand the minority classes in the real sample data, improve the class balance of the training sample set, and thus improve the performance of the trained medical insurance fraud detection model, achieving the effect of improving the accuracy of medical insurance fraud detection. Brief Description of the Drawings
[0016] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It is a flowchart of the medical insurance fraud detection method based on the particle swarm optimization generative adversarial network provided by the present invention.
[0018] Figure 2 It is a flowchart of the model training process in the medical insurance fraud detection method based on the particle swarm optimization generative adversarial network provided by the present invention.
[0019] Figure 3 It is a structural diagram of the medical insurance fraud detection device based on the particle swarm optimization generative adversarial network provided by the present invention.
[0020] Figure 4 It is a structural diagram of the electronic device provided by the present invention. Detailed Embodiments
[0021] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0022] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0023] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0024] It should be further understood that the term "and / or" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0025] As used in this specification and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrases "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.
[0026] The following Figure 1 - Figure 2 describes the medical insurance fraud detection method based on the particle swarm optimization generative adversarial network provided by the present invention. As Figure 1 shown, the medical insurance fraud detection method based on the particle swarm optimization generative adversarial network includes the steps of: S110. Input random noise into the generator of the trained generative adversarial network to obtain synthetic sample data output by the generator. The synthetic sample data includes synthetic sample data to be detected and fraud labels corresponding to the synthetic sample data to be detected. The generative adversarial network also includes a discriminator. The training data of the generative adversarial network includes real sample data. The real sample data includes real sample data to be detected and fraud labels corresponding to the real sample data to be detected; S120. Based on the first particle swarm optimization algorithm, screen the synthetic sample data to obtain a plurality of screened synthetic sample data. In the first particle swarm optimization algorithm, each particle corresponds to a sample set. The sample set includes a plurality of synthetic sample data. The fitness value of each particle is determined based on the global diversity of the sample set corresponding to the particle and the local density of each synthetic sample data in the sample set; S130. Combine the plurality of screened synthetic sample data and the plurality of real sample data to obtain a training sample set, and train a classification model based on the training sample set to obtain a medical insurance fraud detection model; S140. Input the data to be detected into the medical insurance fraud detection model to obtain the fraud detection result output by the medical insurance fraud detection model.
[0027] Medical insurance fraud refers to the act of fabricating false medical data to claim compensation from medical insurance institutions (such as insurance companies selling commercial medical insurance, etc.). Medical insurance fraud detection aims to identify whether there is false data in medical insurance claim data. That is to say, medical insurance fraud detection can be regarded as a binary classification problem, which divides the medical data used for claims into two categories: with false data and without false data. The medical insurance fraud detection method based on particle swarm optimization generative adversarial network provided by the present invention trains the generative adversarial network by using real sample data, so that the generator in the generative adversarial network can generate synthetic sample data, and through the particle swarm optimization algorithm, screens the synthetic sample data generated by the generator to obtain screening synthetic sample data with better global diversity and local density, which together with the real sample data constitute a training sample set to train the classification model. In this way, the minority class in the real sample data can be expanded, the class balance of the training sample set can be improved, and thus the performance of the trained medical insurance fraud detection model can be improved, achieving the effect of improving the accuracy of medical insurance fraud detection.
[0028] Construct real sample data based on real medical insurance claim data, and the medical insurance fraud labels in the real sample data are obtained through manual annotation. The medical insurance claim data includes many different types of features, such as drug combinations, diagnosis codes, etc. In one possible implementation, all types of features in the medical insurance claim data can be used for medical insurance fraud detection. However, this method will lead to high dimensionality and sparsity of the data that the model needs to analyze, resulting in the model being easily interfered by noise, low training efficiency and easy overfitting, and it is difficult to capture the implicit association patterns of fraud behaviors. In another possible implementation of the method provided by the present invention, a large number of redundant features included in the medical insurance claim data are screened. Specifically, before inputting random noise into the generator of the trained generative adversarial network, it includes: Obtain real medical insurance data, which includes various types of features; Use the second particle swarm optimization algorithm to determine the target feature type among various types of features. Among them, the particles in the second particle swarm optimization algorithm correspond to a set of feature types, and the set of feature types includes at least one feature type. The fitness value of the particles in the second particle swarm optimization algorithm reflects the discriminative ability and independence of the feature types in the set of feature types corresponding to the particles; Construct real sample data based on the target feature type.
[0029] By using the second particle swarm optimization algorithm to screen the feature types, determine the target feature type to construct real sample data, dynamically balance the feature discriminability and redundancy consistency, enhance the robustness of feature selection under high-dimensional coefficient data, and solve the problem of insufficient semantic expression ability in complex data.
[0030] Such asFigure 2 As shown in Figure 2 , when using the second particle swarm optimization algorithm to screen feature types, the particle swarm is first initialized: the size of the particle swarm is defined as N particles, and the position of each particle represents a set of feature types, which can be represented by a binary encoding. The number of digits in this encoding represents the total number of feature types in the medical insurance data. When the position corresponding to a certain feature type is 1, it means that the feature type is selected, and 0 means that the feature type is not selected. Randomly initialize the positions and velocities of the particles, and set the maximum number of iterations. In each iteration, for each particle, calculate the fitness value, and perform retention or update processing on the particle according to the fitness value of the particle.
[0031] In the second particle swarm optimization algorithm, the fitness value of a particle reflects the discriminative ability and independence of the feature types in the set of feature types corresponding to the particle. Specifically, the fitness value of a particle in the second particle swarm optimization algorithm is determined based on the first formula, and the first formula is: ; where represents the fitness value of the particle corresponding to the set of feature types X, , , are weight coefficients, represents the mutual information between X and the fraud label set Y, , n is the total number of samples, represents the value vector of the feature types of the i-th sample in the set of feature types X, is the fraud label of the i-th sample, represents the average value vector of the feature types of all samples in the set of feature types X, represents the mean of the fraud labels of all samples, represents the number of feature types in the set of feature types X, and are the standard deviation vectors of X and Y respectively.
[0032] The samples in the above formula refer to the real medical insurance data without feature selection. Each real medical insurance data includes the feature values of various feature types and the corresponding fraud labels. For each feature type set, since the feature types are selected in the feature type set, the value vectors of the feature types of each sample in different feature type sets are different. For example, assume that there are a total of 4 feature types: A, B, C, and D, and there are 2 samples: A and B. The feature type set C1 only includes the feature types A, B, and C. The values of A, B, C, and D in sample A are a, b, c, and d respectively, and the values of A, B, C, and D in sample B are e, f, g, and h respectively. Then the value vector of the feature types of sample A in the feature type set C1 is (a, b, c), and the value vector of the feature types of sample B in the feature type set C1 is (e, f, g). The average value vector of the feature types of all samples in the feature set C1 is ((a + e) / 2, (b + f) / 2, (c + g) / 2). Similarly, if the feature type set C2 includes the feature types A and B, then the value vector of the feature types of sample A in the feature type set C2 is (a, b), and so on.
[0033] In each iteration, based on the fitness value of the particle, update the velocity and position of the particle. The particle velocity update formula is: ; The particle position update formula is: ; In the particle velocity update formula, the superscript t represents the value in the t-th iteration. The position of each particle is a binary vector. 0 indicates that the corresponding feature type at this position is not selected, and 1 indicates that the corresponding feature type at this position is selected; w is the inertia weight, which is used to balance global search and local optimization. 、 are the learning factors, which are used to control the degree to which the particle follows the individual optimum and the swarm optimum respectively; 、 are random numbers within [0, 1], which ensure the diversity of the feature type set through random perturbation; represents the feature type set (binary vector) with the highest fitness found by the particle during the iteration process; is the feature type set (binary vector) with the highest fitness found by the entire particle swarm during the iteration process; A probability value between 0 and 1 can be obtained, which is used to determine the position update of the particle.
[0034] After the iteration is completed, select the feature type set corresponding to the globally optimal particle for constructing real sample data. Optimize the multi-objective optimization strategy iteration to screen the feature type set through the second particle algorithm, dynamically balance the feature discrimination ability and redundancy suppression, and significantly enhance the semantic expression ability of high-dimensional sparse data. As Figure 2 shown, after constructing multiple real sample data, use these real sample data to train the generative adversarial network so that the generator in the generative adversarial network can generate synthetic sample data similar to the real sample data.
[0035] The generative adversarial network (WGAN) includes a generator and a discriminator (which can also be called the discriminator in Figure 2 ). The input of the generator is random noise, such as a 100-dimensional random noise vector z~N(0,1), and the output is the node feature , and the node feature includes the feature values corresponding to each type in the target feature type. The input of the discriminator is the node feature, and the data input to the discriminator is generated by the generator or is real. The output of the discriminator is a real number score, and this score reflects whether the data input to the discriminator is real or generated. The network architectures of the discriminator and the generator can adopt existing network structures, such as a 3-layer fully connected layer.
[0036] The training objective of the generator is to generate synthetic sample data as close as possible to the real sample data, and the training objective of the discriminator is to maximize the score difference between the real sample data and the sample data generated by the generator. According to these training objectives, the loss functions of the generator and the discriminator can be constructed.
[0037] Specifically, the loss function of the generator should be constructed based on minimizing the average score of the discriminator for the synthetic samples. The loss function of the generator can be expressed as: ; where: represents the expected value of the noise vector z following the distribution P z (z) (such as the standard normal distribution), represents the score of the discriminator for the samples generated by the generator G based on the noise z.
[0038] The loss function of the discriminator should be constructed based on maximizing the difference between the real sample score and the synthetic sample score. The loss function of the discriminator can be expressed as: ; where, represents the score of the discriminator for the samples generated by the real sample x, represents the expected value of the distribution followed by the real sample.
[0039] As shown Figure 2 in the figure, the training iteration loop process of the generative adversarial network includes: 1. Generate noise samples.
[0040] 2. Generate fake data.
[0041] 3. Calculate the discriminator loss .
[0042] 4. Use the gradient descent method to update the parameters of the discriminator according to the loss .
[0043] 5. Calculate the total loss of the generator .
[0044] 6. Use the gradient descent method to update the parameters of the generator according to the loss of the generator.
[0045] 7. Repeat the above steps until the preset number of iterations is reached or other stopping conditions are met.
[0046] After the training of the generative adversarial network is completed, the generator therein is used to generate synthetic sample data. In the method provided by the present invention, the generated synthetic sample data is also screened to obtain screened synthetic sample data with better global diversity and local density, so as to improve the data quality of the training data set of the medical insurance fraud detection model. When generating synthetic sample data, the fraud label corresponding to the synthetic sample data is the minority class in the real sample data. In the medical insurance claim data, the data labeled as fraud is the minority class, and the number of synthetic sample data is N times the difference between the number of majority class and minority class samples in the real sample data, where N is greater than or equal to 2, so that the minority class samples can be fully expanded. Specifically, the generator can be used to generate N gen =(N maj -N min ) 2 synthetic sample data, where N maj and N min respectively represent the number of majority class and minority class samples in the real sample data. Using the first particle swarm optimization algorithm, the synthetic sample data is screened to obtain multiple screened synthetic sample data. In the first particle swarm optimization algorithm, each particle corresponds to a sample set, and the sample set includes multiple of the synthetic sample data output by the generator. The position of each particle can be represented by a binary vector, the number of bits of the binary vector is the same as the total number of synthetic sample data, and the value of each bit is 0 or 1, which is used to indicate whether the synthetic sample data corresponding to this bit is selected. The specific process of the first particle swarm optimization algorithm is described below.
[0047] First, initialize the particle swarm. After that, for each particle in each iteration, calculate its corresponding fitness value. In the method provided by the present invention, the fitness value of the particle in the first particle swarm optimization algorithm reflects the global diversity of the sample set corresponding to the particle and the local density of each synthetic sample data in the sample set.
[0048] Specifically, the formula for calculating the local density in the fitness value of the particle in the first particle swarm optimization algorithm is: ; ; Among them, is the local density of the synthetic sample data s, is the estimated probability density of the synthetic sample data s, k is the total number of synthetic sample data, represents the Euclidean distance between the synthetic sample data s and the synthetic sample data ; is the standard deviation of the real sample data; d is the feature dimension in the synthetic sample data.
[0049] The formula for calculating the global diversity in the fitness value of the particle in the first particle swarm optimization algorithm is: ; Among them, represents the global diversity of the sample set corresponding to the particle, represents the 1-Wasserstein distance between the distribution of the sample set and the distribution of the real sample data.
[0050] In each iteration, update the velocity and position of the particle according to the fitness value of the particle, and select the sample set with the highest fitness S opt .
[0051] In the synthetic sample data screening stage, introduce the particle swarm optimization algorithm to perform density perception and diversity evaluation on the synthetic sample pool, and dynamically select the optimal sample subset through the weighted fitness function of local density (based on kernel density estimation) and global diversity (based on Wasserstein distance), covering the multimodal distribution while avoiding redundancy.
[0052] The method provided by the present invention realizes the collaborative optimization of the feature space and the quality of synthetic samples through the global search ability and dynamic feedback of the particle swarm optimization algorithm in both the feature type selection and synthetic sample screening stages, effectively solving the problems of noise interference and distribution deviation in high-dimensional sparse medical insurance data.
[0053] After determining the screened synthetic data, the screened synthetic data and the real sample data are combined into a training sample set, and the classification model is trained based on the training sample set to obtain a medical insurance fraud detection model.
[0054] After training the medical insurance fraud detection model, the data to be detected is input into the medical insurance fraud detection model, and the medical insurance fraud detection result output by the medical insurance fraud detection model is obtained. The feature type of the data to be detected input into the medical insurance fraud detection model should be the same as the feature type of the data in the training sample set used to train the medical insurance fraud detection model.
[0055] Next, the medical insurance fraud detection device based on the particle swarm optimization generative adversarial network provided by the present invention will be described. The medical insurance fraud detection device based on the particle swarm optimization generative adversarial network described below can be mutually corresponding and referred to the medical insurance fraud detection method based on the particle swarm optimization generative adversarial network described above. As Figure 3 shown, the medical insurance fraud detection device based on the particle swarm optimization generative adversarial network provided by the present invention includes a sample synthesis module 310, a sample screening module 320, a training module 330, and a detection module 340. Among them: The sample synthesis module 310 is used to input random noise into the generator of the trained generative adversarial network to obtain the synthetic sample data output by the generator. The synthetic sample data includes synthetic sample data to be detected and the fraud label corresponding to the synthetic sample data to be detected. The generative adversarial network also includes a discriminator. The training data of the generative adversarial network includes real sample data, and the real sample data includes real sample data to be detected and the fraud label corresponding to the real sample data to be detected; The sample screening module 320 is used to screen the synthetic sample data based on the first particle swarm optimization algorithm to obtain a plurality of screened synthetic sample data. In the first particle swarm optimization algorithm, each particle corresponds to a sample set, the sample set includes a plurality of synthetic sample data, and the fitness value of each particle is determined based on the global diversity of the sample set corresponding to the particle and the local density of each synthetic sample data in the sample set; The training module 330 is used to combine a plurality of screened synthetic sample data and a plurality of real sample data to obtain a training sample set, and train the classification model based on the training sample set to obtain a medical insurance fraud detection model; The detection module 340 is used to input the data to be detected into the medical insurance fraud detection model to obtain the fraud detection result output by the medical insurance fraud detection model.
[0056] Figure 4 Illustrates a schematic diagram of the physical structure of an electronic device, as Figure 4As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may invoke the logical instructions in the memory 430 to execute a medical insurance fraud detection method based on a particle swarm optimization generative adversarial network. The medical insurance fraud detection method based on a particle swarm optimization generative adversarial network includes: inputting random noise into the generator of a trained generative adversarial network to obtain synthetic sample data output by the generator. The synthetic sample data includes synthetic sample data to be detected and fraud labels corresponding to the synthetic sample data to be detected. The generative adversarial network further includes a discriminator. The training data of the generative adversarial network includes real sample data. The real sample data includes real sample data to be detected and fraud labels corresponding to the real sample data to be detected; screening the synthetic sample data based on a first particle swarm optimization algorithm to obtain a plurality of screened synthetic sample data. In the first particle swarm optimization algorithm, each particle corresponds to a sample set, and the sample set includes a plurality of synthetic sample data. The fitness value of each particle is determined based on the global diversity of the sample set corresponding to the particle and the local density of each synthetic sample data in the sample set; combining the plurality of screened synthetic sample data and the plurality of real sample data to obtain a training sample set, and training a classification model based on the training sample set to obtain a medical insurance fraud detection model; inputting the data to be detected into the medical insurance fraud detection model to obtain a fraud detection result output by the medical insurance fraud detection model.
[0057] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0058] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the medical insurance fraud detection method based on the particle swarm optimization generative adversarial network provided by the above-mentioned various methods. The medical insurance fraud detection method based on the particle swarm optimization generative adversarial network includes: inputting random noise into the generator of the trained generative adversarial network to obtain synthetic sample data output by the generator. The synthetic sample data includes synthetic sample data to be detected and fraud labels corresponding to the synthetic sample data to be detected. The generative adversarial network also includes a discriminator. The training data of the generative adversarial network includes real sample data, and the real sample data includes real sample data to be detected and fraud labels corresponding to the real sample data to be detected; screening the synthetic sample data based on the first particle swarm optimization algorithm to obtain a plurality of screened synthetic sample data. In the first particle swarm optimization algorithm, each particle corresponds to a sample set, and the sample set includes a plurality of synthetic sample data. The fitness value of each particle is determined based on the global diversity of the sample set corresponding to the particle and the local density of each synthetic sample data in the sample set; combining the plurality of screened synthetic sample data and the plurality of real sample data to obtain a training sample set, and training a classification model based on the training sample set to obtain a medical insurance fraud detection model; inputting the data to be detected into the medical insurance fraud detection model to obtain a fraud detection result output by the medical insurance fraud detection model.
[0059] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the medical insurance fraud detection method based on a particle swarm optimization generative adversarial network provided by the above-mentioned various methods. The medical insurance fraud detection method based on a particle swarm optimization generative adversarial network includes: inputting random noise into the generator of a trained generative adversarial network to obtain synthetic sample data output by the generator. The synthetic sample data includes synthetic sample data to be detected and fraud labels corresponding to the synthetic sample data to be detected. The generative adversarial network further includes a discriminator. The training data of the generative adversarial network includes real sample data. The real sample data includes real sample data to be detected and fraud labels corresponding to the real sample data to be detected; based on the first particle swarm optimization algorithm, screening the synthetic sample data to obtain a plurality of screened synthetic sample data. In the first particle swarm optimization algorithm, each particle corresponds to a sample set, and the sample set includes a plurality of synthetic sample data. The fitness value of each particle is determined based on the global diversity of the sample set corresponding to the particle and the local density of each synthetic sample data in the sample set; combining the plurality of screened synthetic sample data and the plurality of real sample data to obtain a training sample set, and training a classification model based on the training sample set to obtain a medical insurance fraud detection model; inputting the data to be detected into the medical insurance fraud detection model to obtain a fraud detection result output by the medical insurance fraud detection model.
[0060] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0061] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A medical insurance fraud detection method based on a particle swarm optimization generative adversarial network, characterized in that Including: Input random noise into the generator of the trained generative adversarial network to obtain synthetic sample data output by the generator. The synthetic sample data includes synthetic sample data to be detected and fraud labels corresponding to the synthetic sample data to be detected. The generative adversarial network further includes a discriminator. The training data of the generative adversarial network includes real sample data. The real sample data includes real sample data to be detected and fraud labels corresponding to the real sample data to be detected; Based on the first particle swarm optimization algorithm, screen the synthetic sample data to obtain multiple screened synthetic sample data. In the first particle swarm optimization algorithm, each particle corresponds to a sample set, and the sample set includes multiple pieces of the synthetic sample data. The fitness value of each particle is determined based on the global diversity of the sample set corresponding to the particle and the local density of each piece of the synthetic sample data in the sample set; Combine multiple pieces of the screened synthetic sample data and multiple pieces of the real sample data to obtain a training sample set, and train a classification model based on the training sample set to obtain a medical insurance fraud detection model; Input the data to be detected into the medical insurance fraud detection model to obtain a fraud detection result output by the medical insurance fraud detection model.
2. The medical insurance fraud detection method based on particle swarm optimization generative adversarial network according to claim 1, wherein Before inputting the random noise into the generator of the trained generative adversarial network, it includes: Obtain real medical insurance data, and the real medical insurance data includes various types of features; Use the second particle swarm optimization algorithm to determine the target feature type among the various types of features. Among them, the particle in the second particle swarm optimization algorithm corresponds to a feature type set, and the feature type set includes at least one feature type. The fitness value of the particle in the second particle swarm optimization algorithm reflects the discrimination ability and independence of the feature types in the feature type set corresponding to the particle; Construct the real sample data based on the target feature type.
3. The medical insurance fraud detection method based on particle swarm optimization generative adversarial network according to claim 2, characterized in that The fitness value of the particle in the second particle swarm optimization algorithm is determined based on the first formula; The first formula is: ; Among them, represents the fitness value of the particle corresponding to the feature type set X, , , are weight coefficients, represents the mutual information between X and the fraud label set Y, , where n is the total number of samples, represents the value vector of the feature type of the i-th sample in the feature type set X, is the fraud label of the i-th sample, represents the average value vector of the feature types of all samples in the feature type set X, represents the mean value of the fraud labels of all samples, represents the number of feature types in the feature type set X, and are the standard deviation vectors of X and Y respectively.
4. The medical insurance fraud detection method based on particle swarm optimization generative adversarial network according to claim 1, characterized in that The formula for calculating the local density in the fitness value of the particle in the first particle swarm optimization algorithm is: ; ; wherein, is the local density of the synthetic sample data s, is the estimated probability density of the synthetic sample data s, k is the total number of the synthetic sample data, represents the Euclidean distance between the synthetic sample data s and the synthetic sample data ; is the standard deviation of the real sample data; d is the feature dimension in the synthetic sample data; is a normalization constant.
5. The medical insurance fraud detection method based on particle swarm optimization generative adversarial network according to claim 1, characterized in that The formula for calculating the global diversity in the fitness value of the particle in the first particle swarm optimization algorithm is: ; Among them, represents the global diversity of the sample set corresponding to the particles, represents the sample set is the 1-Wasserstein distance between the distribution of the sample set and the distribution of the true sample data.
6. The medical insurance fraud detection method based on particle swarm optimization generative adversarial network according to claim 1, wherein The number of the synthetic sample data is N times the difference between the number of samples of the majority class and the minority class in the real sample data, and N is greater than or equal to 2.
7. A medical insurance fraud detection device based on a particle swarm optimization generative adversarial network, characterized in that, Including: A sample synthesis module for inputting random noise into the generator of the trained generative adversarial network to obtain synthetic sample data output by the generator. The synthetic sample data includes synthetic sample data to be detected and fraud labels corresponding to the synthetic sample data to be detected. The generative adversarial network further includes a discriminator. The training data of the generative adversarial network includes real sample data. The real sample data includes real sample data to be detected and fraud labels corresponding to the real sample data to be detected; A sample screening module, which is used to screen the synthetic sample data based on the first particle swarm optimization algorithm to obtain multiple screened synthetic sample data. In the first particle swarm optimization algorithm, each particle corresponds to a sample set, and the sample set includes multiple synthetic sample data. The fitness value of each particle is determined based on the global diversity of the sample set corresponding to the particle and the local density of each synthetic sample data in the sample set; A training module, which is used to combine multiple screened synthetic sample data and multiple real sample data to obtain a training sample set, and train a classification model based on the training sample set to obtain a medical insurance fraud detection model; A detection module, which is used to input the data to be detected into the medical insurance fraud detection model to obtain a fraud detection result output by the medical insurance fraud detection model.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the medical insurance fraud detection method based on the particle swarm optimization generative adversarial network according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the medical insurance fraud detection method based on the particle swarm optimization generative adversarial network according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the medical insurance fraud detection method based on the particle swarm optimization generative adversarial network according to any one of claims 1 to 6.
Citation Information
Patent Citations
Training method, device and equipment of medical insurance fraud prediction network and storage medium
CN109636061A
quantum optimization parameter adjustment method for distributed deep learning under a Spark framework
CN109871995A
AdvGAN-based adversarial sample generation method
CN115510986A
Drainage basin multi-point water level prediction and early warning method based on generative adversarial network
CN115688579A
Deep learning inverse design method based on high-quality sampling and wavelet transform
CN116796637A