Imbalanced data enhancement method, device and equipment for medical insurance fraud detection and storage medium
Patent Information
- Application Number
- CN202611274087.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-21
- Publication Date
- 2026-09-22
AI Technical Summary
[0005]本发明提供一种面向医保欺诈检测的不平衡数据增强方法、装置、设备和存储介质,用以解决现有技术中医保欺诈检测中少数类欺诈样本不足、模型难以识别稀有违规模式的缺陷,实现提高医保欺诈样本生成质量、增强医保欺诈检测模型对少数类违规行为的识别能力
[0015] The present invention provides an imbalanced data augmentation method, apparatus, device, and storage medium for medical insurance fraud detection. After generating candidate minority class samples through a conditional variational autoencoder, low-quality samples are eliminated based on quality assessment results to ensure the authenticity of the first augmented sample set. The overall distribution of this sample set is used as feedback input, and the generation direction is continuously adjusted and the overall distribution is iteratively optimized through a dynamic feedback optimization mechanism. This allows the second augmented sample set to maintain authenticity while improving diversity. By fusing the two sample sets, the quality and diversity of the generated samples are taken into account. This effectively alleviates the shortcomings of existing technologies, such as unstable quality due to lack of structural constraints in the generated samples, pattern collapse due to fixed generation strategies, and insufficient diversity. As a result, the recognition performance of classification models on imbalanced data is improved.
Smart Images

Figure CN122796531A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical insurance data processing and intelligent risk control technology, and in particular to an imbalanced data augmentation method, apparatus, device, and storage medium for medical insurance fraud detection. Background Technology
[0002] In scenarios involving medical insurance fund supervision and fraud detection, the number of normal settlement records far exceeds the number of fraudulent or irregular records. This leads to a severe bias in model training towards the majority class, making it difficult to effectively identify the minority class. To address this, researchers have proposed data augmentation methods to expand the minority class sample.
[0003] Existing data augmentation methods mainly fall into three categories: traditional resampling methods, such as Synthetic Minority Over-sampling Technique (SMOTE) and Adaptive Synthetic Sampling (ADASYN), which rely on local linear interpolation, making it difficult to characterize complex nonlinear distributions and easily introducing noise into high-dimensional data; cost-sensitive learning methods, such as Cost-Sensitive Support Vector Machine (CS-SVM), which essentially still rely on discriminative models and cannot alleviate imbalance at the data generation level; and deep generative model methods, such as Conditional Tabular GAN (CTGAN), Contrastive Tabular Variational Autoencoder (CTVAE), and Tabular Denoising Diffusion Probabilistic Model (TabDDPM), which generate samples through adversarial training, variational inference, or diffusion processes, but suffer from problems such as training instability, pattern collapse, and lack of dynamic adjustment for fixed policies.
[0004] Medical insurance data typically exhibits characteristics such as high-dimensional tabular features, complex cost items, heterogeneous medical treatment behaviors, scarcity of fraud samples, and concealed fraud patterns. Existing generation methods, if they do not consider the local structure and distribution boundaries of medical insurance fraud samples, are prone to generating samples that deviate from the actual medical insurance fraud patterns. Summary of the Invention
[0005] This invention provides an imbalanced data augmentation method, apparatus, device, and storage medium for medical insurance fraud detection, which addresses the shortcomings of existing medical insurance fraud detection technologies, such as insufficient minority fraud samples and difficulty in identifying rare violation patterns. This invention aims to improve the quality of medical insurance fraud sample generation and enhance the ability of medical insurance fraud detection models to identify minority violations.
[0006] This invention provides a method for imbalanced data augmentation for medical insurance fraud detection, comprising the following steps: Obtain an imbalanced dataset for medical insurance fraud detection; the imbalanced dataset contains majority class samples and minority class samples; The minority class samples are input into the conditional variational autoencoder to obtain the candidate minority class samples output by the conditional variational autoencoder. Based on the quality assessment results of the candidate minority class samples, candidate minority class samples that do not meet the preset quality standards are removed to obtain the first enhanced sample set; The first enhanced sample set is used as the initial sample pool, the distribution state features of the current sample pool are used as the environment state, the feature vector is used as the action, and the update of the generation strategy is guided by a preset reward function; the feature vector is used to represent the generated sample. The process of generating samples based on the current generation strategy, updating the current sample pool with the generated samples, calculating the environment state and reward of the updated current sample pool, and updating the generation strategy is repeated until a preset stopping condition is met. The current sample pool that reaches the preset stopping condition will be used as the second enhanced sample set. The first enhanced sample set and the second enhanced sample set are merged to obtain the enhanced sample set.
[0007] According to the present invention, an imbalanced data augmentation method for detecting medical insurance fraud is provided, wherein the conditional variational autoencoder includes an encoder and a decoder; the conditional variational autoencoder is trained in the following manner: The encoder takes the sample features and class labels from the training set as inputs and outputs the mean and log-variance of the latent space. The mean and the log-variance are reparameterized and sampled to obtain latent variables; The latent variables and the category labels are input into the decoder, and the reconstructed sample is output. The reconstruction loss is calculated based on the original samples and the reconstructed samples in the training set, and the KL divergence loss is calculated based on the mean and the log-variance. The parameters of the conditional variational autoencoder are updated using the weighted sum of the reconstruction loss and the KL divergence loss as the total loss, thus obtaining the conditional variational autoencoder.
[0008] According to the imbalanced data augmentation method for medical insurance fraud detection provided by the present invention, in the prototype density-aware denoising stage, the conditional variational autoencoder is used for: The minority class samples are input into the encoder to extract the latent representation set of the minority class samples in the latent space; Clustering the potential representation set yields multiple prototype centers; Random perturbations are added to the neighborhood of each of the prototype centers to generate candidate potential vectors; The candidate latent vectors are input into the decoder to reconstruct and generate the candidate minority class samples.
[0009] According to the present invention, an imbalanced data augmentation method for detecting medical insurance fraud is provided, wherein the first augmented sample set is obtained by removing candidate minority class samples that do not meet preset quality standards based on the quality assessment results of the candidate minority class samples, including: Calculate the neighbor density between each candidate minority class sample and the minority class sample to obtain the density score of each candidate minority class sample; The threshold is set at the preset quantile of the density score of each of the candidate minority class samples. Candidate minority class samples with density scores below the threshold are removed to obtain the first enhanced sample set.
[0010] According to the present invention, an imbalanced data augmentation method for detecting medical insurance fraud includes calculating the proximity density between each candidate minority class sample and the minority class sample to obtain a density score for each candidate minority class sample, comprising: Calculate the average K-nearest neighbor distance between each of the candidate minority class samples and the minority class sample; The reciprocal of the average distance of the K nearest neighbors is used as the density score of each candidate minority class sample.
[0011] According to the imbalanced data augmentation method for medical insurance fraud detection provided by the present invention, after fusing the first augmented sample set and the second augmented sample set to obtain the augmented sample set, the method further includes: The augmented sample set is merged with the imbalanced dataset to obtain the augmented training set; The downstream classification model is trained using the enhanced training set.
[0012] The present invention also provides an imbalanced data augmentation device for detecting medical insurance fraud, comprising the following modules: The data acquisition module is used to acquire an imbalanced dataset for medical insurance fraud detection; the imbalanced dataset contains majority class samples and minority class samples. The candidate sample generation module is used to input the minority class samples into the conditional variational autoencoder to obtain the candidate minority class samples output by the conditional variational autoencoder. The first enhanced sample set generation module is used to remove candidate minority samples that do not meet the preset quality standards based on the quality assessment results of the candidate minority samples, and obtain the first enhanced sample set. The second enhanced sample set generation module is used to construct a dynamic feedback optimization mechanism. Taking the overall distribution state of the first enhanced sample set as the feedback input, the generation direction is continuously adjusted through the dynamic feedback optimization mechanism to iteratively optimize the overall distribution of the first enhanced sample set and obtain the second enhanced sample set. The fusion module is used to fuse the first enhanced sample set and the second enhanced sample set to obtain an enhanced sample set.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the imbalanced data augmentation method for medical insurance fraud detection as described above.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the imbalanced data augmentation method for medical insurance fraud detection as described above.
[0015] The present invention provides an imbalanced data augmentation method, apparatus, device, and storage medium for medical insurance fraud detection. After generating candidate minority class samples through a conditional variational autoencoder, low-quality samples are eliminated based on quality assessment results to ensure the authenticity of the first augmented sample set. The overall distribution of this sample set is used as feedback input, and the generation direction is continuously adjusted and the overall distribution is iteratively optimized through a dynamic feedback optimization mechanism. This allows the second augmented sample set to maintain authenticity while improving diversity. By fusing the two sample sets, the quality and diversity of the generated samples are taken into account. This effectively alleviates the shortcomings of existing technologies, such as unstable quality due to lack of structural constraints in the generated samples, pattern collapse due to fixed generation strategies, and insufficient diversity. As a result, the recognition performance of classification models on imbalanced data is improved. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1This is a flowchart illustrating the imbalanced data augmentation method for detecting medical insurance fraud provided by the present invention.
[0018] Figure 2 This is a flowchart illustrating the unbalanced table data augmentation method based on prototype-guided generation and PPO dynamic optimization provided by the present invention.
[0019] Figure 3 This is a schematic diagram of the unbalanced data augmentation device for detecting medical insurance fraud provided by the present invention.
[0020] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0022] The imbalanced data augmentation method for medical insurance fraud detection provided in this invention is applicable to imbalanced tabular data processing scenarios in medical insurance fraud detection where fraud samples are scarce and the proportion of normal samples is too high.
[0023] The following is combined Figures 1-4 The present invention describes an imbalanced data augmentation method, apparatus, device, and storage medium for medical insurance fraud detection.
[0024] Figure 1 This is a flowchart illustrating the imbalanced data augmentation method for detecting medical insurance fraud provided by the present invention, as shown below. Figure 1 As shown, the method includes the following: Step 101: Obtain the imbalanced dataset for medical insurance fraud detection.
[0025] An imbalanced dataset is a labeled dataset where the class distribution is significantly skewed, with the minority class having a much smaller number of samples than the majority class. For example, an imbalanced dataset might be a medical insurance business dataset, which could include medical insurance settlement details, medical service items, drug costs, medical institution information, insured person's medical records, reimbursement amounts, diagnosis codes, etc.
[0026] In the scenario of medical insurance fraud detection, the majority of samples correspond to normal medical insurance settlement / compliant reimbursement records, while the minority of samples correspond to medical insurance fraud, illegal reimbursement, abnormal settlement, or fraudulent insurance records.
[0027] The majority class refers to the samples corresponding to the dominant and densely distributed category in a dataset. Taking fraud detection as an example, majority class samples are normal medical insurance settlement samples, compliant reimbursement samples, or normal medical treatment records. This type of sample dominates the dataset, typically accounting for over 90% of the total. In the Kaggle medical insurance fraud detection dataset, normal samples account for approximately 91%; in the A08 service outsourcing competition dataset, normal samples account for approximately 95%. Because the majority class samples are abundant, the classification model can fully learn their feature distribution during training, thus typically achieving higher accuracy in identifying this type of sample.
[0028] Minority class samples refer to samples belonging to a category that is scarce and sparsely distributed in a dataset. Taking fraud detection as an example, minority class samples are medical insurance fraud samples, i.e., records of fraudulent user behavior. These samples are extremely rare in the dataset, typically accounting for less than 10% of the total, and in some extreme scenarios, even as low as 1% to 5%. In the Kaggle medical insurance fraud detection dataset, fraudulent samples account for approximately 9%; in the A08 service outsourcing competition dataset, fraudulent samples account for only about 5%. Due to the severe shortage of minority class samples, classification models struggle to learn their complete feature distribution and boundary information, resulting in insufficient model recognition ability for this type of sample and a tendency to misclassify minority class samples as majority class samples.
[0029] In one embodiment, raw data is collected from a business system or a public data source. This raw data includes sample features and corresponding category labels. The collected raw data undergoes preprocessing, which includes data cleaning, feature transformation, feature filtering, feature normalization, and data partitioning. Data cleaning handles missing and outlier values; feature transformation encodes non-numerical features into numerical features; feature filtering removes features irrelevant to the classification task or that are redundant; and feature normalization eliminates the impact of differences in the units of measurement between different features on model training. Stratified sampling is used to partition the preprocessed data into training and testing sets. Stratified sampling ensures that the majority and minority class proportions in both sets remain consistent, thus guaranteeing the reliability of the evaluation results. The partitioned training set is used for training the conditional variational autoencoder and generating samples, while the testing set is used for the final evaluation of the downstream classification model's performance.
[0030] Through the above steps, a feature matrix containing samples was obtained. and corresponding category tag set An imbalanced dataset in which the majority class label is 0 and the minority class label is 1. They represent the 1st, 2nd, and so on. The feature vector of each sample They represent the 1st, 2nd, and so on. The category label of each sample, This represents the total number of samples in the dataset.
[0031] Step 102: Input the minority class samples into the conditional variational autoencoder to obtain the candidate minority class samples output by the conditional variational autoencoder.
[0032] A Conditional Variational Autoencoder (CVAE) is used to learn the latent distribution of minority class samples and generate candidate minority class samples. Specifically, a CVAE is a generative model built on the basis of a standard variational autoencoder by introducing class-conditional control. Its network structure consists of two parts: an encoder and a decoder. The encoder is responsible for mapping the input sample and its corresponding class label to the latent space, and outputting the mean and log-variance of the sample in the latent space, thereby determining a probability distribution. The decoder is responsible for reconstructing the sample by combining the latent variables sampled from this distribution with the class label.
[0033] The real samples labeled as minority class in the training set are used as input to the conditional variational autoencoder. The encoder maps the samples to a low-dimensional representation in the latent space, and the decoder reconstructs them into new samples that are similar to the original sample distribution but have certain differences. These new samples are used as candidate minority class samples.
[0034] Step 103: Based on the quality assessment results of the candidate minority class samples, remove the candidate minority class samples that do not meet the preset quality standards to obtain the first enhanced sample set.
[0035] It should be understood that steps 102-103 correspond to the first stage in the imbalanced data augmentation method for medical insurance fraud detection, namely the prototype density-aware denoising stage.
[0036] The first augmented sample set refers to the high-quality augmented sample set obtained in the first stage after generating candidate samples around the minority class potential prototypes and filtering them through density-aware denoising.
[0037] The quality of each candidate minority class sample generated by decoding is evaluated. The density score of the candidate minority class sample in the neighborhood of the real minority class sample is used as a quantitative indicator to measure its credibility. Low-quality samples with density scores below the preset quantile threshold are identified as abnormal samples that deviate from the real distribution and are removed. High-quality samples with density scores above the threshold are retained. The retained candidate minority class samples constitute the first enhanced sample set.
[0038] Step 104: Construct a dynamic feedback optimization mechanism. Using the overall distribution state of the first enhanced sample set as feedback input, continuously adjust the generation direction through the dynamic feedback optimization mechanism to iteratively optimize the overall distribution of the first enhanced sample set and obtain the second enhanced sample set.
[0039] It should be understood that step 104 corresponds to the second stage in the imbalanced data augmentation method for medical insurance fraud detection, namely the PPO (Proximal Policy Optimization) dynamic feedback optimization stage. This stage should further optimize the overall distribution of the sample pool based on the first augmented sample set output from the first stage.
[0040] The second enhanced sample set refers to a new sample set obtained by iteratively optimizing the overall distribution of the sample pool through a dynamic feedback optimization mechanism based on the first enhanced sample set. Its core feature is that the sample distribution is more balanced and the coverage is more comprehensive while maintaining the authenticity of the samples.
[0041] The first enhanced sample set is used as the initial sample pool. The distribution characteristics of the sample pool are used as feedback signals to be input into the preset optimization strategy network, which drives the strategy network to continuously output new generation directions. In the cyclic iteration of sample generation and state feedback, the overall distribution of the sample pool gradually tends to the true minority class distribution. Finally, the sample pool that reaches the preset stopping condition is output as the second enhanced sample set.
[0042] Step 105: Merge the first enhanced sample set and the second enhanced sample set to obtain an enhanced sample set.
[0043] All samples from the first augmented sample set are merged with all samples from the second augmented sample set to form a unified sample set that combines realism and diversity. This unified sample set is then used as the final augmented sample set output and combined with the original training set to train the downstream classification model.
[0044] It should be understood that the first enhanced sample set originates from density-aware denoising screening. Its core advantage lies in fidelity; each sample in the set has passed quality verification based on its proximity density to the real minority class samples, and is located within the high-density reliable region of the real minority class distribution, exhibiting high consistency with the real minority class distribution. However, the samples retained after density screening are mainly concentrated in the high-density core region, with relatively limited coverage of sparse and boundary regions. The second enhanced sample set originates from PPO dynamic feedback optimization. Its core advantages lie in diversity and coverage. By continuously monitoring the sample pool distribution and adjusting the generation direction during iteration, the sample distribution gradually expands towards sparse and boundary regions, resulting in wider coverage. However, the second enhanced sample set is gradually expanded based on the first enhanced sample set. The optimization process starts with the first enhanced sample set; therefore, the second enhanced sample set includes samples from the first enhanced sample set and adds new samples on top of that. In other words, the first enhanced sample set is a subset of the second enhanced sample set, and the second enhanced sample set has more samples than the first enhanced sample set.
[0045] In one embodiment, the PPO-generated samples are fused with density-aware denoised samples: ; In the formula, This represents the final augmented sample set. This indicates that the PPO generates a sample set, which is the second enhanced sample set fusion. This represents the set of samples retained after density-aware denoising, i.e., the first enhanced sample set; This represents the union operation of sets.
[0046] The imbalanced data augmentation method for medical insurance fraud detection provided in this invention generates candidate minority class samples through a conditional variational autoencoder. Low-quality samples are then removed based on quality assessment results to ensure the authenticity of the first augmented sample set. Using the overall distribution of this sample set as feedback input, a dynamic feedback optimization mechanism continuously adjusts the generation direction and iteratively optimizes the overall distribution, enabling the second augmented sample set to maintain authenticity while increasing diversity. By fusing the two sample sets, the method simultaneously considers the quality and diversity of the generated samples, effectively alleviating the shortcomings of existing technologies where the lack of structural constraints in generated samples leads to unstable quality, fixed generation strategies lead to pattern collapse, and insufficient diversity. This improves the recognition performance of classification models on imbalanced data.
[0047] Based on the above embodiments, the conditional variational autoencoder includes an encoder and a decoder; the conditional variational autoencoder is trained in the following manner: The encoder takes the sample features and class labels from the training set as inputs and outputs the mean and log-variance of the latent space. The mean and the log-variance are reparameterized and sampled to obtain latent variables; The latent variables and the category labels are input into the decoder, and the reconstructed sample is output. The reconstruction loss is calculated based on the original samples and the reconstructed samples in the training set, and the KL divergence loss is calculated based on the mean and the log-variance. The parameters of the conditional variational autoencoder are updated using the weighted sum of the reconstruction loss and the KL divergence loss as the total loss, thus obtaining the conditional variational autoencoder.
[0048] By fusing raw features with category-conditional information at the input, the encoder can simultaneously perceive the content features and category attributes of samples when constructing the latent space, thereby establishing a category-oriented latent representation. Specifically, for each sample in the training set, its raw data contains two parts of information: sample features... This describes the values of the sample across each feature dimension; category label. The category label is used to identify the category to which the sample belongs. Since category labels are discrete scalar values, directly inputting them into a neural network is insufficient to effectively express the semantic relationships between categories. Therefore, this embodiment of the invention first uses an embedding layer to embed the category labels... Mapped to low-dimensional, dense class embedding vectors : ; In the formula, Indicates category label through The class embedding vector obtained after layer mapping Indicates category label Embedded representation is performed. The embedding layer, as a trainable network layer, adaptively learns the optimal vector representation for each class during training.
[0049] Sample features With category embedding vector Perform a concatenation operation to form a joint input vector. , This represents the concatenated encoder input vector. Indicates will and The concatenated vector contains both complete feature information and category guidance information of the sample, and serves as the input to the encoder. This concatenated vector is then input into the encoder network, which consists of multiple stacked fully connected layers. In a preferred embodiment, a three-layer fully connected structure is used, with dimensions of 128, 64, and 32 respectively. Each layer is followed by an activation function to introduce nonlinear transformation capabilities.
[0050] The encoder performs layer-by-layer feature extraction and dimensionality compression on the joint input vector through the aforementioned multi-layer fully connected network, ultimately outputting two key parameter vectors: the latent space mean vector and the latent space mean vector. Sum of logarithmic variance vector : ; ; In the formula, Represents the latent space mean vector; Represents the latent space variance vector The logarithm of , i.e., the log-variance vector; Represents the latent space standard deviation vector. Indicates output The mapping function, Indicates output The mapping function. Wherein, the latent space mean vector. It determines the central location of the latent variables in space, the log-variance vector This determines the extent to which the latent variables are distributed around the center.
[0051] Through the above process, each training sample is mapped to a Gaussian distribution in the latent space, rather than a single fixed point. This probabilistic mapping method gives the latent space good continuity and local smoothness. Sampling in the neighborhood of the latent space can generate effective samples that are similar to but not exactly the same as the input samples.
[0052] After the encoder outputs the mean and log-variance, instead of directly performing non-differentiable random sampling from this distribution, an independent noise source is introduced to decouple the randomness and differentiability of the sampling. This allows the sampling operation to participate in gradient backpropagation, thus supporting end-to-end training of the entire conditional variational autoencoder. Specifically, during training, the latent space mean vector output by the encoder... Sum of logarithmic variance vector A multidimensional Gaussian distribution is jointly defined in the latent space in order to obtain latent variables that can be used by the decoder. Sampling is required from this Gaussian distribution. However, direct random sampling from this distribution is non-differentiable in the computational graph, and the gradient cannot propagate back to the encoder through the sampling nodes, thus preventing encoder optimization. To address this issue, this embodiment of the invention employs reparameterization to decompose the sampling process into two parts: deterministic computation and random noise.
[0053] Specifically, assuming latent variables in the latent space Obey For the mean, Given a Gaussian distribution with variance, the reparameterization sampling technique represents the sampling process as follows: ; ; In the formula, Represents a latent vector. This indicates element-wise multiplication. Represents random noise variables. express It follows a standard normal distribution. Optionally, during the training phase, it is resampled at each forward propagation. This allows the same input sample to be mapped to different latent variables in the latent space, enhancing the model's generalization ability. During the generation phase, the variables can be fixed. Alternatively, latent variables can be generated directly by adding perturbations to the prototype center, thus achieving flexible and controllable sample generation.
[0054] The process involves restoring the low-dimensional representation in the latent space to a high-dimensional sample in the original feature space, while using class labels as conditional controls to reconstruct the class attributes of the samples, enabling the decoder to generate corresponding samples based on the specified class. Specifically, latent variables... It is a low-dimensional vector sampled from the latent space distribution output by the encoder. It represents the compressed representation of the input sample in the latent space. This latent variable contains the core feature information of the input sample, but it is difficult to determine which category the reconstructed sample should belong to based on this information alone. Therefore, it is necessary to re-inject the category label information into the decoder.
[0055] Similar to the encoder, the category labels on the decoder also need to be mapped to category embedding vectors through an embedding layer. latent variables With the category embedding vector Perform a concatenation operation to form a joint input vector. This concatenated vector contains both the latent representation information and category-guided information of the sample, and serves as the input to the decoder.
[0056] splicing vectors The input decoder network has a structure symmetrical to the encoder network. In a preferred embodiment, it employs a three-layer fully connected structure with dimensions of 32, 64, and 128 respectively. Each layer is followed by an activation function to introduce nonlinear transformation capability. The decoder progressively restores the low-dimensional representation in the latent space to the high-dimensional representation in the original feature space through layer-by-layer inverse transformation. During training, the decoder network continuously learns the mapping relationship from the latent space to the original feature space, making the reconstructed sample as close as possible to the input sample at the feature level. The final layer of the decoder outputs the reconstructed sample. The dimensions and features of the reconstructed sample With consistent dimensions, reconstruct the sample It is a sample that the decoder reconstructs based on latent variables and class labels. Its position in the feature space is determined by the latent variables, and its class assignment is determined by the class label.
[0057] During the training phase, two independent loss terms are constructed for the conditional variational autoencoder to evaluate the reconstruction capability of the decoder and the quality of the latent spatial distribution of the encoder output, respectively. These two terms jointly guide the direction of model parameter optimization. Specifically, the reconstruction loss measures the similarity between the reconstructed samples output by the decoder and the original samples input by the encoder, reflecting the model's ability to preserve information from the original samples. This embodiment of the invention uses mean squared error as the reconstruction loss function, for the first... Each sample, its original sample features With the corresponding reconstructed sample The sum of squared feature-wise differences between samples is taken as the reconstruction error of that sample. The average reconstruction error of all samples across the entire training set is then used to obtain the overall reconstruction loss. ; In the formula, Indicates the reconstruction loss. Indicates the number of samples. Indicates the first Individual sample features, Indicates the first One reconstructed sample, This represents the sum of reconstruction errors over all samples. Where, reconstruction loss... The smaller the value, the closer the reconstructed sample output by the decoder is to the original sample. This indicates that the encoder retains more effective information during the compression process, and the decoder can accurately restore the latent variables to the original features.
[0058] KL divergence loss is used to measure the degree of difference between the latent space distribution of the encoder output and the standard normal distribution, imposing constraints on the regularity of the latent space. In variational autoencoders, it is generally desirable for the latent space distribution of the encoder output to be close to the standard normal distribution, so that the latent space has continuity and sampleability, that is, any sampling point in the latent space can generate reasonable and effective samples after being decoded, without large-area holes.
[0059] The encoder outputs a Gaussian distribution for each input sample, derived from the latent space mean vector. Sum of logarithmic variance vector It is determined that the KL divergence loss calculates the difference between this distribution and the standard normal distribution N(0,1). For the For each sample, the KL divergence loss is: : ; The smaller the KL divergence loss, the closer the latent space distribution of the encoder output is to the standard normal distribution, and the more regular the latent space is. The role of this constraint is to ensure that effective samples can be generated after decoding two adjacent latent variables.
[0060] Reconstructing loss and KL divergence loss The total loss function is constructed by weighting and summing the results according to preset weight coefficients. The expression for the total loss function is as follows: ; ; In the formula, This represents the total CVAE loss. Indicates the weight of the KL divergence loss. express The maximum value, Indicates the current training round. This indicates the KL annealing cycle.
[0061] By adjusting The value of can control the degree to which the model emphasizes reconstruction quality and latent space regularity during training. The larger the value, the more the model emphasizes the regularity of the latent space; The smaller the value, the more the model focuses on the quality of the reconstructed samples. In this embodiment of the invention, the KL annealing strategy is used. The value starts from 0 and gradually increases to the preset maximum value as the number of training rounds increases, so that the model prioritizes learning reconstruction ability in the early stage of training and gradually strengthens the regularity of the latent space in the later stage.
[0062] To minimize the total CVAE loss The model uses this as an optimization objective during training. For example, in each forward propagation, the model calculates the total loss value for the current batch of samples. This loss value reflects the overall fit of the model to the training set under the current parameters. A larger total loss value indicates a larger reconstruction error or a more irregular latent space distribution; a smaller total loss value indicates a better balance between the two types of losses. Simultaneously, an adaptive moment estimation optimizer updates the network parameters based on the calculated gradient values. For example, the Adam optimizer adaptively adjusts the learning rate of each parameter by estimating the first and second moments of the gradient, ensuring stable and efficient parameter updates at different training stages. This avoids training oscillations caused by an excessively large fixed learning rate or slow convergence caused by an excessively small one. The process of calculating the total loss during forward propagation, calculating the gradient during backpropagation, and updating the parameters by the optimizer is repeated until the preset number of training epochs is reached. As the number of training epochs increases, the total loss value gradually decreases and tends to converge, indicating that the model has fully learned effective feature representations and latent space structures from the training data. When training stops, the currently saved encoder and decoder network parameters are the final weights of the trained conditional variational autoencoder model.
[0063] This invention introduces category labels as conditional variables, enabling similar samples to cluster in the latent space and dissimilar samples to separate from each other, thus establishing a category-oriented structured latent space. At the same time, the synergistic optimization of reconstruction loss and KL divergence loss takes into account both the generation quality and the regularity of the latent space.
[0064] Based on the above embodiments, in the prototype density-aware denoising stage, the conditional variational autoencoder is used for: The minority class samples are input into the encoder to extract the latent representation set of the minority class samples in the latent space; Clustering the potential representation set yields multiple prototype centers; Random perturbations are added to the neighborhood of each of the prototype centers to generate candidate potential vectors; The candidate latent vectors are input into the decoder to reconstruct and generate the candidate minority class samples.
[0065] After the Conditional Variational Autoencoder (CVA) is trained, the trained encoder is used to map the minority class samples in the original feature space to the latent space, obtaining its compressed low-dimensional representation. Specifically, after the CVA completes pre-training, the encoder's parameters are fixed, and it has learned to map input samples to an effective representation in the latent space. At this point, all minority class samples in the training set are input into the encoder one by one. For each minority class sample... The encoder performs forward computation through a multi-layer fully connected network, outputting its mean vector in the latent space. The mean vectors of all minority class samples in the latent space constitute the minority class latent representation set. ,in This represents the total number of minority class samples. This set resides in a low-dimensional latent space, where each element corresponds to a compressed representation of a real minority class sample, preserving the class semantic features of the original sample while removing redundant information irrelevant to classification.
[0066] After extracting the latent representation set of minority class samples, a clustering algorithm is used to identify the intrinsic substructure of the minority class in the latent space. Multiple prototype centers are used to represent the core regions of the minority class distribution, thereby constructing a generative anchor point system oriented towards multiple sub-patterns. Specifically, the latent representation set... In Each data point is represented by a mean vector, and each data point corresponds to the position of a true minority class sample in the latent space. Since different fraudulent behavior patterns differ significantly in the original feature space, they naturally form different clustering regions after mapping to the latent space. The goal of the clustering algorithm is to automatically divide these data points into several dense clusters without prior labels, with each cluster corresponding to a sub-pattern.
[0067] The K-Means clustering algorithm is used to cluster the above data points. The K-Means algorithm divides the data points into K clusters by iteratively optimizing the sum of squared intra-cluster distances and determines the centroid of each cluster. The algorithm process is as follows: From... K data points are randomly selected as initial cluster centers; each data point is assigned to the nearest cluster center to form K clusters; the mean of each cluster is recalculated as the new cluster center; the assignment and update steps are repeated until the cluster centers no longer change or the preset maximum number of iterations is reached.
[0068] It should be understood that the determination of the number of clusters K directly affects the granularity of prototype extraction. If K is too small, multiple sub-patterns are forcibly merged into a single cluster, and the prototype center cannot accurately represent any sub-pattern, resulting in incomplete sample coverage. If K is too large, the same sub-pattern is split into multiple clusters, and the prototype centers are too close together, leading to redundant and homogeneous generated samples. In this embodiment of the invention, K is set to 5. This value sufficiently covers multiple sub-patterns within a few classes while avoiding redundancy caused by excessive subdivision.
[0069] After clustering, K prototype centers are obtained. , These represent the 1st, 2nd to Kth prototype centers, respectively. Each prototype center represents the core location of a high-density region in the latent space, which corresponds to the typical location of a sub-pattern of a minority class in the latent space.
[0070] After a plurality of prototype centers are obtained through clustering, sampling is no longer directly performed from the standard normal distribution. Instead, each prototype center is used as an anchor point for local sampling in its neighborhood, so that the generated latent vectors are anchored in the latent space region where the real minority classes are located, thereby ensuring that the decoded samples have real class attributes. Specifically, one prototype center is randomly selected from K prototype centers as the anchor point for current generation. Each prototype center represents the typical position of a sub-pattern of a minority class in the latent space, and selecting different prototype centers will guide the generation of samples of different sub-patterns. The prototype center can be selected by uniform random selection, which ensures that the generation frequency of each sub-pattern is roughly equivalent; it can also be selected by weighting according to the sample density of each cluster, and the higher the density of a region is, the more samples will be generated there.
[0071] Random perturbation is added in the neighborhood of the selected prototype center to generate candidate latent vectors , and the mathematical expression of the perturbation is: , where represents a new latent vector obtained after adding perturbation, that is, a candidate latent vector; represents a Gaussian noise intensity coefficient; represents standard normal distribution noise. Different from the direct sampling method from standard normal distribution , in the embodiment of the present invention, the prototype center is used as a reference position, and the perturbation only makes the sampling point offset in the local area near the prototype center instead of random walking in the whole latent space.
[0072] It should be understood that the value of determines the diversity range of generated samples, wherein if the value is too large, the sampling point may be far away from the aggregation region of the latent space of the real minority classes, resulting in that the decoded generated samples deviate from the distribution of the minority classes and cause class ambiguity or semantic shift; if the value is too small, all generated samples are highly concentrated near the prototype center with insufficient diversity, which is equivalent to repeated copying of the original samples and cannot effectively expand new feature patterns. In the embodiment of the present invention, the value is 0.1, which can provide sufficient local diversity on the premise of maintaining the authenticity of the samples.
[0073] The candidate latent vector is obtained by adding random perturbation in the neighborhood of the prototype center The resulting vector has the same dimension as the latent space, representing a sampling point located within the true minority class distribution region. However, this latent vector is only a low-dimensional numerical vector, not a complete sample, and cannot be directly used as the result of data augmentation. The function of the decoder is to map this low-dimensional latent representation back to the high-dimensional original feature space, restoring a true sample with semantic meaning.
[0074] Since the decoder is trained using the concatenation of latent variables and class embedding vectors as input during the training phase, the generation phase requires the candidate latent vectors to be used as input. With the corresponding category embedding vector The vectors are concatenated to form a joint input vector. The category embedding vector should be consistent with the category embedding used during training for the minority class, ensuring that the decoder reconstructs samples with the correct category conditions and generates samples of the specified class. In this embodiment of the invention, the category label is set to the minority class, so that the samples reconstructed by the decoder belong to the minority class.
[0075] splicing vectors The input decoder network consists of multiple stacked fully connected layers. In a preferred embodiment, it employs a three-layer structure symmetrical to the encoder, with dimensions of 32, 64, and 128 respectively. Each layer progressively maps the low-dimensional latent representation back to the high-dimensional feature space through linear transformations and non-linear activation functions. With layer-by-layer upsampling, the vector dimension gradually expands, and the sample information becomes increasingly rich. Finally, the decoder's output layer outputs a vector with the same dimension as the original input feature space. That is, the candidate minority class samples generated by reconstruction.
[0076] It should be understood that the candidate minority class samples generated by the decoder have similar feature value patterns to the real minority class samples at the feature level, but are not exactly the same as any real sample. The number of generated candidate samples is not limited by the number of original samples. By performing multiple perturbation sampling on the same prototype center, any number of different candidate latent vectors can be generated, which are then reconstructed by the decoder to obtain the corresponding number of candidate minority class samples.
[0077] This invention extracts multiple prototype centers through clustering to anchor various sub-patterns of the minority class. Within the neighborhood of each prototype center, candidate latent vectors are generated by perturbation and then decoded and reconstructed. This anchors the generated samples to the real minority class distribution area while covering its diverse sub-pattern structure. This effectively avoids the problems of generated samples deviating from the real distribution, class overlap, and semantic shift, and improves the class fidelity and structural diversity of the generated samples.
[0078] Based on the above embodiments, the step of removing candidate minority class samples that do not meet the preset quality standards according to the quality assessment results of the candidate minority class samples to obtain the first enhanced sample set includes: Calculate the neighbor density between each candidate minority class sample and the minority class sample to obtain the density score of each candidate minority class sample; The threshold is set at the preset quantile of the density score of each of the candidate minority class samples. Candidate minority class samples with density scores below the threshold are removed to obtain the first enhanced sample set.
[0079] For each candidate minority class sample, find the K nearest real minority class samples in the real minority class sample space, calculate the average of these K distances as the proximity index between the candidate minority class sample and the real minority class distribution, and then use the reciprocal of the average distance as the density score of the candidate minority class sample to quantify its clustering degree within the real minority class distribution area.
[0080] In one embodiment, the density score of each candidate minority class sample is obtained by calculating the average K-nearest neighbor distance between each candidate minority class sample and the minority class sample; and using the reciprocal of the average K-nearest neighbor distance as the density score of each candidate minority class sample.
[0081] Specifically, for each candidate minority class sample generated by the conditional variational autoencoder, it is placed in the original feature space, and the distance between it and each real minority class sample is calculated. In this embodiment of the invention, the distance metric is Euclidean distance. Calculate its relationship with the first A real minority sample Euclidean distance between The distance sequence between the candidate minority class sample and all real minority class samples is calculated one by one. This distance sequence is then sorted in ascending order, and the K smallest distance values are selected. These are the K nearest neighbors of the candidate minority class sample in the feature space. The average of these K smallest distances is then used as the candidate minority class sample's distance. Mean nearest neighbor distance to the true minority class distribution: ; In the formula, Indicates the first The average K-nearest neighbor distance of each generated sample This represents the number of nearest neighbors. It should be understood that the smaller the average distance of K-nearest neighbors, the more real minority class neighbors the candidate minority sample has around it. The closer the candidate minority sample is to these neighbors, the higher the density of minority class samples in the region. The larger the average distance of K-nearest neighbors, the fewer real minority class neighbors the candidate minority sample has around it and the farther away they are. The candidate minority sample is located in a low-density region where minority class samples are sparse or in an area where the two classes overlap.
[0082] Average distance of K nearest neighbors Take the reciprocal and add a very small constant. To prevent the denominator from being zero, candidate minority class samples are obtained. Density score: ; In the formula, Indicates the first Density score of each generated sample, This indicates the prevention of extremely small positive numbers with a denominator of 0.
[0083] By taking the reciprocal, candidate minority samples with smaller average distances receive higher density scores, while those with larger average distances receive lower density scores. This results in a positive correlation between density scores and sample quality. Samples with high density scores are located in the high-density, reliable region of the minority class, while samples with low density scores are located in the low-density, unreliable region. This makes it easier to use a single density score as a criterion for subsequent screening.
[0084] After sorting the density scores of all candidate minority class samples in ascending order, the density score value at the preset percentile position is taken as the screening threshold. Candidate minority class samples with density scores higher than the threshold are retained, and candidate minority class samples with density scores lower than the threshold are removed.
[0085] Specifically, the density scores of all candidate minority class samples are sorted in ascending order to obtain an ordered sequence. The density score reflects the proximity of each candidate minority class sample to the true minority class distribution; a higher score indicates a more credible sample, while a lower score indicates a more suspicious sample. Within the ordered sequence, a threshold position is determined based on a preset quantile percentage. For example, if the preset quantile is 20%, the threshold position is N×20%, and the density score at that position is taken as the threshold. Different percentiles correspond to different levels of screening stringency: lower percentiles result in lower thresholds, more lenient screening conditions, and more retained samples, but may also include more low-quality samples; higher percentiles result in higher thresholds, more stringent screening conditions, and fewer retained samples, but with better sample quality. In this embodiment, a 20% percentile is used as the threshold, meaning that the top 80% of candidate minority class samples with the highest density scores are retained, ensuring a sufficient number of samples while maintaining quality.
[0086] The density score threshold is used as the quantitative basis for the preset quality standard. The density score of each candidate minority class sample is compared with the threshold: candidate minority class samples with a density score higher than the threshold are determined to meet the preset quality standard and are retained; candidate minority class samples with a density score lower than the threshold are determined to be low-quality anomalous samples and are removed. For example, samples that meet the following conditions are retained: ,satisfy The generated samples are retained if they meet the criteria, otherwise they are discarded. Because the threshold is dynamically calculated based on the density score distribution of all candidate minority class samples in the current batch, rather than being fixed in advance, it can automatically adapt to differences in density score values under different data distributions.
[0087] All candidate minority class samples with density scores higher than the threshold are retained. The retained high-quality candidate minority class samples constitute the first augmentation sample set. Each sample in this sample set has passed the neighbor density quality verification and has a high degree of consistency with the real minority class distribution, thus possessing the credibility to serve as a minority class augmentation sample.
[0088] The embodiments of the present invention effectively eliminate low-quality samples and boundary samples that deviate from the true minority class distribution by using proximity density calculation and quantile threshold screening, ensuring that the first enhanced sample set is highly consistent with the true minority class distribution.
[0089] Based on the above embodiments, the construction of a dynamic feedback optimization mechanism, using the overall distribution state of the first enhanced sample set as feedback input, continuously adjusts the generation direction through the dynamic feedback optimization mechanism, iteratively optimizes the overall distribution of the first enhanced sample set, and obtains a second enhanced sample set, includes: Use the first enhanced sample set as the initial sample pool; The current sample pool distribution characteristics are used as the environment state, the feature vector is used as the action, and the preset reward function guides the update of the generation strategy; the feature vector is used to represent the generated sample. The process of generating samples based on the current generation strategy, updating the current sample pool with the generated samples, calculating the environment state and reward of the updated current sample pool, and updating the generation strategy is repeated until a preset stopping condition is met. The current sample pool that reaches the preset stopping condition is used as the second enhanced sample set.
[0090] The sample pool is a dynamic data structure used to store all available augmented samples at the current time. At the start of the optimization process, the initial content of the sample pool consists of all samples from the first augmented sample set. These samples have passed the quality verification of density-aware denoising and have high fidelity; therefore, the initial state of the sample pool is reliable. For example, constructing the sample pool: ; In the formula, Indicates the sample pool; Indicates the 1st, 2nd and so on. One generated sample, This indicates the current number of steps generated.
[0091] It should be understood that using the first augmented sample set as the initial sample pool gives the optimization process a high-quality starting point. Existing reinforcement learning schemes usually start optimization from an empty pool or random samples, requiring a lot of exploration in the early stages to generate effective samples. However, the sample pool of this embodiment is filled with high-fidelity samples that have been verified for quality before optimization begins, so that the optimization process can be finely adjusted based on a relatively good initial distribution, rather than exploring from scratch.
[0092] The sample pool is dynamically changing during the optimization process. In each iteration, the policy network generates actions based on the current distribution of the sample pool, producing a batch of new samples. These new samples are added to the sample pool, continuously expanding and updating it. Simultaneously, the distribution of the sample pool also evolves. With the addition of new samples, statistical characteristics such as the mean, variance, KL divergence, and information entropy of the sample pool change accordingly. These changes serve as the input state for the next iteration. Therefore, the sample pool is both the starting point of the optimization process and the direct object of state awareness and action interaction in each iteration. Its evolution from the initial state to the final state fully records the optimization effect and path.
[0093] Through the above settings, the sample pool organically connects the first enhanced sample set with the dynamic feedback optimization mechanism, enabling the optimization process to iterate from a high-quality starting point, rather than blindly searching from scratch.
[0094] The reinforcement learning framework is applied to the sample generation optimization process. The environmental state is described by the overall distribution statistics of the current sample pool. The action is directly output in the form of the feature vector of the generated sample. A preset reward function is used to evaluate the quality of each generation and generate a feedback signal. This feedback signal drives the policy network parameters to be updated in a better direction through the reinforcement learning algorithm, so that the policy network gradually generates higher quality samples in subsequent decisions.
[0095] Specifically, the environmental state is the basis for the policy network's decision-making and describes the overall distribution of the current sample pool. This embodiment of the invention combines multiple statistics of the current sample pool into a state vector. ,in, It represents the mean of all samples in the sample pool across each feature dimension, reflecting the central location of the sample distribution; It represents the variance of the sample pool, reflecting the degree of dispersion of the sample; The KL divergence represents the difference between the sample pool distribution and the true minority class distribution, reflecting the degree of deviation between the two. The information entropy of the sample pool reflects the level of sample diversity. These five dimensions represent the current training progress and reflect the stage of optimization. They characterize the distribution quality of the sample pool from different perspectives, enabling the policy network to fully perceive the current state.
[0096] Actions are the decision outputs of the policy network based on the current state, used to indicate the generation direction. In this embodiment of the invention, actions are output by the policy network based on the current state information, used to represent the feature vectors of the samples to be generated. Specifically, the action space is defined as... 3D continuous space: ,in, Represents the space of continuous actions. Indicates the first The action vector output at each time step This represents the feature dimension of the sample. Each action vector corresponds to the feature representation of a candidate minority class sample and can be added to the current sample generation set to update the state information of the next time step. Compared with the scheme of outputting policy parameter adjustment, directly outputting feature vectors makes the meaning of actions more intuitive. New samples can be directly added to the sample pool and immediately used for the next round of state updates without additional parameter transformation steps, simplifying the environment interaction logic.
[0097] The reward function is used to quantitatively evaluate the actions output by the policy network, generating feedback signals to drive policy network updates. The design goal of the reward function aligns with the optimization goal, guiding the generated samples to achieve a good balance between realism and diversity. In this embodiment of the invention, multiple sub-reward items are weighted and combined into a total reward: ; ; ; ; ; in, Indicates the total reward. This indicates the KL dispersion reward. Represents entropy reward, Indicates coverage bonus, Represents KL regularization terms. This represents the weight coefficient of the regularization term; This represents the KL divergence before adding the currently generated sample; This represents the KL divergence after adding the currently generated sample; This represents the entropy value before the current sample was added. This represents the entropy value after incorporating the currently generated sample; Indicates the generation of sample pools Variance across each feature dimension This indicates that the variances of each dimension are averaged. Indicates the weight of the regularization term. This represents the sample distribution generated by CVAE. This indicates the distribution of PPO generated samples. This represents the KL divergence between the two.
[0098] Total Rewards The quality of generated actions is comprehensively evaluated from four dimensions: KL divergence reward. The reward measures the reduction in the difference between the sample pool distribution and the true minority class distribution; the smaller the difference, the higher the reward. (Entropy reward) The reward measures the increase in sample pool diversity; the greater the diversity, the higher the reward. (Coverage reward) The wider the coverage of the sample pool in the feature space, the higher the reward; KL regularization term. The constraint generation policy does not deviate from the prior distribution of the conditional variational autoencoder; the smaller the deviation, the higher the reward. Through the design of a multi-objective reward function, the policy network is guided to achieve a balance between realism and diversity in the samples it generates.
[0099] The policy network employs a reinforcement learning algorithm for parameter updates. In each round of interaction, the policy network generates actions based on the current state, and the environment updates the sample pool and returns a reward based on the actions. The algorithm uses the collected state-action-reward data to calculate an advantage estimate and updates the policy network parameters by maximizing the accumulated expected reward. As the policy network is continuously updated, the quality of its generated samples gradually improves, and the overall distribution of the sample pool is continuously optimized.
[0100] In one embodiment, the PPO algorithm is used to update the policy network, and the objective function is: ; In the formula, This represents the PPO objective function. Expressing expectations, Indicates the first The probability ratio of the new and old strategies, Indicates the first Step advantage function, This represents the clipping function. This represents the PPO pruning threshold, in this embodiment... .
[0101] In each iteration, the policy network generates new samples based on the current sample pool state and adds them to the sample pool. It calculates the updated state characteristics and reward value of the sample pool and uses the reward signal to update the policy network parameters so that the policy network can generate higher quality samples in the next iteration. This process is repeated until a preset stopping condition is met, and all samples in the sample pool at the time of stopping are output as the second augmented sample set.
[0102] This invention uses a first enhanced sample set as the initial sample pool, and uses the distribution state of the sample pool as a feedback signal to drive the policy network to dynamically adjust the generation direction. Through multiple iterations, the overall distribution of the sample pool is continuously optimized, so that the generated samples gradually expand to sparse and boundary regions while maintaining authenticity. This effectively avoids the pattern collapse and sample homogenization problems caused by fixed generation strategies and lack of dynamic adjustment in the prior art, and achieves a continuous improvement in the diversity and distribution coverage of generated samples.
[0103] Based on the above embodiments, after fusing the first enhanced sample set and the second enhanced sample set to obtain the enhanced sample set, the method further includes: The augmented sample set is merged with the imbalanced dataset to obtain the augmented training set; The downstream classification model is trained using the enhanced training set.
[0104] The downstream classification model is a medical insurance fraud detection classification model, used to identify potential fraud, violations, or abnormal behaviors based on the characteristic information of medical insurance settlement records. Specifically, the final enhanced sample set is merged with the original imbalanced training set to form an enhanced training set with improved class distribution. This enhanced training set is then used as training data to train the medical insurance fraud detection classification model's parameters, improving its ability to identify minority fraud samples in a more balanced data environment.
[0105] Specifically, the original training set contains original majority class samples and minority class samples, with a sufficient number of majority class samples and a scarce number of minority class samples; the augmented sample set contains minority class augmented samples generated through the aforementioned steps of this invention, which have undergone quality screening and dynamic optimization. After merging the two, the number of minority class samples in the new training set is expanded, and the ratio of majority class to minority class samples is effectively improved, so that the class distribution of the training set changes from a severely imbalanced state to a relatively balanced state.
[0106] The downstream classification model is trained using the augmented training set obtained after merging. The downstream classification model can employ various standard classifier architectures. In this embodiment, a multilayer perceptron is used as the downstream classification model. Its structure is as follows: the input layer receives the feature vectors of the augmented training set samples; two hidden layers (128-dimensional and 64-dimensional) sequentially undergo nonlinear transformations to extract high-level features; the output layer uses the sigmoid activation function, outputting the probability value of a sample belonging to the minority class. During training, binary cross-entropy loss is used as the optimization objective. By minimizing the classification loss, the classification model learns an effective mapping from input features to class labels. ; In the formula, Represents classification loss, Indicates the true label, This represents the probability that the model predicts the value as fraud.
[0107] Since the minority class samples in the training set have been sufficiently enhanced, the classification model no longer faces the serious class imbalance problem in the original dataset. During training, it can learn the discrimination boundaries of the majority and minority classes simultaneously, effectively avoiding the problem of biased prediction of the majority class due to insufficient minority class samples.
[0108] The trained downstream classification model was evaluated on a reserved test set. The classification performance of the model was quantified by three indicators: F1-score, PR-AUC (Precision-Recall Area Under the Curve), and MCC (Matthews Correlation Coefficient) to verify that the imbalanced data augmentation method for medical insurance fraud detection proposed in this invention can effectively improve the classification model's ability to identify minority class samples in imbalanced data scenarios.
[0109] This invention combines the augmented sample set with the original imbalanced dataset, effectively expanding the number of minority class samples in the training set and making the class distribution more balanced. On this basis, the downstream classification model trained can simultaneously learn the feature patterns of both the majority and minority classes, effectively alleviating the problem of the model being biased towards the majority class caused by the original imbalanced data and improving the classification model's ability to identify minority class samples.
[0110] To further explain the imbalanced data augmentation method for medical insurance fraud detection proposed in this invention, please refer to the following embodiments.
[0111] This invention specifically proposes a method for enhancing imbalanced table data based on prototype-guided generation and dynamic PPO optimization, referencing... Figure 2The method consists of three stages: the first stage is the prototype density-aware denoising stage, which is used to generate and screen high-quality minority class samples; the second stage is the PPO dynamic feedback optimization stage, which is used to further optimize the sample distribution based on the first enhanced sample set; and the third stage is the enhanced sample fusion and classification model training stage, which is used to form the final enhanced training set and train the downstream classification model.
[0112] Before the first stage, CVAE pre-training is required as a basic model preparation process: Initialize the network parameters of the conditional variational autoencoder; input the sample features and class labels from the training set into the encoder to calculate the mean and variance of the latent variable distribution; reparameterize the mean and variance to obtain the latent variables; input the latent variables and class labels into the decoder to reconstruct and generate samples; calculate the reconstruction loss and KL divergence loss, and sum them by weight to obtain the total CVAE loss; update the CVAE network parameters based on the total loss; determine whether the maximum number of training rounds has been reached. If not, repeat the above process; if it has been reached, output the pre-trained CVAE weights.
[0113] Prototype density-aware denoising stage: The minority class latent space representation extracted by the pre-trained CVAE encoder is clustered to obtain multiple prototype centers; random perturbations are added to the neighborhood of each prototype center to generate candidate latent vectors, which are then reconstructed by the decoder to generate candidate minority class samples; the neighborhood density between each candidate minority class sample and the real minority class sample is calculated to obtain the density score of each candidate minority class sample; using the preset quantile of the density score of each candidate minority class sample as a threshold, candidate minority class samples with density scores lower than the threshold are removed to obtain the first enhanced sample set, i.e., the denoised enhanced samples.
[0114] The PPO dynamic feedback optimization phase involves: using the first augmented sample set as the initial sample pool to initialize the MDP dynamic environment; extracting the distribution state features of the current sample pool as the environment state, including the sample pool mean, variance, KL divergence with the true minority class distribution, information entropy, and current training progress; the CVAE-Actor network generates action samples based on the current environment state, with actions as feature vectors to represent the new samples to be generated; calculating multi-objective reward values based on the generated samples, including KL divergence reward, entropy reward, coverage reward, and KL regularization term, to evaluate the quality of the generated samples; updating the policy network parameters using the PPO algorithm based on the calculated reward values; determining whether the maximum number of training rounds has been reached; if not, repeating the loop of extracting the environment state, generating action samples, calculating rewards, and updating PPO parameters; if the maximum number of training rounds has been reached, using the current sample pool as the PPO-optimized sample set, i.e., the second augmented sample set, i.e., the PPO-optimized samples.
[0115] Enhanced Sample Fusion and Classification Model Training Phase: The first and second enhanced sample sets are fused together and then merged with the original imbalanced dataset to form an enhanced training set; the MLP classification model is trained using the enhanced training set. The MLP classification model includes an input layer, a hidden layer, and an output layer. The output layer uses the Sigmoid activation function and is optimized using binary cross-entropy loss; the trained classification model is evaluated on the test set, and the classification results are output.
[0116] This invention provides structured guidance for the potential representation space of minority classes through prototype structure constraints and a density-aware cleanup mechanism. Generated samples are first generated under constraints around representative minority class prototypes, and outlier samples are filtered using neighborhood density evaluation, thereby reducing deviations of generated samples from the true distribution. Compared to methods such as CTGAN-ENN, CTVAE, and CTAB-GAN+, this invention effectively reduces class overlap, semantic shift, and the proportion of noisy samples, making the generated samples more consistent with the distribution characteristics of real minority class data, improving the quality of generated data and the training effect of subsequent classification models.
[0117] By constructing a dynamic generation and control mechanism based on reinforcement learning, and jointly building feedback signals using metrics such as KL divergence, information entropy, and coverage, the distribution of the generated sample population is continuously optimized, guiding the generation process to cover more potential patterns and feature regions. Compared with existing fixed generation strategy methods, this invention can effectively avoid the problem of generated samples being concentrated in local high-density regions, improve the expressive power of different sub-patterns and sparse regions within minority classes, thereby enhancing the diversity and representativeness of generated samples.
[0118] By continuously evaluating the difference between the generated samples and the true minority class distribution through a dynamic feedback mechanism, and utilizing coverage rewards to encourage the generated samples to expand into regions with higher information content in the feature space, the generated samples can more fully cover the area near the classification boundary. Compared with TabDDPM and traditional CVAE generation mechanisms, the enhanced samples generated by this invention can provide the classifier with more training information with discriminative value, improving the model's ability to identify complex boundary samples and rare fraudulent patterns.
[0119] The imbalanced data augmentation device for medical insurance fraud detection provided by the present invention will be described below. The imbalanced data augmentation device for medical insurance fraud detection described below can be referred to in correspondence with the imbalanced data augmentation method for medical insurance fraud detection described above.
[0120] refer to Figure 3 The present invention provides an imbalanced data augmentation device for detecting medical insurance fraud, comprising: Data acquisition module 301 is used to acquire an imbalanced dataset for medical insurance fraud detection; the imbalanced dataset contains majority class samples and minority class samples; The candidate sample generation module 302 is used to input the minority class samples into the conditional variational autoencoder to obtain the candidate minority class samples output by the conditional variational autoencoder. The first enhanced sample set generation module 303 is used to remove candidate minority samples that do not meet the preset quality standards based on the quality assessment results of the candidate minority samples, and obtain the first enhanced sample set. The second enhanced sample set generation module 304 is used to construct a dynamic feedback optimization mechanism. Taking the overall distribution state of the first enhanced sample set as the feedback input, the generation direction is continuously adjusted through the dynamic feedback optimization mechanism to iteratively optimize the overall distribution of the first enhanced sample set and obtain the second enhanced sample set. The fusion module 305 is used to fuse the first enhanced sample set and the second enhanced sample set to obtain an enhanced sample set.
[0121] The imbalanced data augmentation device for medical insurance fraud detection provided in this invention generates candidate minority class samples through a conditional variational autoencoder. Based on quality assessment results, low-quality samples are eliminated to ensure the authenticity of the first augmented sample set. Using the overall distribution of this sample set as feedback input, a dynamic feedback optimization mechanism continuously adjusts the generation direction and iteratively optimizes the overall distribution, enabling the second augmented sample set to maintain authenticity while increasing diversity. By fusing the two sample sets, both the quality and diversity of the generated samples are considered, effectively alleviating the shortcomings of existing technologies where the lack of structural constraints in generated samples leads to unstable quality, fixed generation strategies lead to pattern collapse, and insufficient diversity. This improves the recognition performance of classification models on imbalanced data.
[0122] In one embodiment, the conditional variational autoencoder includes an encoder and a decoder; the conditional variational autoencoder is trained in the following manner: The encoder takes the sample features and class labels from the training set as inputs and outputs the mean and log-variance of the latent space. The mean and the log-variance are reparameterized and sampled to obtain latent variables; The latent variables and the category labels are input into the decoder, and the reconstructed sample is output. The reconstruction loss is calculated based on the original samples and the reconstructed samples in the training set, and the KL divergence loss is calculated based on the mean and the log-variance. The parameters of the conditional variational autoencoder are updated using the weighted sum of the reconstruction loss and the KL divergence loss as the total loss, thus obtaining the conditional variational autoencoder.
[0123] In one embodiment, during the prototype density-aware denoising stage, the conditional variational autoencoder is used to: The minority class samples are input into the encoder to extract the latent representation set of the minority class samples in the latent space; Clustering the potential representation set yields multiple prototype centers; Random perturbations are added to the neighborhood of each of the prototype centers to generate candidate potential vectors; The candidate latent vectors are input into the decoder to reconstruct and generate the candidate minority class samples.
[0124] In one embodiment, the first enhanced sample set generation module 303 is further configured to: Calculate the neighbor density between each candidate minority class sample and the minority class sample to obtain the density score of each candidate minority class sample; The threshold is set at the preset quantile of the density score of each of the candidate minority class samples. Candidate minority class samples with density scores below the threshold are removed to obtain the first enhanced sample set.
[0125] In one embodiment, the first enhanced sample set generation module 303 is further configured to: Calculate the average K-nearest neighbor distance between each of the candidate minority class samples and the minority class sample; The reciprocal of the average distance of the K nearest neighbors is used as the density score of each candidate minority class sample.
[0126] In one embodiment, the second enhanced sample set generation module 304 is further configured to: Use the first enhanced sample set as the initial sample pool; The current sample pool distribution characteristics are used as the environment state, the feature vector is used as the action, and the preset reward function guides the update of the generation strategy; the feature vector is used to represent the generated sample. The process of generating samples based on the current generation strategy, updating the current sample pool with the generated samples, calculating the environment state and reward of the updated current sample pool, and updating the generation strategy is repeated until a preset stopping condition is met. The current sample pool that reaches the preset stopping condition is used as the second enhanced sample set.
[0127] In one embodiment, the fusion module 305 is further configured to: The augmented sample set is merged with the imbalanced dataset to obtain the augmented training set; The downstream classification model is trained using the enhanced training set.
[0128] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute an imbalanced data augmentation method for medical insurance fraud detection. This method includes: acquiring an imbalanced dataset for medical insurance fraud detection; the imbalanced dataset containing majority class samples and minority class samples; inputting the minority class samples into a conditional variational autoencoder to obtain candidate minority class samples output by the conditional variational autoencoder; based on the quality assessment results of the candidate minority class samples, removing candidate minority class samples that do not meet preset quality standards to obtain a first augmented sample set; constructing a dynamic feedback optimization mechanism, using the overall distribution state of the first augmented sample set as feedback input, continuously adjusting the generation direction through the dynamic feedback optimization mechanism, iteratively optimizing the overall distribution of the first augmented sample set to obtain a second augmented sample set; and merging the first augmented sample set and the second augmented sample set to obtain an augmented sample set.
[0129] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0130] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the imbalanced data augmentation method for medical insurance fraud detection provided by the above methods. The method includes: acquiring an imbalanced dataset for medical insurance fraud detection; the imbalanced dataset contains majority class samples and minority class samples; inputting the minority class samples into a conditional variational autoencoder to obtain candidate minority class samples output by the conditional variational autoencoder; removing candidate minority class samples that do not meet preset quality standards based on the quality assessment results of the candidate minority class samples to obtain a first augmented sample set; constructing a dynamic feedback optimization mechanism, using the overall distribution state of the first augmented sample set as feedback input, continuously adjusting the generation direction through the dynamic feedback optimization mechanism, iteratively optimizing the overall distribution of the first augmented sample set to obtain a second augmented sample set; and merging the first augmented sample set and the second augmented sample set to obtain an augmented sample set.
[0131] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements an imbalanced data augmentation method for medical insurance fraud detection provided by the above methods. This method includes: acquiring an imbalanced dataset for medical insurance fraud detection; the imbalanced dataset containing majority class samples and minority class samples; inputting the minority class samples into a conditional variational autoencoder to obtain candidate minority class samples output by the conditional variational autoencoder; based on the quality assessment results of the candidate minority class samples, removing candidate minority class samples that do not meet preset quality standards to obtain a first augmented sample set; constructing a dynamic feedback optimization mechanism, using the overall distribution state of the first augmented sample set as feedback input, continuously adjusting the generation direction through the dynamic feedback optimization mechanism, iteratively optimizing the overall distribution of the first augmented sample set to obtain a second augmented sample set; and merging the first augmented sample set and the second augmented sample set to obtain an augmented sample set.
[0132] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0133] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for imbalanced data augmentation for medical insurance fraud detection, characterized in that, include: Obtain an imbalanced dataset for medical insurance fraud detection; the imbalanced dataset contains majority class samples and minority class samples; The minority class samples are input into the conditional variational autoencoder to obtain the candidate minority class samples output by the conditional variational autoencoder. Based on the quality assessment results of the candidate minority class samples, candidate minority class samples that do not meet the preset quality standards are removed to obtain the first enhanced sample set; The first enhanced sample set is used as the initial sample pool, the distribution state features of the current sample pool are used as the environment state, the feature vector is used as the action, and the preset reward function is used to guide the update of the generation strategy. The feature vector is used to represent the generated sample; The process of generating samples based on the current generation strategy, updating the current sample pool with the generated samples, calculating the environment state and reward of the updated current sample pool, and updating the generation strategy is repeated until a preset stopping condition is met. The current sample pool that reaches the preset stopping condition will be used as the second enhanced sample set. The first enhanced sample set and the second enhanced sample set are merged to obtain the enhanced sample set.
2. The imbalanced data augmentation method for medical insurance fraud detection according to claim 1, characterized in that, The conditional variational autoencoder includes an encoder and a decoder; the conditional variational autoencoder is trained in the following manner: The encoder takes the sample features and class labels from the training set as inputs and outputs the mean and log-variance of the latent space. The mean and the log-variance are reparameterized and sampled to obtain latent variables; The latent variables and the category labels are input into the decoder, and the reconstructed sample is output. The reconstruction loss is calculated based on the original samples and the reconstructed samples in the training set, and the KL divergence loss is calculated based on the mean and the log-variance. The parameters of the conditional variational autoencoder are updated using the weighted sum of the reconstruction loss and the KL divergence loss as the total loss, thus obtaining the conditional variational autoencoder.
3. The imbalanced data augmentation method for medical insurance fraud detection according to claim 2, characterized in that, In the prototype density-aware denoising stage, the conditional variational autoencoder is used for: The minority class samples are input into the encoder to extract the latent representation set of the minority class samples in the latent space; Clustering the potential representation set yields multiple prototype centers; Random perturbations are added to the neighborhood of each of the prototype centers to generate candidate potential vectors; The candidate latent vectors are input into the decoder to reconstruct and generate the candidate minority class samples.
4. The imbalanced data augmentation method for medical insurance fraud detection according to claim 1, characterized in that, The first enhanced sample set is obtained by removing candidate minority class samples that do not meet the preset quality standards based on the quality assessment results of the candidate minority class samples, including: Calculate the neighbor density between each candidate minority class sample and the minority class sample to obtain the density score of each candidate minority class sample; The threshold is set at the preset quantile of the density score of each of the candidate minority class samples. Candidate minority class samples with density scores below the threshold are removed to obtain the first enhanced sample set.
5. The imbalanced data augmentation method for medical insurance fraud detection according to claim 4, characterized in that, The calculation of the proximity density between each candidate minority class sample and the minority class sample to obtain the density score of each candidate minority class sample includes: Calculate the average K-nearest neighbor distance between each of the candidate minority class samples and the minority class sample; The reciprocal of the average distance of the K nearest neighbors is used as the density score of each candidate minority class sample.
6. The imbalanced data augmentation method for medical insurance fraud detection according to claim 1, characterized in that, After fusing the first enhanced sample set and the second enhanced sample set to obtain the enhanced sample set, the method further includes: The augmented sample set is merged with the imbalanced dataset to obtain the augmented training set; The downstream classification model is trained using the enhanced training set.
7. An imbalanced data augmentation device for medical insurance fraud detection, characterized in that, include: The data acquisition module is used to acquire an imbalanced dataset for medical insurance fraud detection; the imbalanced dataset contains majority class samples and minority class samples. The candidate sample generation module is used to input the minority class samples into the conditional variational autoencoder to obtain the candidate minority class samples output by the conditional variational autoencoder. The first enhanced sample set generation module is used to remove candidate minority samples that do not meet the preset quality standards based on the quality assessment results of the candidate minority samples, and obtain the first enhanced sample set. The second enhanced sample set generation module is used to take the first enhanced sample set as the initial sample pool, take the distribution state features of the current sample pool as the environment state, take the feature vector as the action, and use a preset reward function to guide the update of the generation strategy. The feature vector is used to represent the generated sample; the process of generating a sample based on the current generation strategy, updating the current sample pool with the generated sample, calculating the environment state and reward of the updated current sample pool, and updating the generation strategy is repeated until a preset stopping condition is reached; The current sample pool that reaches the preset stopping condition will be used as the second enhanced sample set. The fusion module is used to fuse the first enhanced sample set and the second enhanced sample set to obtain an enhanced sample set.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the imbalanced data augmentation method for medical insurance fraud detection as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the imbalanced data augmentation method for medical insurance fraud detection as described in any one of claims 1 to 6.