Sample generation method and device, computer equipment and readable storage medium

By performing structural analysis and dynamic weight generation on minority class samples, combined with category boundary screening, the problems of noise sensitivity and distribution deviation in traditional methods are solved. The generated samples significantly improve the accuracy and reliability of the model in financial fraud detection and medical diagnosis.

CN120744504APending Publication Date: 2025-10-03CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510897331.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing minority class sample generation methods cannot meet the stringent reliability requirements of practical applications in high-dimensional data scenarios. Traditional methods do not consider the intrinsic distribution structure of minority class samples. Blind interpolation may introduce noise samples across class boundaries and ignore the constraints of majority class samples on the generation process, causing the generated samples to deviate from the true minority class distribution, which cannot meet the needs of scenarios such as financial fraud detection and medical health diagnosis.

Method used

By obtaining the original data set and separating it into minority and majority class sample sets, and then standardizing it, the silhouette coefficient is used to determine the optimal number of clusters. The minority class samples are clustered, and the dynamic weights are calculated to generate candidate samples. Strict screening is performed based on the category boundaries to retain samples that are close to the minority class and far away from the majority class.

Benefits of technology

The generated samples are more consistent with real fraud patterns and the distribution of rare disease characteristics, reducing the model's missed reporting rate and misjudgment risk, and improving the accuracy and reliability of financial fraud detection and medical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744504A_ABST
    Figure CN120744504A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, and provides a sample generation method and device, computer equipment and a readable storage medium, and the method comprises the steps: obtaining a to-be-generated original data set, and separating the original data set into a minority class sample set and a majority class sample set; for the minority class samples, by calculating a contour coefficient under each clustering number, determining an optimal clustering number enabling the contour coefficient to be maximum, and clustering the minority class samples based on the optimal clustering number to obtain a plurality of minority class sample clusters so as to generate a plurality of candidate samples; and respectively acquiring a first average nearest neighbor distance from each candidate sample to the minority class sample and a second average nearest neighbor distance from each candidate sample to the majority class sample, so as to determine the target sample as a newly generated minority class sample. According to the method, samples conforming to a real fraud mode can be accurately generated in financial fraud detection, rare disease feature distribution can be effectively simulated in medical diagnosis, and the early screening accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of machine learning technology, and in particular to a sample generation method, apparatus, computer device, and readable storage medium. Background Art

[0002] In the fields of machine learning and data mining, data imbalance is a key challenge that hinders model performance. This problem is particularly prominent in scenarios such as financial fraud detection and healthcare diagnosis, where the identification of minority samples (such as fraudulent transactions and rare disease cases) is crucial. In the financial sector, for example, fraudulent transactions only account for a tiny fraction of the massive amount of transaction data. Traditional models are prone to "class bias" due to the lack of minority samples, leading to the risk of missed detections or a sharp increase in the cost of misjudgment. Furthermore, in healthcare, data on patients with rare diseases is scarce, and diagnostic models trained on imbalanced data can miss key features, delaying diagnosis and treatment.

[0003] Existing mainstream minority sample generation methods (such as SMOTE and its variants) generate new samples by interpolating neighboring samples. However, these methods suffer from significant flaws: First, they fail to consider the intrinsic distribution structure of minority samples. This blind interpolation can introduce noise samples that cross class boundaries (e.g., misclassifying the transition area between fraudulent and legitimate transactions as fraudulent). Second, they ignore the constraints imposed by majority class samples on the generation process, causing generated samples to deviate from the true minority class distribution or approach dense majority class areas, exacerbating model confusion. This is particularly true in high-dimensional data scenarios (such as the multi-dimensional characteristics of financial transactions and the high-dimensional parameters of medical images). Traditional methods suffer from the "curse of dimensionality," leading to a sharp decline in generated sample quality and failing to meet the stringent reliability requirements of practical applications.

[0004] Therefore, a sample generation method is urgently needed to solve at least one of the above problems. Summary of the Invention

[0005] The present application provides a sample generation method, apparatus, computer device and readable storage medium, which aims to solve the data imbalance problem in the field of machine learning and data mining, which is one of the key challenges that restrict model performance, especially in scenarios such as financial fraud detection and medical health diagnosis that have extremely high recognition requirements for minority samples (such as fraudulent transactions and rare disease cases).

[0006] In a first aspect, the present application provides a sample generation method, comprising: Obtaining an original data set to be generated, separating the original data set into a minority class sample set and a majority class sample set, and performing standardization on the minority class sample set and the majority class sample set to obtain standardized minority class samples and majority class samples; For the minority class samples, determining an optimal number of clusters that maximizes the silhouette coefficient by calculating the silhouette coefficient under each number of clusters, and clustering the minority class samples based on the optimal number of clusters to obtain a plurality of minority class sample clusters; Obtaining the nearest neighbor distance between the minority class sample corresponding to the minority class sample cluster and the majority class sample, calculating a dynamic weight according to the nearest neighbor distance, and generating a plurality of candidate samples according to the dynamic weight; Obtain a first average nearest neighbor distance from each candidate sample to the minority class sample and a second average nearest neighbor distance from each candidate sample to the majority class sample, and use the candidate sample whose first average nearest neighbor distance is smaller than the second average nearest neighbor distance as a target sample, and use the target sample as a newly generated minority class sample.

[0007] In some embodiments, separating the original data set into a minority class sample set and a majority class sample set includes: counting the number of samples corresponding to each class label according to the class labels of the samples in the original data set, and dividing the samples corresponding to the class whose sample number ratio is less than a preset ratio threshold into the minority class sample set; and dividing the samples corresponding to the class whose sample number ratio is greater than or equal to the preset ratio threshold into the majority class sample set.

[0008] In some embodiments, the standardization processing of the minority class sample set and the majority class sample set includes: for each feature dimension, calculating the feature mean and feature standard deviation of the overall data set after the minority class sample set and the majority class sample set are merged; based on the feature mean and feature standard deviation, standardizing each sample in the minority class sample set and the majority class sample set so that the mean corresponding to each feature dimension after standardization is 0 and the standard deviation is 1, thereby obtaining the standardized minority class samples and majority class samples.

[0009] In some embodiments, clustering the minority class samples based on the optimal number of clusters to obtain multiple minority class sample clusters includes: inputting the optimal number of clusters and the minority class samples into a preset Gaussian mixture model, the Gaussian mixture model uses the optimal number of clusters as the number of components, performing probability density fitting and clustering on the minority class samples, so that each of the minority class samples is assigned to a corresponding sample cluster, and obtaining multiple minority class sample clusters.

[0010] In some embodiments, obtaining the nearest neighbor distance from the minority class sample corresponding to the minority class sample cluster to the majority class sample includes: constructing a KD tree for the minority class sample and the majority class sample respectively, and quickly searching the nearest neighbor distance from each minority class sample to the majority class sample through the KD tree.

[0011] In some embodiments, before obtaining the first average nearest neighbor distance of each candidate sample to the minority class sample and the second average nearest neighbor distance to the majority class sample, the method further includes: calculating the mean and maximum sample distance of the minority class samples corresponding to each minority class sample cluster; using the mean as the center of an envelope circle and the maximum sample distance as the radius to construct an envelope circle; filtering out candidate samples that are not within the envelope circle and retaining the candidate samples within the envelope circle.

[0012] In some embodiments, after the target sample is used as a newly generated minority class sample, the method further includes: denormalizing the target sample from the normalized space back to the original data space, and merging the target sample with the original data set to form a balanced new data set.

[0013] In a second aspect, the present application provides a sample generation device, comprising: a data acquisition unit, configured to acquire an original data set to be generated, separate the original data set into a minority class sample set and a majority class sample set, and perform standardization on the minority class sample set and the majority class sample set to obtain standardized minority class samples and majority class samples; a clustering calculation unit, configured to determine, for the minority class samples, an optimal number of clusters that maximizes the silhouette coefficient by calculating the silhouette coefficient under each number of clusters, and cluster the minority class samples based on the optimal number of clusters to obtain a plurality of minority class sample clusters; a distance acquisition unit, configured to acquire the nearest neighbor distance between the minority class sample corresponding to the minority class sample cluster and the majority class sample, calculate a dynamic weight according to the nearest neighbor distance, and generate a plurality of candidate samples according to the dynamic weight; The sample generation unit is used to respectively obtain a first average nearest neighbor distance from each candidate sample to the minority class sample and a second average nearest neighbor distance from each candidate sample to the majority class sample, and use the candidate sample whose first average nearest neighbor distance is smaller than the second average nearest neighbor distance as a target sample, and use the target sample as a newly generated minority class sample.

[0014] In a third aspect, the present application further provides a computer device, comprising: memory and processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the steps of the sample generation method described in the first aspect above when executing the computer program.

[0015] In a fourth aspect, the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the steps of the sample generation method described in the first aspect above.

[0016] The present invention provides a sample generation method, apparatus, computer device, and readable storage medium. This method fundamentally addresses the blindness and noise sensitivity of traditional methods by performing structural analysis on minority class samples (adaptively determining the optimal number of clusters and dividing the sample clusters), introducing distance constraints on majority class samples (dynamically weighting the generation direction), and performing strict screening based on class boundaries (retaining samples close to the minority class and far from the majority class). This method can accurately generate samples that conform to real fraud patterns in financial fraud detection, reducing the model's false negative rate; and effectively simulate the characteristic distribution of rare diseases in medical diagnosis, improving the accuracy of early screening. This method offers significant technological advancement and practical application value.

[0017] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 This is a schematic flow chart of the steps of a sample generation method provided in one embodiment of the present application; Figure 2 is a structural diagram of a sample generation device provided in one embodiment of the present application; Figure 3 This is a schematic block diagram of the structure of a computer device provided in one embodiment of the present application.

[0020] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. DETAILED DESCRIPTION

[0021] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0022] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0023] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish between identical or similar items having substantially the same functions and effects. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or order of execution, and that terms such as "first" and "second" do not necessarily define differences.

[0024] It should be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0025] It will also be understood that the term "and / or" as used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0026] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.

[0027] In the fields of machine learning and data mining, data imbalance is a key challenge that hinders model performance. This problem is particularly prominent in scenarios such as financial fraud detection and healthcare diagnosis, where the identification of minority samples (such as fraudulent transactions and rare disease cases) is crucial. In the financial sector, for example, fraudulent transactions only account for a tiny fraction of the massive amount of transaction data. Traditional models are prone to "class bias" due to the lack of minority samples, leading to the risk of missed detections or a sharp increase in the cost of misjudgment. Furthermore, in healthcare, data on patients with rare diseases is scarce, and diagnostic models trained on imbalanced data can miss key features, delaying diagnosis and treatment.

[0028] Existing mainstream minority sample generation methods (such as SMOTE and its variants) generate new samples by interpolating neighboring samples. However, these methods suffer from significant flaws: First, they fail to consider the intrinsic distribution structure of minority samples. This blind interpolation can introduce noise samples that cross class boundaries (e.g., misclassifying the transition area between fraudulent and legitimate transactions as fraudulent). Second, they ignore the constraints imposed by majority class samples on the generation process, causing generated samples to deviate from the true minority class distribution or approach dense majority class areas, exacerbating model confusion. This is particularly true in high-dimensional data scenarios (such as the multi-dimensional characteristics of financial transactions and the high-dimensional parameters of medical images). Traditional methods suffer from the "curse of dimensionality," leading to a sharp decline in generated sample quality and failing to meet the stringent reliability requirements of practical applications.

[0029] Therefore, a sample generation method is urgently needed to solve at least one of the above problems.

[0030] To resolve the above issues, please refer to Figure 1 , Figure 1 This is a schematic flow chart of a sample generation method provided in one embodiment of the present application. This sample generation method can be implemented by a computer device, which can be deployed on a single server or a server cluster. It can also be deployed on a handheld terminal, a laptop computer, a wearable device, or a robot.

[0031] It should be noted that the acquisition of any information mentioned in the provided method complies with relevant regulations and is carried out with the user's consent, and will not infringe on the user's privacy or violate relevant laws and regulations.

[0032] To solve the above problems, please refer to Figure 1 Specifically, Figure 1 As shown, the provided sample generation method includes steps S101 to S104. The details are as follows: Step S101: Obtain an original data set to be generated, separate the original data set into a minority class sample set and a majority class sample set, and perform standardization on the minority class sample set and the majority class sample set to obtain standardized minority class samples and majority class samples.

[0033] Specifically, based on the class labels of the original dataset, the samples are divided into the minority class (the class with a low sample size) and the majority class (the class with a high sample size). The features of the two classes of samples are then normalized to eliminate the impact of dimensional differences in the different feature dimensions. Specifically, the mean and standard deviation of the features of the combined overall dataset are calculated, and each sample is Z-score normalized to ensure that the features of each dimension follow a standard normal distribution with a mean of 0 and a standard deviation of 1.

[0034] For example, in financial scenarios, data separation involves using transaction data as an example, labeling it "fraud" (minority class) and "normal" (majority class), separating the two classes of samples based on a preset threshold (e.g., fraud sample ratio <1%). Standardization involves combining all samples to calculate the global mean and standard deviation of multiple features, such as transaction amount (in the tens of thousands of yuan), transaction time interval (in seconds), and geographic latitude and longitude (floating point). This converts each sample's feature value into a standardized score, preventing high-value features like amount from dominating the model.

[0035] For example, in healthcare scenarios, data separation involves labeling disease diagnosis data as "rare diseases" (minority category) and "non-rare diseases" (majority category), separating case data based on the rare disease incidence threshold (e.g., <0.01%) in medical statistics. Standardization involves combining samples from both categories to calculate global statistics based on features such as age (years), blood pressure (mmHg), and blood markers (e.g., white blood cell count × 10^9 / L). This standardization eliminates scale differences between features of different dimensions, such as age and blood pressure, facilitating subsequent clustering and distance calculations.

[0036] This step eliminates dimensional differences between amounts and time in financial transactions and between age and test indicators in medical data, preventing high-value features from interfering with clustering and distance calculations, thus ensuring fairness in subsequent analyses. This standardization approach is applicable to both high-dimensional financial transaction features and multimodal medical physiological indicators, providing a foundation for unified processing across diverse scenarios and enhancing the method's universality.

[0037] Step S102: For the minority class samples, the silhouette coefficient under each cluster number is calculated to determine the optimal cluster number that maximizes the silhouette coefficient, and the minority class samples are clustered based on the optimal cluster number to obtain multiple minority class sample clusters.

[0038] Specifically, for the standardized minority class samples, the team enumerated different numbers of clusters (e.g., k = 2 to 10) and calculated the corresponding silhouette coefficient (Silhouette Coefficient, a measure of the closeness of sample clustering and the degree of separation between clusters) for each k value. The k with the largest silhouette coefficient was selected as the optimal number of clusters. Probabilistic clustering of minority class samples was performed based on a Gaussian mixture model (GMM), assigning each sample to a corresponding cluster with probability. This revealed substructure differences within the minority class (e.g., different fraud patterns and rare disease subtypes).

[0039] In financial scenarios, for example, the silhouette coefficient is calculated based on characteristics of fraud samples, such as transaction device type, IP address geographic entropy, and transaction time periodicity, to determine the silhouette coefficient for different numbers of clusters. For example, when k=3, it can distinguish between three fraud patterns: "counterfeit card transactions," "phishing," and "internal fraud." GMM clustering assumes that each fraud pattern follows a Gaussian distribution. Using the EM algorithm to fit the mean and covariance matrix of each cluster, each fraud sample is assigned to the most likely cluster based on posterior probability. For example, a sample has an 80% probability of belonging to the "counterfeit card transaction cluster."

[0040] For example, in healthcare scenarios, silhouette coefficient calculations are performed using features such as gene expression data and clinical symptom scores for rare disease patients, with a range of k = 2 to 5. When k = 2, early-onset and late-onset subtypes of the disease can be distinguished, corresponding to different pathological mechanisms. GMM clustering explicitly separates subtype characteristics by fitting Gaussian distribution parameters to each subtype. For example, the feature mean of the early-onset subtype tends to favor "younger-onset with abnormalities in specific indicators," while the late-onset subtype tends to favor "middle-aged and older-onset with abnormalities in another set of indicators."

[0041] This step breaks the traditional assumption of uniform distribution of minority classes. Clustering captures the multimodal nature of financial fraud and the multi-subtype diversity of rare diseases, ensuring that the generated samples more closely align with the distribution characteristics of the actual subclusters and avoiding invalid cross-modal interpolation. The silhouette coefficient automatically determines the optimal number of clusters, eliminating the need for manual parameter adjustment and adapting to the complex internal structure of minority classes in different scenarios (e.g., variable fraud patterns and unknown rare disease subtypes).

[0042] Step S103: Obtain the nearest neighbor distance between the minority class sample and the majority class sample corresponding to the minority class sample cluster, calculate the dynamic weight according to the nearest neighbor distance, and generate multiple candidate samples according to the dynamic weight.

[0043] Specifically, for each minority class sample cluster, the KD tree is used to quickly search the nearest neighbor distance of each sample in the cluster to the majority class sample, and the dynamic weight is calculated based on the principle of "the farther the distance, the higher the weight" (such as weight = 1 / (1+e^(-d)), where d is the nearest neighbor distance). The sampling density of the kernel density estimation (KDE) is weighted and adjusted to generate candidate samples that are biased away from the dense area of ​​the majority class.

[0044] For example, in financial scenarios, the nearest neighbor search uses a KD tree to find the 10 closest samples among normal transaction samples for each sample in the "counterfeit card transaction cluster," and calculates the Euclidean distance as the nearest neighbor distance d. (A larger d indicates that the counterfeit card sample is further away from the core area of ​​normal transactions.) Dynamic weighting and KDE sampling assign higher weights to samples with larger d, guiding KDE to generate more candidate samples near them. For example, it can generate fraud candidate samples in areas characterized by "midnight, out-of-town, large-value transactions" (areas with fewer normal transactions) while avoiding noise generation in areas with dense normal transactions such as "weekday, local, small-value transactions."

[0045] For example, in healthcare scenarios, nearest neighbor search calculates the nearest neighbor distance d (e.g., Mahalanobis distance based on gene expression profiles) from each patient sample in an early-onset rare disease cluster to a healthy sample. A larger d indicates that the patient's characteristics deviate further from those of a healthy population. Dynamic weighting and KDE sampling assign higher weights to cases with larger d values, generating candidate samples in the feature space characterized by high expression of specific genes during adolescence. This avoids generating invalid samples in regions with characteristics similar to those of a healthy population (e.g., regions with normal gene expression levels).

[0046] By introducing distance constraints on majority class samples, generated samples are more likely to be in the "pure regions" of the minority class (e.g., the fraud feature space far from normal transactions, and the disease feature space far from healthy people), reducing noise generated across class boundaries. KD trees are used to address the "curse of dimensionality" problem in high-dimensional financial transaction features (e.g., 50+ dimensions) and medical genomic data (tens of thousands of features), ensuring the efficiency and accuracy of nearest neighbor searches.

[0047] Step S104. Obtain the first average nearest neighbor distance of each candidate sample to the minority class sample and the second average nearest neighbor distance to the majority class sample respectively, and take the candidate sample whose first average nearest neighbor distance is smaller than the second average nearest neighbor distance as the target sample, and take the target sample as the newly generated minority class sample.

[0048] Specifically, for each candidate sample, the average nearest neighbor distance to all samples of the minority class (the first average distance, which measures the closeness to the minority class) and the average nearest neighbor distance to all samples of the majority class (the second average distance, which measures the distance from the majority class) are calculated respectively. Only samples with the first average distance less than the second average distance are retained to ensure that the generated samples are closer to the minority class distribution and away from the majority class dense area.

[0049] For example, in financial scenarios, distance calculation involves calculating the mean Euclidean distance (first average distance) from a candidate fraud sample to all real fraud samples, as well as the mean Euclidean distance (second average distance) from all normal transaction samples. The screening logic involves retaining a candidate sample as valid if its first average distance is less than its second average distance, indicating that it more closely matches the distribution of real fraud samples and is further away from the core area of ​​normal transactions. Conversely, if it is closer to normal transactions (e.g., if its second average distance is smaller), it is filtered out (for example, excluding candidate samples with the "small local transactions on workdays" characteristic, as they are closer to normal transaction patterns).

[0050] For example, in healthcare scenarios, distance calculation involves calculating the mean Mahalanobis distance (taking into account feature correlation) between a candidate rare disease sample and all actual rare disease patient samples, as well as the mean Mahalanobis distance to healthy samples. The screening logic then retains only samples that are closer to rare disease patients (with a smaller first mean distance) and further away from healthy individuals (with a larger second mean distance), excluding candidate samples that are close to healthy indicators but slightly abnormal (to avoid generating noisy samples that could be misdiagnosed as healthy).

[0051] Through a two-way screening process of "close to the minority class and away from the majority class," generated samples are ensured to lie within clear class boundaries, avoiding the "cross-boundary noise" (such as seemingly normal fraud samples and seemingly healthy disease samples) that can confound the model in traditional methods. Quantitative screening criteria (distance comparison) replace subjective thresholds, adapting to the complexity of class boundaries in different scenarios. This significantly improves the reliability of generated samples and reduces the risk of missed or misclassified classifications, particularly in scenarios like financial fraud, where small samples carry high risk, and medical diagnosis, where the cost of misdiagnosis is extremely high.

[0052] In financial scenarios, the provided method solves the problem of "scarce fraud samples + changeable patterns", generates samples that fit the real fraud sub-patterns and are far away from normal transaction-intensive areas, and improves the anti-fraud model's ability to identify new types of fraud; in medical scenarios, it solves the problem of "insufficient rare disease data + subtype differences", generates samples that focus on areas with significant disease characteristics, and assists the model in capturing key indicators for early diagnosis, reducing missed diagnosis and delays.

[0053] In some embodiments, separating the original data set into a minority class sample set and a majority class sample set includes: counting the number of samples corresponding to each class label according to the class labels of the samples in the original data set, and dividing the samples corresponding to the class whose sample number ratio is less than a preset ratio threshold into the minority class sample set; and dividing the samples corresponding to the class whose sample number ratio is greater than or equal to the preset ratio threshold into the majority class sample set.

[0054] By counting the number of samples in each category according to the category labels of the original data set, the minority class and the majority class are divided according to a preset ratio threshold (such as 5%): the categories with a sample ratio less than the threshold are classified into the minority class sample set, and the categories with a sample ratio greater than or equal to the threshold are classified into the majority class sample set.

[0055] For example, in a financial scenario, the proportion of samples labeled "fraud" in the statistical transaction data (e.g., 0.3%) is lower than the preset threshold (e.g., 1%), so all "fraud" labeled samples are classified as the minority class sample set; the "normal transaction" labeled samples account for 99.7% and are classified as the majority class sample set.

[0056] For example, in medical and health scenarios, the proportion of samples labeled "rare diseases" in statistical disease diagnosis data (such as 0.05%) is lower than the preset threshold of the medical scenario (such as 0.1%), and is classified as the minority class; the proportion of samples labeled "non-rare diseases" is 99.95%, which is classified as the majority class.

[0057] By quantifying the degree of class imbalance through preset ratio thresholds, we avoid subjective classifications and adapt to the objective data distribution characteristics of financial (where fraud rates are extremely low) and medical (where rare diseases are strictly defined). The thresholds can be adjusted based on domain requirements (e.g., 0.5% for finance and 0.1% for medical), flexibly adapting to the definition of "minority classes" in different scenarios and enhancing the method's universality.

[0058] In some embodiments, the standardization processing of the minority class sample set and the majority class sample set includes: for each feature dimension, calculating the feature mean and feature standard deviation of the overall data set after the minority class sample set and the majority class sample set are merged; based on the feature mean and feature standard deviation, standardizing each sample in the minority class sample set and the majority class sample set so that the mean corresponding to each feature dimension after standardization is 0 and the standard deviation is 1, thereby obtaining the standardized minority class samples and majority class samples.

[0059] For the combined dataset of minority and majority class samples, the mean and standard deviation of each feature dimension are calculated, and the samples are converted to a standard normal distribution with a mean of 0 and a standard deviation of 1 through Z-score normalization.

[0060] For example, in a financial scenario, by combining fraudulent and normal transaction samples, the global mean (e.g., 500 yuan) and standard deviation (e.g., 2,000 yuan) of the "transaction amount" dimension are calculated. The transaction amount of a fraudulent sample (3,000 yuan) is standardized to (3,000-500) / 2,000 = 1.25, eliminating the dimensional difference with low-value domain features such as "transaction frequency" (mean 10 times, standard deviation 5 times).

[0061] For example, in a healthcare scenario, by combining rare disease and healthy samples, the global mean (120 mmHg) and standard deviation (15 mmHg) of the "systolic blood pressure" dimension are calculated. The systolic blood pressure (150 mmHg) of a patient with a rare disease is standardized to (150-120) / 15=2, ensuring that it is analyzed in the same metric space as features such as "age" (mean 40 years, standard deviation 15 years).

[0062] This approach avoids analytical bias caused by high-value features like "amount" dominating distance calculations in financial transactions (e.g., Euclidean distance being dominated by the amount dimension), or by the different scales of "blood pressure" and "age" in medical data. The standardized data meets the feature scale requirements of Gaussian mixture models (GMMs) and KD tree searches, improving the accuracy of clustering and distance calculations. This is particularly crucial in scenarios involving high-dimensional financial features (e.g., 50+ dimensions) and multimodal medical data (physiological indicators + imaging features).

[0063] In some embodiments, clustering the minority class samples based on the optimal number of clusters to obtain multiple minority class sample clusters includes: inputting the optimal number of clusters and the minority class samples into a preset Gaussian mixture model, the Gaussian mixture model uses the optimal number of clusters as the number of components, performing probability density fitting and clustering on the minority class samples, so that each of the minority class samples is assigned to a corresponding sample cluster, and obtaining multiple minority class sample clusters.

[0064] By taking the optimal number of clusters as the number of components of the Gaussian mixture model (GMM), probability density fitting is performed on the minority class samples, the mean and covariance matrix of each component are estimated by the EM algorithm, and each sample is assigned to the most likely cluster according to the posterior probability.

[0065] For example, in a financial scenario, if the optimal number of clusters is three, GMM assumes three fraud patterns (such as "counterfeit card transactions," "phishing," and "internal fraud") and fits three Gaussian distributions to fraud samples based on characteristics such as "transaction device uniqueness" and "IP address variability." The posterior probability of a sample belonging to the "counterfeit card transaction cluster" is 0.7, and it is assigned to that cluster.

[0066] For example, in a healthcare scenario, if the optimal number of clusters is 2, GMM distinguishes between "early-onset" and "late-onset" subtypes of rare diseases, fitting two Gaussian distributions to the features "age of onset" and "specific gene expression levels." The posterior probability of a patient with a rare disease belonging to the "early-onset subtype cluster" is 0.85, and the patient is assigned to this cluster.

[0067] By breaking the traditional assumption of a single distribution for minority classes, this approach can identify multiple modes of financial fraud (such as device fraud and data fraud) and multiple subtypes of rare medical diseases (e.g., with different pathological mechanisms), ensuring that generated samples more closely align with the characteristic distribution of true subclusters. By using posterior probabilities to handle ambiguous sample attribution (e.g., a sample may possess some characteristics of two fraud modes simultaneously), this approach avoids the rigid division of clusters and improves the reliability of clustering results. This approach is particularly suitable for complex scenarios with overlapping symptoms in medical data.

[0068] In some embodiments, obtaining the nearest neighbor distance from the minority class sample corresponding to the minority class sample cluster to the majority class sample includes: constructing a KD tree for the minority class sample and the majority class sample respectively, and quickly searching the nearest neighbor distance from each minority class sample to the majority class sample through the KD tree.

[0069] This embodiment constructs a KD tree (a high-dimensional spatial index structure) for the minority class and majority class samples respectively, uses the KD tree to quickly search for the nearest neighbor sample of each minority class sample in the majority class, and calculates the Euclidean distance or Mahalanobis distance as the nearest neighbor distance.

[0070] For example, in financial scenarios, a KD tree is constructed for each of the standardized fraud samples (minority class) and normal transaction samples (majority class). For a fraud sample, the KD tree is used to search for the 10 nearest neighbors among the normal transaction samples. The minimum Euclidean distance is taken as the nearest neighbor distance from the fraud sample to the majority class (reflecting its closeness to normal transactions).

[0071] For example, in a healthcare scenario, a KD tree is constructed for samples of patients with rare diseases (minority class) and healthy samples (majority class). For a sample of a patient with a rare disease, the five nearest neighbor samples in the healthy sample are searched, and the Mahalanobis distance (taking feature correlation into account) is calculated as the nearest neighbor distance to measure the degree of difference between its features and those of the healthy population.

[0072] By addressing the "curse of dimensionality" problem in financial transactions (over 50 dimensions) and medical genomics data (tens of thousands of dimensions), the KD tree achieves near-linear search time complexity, far outperforming the exponential complexity of brute-force searches and ensuring computational efficiency for large datasets. Spatial indexing rapidly locates nearest neighbor samples, providing accurate distance information for subsequent dynamic weight calculations and boundary screening, thus avoiding the computational bottlenecks of traditional brute-force searches in high-dimensional scenarios.

[0073] In some embodiments, before obtaining the first average nearest neighbor distance of each candidate sample to the minority class sample and the second average nearest neighbor distance to the majority class sample, the method further includes: calculating the mean and maximum sample distance of the minority class samples corresponding to each minority class sample cluster; using the mean as the center of an envelope circle and the maximum sample distance as the radius to construct an envelope circle; filtering out candidate samples that are not within the envelope circle and retaining the candidate samples within the envelope circle.

[0074] For each minority class sample cluster, the mean of the samples in the cluster (as the center of the envelope circle) and the maximum distance from the sample to the center of the circle (as the radius) are calculated to construct a hypersphere (envelope circle). The candidate samples that are not in the sphere are filtered out, and only the samples in the sphere are retained for subsequent screening.

[0075] For example, in a financial scenario, for a "counterfeit card transaction cluster," the feature mean of all fraudulent samples within the cluster is calculated (for example, the mean vector for "midnight transactions + remote devices + large amounts"). The Euclidean distance from each sample to the mean is calculated, and the maximum value is used as the radius to construct an enveloping circle. Candidate samples generated that meet the "distance to the mean ≤ radius" condition are retained (for example, candidate samples with the "daytime local small amount transactions" feature are excluded because they fall outside the cluster's feature range).

[0076] For example, in a healthcare scenario, for an "early-onset rare disease cluster," the mean characteristic value of patient samples within the cluster (e.g., the mean vector for "adolescent age + high expression of gene A") is calculated, and the maximum Mahalanobis distance is used as the radius to construct an enveloping circle. Candidate samples within the circle (e.g., "under 20 years old + gene A expression ≥ mean - 2σ") are retained, while invalid samples such as "middle-aged or elderly + normal gene A expression" are excluded.

[0077] By ensuring that generated samples lie within the inherent feature space of the minority cluster (e.g., the typical feature range of the financial fraud cluster or the core symptom range of rare medical disease subtypes), we avoid generating cross-cluster or extremely abnormal noise samples (e.g., seemingly normal fraud samples or cases that fall outside the known subtype characteristics). By using envelope filtering, we can bring candidate samples closer to the true distribution of the minority cluster. This, particularly in medical scenarios, prevents the generation of unreasonable samples that "deviate from known disease characteristics" and reduces the risk of the model misjudging abnormal noise.

[0078] In some embodiments, after the target sample is used as a newly generated minority class sample, the method further includes: denormalizing the target sample from the normalized space back to the original data space, and merging the target sample with the original data set to form a balanced new data set.

[0079] The filtered target samples are converted from the standardized space back to the original data space through denormalization (x=x′×σ+μ) and merged with the original dataset to form a new dataset with balanced categories.

[0080] For example, in a financial scenario, the standardized target fraud samples are denormalized to the original scale (e.g., the standardized value of 1.25 is converted to 1.25×2000+500=3000 yuan) using the mean (500 yuan) and standard deviation (2000 yuan) of the "transaction amount" dimension calculated in Example 2, and then merged with the original transaction data to increase the number of fraud samples to a ratio close to that of normal transactions (e.g., 1:10).

[0081] For example, in a healthcare scenario, the standardized target rare disease samples are denormalized to their original values ​​(e.g., the standardized value 2 is converted to 2×15+120=150 mmHg) using the mean (120 mmHg) and standard deviation (15 mmHg) of the "systolic blood pressure" dimension, and then merged with the original case data to balance the number of rare disease and non-rare disease samples and improve the balance of the diagnostic model training data.

[0082] Denormalization ensures that the feature values ​​of generated samples conform to their original business meanings (e.g., the actual scale of financial transaction amounts or the clinical units of medical indicators). This facilitates subsequent model training and business interpretation, avoiding analytical bias caused by data scale distortion. By merging generated samples with the original data, the proportion of minority class samples is directly increased, providing more valid training samples for financial anti-fraud models (reducing missed detections) and a more balanced feature distribution for medical diagnosis models (reducing misdiagnoses), forming a complete solution from generation to application.

[0083] In some embodiments, the present application proposes a proximity-based synthetic over-sampling (PBSS) algorithm. Data imbalance is a common and thorny problem in machine learning. When the number of samples in one class is significantly smaller than that in other classes, traditional machine learning algorithms often exhibit poor learning and prediction performance for minority class samples. This is because most algorithms tend to favor the more numerous class, significantly reducing their ability to identify and classify minority class samples. SMOTE (Synthetic Minority Over-sampling Technique) is an interpolation-based oversampling technique. Its core idea is to balance the dataset by artificially synthesizing new minority class samples. The specific steps are as follows: For each minority class sample, its k nearest neighbors are found. Random linear interpolation is performed between the original sample and its nearest neighbors to generate new synthetic samples, thereby increasing the number of minority class samples. SMOTE addresses data imbalance and avoids the risk of overfitting caused by simply duplicating samples, thereby increasing the diversity of minority class samples and improving the generalization ability of machine learning models on imbalanced datasets.

[0084] The biggest flaw of the traditional SMOTE algorithm is its overly simplistic and blind interpolation method. It simply selects the nearest neighbors based on Euclidean distance and performs linear interpolation between these samples. This approach has serious limitations: it fails to account for the local distribution of samples; it may generate unreasonable samples near class boundaries; and it performs poorly on high-dimensional data. The SMOTE algorithm is highly sensitive to noise and outliers in the data: noisy minority class samples may be incorrectly used to generate new samples; outliers may cause the generated synthetic samples to deviate from the true data distribution; and it lacks an effective mechanism for screening sample quality. As the feature dimensionality increases, the performance of the SMOTE algorithm degrades dramatically: in high-dimensional space, the concept of "nearest neighbor" becomes unreliable; distance calculations between samples become meaningless; interpolated samples may be distributed in a sparse high-dimensional space. Traditional SMOTE also has significant shortcomings in handling class boundaries: it cannot effectively distinguish complex boundaries between classes; it may generate inappropriate samples in areas of class overlap; and it lacks a deep understanding of class distribution. The SMOTE algorithm has high computational overhead: it requires finding the k nearest neighbors for each minority class sample; it is computationally inefficient on large datasets; and its computational complexity increases exponentially with the number of samples and dimensionality. The SMOTE algorithm also lacks flexibility in determining the oversampling ratio: it is difficult to precisely control the number of generated samples; excessive or insufficient oversampling can affect model performance; and it lacks an adaptive sampling strategy. Traditional SMOTE fails to fully account for: the complex distribution characteristics of different classes; nonlinear relationships between samples; and differences in local data density.

[0085] The PBSS algorithm provided is an advanced and intelligent method for generating minority class samples. It significantly improves the performance of the traditional SMOTE algorithm through multi-level data processing and intelligent sampling. It specifically includes the following steps: 1. Data Separation and Preprocessing: In the first step of the algorithm, the input data is divided into minority and majority class samples. This separation allows the algorithm to focus on generating minority class samples, ensuring that the generated new samples better represent the characteristics of the minority class. Furthermore, the algorithm normalizes the data to eliminate scale differences between features, thereby improving the accuracy of subsequent clustering and distance calculations. This preprocessing step lays the foundation for subsequent clustering and sample generation.

[0086] 2. Automatically select the optimal number of clusters: The algorithm automatically selects the optimal number of clusters by calculating the Silhouette Score. By trying different numbers of clusters, the algorithm evaluates the effectiveness of each clustering scheme and selects the optimal number of clusters. This process not only improves clustering effectiveness but also ensures the diversity and representativeness of the generated samples, avoiding the degradation of sample quality caused by an inappropriate choice of cluster number.

[0087] 3. Clustering using a Gaussian mixture model: After determining the optimal number of clusters, the algorithm uses a Gaussian mixture model (GMM) to cluster minority samples. GMMs capture the complex distribution characteristics of data and perform better than traditional K-means clustering when handling clusters of varying shapes and sizes. This choice ensures that the generated new samples better reflect the true distribution of minority samples.

[0088] 4. KD Tree Construction and Nearest Neighbor Search: The algorithm constructs a KD tree to accelerate nearest neighbor search. This data structure efficiently finds distances between samples. Using the KD tree, the algorithm can quickly find the distance from each minority class sample to the nearest majority class sample, providing important information for subsequent sample generation. This optimization significantly improves the algorithm's computational efficiency, especially when working with large datasets.

[0089] 5. Kernel Density Estimation and Weighted Sampling: When generating new samples, the algorithm incorporates kernel density estimation (KDE) and dynamically calculates weights based on the distance between the sample and the nearest majority class sample. Samples with greater distances receive higher weights when generating new samples. This mechanism ensures that the generated new samples are closer to minority class samples and further away from majority class samples. This weighted sampling strategy effectively improves the quality of new samples and reduces noise in the generated samples.

[0090] 6. Constructing an Envelope and Sample Filtering: The algorithm constructs an envelope by calculating the mean and maximum distance of minority class samples to limit the range of new sample generation. During the sampling process, the algorithm filters out samples outside the envelope to ensure that new samples generated fall within the distribution range of minority class samples. This step effectively avoids the risk of generating unreasonable samples and further improves sample validity.

[0091] 7. Generating New Samples and Denormalizing: After successfully generating new samples, the algorithm denormalizes them from the normalized space back to the original data space. This process ensures that the generated new samples can be seamlessly integrated with the original dataset, maintaining data consistency and usability. Finally, the algorithm merges the newly generated samples with the original data to form a new feature matrix and label array.

[0092] The improved distance-based SMOTE algorithm significantly improves the performance of the traditional SMOTE algorithm by introducing clustering, KD trees, kernel density estimation, and weighted sampling techniques. By automatically selecting the optimal number of clusters and dynamically calculating sample weights, the algorithm generates higher-quality minority class samples while reducing the generation of noisy and unreasonable samples. Furthermore, the algorithm excels at handling high-dimensional data and complex class boundaries, effectively addressing data imbalance. Overall, this algorithm not only improves the quality of minority class sample generation but also enhances the generalization ability of machine learning models on imbalanced datasets, providing a more reliable solution for practical applications.

[0093] This improved distance-based SMOTE algorithm demonstrates superior performance across multiple dimensions in addressing data imbalance. By thoroughly analyzing the provided experimental data, we can fully evaluate the algorithm's innovative value and practical effectiveness.

[0094] Traditional SMOTE and non-oversampling methods have significant limitations in processing minority samples. By introducing kernel density estimation and a dynamic weighting mechanism, the improved algorithm significantly improves the quality of new sample generation. Specifically, this more accurately captures the distribution characteristics of minority samples, reduces the generation of noisy samples, and enhances sample representativeness and diversity.

[0095] At the same time, the algorithm significantly improved the performance of different classifiers: the F1 score of the Naive Bayes classifier increased from 0.981615 to 0.990874; the precision of the LGBM classifier increased from 0.981682 to 0.981912; and the recall rate of the MLP classifier increased significantly from 0.000000 to 0.027523.

[0096] The algorithm's innovations lie in: automatically selecting the optimal number of clusters; a distance-based weighted sampling strategy; and generating samples using envelope constraints. These mechanisms enable the algorithm to adapt to more complex data distributions and improve the model's generalization capabilities. By calculating the distance between a sample and the majority class sample and assigning a weight to each sample, the generated new samples are more meaningful. This mechanism significantly improves the targetedness and effectiveness of sample generation. By introducing kernel density estimation techniques, it can more accurately simulate the probability distribution of minority class samples, and compared to traditional linear interpolation methods, the generated samples are closer to the true data distribution. The algorithm automatically selects the optimal number of clusters using the silhouette coefficient, overcoming the limitations of traditional methods that manually set the number of clusters, making the sample generation process more intelligent and precise. For highly imbalanced datasets such as financial fraud detection, rare disease diagnosis, and industrial defect detection, our algorithm has significant practical significance: it significantly improves the recognition accuracy of minority class samples, reduces the risk of misclassification, and enhances the performance of machine learning models on extremely imbalanced datasets.

[0097] By comparing experimental data, the algorithm has made significant progress in multiple key indicators: the precision rate has increased by an average of about 2-3 percentage points; the recall rate has increased from close to 0 to 0.02-0.09; and the F1 score has increased by about 1% overall.

[0098] In general, this improved distance-based SMOTE algorithm is not only innovative in theory, but also demonstrates excellent performance in practical applications, providing a more intelligent and efficient solution to the data imbalance problem.

[0099] See also Figure 2 As shown, Figure 2 2 is a schematic diagram of the structure of a sample generation device 200 provided in an embodiment of the present application. The sample generation device 200 is configured to execute the steps of the sample generation method described in each of the above embodiments. The sample generation device 200 can be a single server or a server cluster, or the sample generation device 200 can be a terminal, such as a handheld terminal, a laptop computer, a wearable device, or a robot.

[0100] like Figure 2 As shown, the sample generating device 200 includes: The data acquisition unit 201 is used to acquire an original data set to be generated, separate the original data set into a minority class sample set and a majority class sample set, and perform normalization on the minority class sample set and the majority class sample set to obtain normalized minority class samples and majority class samples; A clustering calculation unit 202 is configured to determine an optimal number of clusters that maximizes the silhouette coefficient for the minority class samples by calculating the silhouette coefficient for each number of clusters, and cluster the minority class samples based on the optimal number of clusters to obtain a plurality of minority class sample clusters; a distance acquisition unit 203 configured to acquire the nearest neighbor distance between the minority class sample corresponding to the minority class sample cluster and the majority class sample, calculate a dynamic weight based on the nearest neighbor distance, and generate a plurality of candidate samples based on the dynamic weight; The sample generation unit 204 is configured to obtain a first average nearest neighbor distance from each candidate sample to the minority class sample and a second average nearest neighbor distance from each candidate sample to the majority class sample, and to use the candidate sample whose first average nearest neighbor distance is smaller than the second average nearest neighbor distance as a target sample, and to use the target sample as a newly generated minority class sample.

[0101] In some embodiments, separating the original data set into a minority class sample set and a majority class sample set includes: counting the number of samples corresponding to each class label according to the class labels of the samples in the original data set, and dividing the samples corresponding to the class whose sample number ratio is less than a preset ratio threshold into the minority class sample set; and dividing the samples corresponding to the class whose sample number ratio is greater than or equal to the preset ratio threshold into the majority class sample set.

[0102] In some embodiments, the standardization processing of the minority class sample set and the majority class sample set includes: for each feature dimension, calculating the feature mean and feature standard deviation of the overall data set after the minority class sample set and the majority class sample set are merged; based on the feature mean and feature standard deviation, standardizing each sample in the minority class sample set and the majority class sample set so that the mean corresponding to each feature dimension after standardization is 0 and the standard deviation is 1, thereby obtaining the standardized minority class samples and majority class samples.

[0103] In some embodiments, clustering the minority class samples based on the optimal number of clusters to obtain multiple minority class sample clusters includes: inputting the optimal number of clusters and the minority class samples into a preset Gaussian mixture model, the Gaussian mixture model uses the optimal number of clusters as the number of components, performing probability density fitting and clustering on the minority class samples, so that each of the minority class samples is assigned to a corresponding sample cluster, and obtaining multiple minority class sample clusters.

[0104] In some embodiments, obtaining the nearest neighbor distance from the minority class sample corresponding to the minority class sample cluster to the majority class sample includes: constructing a KD tree for the minority class sample and the majority class sample respectively, and quickly searching the nearest neighbor distance from each minority class sample to the majority class sample through the KD tree.

[0105] In some embodiments, before obtaining the first average nearest neighbor distance of each candidate sample to the minority class sample and the second average nearest neighbor distance to the majority class sample, the method further includes: calculating the mean and maximum sample distance of the minority class samples corresponding to each minority class sample cluster; using the mean as the center of an envelope circle and the maximum sample distance as the radius to construct an envelope circle; filtering out candidate samples that are not within the envelope circle and retaining the candidate samples within the envelope circle.

[0106] In some embodiments, after the target sample is used as a newly generated minority class sample, the method further includes: denormalizing the target sample from the normalized space back to the original data space, and merging the target sample with the original data set to form a balanced new data set.

[0107] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the sample generation device and each module described above can refer to the corresponding processes in the sample generation method embodiments described in the above embodiments, and will not be repeated here.

[0108] The above-mentioned sample generation method can be implemented in the form of a computer program. The computer program can be used in Figure 2 Run on the device shown.

[0109] See also Figure 3 , Figure 3 1 is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. The computer device includes a processor, a memory, and a network interface connected via a device bus, wherein the memory may include a storage medium and an internal memory.

[0110] The storage medium can store an operating device and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any sample generation method.

[0111] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.

[0112] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any sample generation method.

[0113] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 3The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the terminal to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0114] It should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0115] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps: Obtaining an original data set to be generated, separating the original data set into a minority class sample set and a majority class sample set, and performing standardization on the minority class sample set and the majority class sample set to obtain standardized minority class samples and majority class samples; For the minority class samples, determining an optimal number of clusters that maximizes the silhouette coefficient by calculating the silhouette coefficient under each number of clusters, and clustering the minority class samples based on the optimal number of clusters to obtain a plurality of minority class sample clusters; Obtaining the nearest neighbor distance between the minority class sample corresponding to the minority class sample cluster and the majority class sample, calculating a dynamic weight according to the nearest neighbor distance, and generating a plurality of candidate samples according to the dynamic weight; Obtain a first average nearest neighbor distance from each candidate sample to the minority class sample and a second average nearest neighbor distance from each candidate sample to the majority class sample, and use the candidate sample whose first average nearest neighbor distance is smaller than the second average nearest neighbor distance as a target sample, and use the target sample as a newly generated minority class sample.

[0116] In some embodiments, separating the original data set into a minority class sample set and a majority class sample set includes: counting the number of samples corresponding to each class label according to the class labels of the samples in the original data set, and dividing the samples corresponding to the class whose sample number ratio is less than a preset ratio threshold into the minority class sample set; and dividing the samples corresponding to the class whose sample number ratio is greater than or equal to the preset ratio threshold into the majority class sample set.

[0117] In some embodiments, the standardization processing of the minority class sample set and the majority class sample set includes: for each feature dimension, calculating the feature mean and feature standard deviation of the overall data set after the minority class sample set and the majority class sample set are merged; based on the feature mean and feature standard deviation, standardizing each sample in the minority class sample set and the majority class sample set so that the mean corresponding to each feature dimension after standardization is 0 and the standard deviation is 1, thereby obtaining the standardized minority class samples and majority class samples.

[0118] In some embodiments, clustering the minority class samples based on the optimal number of clusters to obtain multiple minority class sample clusters includes: inputting the optimal number of clusters and the minority class samples into a preset Gaussian mixture model, the Gaussian mixture model uses the optimal number of clusters as the number of components, performing probability density fitting and clustering on the minority class samples, so that each of the minority class samples is assigned to a corresponding sample cluster, and obtaining multiple minority class sample clusters.

[0119] In some embodiments, obtaining the nearest neighbor distance from the minority class sample corresponding to the minority class sample cluster to the majority class sample includes: constructing a KD tree for the minority class sample and the majority class sample respectively, and quickly searching the nearest neighbor distance from each minority class sample to the majority class sample through the KD tree.

[0120] In some embodiments, before obtaining the first average nearest neighbor distance of each candidate sample to the minority class sample and the second average nearest neighbor distance to the majority class sample, the method further includes: calculating the mean and maximum sample distance of the minority class samples corresponding to each minority class sample cluster; using the mean as the center of an envelope circle and the maximum sample distance as the radius to construct an envelope circle; filtering out candidate samples that are not within the envelope circle and retaining the candidate samples within the envelope circle.

[0121] In some embodiments, after the target sample is used as a newly generated minority class sample, the method further includes: denormalizing the target sample from the normalized space back to the original data space, and merging the target sample with the original data set to form a balanced new data set.

[0122] The present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor implements the steps of the sample generation method described in the first aspect above.

[0123] The computer-readable storage medium may be an internal storage unit of the computer device described in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc., equipped on the computer device.

[0124] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A sample generation method, characterized in that: include: Obtaining an original data set to be generated, separating the original data set into a minority class sample set and a majority class sample set, and performing standardization on the minority class sample set and the majority class sample set to obtain standardized minority class samples and majority class samples; For the minority class samples, determining an optimal number of clusters that maximizes the silhouette coefficient by calculating the silhouette coefficient under each number of clusters, and clustering the minority class samples based on the optimal number of clusters to obtain a plurality of minority class sample clusters; Obtaining the nearest neighbor distance between the minority class sample corresponding to the minority class sample cluster and the majority class sample, calculating a dynamic weight according to the nearest neighbor distance, and generating a plurality of candidate samples according to the dynamic weight; Obtain a first average nearest neighbor distance from each candidate sample to the minority class sample and a second average nearest neighbor distance from each candidate sample to the majority class sample, and use the candidate sample whose first average nearest neighbor distance is smaller than the second average nearest neighbor distance as a target sample, and use the target sample as a newly generated minority class sample.

2. The method according to claim 1, characterized in that The separating the original data set into a minority class sample set and a majority class sample set comprises: According to the category labels of the samples in the original data set, the number of samples corresponding to each category label is counted, and the samples corresponding to the category whose sample number ratio is less than a preset ratio threshold are divided into the minority class sample set; The samples corresponding to the categories whose sample quantity ratio is greater than or equal to the preset ratio threshold are divided into the majority class sample set.

3. The method according to claim 1, characterized in that The standardization processing of the minority class sample set and the majority class sample set includes: For each feature dimension, calculate the feature mean and feature standard deviation of the overall data set after the minority class sample set and the majority class sample set are merged; Based on the feature mean and feature standard deviation, each sample in the minority class sample set and the majority class sample set is standardized so that the mean corresponding to each feature dimension after standardization is 0 and the standard deviation is 1, thereby obtaining the standardized minority class samples and majority class samples.

4. The method according to claim 1, wherein The clustering of the minority class samples based on the optimal number of clusters to obtain a plurality of minority class sample clusters includes: The optimal number of clusters and the minority class samples are input into a preset Gaussian mixture model. The Gaussian mixture model uses the optimal number of clusters as the number of components, performs probability density fitting and clustering on the minority class samples, and assigns each minority class sample to a corresponding sample cluster to obtain a plurality of minority class sample clusters.

5. The method according to claim 1, wherein The obtaining of the nearest neighbor distance between the minority class sample corresponding to the minority class sample cluster and the majority class sample includes: A KD tree is constructed for the minority class samples and the majority class samples respectively, and the nearest neighbor distance from each minority class sample to the majority class sample is quickly searched through the KD tree.

6. The method according to claim 1, wherein Before respectively obtaining the first average nearest neighbor distance from each candidate sample to the minority class sample and the second average nearest neighbor distance from each candidate sample to the majority class sample, the method further includes: Calculating the mean and maximum sample distance of the minority class samples corresponding to each of the minority class sample clusters; Taking the mean as the center of the envelope circle and the maximum sample distance as the radius to construct the envelope circle; The candidate samples that are not within the envelope circle are filtered out, and the candidate samples that are within the envelope circle are retained.

7. The method according to claim 1, characterized in that After taking the target sample as a newly generated minority class sample, the method further includes: The target samples are denormalized from the normalized space back to the original data space and merged with the original data set to form a balanced new data set.

8. A sample generating device, characterized in that: include: a data acquisition unit, configured to acquire an original data set to be generated, separate the original data set into a minority class sample set and a majority class sample set, and perform standardization on the minority class sample set and the majority class sample set to obtain standardized minority class samples and majority class samples; a clustering calculation unit, configured to determine, for the minority class samples, an optimal number of clusters that maximizes the silhouette coefficient by calculating the silhouette coefficient under each number of clusters, and cluster the minority class samples based on the optimal number of clusters to obtain a plurality of minority class sample clusters; a distance acquisition unit, configured to acquire the nearest neighbor distance between the minority class sample corresponding to the minority class sample cluster and the majority class sample, calculate a dynamic weight according to the nearest neighbor distance, and generate a plurality of candidate samples according to the dynamic weight; The sample generation unit is used to respectively obtain a first average nearest neighbor distance from each candidate sample to the minority class sample and a second average nearest neighbor distance from each candidate sample to the majority class sample, and use the candidate sample whose first average nearest neighbor distance is smaller than the second average nearest neighbor distance as a target sample, and use the target sample as a newly generated minority class sample.

9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • High-salt-tolerance rice breeding data management method and system based on clustering processing

    CN121030284A

  • Endoscope image data enhancement method and system based on adversarial network

    CN121304471A