Sample data processing method and computer program product

By constructing and adjusting the sample data set, the problem of sample data imbalance was solved, and more accurate model training and prediction effects were achieved.

CN120670840APending Publication Date: 2025-09-19AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510687767.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In existing technologies, the imbalance of sample data leads to undesirable model training results. In particular, in binary classification problems, the number of samples in some categories is far less than that in other categories, resulting in inaccurate model predictions for minority categories.

Method used

Construct positive sample data sets and negative sample data sets, sample the majority class data in a targeted manner through clustering and sampling rate adjustment, and construct the target sample data set to enhance data balance.

Benefits of technology

By building a balanced sample data set, the accuracy and generalization ability of model training can be improved, the distribution of sample data can be optimized, and the quality of sample data can be enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670840A_ABST
    Figure CN120670840A_ABST
Patent Text Reader

Abstract

The invention discloses a sample data processing method and a computer program product. The method comprises the following steps: constructing a positive sample data set and a negative sample data set; determining a majority class sample set and a minority class sample set in the positive sample data set and the negative sample data set according to the sample number, and determining the total sampling rate of the majority class sample set according to the sample number corresponding to the majority class sample set and the minority class sample set; clustering a plurality of initial sample data in the majority class sample set to obtain a plurality of majority class sample clusters, and determining the target sampling rate of each determined majority class sample cluster according to the distance between the majority class sample cluster and the minority class sample set and the total sampling rate; sampling the initial sample data in the majority class sample clusters according to the target sampling rates corresponding to the majority class sample clusters to obtain multiple pieces of sampling sample data, and constructing a target sample data set according to the multiple pieces of sampling sample data and the minority class sample set. The data balance of the sample data is realized, and the quality of the sample data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a sample data processing method and a computer program product. Background Art

[0002] Sample data is the core foundation of the model training process. Its role runs through the entire model development process and directly affects the performance and generalization ability of the model.

[0003] In related technologies, sample data is usually generated based on limited data, and the sample data often suffers from problems such as data type imbalance. Data imbalance refers to the situation in which the number of sample data of different categories in a dataset varies greatly, resulting in the number of samples of some categories being far smaller than the number of samples of other categories. Generally speaking, data imbalance refers to a binary classification problem in which the number of samples of one category is far smaller than the number of samples of the other category. In this case, the model may be more inclined to predict the types with more samples because they appear to be more common. This also causes the model to have inaccurate predictions for types with fewer samples. In other words, it is difficult to obtain a data model that meets the requirements using existing sample data. Summary of the Invention

[0004] The present invention provides a sample data processing method and a computer program product to solve the problems of insufficient and unbalanced sample data in related technologies.

[0005] According to one aspect of the present invention, a sample data processing method is provided, the method comprising:

[0006] Constructing a positive sample data set and a negative sample data set; wherein the number of samples of the initial sample data contained in the positive sample data set and the negative sample data set is different;

[0007] Determine the majority class sample set and the minority class sample set in the positive sample data set and the negative sample data set according to the number of samples, and determine the total sampling rate of the majority class sample set according to the number of samples corresponding to the majority class sample set and the number of samples corresponding to the minority class sample set;

[0008] Clustering the plurality of initial sample data contained in the majority class sample set to obtain a plurality of majority class sample clusters, and determining a target sampling rate for each majority class sample cluster according to the distance between each majority class sample cluster and the minority class sample set and the total sampling rate;

[0009] The initial sample data in the majority class sample cluster is sampled according to the target sampling rate corresponding to each majority class sample cluster to obtain multiple sampled sample data, and a target sample data set is constructed based on the multiple sampled sample data and the initial sample data contained in the minority class sample set.

[0010] According to another aspect of the present invention, there is provided a sample data processing apparatus, the apparatus comprising:

[0011] A sample data set construction module is used to construct a positive sample data set and a negative sample data set; wherein the number of samples of the initial sample data contained in the positive sample data set and the negative sample data set is different;

[0012] a total sampling rate determination module, configured to determine the majority class sample set and the minority class sample set in the positive sample data set and the negative sample data set according to the number of samples, and determine the total sampling rate of the majority class sample set according to the number of samples corresponding to the majority class sample set and the number of samples corresponding to the minority class sample set;

[0013] a target sampling rate determination module, configured to cluster the plurality of initial sample data contained in the majority class sample set to obtain a plurality of majority class sample clusters, and determine a target sampling rate for each majority class sample cluster based on the distance between each majority class sample cluster and the minority class sample set and the total sampling rate;

[0014] A target sample data set construction module is used to sample the initial sample data in the majority class sample cluster according to the target sampling rate corresponding to each majority class sample cluster, obtain multiple sampled sample data, and construct a target sample data set based on the multiple sampled sample data and the initial sample data contained in the minority class sample set.

[0015] According to another aspect of the present invention, an electronic device is provided, comprising:

[0016] at least one processor; and

[0017] a memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform a sample data processing method according to any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement a sample data processing method according to any embodiment of the present invention when executed.

[0020] According to another aspect of the present invention, an embodiment of the present disclosure further provides a computer program product, including a computer program, which implements a sample data processing method as described in any one of the embodiments of the present disclosure when executed by a processor.

[0021] The technical solution of the embodiment of the present invention is, first, by constructing a positive sample data set and a negative sample data set. Since the sample numbers of the initial sample data contained in the positive sample data set and the negative sample data set are usually different, that is, the distribution of each type of sample data is uneven, this sample data set construction method is more in line with the actual scenario, reducing the difficulty of constructing the sample data set. However, when the unbalanced sample data is used for model training, it is often easy to cause the model training results to not meet expectations. Then, by determining the majority class sample set and the minority class sample set in the positive sample data set and the negative sample data set according to the sample number, and determining the total sampling rate of the majority class sample set according to the sample number corresponding to the majority class sample set and the sample number corresponding to the minority class sample set, it is possible to distinguish the majority class sample and minority class sample data corresponding to the positive and negative sample data sets, and the total sampling rate of the majority class samples can be determined quickly and easily. Then, by clustering the multiple initial sample data contained in the majority class sample set, multiple majority class sample clusters are obtained. According to the distance between each majority class sample cluster and the minority class sample set and the total sampling rate, the target sampling rate of each majority class sample cluster is determined, thereby achieving reclassification of the majority class sample data set and specifically determining the sampling rates of different majority class sample clusters to obtain a more reasonable number of sample collections. Finally, by sampling the initial sample data in the majority class sample cluster according to the target sampling rate corresponding to each majority class sample cluster, multiple sampled sample data are obtained. The target sample data set is constructed based on the multiple sampled sample data and the initial sample data contained in the minority class sample set. By differentially sampling different majority class sample clusters, the collected sample data can cover more scenarios, further optimize the distribution of sample data, and obtain sample data with a more balanced data type distribution. Furthermore, by combining with the initial sample data in the minority class sample set, comprehensive and balanced sample data is finally obtained, thereby effectively enhancing the quality of the sample data.

[0022] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1A is a flowchart of a sample data processing method provided in accordance with the first embodiment of the present invention;

[0025] Figure 1B is a schematic diagram of a process for constructing a target sample data set that can be used in a sample data processing method according to an embodiment of the present invention, provided in accordance with the first embodiment of the present invention;

[0026] Figure 2 is a flow chart of a sample data processing method provided in accordance with the second embodiment of the present invention;

[0027] Figure 3A is a flow chart of a sample data processing method provided according to the third embodiment of the present invention;

[0028] Figure 3B 1 is a schematic diagram of a generative adversarial network training process that can be used in a sample data processing method according to a third embodiment of the present invention;

[0029] Figure 4 This is a schematic structural diagram of a sample data processing device provided according to a fourth embodiment of the present invention;

[0030] Figure 5 It is a structural diagram of an electronic device for implementing a sample data processing method provided in the fifth embodiment of the present invention. DETAILED DESCRIPTION

[0031] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "initial", "target", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0033] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0034] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0035] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0036] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0037] Example 1

[0038] Figure 1A This is a flow chart of a sample data processing method provided in the first embodiment of the present invention. This embodiment is applicable to the scenario of generating training sample data. The method can be executed by a sample data processing device. The sample data processing device can be implemented in the form of hardware and / or software. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, PC or server. Figure 1A As shown, the method may specifically include:

[0039] S110 , constructing a positive sample dataset and a negative sample dataset; wherein the number of samples of initial sample data included in the positive sample dataset and the negative sample dataset is different.

[0040] The number of samples in the initial sample data contained in the positive sample data set and the negative sample data set is different, indicating that there is an imbalance problem in the initial sample data.

[0041] In the embodiments of the present invention, the positive sample dataset and the negative sample dataset can be relative terms, used to distinguish datasets corresponding to sample data of different categories. In other words, the positive sample dataset and the negative sample dataset can be understood as sample datasets corresponding to different types in a binary classification scenario. For example, the positive sample dataset refers to a dataset consisting of samples of the "target category" or "category of interest"; the negative sample dataset generally refers to a dataset consisting of samples that do not belong to the "target category" or "category of interest" in a classification task.

[0042] As an optional implementation scheme of the embodiments of the present disclosure, the initial sample data can be real sample data obtained through a data acquisition method. Real sample data can be understood as sample data generated in various practical application scenarios. The acquired real sample data may include positive sample data and / or negative sample data. In some embodiments, positive sample data can be obtained by processing negative sample data. Similarly, positive and negative sample data can be obtained by processing positive sample data.

[0043] As another optional implementation scheme of the embodiment of the present disclosure, the initial sample data includes multiple real sample data and multiple simulated sample data (synthetic data or fake data) generated by the sample generation model. The sample generation model can be a pre-trained machine learning model for generating simulated sample data. Specifically, the technical solution of the present invention can enhance the quality of the simulated sample data and enrich the quantity of sample data by constructing positive sample data sets and negative sample data sets using real data and model synthetic data, thereby constructing a positive sample data set and a negative sample data set with sufficient data volume for use in subsequent data processing processes.

[0044] Alternatively, the sample generation model can be obtained by training the generator using a generative adversarial network. The core concept of a generative adversarial network is the zero-sum game in game theory, which frames the generation problem as a competition between two network models: a discriminator and a generator. The generator attempts to generate synthetic data similar to real data, while the discriminator attempts to distinguish between generated data and real data. The former strives to generate data closer to real data, while the latter, in turn, strives to more perfectly distinguish between real and generated data. In this way, the two networks improve through competition, and as they continue to compete, the generator network eventually outputs increasingly better data, approaching real data, thereby generating the desired data.

[0045] Taking user credit assessment as an example, data imbalance is a common problem in credit assessment. Typically, the number of samples with poor credit is far smaller than the number of samples with good credit, meaning that there are relatively few examples of credit default, which is consistent with real-world applications. For example, in a dataset of overdue credit card payments, if the number of overdue cases only accounts for 1% of the total cases, then the dataset is considered imbalanced. Data processing methods for related data often suffer from data limitations and difficulties handling nonlinear relationships. Generative adversarial network-based sample generation models offer new opportunities for improving credit assessments, generating synthetic data to enhance assessment accuracy and support anti-fraud analysis. Using sample generation models to generate synthetic user data and combining it with traditional data for credit assessment can effectively improve assessment accuracy, especially in data-limited situations. This effectively addresses the data limitations of traditional credit assessment methods. Furthermore, the data generated by sample generation models based on generative adversarial networks can provide decision support, assisting in credit grant or rejection decisions and implementing appropriate anti-fraud measures to support credit assessments. Generative adversarial networks, as powerful deep learning models, possess strong high-dimensional data modeling capabilities and can better capture the complex relationships and underlying patterns in user data. Compared with traditional statistical models and some machine learning methods, generative adversarial networks perform better in modeling nonlinear relationships and can more accurately capture the complex patterns of users in credit scenarios.

[0046] In one embodiment, the sample generation model optionally utilizes a generative adversarial network to train a generator; the loss function of the discriminator in the generative adversarial network, which is trained adversarially against the generator, includes a modulation factor; the modulation factor is used to control the degree of weight decay for difficult and easy samples. Difficult and easy samples refer to sample data that is difficult to distinguish between real and simulated samples. By adding the modulation factor to the discriminator, the discriminator's loss output can be adjusted, allowing the discriminator to pay more attention to difficult samples and improve the quality of the generated data.

[0047] Specifically, during the training process of a generative adversarial network, the convergence condition of the discriminator can be adjusted by including a modulation factor in the discriminator's loss function, thereby affecting the overall training process of the generative adversarial network. The additional modulation factor in the technical solution of the present invention can achieve better training results for the generative adversarial network, thereby outputting better simulated sample data and improving the data quality of the generative adversarial network output.

[0048] In one embodiment, optionally, the loss function of the discriminator is:

[0049]

[0050] in, is the discriminant loss corresponding to the discriminator, x i is the i-th real sample data; D(x i ) is the discrimination result of the discriminator on the i-th real sample data; The i-th simulated sample data generated by the sample generation model; is the discrimination result of the discriminator on the i-th simulated sample data; N is the total number of the real sample data, and , the total number of the simulated sample data.

[0051] As an optional example of the embodiment of the present invention, the sample generation model can be trained in the following way:

[0052] (1) Initialize the network parameters that need to be optimized for the generator G and the discriminator D, that is, the model learning parameters, which are expressed as θ g and θ d ;

[0053] (2) Assuming that the number of initial sample data in the negative sample data set is N, the set of true negative sample data is {x 1 ,x 2 ,...,x N}; At the same time, randomly generate N noise samples z, then the generated negative sample set is {z 1 ,z 2 ,...,z N}. For the t-th data sample pair (x t ,z t ), the corresponding output y t The values ​​are 1 and 0, indicating that the probability of judging the real sample as true is 1, and the probability of judging the generated sample as true is 0.

[0054] (3) The training process of the two network models follows the principle of fixing the other while training one network to ensure that the two models reach a Nash equilibrium. The number of training times δ is set. For each round of training, the two networks are learned as follows:

[0055] Discriminator training: Fix the generator G and train the discriminator D. First, pass the noise sample through the generator G to obtain the generated sample (simulated sample data) as follows Right now The learning objective function of D is as follows:

[0056]

[0057] The randomly generated data and real data in step (2) are used as the input of the discriminator model. In the above formula, -logD(x) represents the conversion of x i The smaller the uncertainty of the true data is, the better. Ideally, the best state is D(x) = 1. Indicates that It is judged that the smaller the uncertainty of the generated data, the better. The greater the probability of judging it as false, the better. i )) γ and The γ in the formula is the modulation factor introduced, which aims to control the degree to which different samples reduce the loss function (cross entropy), making the cross entropy value as small as possible, that is, the smaller the distance between the predicted distribution and the true distribution, the better. This allows the model to pay more attention to difficult-to-separate samples and iterate as quickly as possible for a large number of easy-to-separate samples. In this way, the reduction of the model's loss function relies more on those difficult-to-separate sample data.

[0058] Furthermore, the discriminator D can be updated using gradient descent d , as shown below:

[0059]

[0060] (3.2) Fix the discriminator D and train the generator G; the training objective function of G is as follows:

[0061]

[0062] Update the parameters θ of G using gradient descent g , as shown below:

[0063]

[0064] The training requirement of this generator is to make the generated data consistent with the real data distribution as much as possible. If log(D(G(z)))=0, that is, D(G(z))=1, the generated data is judged to be natural (real) data.

[0065] (4) The convergence condition is reached and the process ends. The parameters of the two network models are alternately optimized by the gradient descent method and the back propagation network is used for parameter optimization. Finally, the two models converge, and the probability distribution of the real data is The distance between the probability distribution P(x) and the generated data As close as possible.

[0066] S120. Determine the majority class sample set and the minority class sample set in the positive sample data set and the negative sample data set according to the number of samples, and determine the total sampling rate of the majority class sample set according to the number of samples corresponding to the majority class sample set and the number of samples corresponding to the minority class sample set.

[0067] The sample size refers to the number of samples in the sample dataset. The majority class sample set refers to the data set corresponding to the class with a larger number of samples in a class-imbalanced dataset. Conversely, the minority class sample set refers to the data set corresponding to the class with a smaller number of samples in a class-imbalanced dataset.

[0068] Specifically, the technical solution of the present invention determines the sample data set with a larger number of positive sample data sets and negative sample data sets as the majority class sample set and determines the other sample data set as the minority class sample set based on the number of sample data contained in the positive sample data set and the negative sample data set.

[0069] The total sampling rate of the majority class sample set can be used to indicate the total sample ratio required to be sampled in the majority class sample set, and can be calculated based on the number of samples corresponding to the majority class sample set and the number of samples corresponding to the minority class sample set.

[0070] Optionally, in one embodiment, the ratio of the number of samples corresponding to the minority sample set to the number of samples corresponding to the majority class sample set may be directly determined as the total sampling rate of the majority class sample set.

[0071] In another embodiment, a ratio of the number of samples corresponding to the minority sample set to the number of samples corresponding to the majority class sample set may be determined, and the total sampling rate of the majority class sample set may be determined based on the ratio and a preset ratio adjustment parameter. The preset ratio adjustment parameter may include a preset ratio adjustment ratio (multiplied by 110%, etc.) or a ratio adjustment amount (e.g., increased by 1%).

[0072] For example, the total sampling rate of the majority class sample set is calculated as follows:

[0073]

[0074] Among them, R represents the total sampling rate of the majority class sample set, Num minority Indicates the number of samples in the minority class sample set, Num majoruty Indicates the number of samples in the majority class sample set.

[0075] S130. Cluster the multiple initial sample data contained in the majority class sample set to obtain multiple majority class sample clusters, and determine the target sampling rate of each majority class sample cluster according to the distance between each majority class sample cluster and the minority class sample set and the total sampling rate.

[0076] The majority class sample cluster is the result of clustering similar or related samples in the initial sample data into the same cluster using a clustering algorithm. The distance between the majority class sample cluster and the minority class sample set can be understood as the similarity between the majority class cluster and the minority class sample set. The target sampling rate refers to the sampling rate when sampling the initial sample data in a single majority class sample cluster.

[0077] In an embodiment of the present invention, the sampling rate corresponding to the majority class sample cluster can be positively correlated with the distance between the majority class sample cluster and the minority class sample set. Due to the need for diversified feature data, the greater the difference between the majority class sample cluster and the minority class sample set, the higher the sampling rate can be, that is, the smaller the similarity or the greater the distance, the higher the sampling rate, and vice versa. Furthermore, the majority class sample clusters can be set to 5, and the distances of different sample clusters can be sorted. According to the distance arrangement corresponding to different sample clusters, different sampling rates are set for different majority class sample clusters from large to small. Sample data is collected for different majority class sample clusters according to the distance arrangement, and the sample data finally sampled meets the requirements.

[0078] Optionally, a plurality of initial sample data in the majority class sample set is subjected to a K-means algorithm to obtain a plurality of majority class sample clusters after classification.

[0079] Specifically, the technical solution of the present invention can cluster the initial sample data based on the initial sample data in the majority class sample set to obtain multiple majority class sample clusters. Then, the distance between each majority class sample cluster and the minority class sample set is determined, and the weight is driven by the distance. The closer the distance between the majority class sample cluster and the minority class, the higher its sampling rate should be (because clusters with close distances are more likely to overlap with the minority class, and more downsampling is required to balance the distribution). Then, combined with the total sampling rate constraint, that is, the sampling rate of all clusters and the total sampling rate that must meet the preset. According to the distance and the total sampling rate, the relative weight of each majority class sample cluster is calculated, and then the sampling rate is allocated, thereby realizing the dynamic allocation of the sampling rates of multiple majority class sample clusters. The technical solution of the present invention can obtain corresponding and reasonable sampling of multiple majority class initial sample data clusters with different characteristics in the above manner, so that the distribution of the majority class initial sample data finally obtained is wider and more appropriate, effectively ensuring the data balance between the majority class initial sample data and the minority class initial sample data.

[0080] In one embodiment, determining the target sampling rate of each majority class sample cluster based on the distance between each majority class sample cluster and the minority class sample set and the total sampling rate includes: determining the distance between each majority class sample cluster and the minority class sample set, and determining the sum of the distances corresponding to multiple majority class sample clusters; determining the initial sampling rate based on the distance corresponding to each majority class sample cluster and the sum of the distances, and determining the target sampling rate based on the initial sampling rate and the total sampling rate.

[0081] As an optional technical solution for the embodiments of the present disclosure, the Euclidean distance can be used to calculate the distance between each majority class sample cluster and the minority class sample set. Specifically, the center of the majority class sample cluster can be determined based on multiple initial sample data in the majority class sample cluster; and the center of the minority class sample set can be determined based on multiple initial sample data in the minority class sample set. Then, the Euclidean distance between the center of the majority class sample cluster and the center of the minority class sample set can be calculated to obtain the distance between the majority class sample cluster and the minority class sample set.

[0082] The initial sampling rate refers to the sampling rate of each majority class sample cluster. Generally, the initial sampling rate is determined as the ratio of the distance between the majority class sample cluster and the minority class sample set to the sum of the distances corresponding to multiple majority class sample clusters.

[0083] Specifically, the distance between each majority class sample cluster and the minority class sample set is calculated, and the sum of the distances corresponding to multiple majority class sample clusters is calculated, so as to determine the initial sampling rate according to the corresponding relationship between the distance of a single majority class sample cluster and the sum of the distances.

[0084] Specifically, the initial sampling rate can be calculated by the following formula:

[0085]

[0086] Among them, R n represents the initial sampling rate corresponding to the nth majority class sample cluster; N is the total number of majority class sample clusters; Represents the distance between the cluster center of the nth majority class sample cluster and the minority class sample.

[0087] In one embodiment, determining the target sampling rate according to the initial sampling rate and the total sampling rate includes: multiplying the initial sampling rate and the total sampling rate to obtain the target sampling rate.

[0088] Specifically, the target sampling rate corresponding to the majority class sample cluster can be calculated by the following formula:

[0089] R c =R×R n ,

[0090] Among them, R c represents the target sampling rate corresponding to the nth majority class sample cluster, R n represents the initial sampling rate corresponding to the nth majority class sample cluster; R represents the total sampling rate corresponding to the majority class sample set.

[0091] By combining the total sampling rate with the target sampling rate, the present invention can deeply integrate the quantitative relationship between the majority class sample set and the minority class sample set, obtaining a more accurate target sampling rate for each majority class sample cluster, thereby achieving reasonable sampling of multiple majority class sample clusters. By calculating the initial sampling rate and combining it with the total sampling rate, a more appropriate actual sampling rate corresponding to the majority class sample cluster can be calculated, helping to reduce the imbalance between the majority class sample data and the minority class sample data.

[0092] S140. Sample the initial sample data in the majority class sample cluster according to the target sampling rate corresponding to each majority class sample cluster to obtain multiple sampled sample data, and construct a target sample data set based on the multiple sampled sample data and the initial sample data included in the minority class sample set.

[0093] The sampled sample data refers to the initial sample data set obtained by sampling multiple majority class sample clusters according to the target sampling rate. The target sample data set is the sample data set obtained by combining the sampled sample data with the initial sample data from the minority class sample set. The target sample data set has a more balanced distribution of majority class sample data and minority class sample data.

[0094] Specifically, the initial sample data in the majority class sample cluster is randomly undersampled according to the target sampling rate corresponding to each majority class sample cluster. The total number of majority class samples sampled is:

[0095] Num majority =Clustre1×R C1 +Clustre2×R C2 +...Clustre n ×R CN ,

[0096] Among them, Num majority Represents the total number of initial sample data sampled from multiple majority class sample clusters, Clustre1 represents the total number of initial sample data in the first majority class sample cluster; R C1 represents the total number of initial sample data in the first majority class sample cluster; Clustre2 represents the total number of initial sample data in the second majority class sample cluster; R C2Represents the total number of initial sample data in the second majority class sample cluster; Clustre N Represents the total number of initial sample data in the Nth majority class sample cluster; R CN Represents the total number of initial sample data in the Nth majority class sample cluster.

[0097] In the embodiment of the present invention, the number of sampled samples corresponding to the majority class sample cluster can be determined based on the total number of initial sample data in the majority class sample cluster and the target sampling rate corresponding to the majority class sample cluster. Then, by randomly sampling the initial sample data in the majority class sample cluster, the same number of initial sample data as the number of sampled samples is obtained, that is, the sampled sample data is obtained. This obtains the majority class sample dataset D majority After completing undersampling, a randomly selected set of balanced data is constructed, namely the target sample data set, which can be:

[0098] D=D majority +D minority ,

[0099] Where D represents the target sample dataset; D majority represents a collection of multiple sample data samples sampled from multiple majority class sample clusters; D minority represents the minority class sample set.

[0100] Specifically, the process of using majority class samples to construct the target sample dataset can also refer to Figure 1B FIG. 1 is a schematic diagram of a process of constructing a target sample data set that can be used in a sample data processing method according to an embodiment of the present invention.

[0101] In one embodiment, after constructing the target sample dataset based on the plurality of sample data and the initial sample data included in the minority class sample set, the method further includes: performing feature extraction on each initial sample data in the target sample dataset to obtain sample feature data, and constructing a training dataset based on the sample feature data. That is, after obtaining a balanced sample dataset, feature engineering techniques are then used to filter out redundant features of the balanced sample data to obtain a high-quality data sample set.

[0102] Specifically, the technical solution of the present invention obtains initial sample data from multiple majority class sample clusters after sampling by sampling different majority class sample clusters according to the target sampling rate. By combining and arranging the sampled sample data with the initial sample data contained in the minority class sample set, a target sample data set can be obtained. In this way, sufficient sample data with balanced data features can be obtained, significantly improving the quality of the sample data. Furthermore, after obtaining the target sample data set, redundant features of the target sample data set can be processed to obtain a target sample data set of even better quality.

[0103] In one embodiment, after constructing the target sample data set based on the multiple sampling sample data and the initial sample data included in the minority class sample set, it also includes: training a pre-established initial classification model based on the target sample data set to obtain a target classification model.

[0104] The initial classification model refers to a pre-built classification model that needs to be trained using the target sample dataset. The target classification model refers to a classification model trained using the target sample dataset that can meet the preset classification requirements.

[0105] Specifically, the target sample data set obtained by constructing the initial sample data can be used to train the pre-established initial classification model to obtain the target classification model when the convergence condition is reached. The convergence condition can usually be reaching a preset number of training times, or reaching a preset loss function threshold, etc. The target sample data set obtained by the technical solution of the present invention can meet a sufficient number of model training times, and can better balance the influence of different types of data in the model training process, so as to obtain a target classification model with excellent classification effect, and effectively achieve the effect of training a high-quality target model based on high-quality sample data.

[0106] The technical solution of the embodiment of the present invention is, first, by constructing a positive sample data set and a negative sample data set. Since the sample numbers of the initial sample data contained in the positive sample data set and the negative sample data set are usually different, that is, the distribution of each type of sample data is uneven, this sample data set construction method is more in line with the actual scenario, reducing the difficulty of constructing the sample data set. However, when the unbalanced sample data is used for model training, it is often easy to cause the model training results to not meet expectations. Then, by determining the majority class sample set and the minority class sample set in the positive sample data set and the negative sample data set according to the sample number, and determining the total sampling rate of the majority class sample set according to the sample number corresponding to the majority class sample set and the sample number corresponding to the minority class sample set, it is possible to distinguish the majority class sample and minority class sample data corresponding to the positive and negative sample data sets, and the total sampling rate of the majority class samples can be determined quickly and easily. Then, by clustering the multiple initial sample data contained in the majority class sample set, multiple majority class sample clusters are obtained. According to the distance between each majority class sample cluster and the minority class sample set and the total sampling rate, the target sampling rate of each majority class sample cluster is determined, thereby achieving reclassification of the majority class sample data set and specifically determining the sampling rates of different majority class sample clusters to obtain a more reasonable number of sample collections. Finally, by sampling the initial sample data in the majority class sample cluster according to the target sampling rate corresponding to each majority class sample cluster, multiple sampled sample data are obtained. The target sample data set is constructed based on the multiple sampled sample data and the initial sample data contained in the minority class sample set. By differentially sampling different majority class sample clusters, the collected sample data can cover more scenarios, further optimize the distribution of sample data, and obtain sample data with a more balanced data type distribution. Furthermore, by combining with the initial sample data in the minority class sample set, comprehensive and balanced sample data is finally obtained, thereby effectively enhancing the quality of the sample data.

[0107] Example 2

[0108] Figure 2A flowchart of a sample data processing method provided for the second embodiment of the present invention, the solution in this embodiment is a refinement of the technical solution for constructing positive sample data sets and negative sample data sets on the basis of the above embodiments. Optionally, the construction of positive sample data sets and negative sample data sets includes: obtaining an initial sample data set; wherein the initial sample data set includes a plurality of initial sample data; the plurality of initial sample data include a plurality of real sample data and a plurality of simulated sample data generated by a sample generation model; respectively determining the expected category label corresponding to each of the initial sample data, and dividing the plurality of initial sample data in the initial sample data set into positive sample data sets and negative sample data sets according to the expected category label corresponding to the initial sample data. For specific implementation methods, please refer to the description of this embodiment. Among them, the technical features that are the same or similar to those in the above embodiments will not be repeated here. As Figure 2 As shown, the method may specifically include:

[0109] S210, obtaining an initial sample data set; wherein the initial sample data set includes a plurality of initial sample data; the plurality of initial sample data includes a plurality of real sample data and a plurality of simulated sample data generated by a sample generation model.

[0110] Real sample data refers to observed or recorded data that can be directly collected from the real world, while simulated sample data refers to artificial data generated through algorithms, models, or rules.

[0111] Specifically, multiple simulated sample data and real sample data generated by the sample generation model can be obtained, and the obtained multiple real sample data and multiple simulated sample data can be combined to construct an initial sample data set, so that a rich amount of initial sample data can be obtained, avoiding the problem of insufficient data. In addition, the training method of the sample generation model can refer to the statement in Example 1. In one embodiment, optionally, the sample generation model includes an encoder and a decoder; the encoder is used to convert the input noise vector into a latent space vector; the decoder is used to convert the latent space vector into simulated sample data.

[0112] The noise vector is a random, high-dimensional vector that is input into the sample generation model to generate simulated sample data. The latent space vector is a low-dimensional, continuous numerical vector that represents the essential characteristics of high-dimensional data (such as images and text).

[0113] In one embodiment, the obtaining of the initial sample data set includes: inputting multiple noise vectors into the sample generation model to obtain multiple simulated sample data; the discriminator: obtaining multiple real sample data, and constructing the initial sample data set based on the multiple real sample data and the multiple simulated sample data.

[0114] Specifically, by inputting multiple noise vectors into the sample generation model, multiple corresponding simulated sample data can be obtained. Furthermore, by acquiring multiple real sample data and combining the multiple simulated sample data with the multiple real sample data, an initial sample dataset can be obtained. This allows for a simple and balanced sample dataset to be obtained, which is sufficient for subsequent sample data screening and model training.

[0115] S220 , respectively determining an expected category label corresponding to each of the initial sample data, and dividing the plurality of the initial sample data in the initial sample data set into a positive sample data set and a negative sample data set according to the expected category label corresponding to the initial sample data.

[0116] The expected class label is a label used to represent the data type of the initial sample data. More specifically, the expected class label can be understood as the data type corresponding to the initial sample data that the trained model is expected to output. The expected class label can be used to distinguish positive samples from negative samples.

[0117] Specifically, the expected category label corresponding to each initial sample data can be determined, and then the data type corresponding to each initial sample data can be determined based on the expected category label, thereby obtaining the positive sample data set and the negative sample data set. It is easy to imagine that by establishing an expected category label for each initial sample data, the sample data set type corresponding to each initial sample data can be quickly and easily determined, thereby achieving a fast and efficient delineation of the positive sample data set and the negative sample data set.

[0118] S230. Determine the majority class sample set and the minority class sample set in the positive sample data set and the negative sample data set according to the number of samples, and determine the total sampling rate of the majority class sample set according to the number of samples corresponding to the majority class sample set and the number of samples corresponding to the minority class sample set.

[0119] S240 , clustering the multiple initial sample data included in the majority class sample set to obtain multiple majority class sample clusters, and determining the target sampling rate of each majority class sample cluster according to the distance between each majority class sample cluster and the minority class sample set and the total sampling rate.

[0120] S250 , sampling the initial sample data in the majority class sample cluster according to the target sampling rate corresponding to each majority class sample cluster to obtain multiple sampled sample data, and constructing a target sample data set based on the multiple sampled sample data and the initial sample data included in the minority class sample set.

[0121] The technical solution of the present invention obtains an initial sample data set; wherein the initial sample data set includes multiple initial sample data; the multiple initial sample data include multiple real sample data and multiple simulated sample data generated by a sample generation model; and a sufficient amount of sample data can be obtained simply and efficiently to construct the initial sample data set. The expected category label corresponding to each of the initial sample data is determined respectively, and the multiple initial sample data in the initial sample data set are divided into positive sample data sets and negative sample data sets according to the expected category label corresponding to the initial sample data. Thus, according to the preset expected category label, the positive sample data set and the negative sample data set of the initial sample data in the initial sample data set are quickly completed, and the label data is effectively utilized to improve the efficiency of data processing.

[0122] Example 3

[0123] Figure 3A A flow chart of a sample data processing method provided for the third embodiment of the present invention. The solution in this embodiment is a preferred embodiment provided on the basis of the above embodiments. For the specific implementation method, please refer to the description of this embodiment. It should be noted that this embodiment takes the application scenario of training a credit assessment model as an example. Among them, the technical features that are the same or similar to the above embodiments are not repeated here. The method may specifically include: Among them, the technical features that are the same or similar to the above embodiments are not repeated here. Figure 3A As shown, the method may specifically include:

[0124] Step 1: Build a generative adversarial network and initialize the network parameters of the generator and discriminator in the generative adversarial network model. Use noise vector samples and real sample data to train the generative adversarial network model. By combining the adjustment degree parameters, obtain a generative adversarial network model that meets the training end conditions. Use the trained generator as the sample generation model. This step can refer to Figure 3B The figure shows a schematic diagram of a generative adversarial network training process that can be used in a sample data processing method according to an embodiment of the present invention.

[0125] Step 2: Input the obtained multiple noise vectors into the generative sample generation model to obtain multiple sequence data (simulated sample data); obtain multiple real sample data, combine the multiple sequence data and the multiple real sample data to obtain an initial sample data set. According to the expected category label corresponding to each initial sample data, determine the positive sample data set and the negative sample data set. Among them, the real data can be divided in a ratio of 7:3, with 70% of the original real data as the training set and 30% of the original real data as the test set.

[0126] Step 3: Calculate the number of different types of samples contained in the positive sample data set and the negative sample data set, and determine the majority class sample set and minority class sample set containing multiple initial sample data respectively.

[0127] Step 4: Cluster the majority class sample set to obtain n classified clusters (majority class sample clusters). Undersample the majority class (majority class sample set) to calculate the total sampling rate of the majority class, and calculate the sampling rate (initial sampling rate) corresponding to each cluster (majority class sample cluster). Finally, based on the sampling rate (initial sampling rate) corresponding to each cluster (majority class sample cluster) and the total sampling rate of the majority class samples, calculate the sample sampling rate (target sampling rate) of each cluster (majority class sample cluster).

[0128] Step 5. Based on the n clusters (majority class sample clusters) after clustering and the calculated sample sampling rate of each cluster (majority class sample cluster), calculate the number of initial sample data samples for each cluster (majority class sample cluster) and the number of all initial sample data samples in the majority class sample set, so as to obtain the majority class sample data set (sampling sample data) after random undersampling selection.

[0129] Step 6: Combine the majority class sample dataset (sampled sample data) after random undersampling selection with the minority class samples (minority class sample set) to obtain a balanced sample dataset (target sample dataset). Furthermore, feature engineering-related techniques can be used to filter out redundant features in the balanced sample dataset (target sample dataset) to obtain a high-quality sample set.

[0130] Step 7: Input the balanced sample dataset or sample set (target sample dataset) into the binary classification model to be trained (such as a credit assessment model), and continuously optimize the model parameters to obtain a trained binary classification model. Finally, use 30% of the original real data to verify the trained binary classification model.

[0131] The technical solution of the present invention, through the above-mentioned embodiment steps, creatively proposes to combine the modulation factor and the generative adversarial network to construct a generative adversarial network model to solve the problem of biased model learning caused by data imbalance, thereby utilizing the generative adversarial network model to output positive sample data sets and negative sample data sets, and by performing sample data rebalancing processing on the positive sample data sets and the negative sample data sets, a target sample data set is obtained, thereby greatly enhancing the training effect of the binary classification model according to the target sample data set.

[0132] Example 4

[0133] Figure 4This is a schematic diagram of the structure of a sample data processing device provided by the fourth embodiment of the present invention. The device can be implemented by software and / or hardware and can be configured in an electronic device. Figure 4 As shown, the apparatus includes: a sample data set construction module 401 , a total sampling rate determination module 402 , a target sampling rate determination module 403 and a target sample data set construction module 404 . Among them, the sample data set construction module 401 is used to construct a positive sample data set and a negative sample data set; wherein the number of samples of the initial sample data contained in the positive sample data set and the negative sample data set is different; the total sampling rate determination module 402 is used to determine the majority class sample set and the minority class sample set in the positive sample data set and the negative sample data set according to the sample number, and determine the total sampling rate of the majority class sample set according to the sample number corresponding to the majority class sample set and the sample number corresponding to the minority class sample set; the target sampling rate determination module 403 is used to cluster the multiple initial sample data contained in the majority class sample set to obtain multiple majority class sample clusters, and determine the target sampling rate of each majority class sample cluster according to the distance between each majority class sample cluster and the minority class sample set and the total sampling rate; the target sample data set construction module 404 is used to sample the initial sample data in the majority class sample cluster according to the target sampling rate corresponding to each majority class sample cluster to obtain multiple sampled sample data, and construct a target sample data set according to the multiple sampled sample data and the initial sample data contained in the minority class sample set.

[0134] The technical solution of the embodiment of the present invention is as follows: first, a positive sample data set and a negative sample data set are constructed by a sample data set construction module 401. Since the number of samples of the initial sample data contained in the positive sample data set and the negative sample data set is usually different, that is, the distribution of each type of sample data is uneven. This sample data set construction method is more in line with the actual scenario and reduces the difficulty of constructing the sample data set. However, when the unbalanced sample data is used for model training, it is often easy to cause the model training results to be inconsistent with expectations. Then, the majority class sample set and the minority class sample set in the positive sample data set and the negative sample data set are determined according to the number of samples through the total sampling rate determination module 402. The total sampling rate of the majority class sample set is determined according to the number of samples corresponding to the majority class sample set and the number of samples corresponding to the minority class sample set. The majority class sample and minority class sample data corresponding to the positive and negative sample data sets can be distinguished, and the total sampling rate of the majority class sample can be determined quickly and easily. Then, the target sampling rate determination module 403 clusters the multiple initial sample data contained in the majority class sample set to obtain multiple majority class sample clusters. Based on the distance between each majority class sample cluster and the minority class sample set and the total sampling rate, the target sampling rate of each majority class sample cluster is determined, thereby achieving reclassification of the majority class sample data set and specifically determining the sampling rates of different majority class sample clusters to obtain a more reasonable number of sample collections. Finally, the target sample data set construction module 404 samples the initial sample data in the majority class sample cluster according to the target sampling rate corresponding to each majority class sample cluster to obtain multiple sampled sample data. The target sample data set is constructed based on the multiple sampled sample data and the initial sample data contained in the minority class sample set. By differentially sampling different majority class sample clusters, the collected sample data can cover more scenarios, further optimize the distribution of the sample data, and obtain sample data with a more balanced data type distribution. Furthermore, by combining with the initial sample data in the minority class sample set, comprehensive and balanced sample data is finally obtained, thereby effectively enhancing the quality of the sample data.

[0135] Based on the above-mentioned optional technical solutions, the sample data set construction module 401 may optionally further include: an acquisition unit and a division unit. The acquisition unit is configured to acquire an initial sample data set; the initial sample data set includes a plurality of initial sample data; the plurality of initial sample data includes a plurality of real sample data and a plurality of simulated sample data generated by a sample generation model; and the division unit is configured to respectively determine an expected category label corresponding to each initial sample data, and divide the plurality of initial sample data in the initial sample data set into a positive sample data set and a negative sample data set according to the expected category label corresponding to the initial sample data.

[0136] Based on the above-mentioned optional technical solutions, optionally, the sample generation model is obtained by training the generator using a generative adversarial network; the loss function of the discriminator in the generative adversarial network that performs adversarial training with the generator includes a modulation factor; the modulation factor is used to control the degree of weight attenuation of difficult and easy samples.

[0137] Based on the above optional technical solutions, optionally, the loss function of the discriminator is:

[0138]

[0139] in, is the discriminant loss corresponding to the discriminator, x i is the i-th real sample data; D(x i ) is the discrimination result of the discriminator on the i-th real sample data; The i-th simulated sample data generated by the sample generation model; is the discrimination result of the discriminator on the i-th simulated sample data; N is the total number of the real sample data, and , the total number of the simulated sample data.

[0140] Based on the above optional technical solutions, optionally, the sample generation model (anti-generative network includes an encoder and a decoder; the encoder is used to convert the input noise vector into a latent space vector; the decoder is used to convert the latent space vector into simulated sample data.

[0141] Based on the above optional technical solutions, the acquisition unit may further include: a model sample data acquisition subunit and an initial sample data set construction subunit. The model sample data acquisition subunit is configured to input multiple noise vectors into the sample generation model to obtain multiple simulated sample data; and the initial sample data set construction subunit is configured to acquire multiple real sample data and construct an initial sample data set based on the multiple real sample data and the multiple simulated sample data.

[0142] Based on the above-mentioned optional technical solutions, optionally, the target sampling rate determination module 403 includes: a total distance determination unit and a target sampling rate determination subunit. The total distance determination unit is configured to respectively determine the distance between each of the majority class sample clusters and the minority class sample set, and to determine the sum of the distances corresponding to the plurality of majority class sample clusters; the target sampling rate determination subunit is configured to determine an initial sampling rate based on the distance corresponding to each of the majority class sample clusters and the sum of the distances, and to determine the target sampling rate based on the initial sampling rate and the total sampling rate.

[0143] Based on the above optional technical solutions, optionally, the target sampling rate determination subunit may further include a target sampling rate calculation subunit, wherein the target sampling rate calculation subunit is configured to multiply the initial sampling rate and the total sampling rate to obtain the target sampling rate.

[0144] Based on the above optional technical solutions, the sample data processing device may further include a model training module. The model training module is configured to, after constructing a target sample dataset based on the plurality of sample data and the initial sample data included in the minority class sample set, train a pre-established initial classification model based on the target sample dataset to obtain a target classification model.

[0145] The sample data processing apparatus provided in the embodiments of the present invention can execute the sample data processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the sample data processing method. For technical details not fully described in this embodiment, please refer to any sample data processing method described in the embodiments of the present invention.

[0146] Example 5

[0147] Figure 5 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0148] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0149] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0150] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a sample data processing method.

[0151] In some embodiments, a sample data processing method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the sample data processing method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform a sample data processing method in any other suitable manner (e.g., by means of firmware).

[0152] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0153] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0154] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0155] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0156] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0157] A computing system may include a client and a server. The client and server are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0158] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication unit 19, or installed from the storage unit 18, or installed from the ROM 12. When the computer program is executed by the processor 11, the above-mentioned functions defined in the method of the embodiment of the present invention are performed.

[0159] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0160] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A sample data processing method, characterized in that: include: Constructing a positive sample data set and a negative sample data set; wherein the number of samples of the initial sample data contained in the positive sample data set and the negative sample data set is different; Determine the majority class sample set and the minority class sample set in the positive sample data set and the negative sample data set according to the number of samples, and determine the total sampling rate of the majority class sample set according to the number of samples corresponding to the majority class sample set and the number of samples corresponding to the minority class sample set; Clustering the plurality of initial sample data contained in the majority class sample set to obtain a plurality of majority class sample clusters, and determining a target sampling rate for each majority class sample cluster according to the distance between each majority class sample cluster and the minority class sample set and the total sampling rate; The initial sample data in the majority class sample cluster is sampled according to the target sampling rate corresponding to each majority class sample cluster to obtain multiple sampled sample data, and a target sample data set is constructed based on the multiple sampled sample data and the initial sample data contained in the minority class sample set.

2. The sample data processing method according to claim 1, characterized in that: The constructing of the positive sample dataset and the negative sample dataset includes: Acquire an initial sample data set; wherein the initial sample data set includes a plurality of initial sample data; the plurality of initial sample data includes a plurality of real sample data and a plurality of simulated sample data generated by a sample generation model; An expected category label corresponding to each of the initial sample data is determined respectively, and a plurality of the initial sample data in the initial sample data set are divided into a positive sample data set and a negative sample data set according to the expected category label corresponding to the initial sample data.

3. The sample data processing method according to claim 2, characterized in that: The sample generation model is obtained by training a generator using a generative adversarial network; the loss function of a discriminator in the generative adversarial network that performs adversarial training with the generator includes a modulation factor; The modulation factor is used to control the degree of weight attenuation of difficult and easy samples.

4. The sample data processing method according to claim 3, characterized in that: The loss function of the discriminator is: in, is the discriminant loss corresponding to the discriminator, x i is the i-th real sample data; D(x i ) is the discrimination result of the discriminator on the i-th real sample data; The i-th simulated sample data generated by the sample generation model; is the discrimination result of the discriminator on the i-th simulated sample data; N is the total number of the real sample data, and , the total number of the simulated sample data.

5. The sample data processing method according to claim 2, characterized in that: The sample generation model includes an encoder and a decoder; the encoder is used to convert an input noise vector into a latent space vector; the decoder is used to convert the latent space vector into simulated sample data.

6. The sample data processing method according to claim 5, characterized in that: The obtaining of the initial sample data set includes: Inputting a plurality of noise vectors into the sample generation model to obtain a plurality of simulated sample data; A plurality of real sample data are obtained, and an initial sample data set is constructed according to the plurality of real sample data and the plurality of simulated sample data.

7. The sample data processing method according to claim 1, characterized in that: The step of determining the target sampling rate of each majority class sample cluster according to the distance between each majority class sample cluster and the minority class sample set and the total sampling rate comprises: Determine the distance between each of the majority class sample clusters and the minority class sample set, and determine the sum of the distances corresponding to multiple majority class sample clusters; An initial sampling rate is determined according to the distance corresponding to each of the majority class sample clusters and the sum of the distances, and a target sampling rate is determined according to the initial sampling rate and the total sampling rate.

8. The sample data processing method according to claim 7, characterized in that: The determining of the target sampling rate according to the initial sampling rate and the total sampling rate includes: The initial sampling rate and the total sampling rate are multiplied to obtain a target sampling rate.

9. The sample data processing method according to claim 1, characterized in that: After constructing the target sample data set according to the plurality of sample data and the initial sample data included in the minority class sample set, the method further includes: The pre-established initial classification model is trained according to the target sample data set to obtain a target classification model.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the sample data processing method according to any one of claims 1 to 9.