Account detection model training method, device, electronic device and storage medium

Through clustering and similarity annotation methods, training and aggregating sub-account detection models are solved, and the accuracy and efficiency of the detection model are improved.

CN119441894BActive Publication Date: 2025-06-06BEIJING TRUSFORT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510019031.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-06-06
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

The prior art faces the problems of data imbalance and insufficient marking in account abnormality detection, which leads to limited effectiveness of traditional models in practical applications and it is difficult to effectively detect potential abnormal behavior.

Method used

By obtaining the sample data set, clustering is performed to generate multiple clusters, the similarity is calculated and labeled for unlabeled accounts in each cluster, the sub-account detection model of each cluster is trained, and these sub-models are aggregated to obtain the target account detection model.

Benefits of technology

Expand positive sample accounts through similarity calculations, solve the problem of data imbalance, improve the training accuracy and efficiency of target account detection models, and thus improve the accuracy of account detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119441894B_ABST
    Figure CN119441894B_ABST
Patent Text Reader

Abstract

The present application provides a training method, device, electronic device and storage medium for an account detection model, the method comprising: obtaining a sample data set, wherein the samples of the sample data set include multiple unlabeled accounts and multiple positive sample accounts; performing clustering processing on each sample in the sample data set to obtain a clustering result, wherein the clustering result includes multiple clusters; labeling the unlabeled accounts in each cluster to obtain a labeling result for each cluster; labeling the unlabeled accounts in each cluster, including: calculating the similarity between each unlabeled account in each cluster and the positive sample account; labeling each unlabeled account in each cluster according to the similarity corresponding to each unlabeled account; training a sub-account detection model corresponding to each cluster according to the labeling result of each cluster and the positive sample account in each cluster; and aggregating the sub-account detection models corresponding to each cluster according to the clustering result to obtain a target account detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a training method, device, electronic device and storage medium for an account detection model. Background Art

[0002] With the booming development of online payment and online banking business, account anomaly detection has become a key link in ensuring the security of user funds and the stable operation of banking business. However, this field faces severe challenges of imbalanced and insufficiently labeled abnormal account data, that is, there are relatively few abnormal cases (positive samples) and it is difficult to fully label them, and unlabeled account data accounts for the vast majority. This data characteristic makes the traditional account detection model limited in practical application and difficult to effectively detect potential abnormal behaviors. Therefore, how to train an accurate and reliable account detection model under the condition of unbalanced data has become a technical problem that needs to be solved urgently. Summary of the invention

[0003] Embodiments of the present application provide a method, device, electronic device, and storage medium for training an account detection model.

[0004] According to a first aspect of the present application, a method for training an account detection model is provided, the method comprising:

[0005] Obtain a sample data set, where samples in the sample data set include multiple unlabeled accounts and multiple positive sample accounts;

[0006] Performing clustering processing on each sample in the sample data set to obtain a clustering result, wherein the clustering result includes a plurality of clusters;

[0007] Labeling the unlabeled accounts in each cluster to obtain the labeling results of each cluster; the labeling of the unlabeled accounts in each cluster includes: calculating the similarity between each unlabeled account in each cluster and the positive sample account; labeling each unlabeled account in each cluster according to the similarity corresponding to each unlabeled account;

[0008] Based on the annotation results of each cluster and the positive sample accounts in each cluster, train the sub-account detection model corresponding to each cluster;

[0009] According to the clustering result, the sub-account detection models corresponding to each cluster are aggregated to obtain a target account detection model.

[0010] According to an embodiment of the present application, clustering each sample in the sample data set includes:

[0011] Get sample features of each sample;

[0012] According to the sample characteristics of each positive sample account, the positive sample accounts in the sample data set are initially clustered to obtain an initial clustering result;

[0013] According to the initial clustering results, determine the cluster center;

[0014] Each unlabeled account is clustered according to the cluster center and the sample characteristics of each unlabeled account.

[0015] According to an embodiment of the present application, obtaining the sample feature of each sample includes:

[0016] Obtaining transaction flow data corresponding to the sample data set;

[0017] Obtaining the transaction flow of each positive sample account and the transaction flow of each unlabeled account from the transaction flow data;

[0018] Extract the feature data to be processed of the corresponding account from each transaction flow, and obtain the feature data to be processed of each positive sample account and each unlabeled account;

[0019] Performing data feature preprocessing on each feature data to be processed to obtain a data feature preprocessing result, wherein the data feature preprocessing includes outlier processing, missing value processing and binning processing;

[0020] The sample features of each positive sample account and the sample features of each unlabeled account are extracted from the data feature preprocessing result, wherein the sample features include user basic features, account basic features and account flow features.

[0021] According to an embodiment of the present application, the calculating of the similarity between each unlabeled account in each cluster and the positive sample account includes:

[0022] For each cluster, obtain k samples adjacent to each unlabeled account;

[0023] Determine the proportion of positive sample accounts in the k samples;

[0024] The proportion of the positive sample accounts is used as the similarity between the unlabeled account and the positive sample account.

[0025] According to an implementation manner of the present application, the labeling result includes an account label for each unlabeled account, and the account label includes an updated negative sample account and an updated positive sample account; accordingly,

[0026] The step of labeling each unlabeled account in each cluster according to the similarity corresponding to each unlabeled account includes:

[0027] Generate a random number uniformly distributed between [0, 1] for each unlabeled account;

[0028] Compare the similarity of each unlabeled account with the random number to obtain the comparison result of each unlabeled account;

[0029] If the comparison result shows that the similarity is greater than or equal to the random number, the corresponding unlabeled account is labeled as an updated positive sample account;

[0030] When the comparison result shows that the similarity is less than the random number, the corresponding unlabeled account is labeled as an updated negative sample account.

[0031] According to an embodiment of the present application, the sub-account detection model corresponding to each cluster is trained according to the labeling results of each cluster and the positive sample accounts in each cluster, including:

[0032] For each cluster, the updated negative sample accounts and updated positive sample accounts shown in the labeling results are combined with the positive sample accounts within the cluster to form a training account set, so as to obtain a training account set corresponding to each cluster;

[0033] According to the sample features of each account in the training account set corresponding to each cluster, a sub-account detection model for each cluster is trained.

[0034] According to an embodiment of the present application, the method further includes:

[0035] Receiving an account to be detected and a transaction flow of the account to be detected;

[0036] Extracting account features of the account to be detected from the transaction flow of the account to be detected;

[0037] Using the target account detection model and account features to detect the account to be detected, and obtain an account detection result;

[0038] The target account detection model and account features are used to detect the account to be detected, and the account detection result is obtained, including:

[0039] Determining the cluster to which the account to be detected belongs according to the account characteristics of the account to be detected and the clustering result;

[0040] The account features of the account to be detected are input into the sub-account detection model corresponding to the cluster to which the account to be detected belongs, to obtain the account detection result.

[0041] According to a second aspect of the present application, a training device for an account detection model is provided, the device comprising:

[0042] An acquisition module, used to acquire a sample data set, wherein the samples of the sample data set include a plurality of unlabeled accounts and a plurality of positive sample accounts;

[0043] A clustering module, used for performing clustering processing on each sample in the sample data set to obtain a clustering result, wherein the clustering result includes a plurality of clusters;

[0044] The labeling module is used to label the unlabeled accounts in each cluster to obtain the labeling results of each cluster; labeling the unlabeled accounts in each cluster includes: calculating the similarity between each unlabeled account in each cluster and the positive sample account; labeling each unlabeled account in each cluster according to the similarity corresponding to each unlabeled account;

[0045] The training module is used to train the sub-account detection model corresponding to each cluster based on the labeling results of each cluster and the positive sample accounts in each cluster;

[0046] The aggregation module is used to aggregate the sub-account detection models corresponding to each cluster according to the clustering result to obtain a target account detection model.

[0047] According to a third aspect of the present application, an electronic device is provided, including:

[0048] at least one processor; and

[0049] a memory communicatively connected to the at least one processor; wherein,

[0050] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the present application.

[0051] According to a fourth aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the method described in the present application.

[0052] The present application discloses a training method, device, electronic device and storage medium for an account detection model, which obtains a sample data set, wherein the samples of the sample data set include multiple unlabeled accounts and multiple positive sample accounts; clusters each sample in the sample data set to obtain a clustering result, wherein the clustering result includes multiple clusters; labels the unlabeled accounts in each cluster to obtain a labeling result for each cluster; labels the unlabeled accounts in each cluster, including: calculating the similarity between each unlabeled account in each cluster and the positive sample account; labeling each unlabeled account in each cluster according to the similarity corresponding to each unlabeled account; training a sub-account detection model corresponding to each cluster according to the labeling result of each cluster and the positive sample accounts in each cluster; and aggregating the sub-account detection models corresponding to each cluster according to the clustering result to obtain a target account detection model. By labeling the unlabeled accounts in the sample data set through similarity calculation, more positive sample accounts can be obtained, the positive sample accounts can be expanded, and the problem of data imbalance can be solved. Therefore, when the target account detection model is subsequently trained, it can help the target account detection model learn more positive sample accounts, that is, abnormal accounts, improve the accuracy and efficiency of the target account detection model training, and thus improve the accuracy of account detection.

[0053] It should be understood that the teachings of the present application are not required to achieve all of the beneficial effects described above, but specific technical solutions can achieve specific technical effects, and other embodiments of the present application can also achieve beneficial effects not mentioned above. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] By reading the detailed description below with reference to the accompanying drawings, the above and other purposes, features and advantages of the exemplary embodiments of the present application will become readily understood. In the accompanying drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, wherein:

[0055] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.

[0056] Figure 1 A schematic diagram of the implementation process of the training method of the account detection model provided in the embodiment of the present application is shown;

[0057] Figure 2 A schematic diagram of the implementation process of the clustering operation of the account detection model training method provided in an embodiment of the present application is shown;

[0058] Figure 3 A schematic diagram of the implementation process of the similarity calculation operation of the account detection model training method provided in an embodiment of the present application is shown;

[0059] Figure 4A schematic diagram of the implementation process of the labeling operation of the account detection model training method provided in an embodiment of the present application is shown;

[0060] Figure 5 An example diagram showing the relationship between k and the number of positive sample account annotations in an embodiment of the present application;

[0061] Figure 6 A schematic diagram showing the composition structure of a training device for an account detection model provided in an embodiment of the present application is shown;

[0062] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0063] In order to make the purpose, features, and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0064] Figure 1 A schematic diagram of the implementation process of the account detection model training method provided in an embodiment of the present application is shown.

[0065] refer to Figure 1 , an embodiment of the present application provides a method for training an account detection model, the method comprising:

[0066] Operation 101 : obtaining a sample data set, where samples of the sample data set include a plurality of unlabeled accounts and a plurality of positive sample accounts.

[0067] In order to manage the accounts, a blacklist and a whitelist are usually configured for the accounts. The blacklist usually includes accounts with abnormalities, and the whitelist usually includes accounts for which it is not determined whether there are abnormalities.

[0068] It can be understood that under normal circumstances, the blacklist includes abnormal accounts, while the whitelist includes normal accounts. However, in the field of bank account detection, it is difficult to determine whether an account belongs to a normal user. Therefore, the whitelist in the embodiment of the present application includes accounts that are not determined to be abnormal.

[0069] After obtaining the blacklist and the whitelist, the accounts in the blacklist can be determined as positive sample accounts, and the accounts in the whitelist can be determined as unlabeled accounts.

[0070] In order to avoid interference of too many negative samples on subsequent model training, determining the accounts in the whitelist as unlabeled accounts may include: removing the accounts in the whitelist that are clearly not abnormal according to the account flow, and obtaining unlabeled accounts. In this way, by reducing negative samples, the subsequent model can focus more on learning the characteristics of positive sample accounts, thereby improving the model's ability to detect abnormal accounts. Among them, negative samples refer to normal accounts.

[0071] Operation 102 : performing clustering processing on each sample in the sample data set to obtain a clustering result, where the clustering result includes a plurality of clusters.

[0072] The transaction flow of an account shows multiple features of the account. According to the differences in the performance of the account in different features, abnormal accounts can be divided into different types of abnormal accounts.

[0073] For example, abnormal accounts may behave differently in terms of transaction amount, transaction frequency, transaction location, transaction time and other characteristics. Based on their performance in different characteristics, abnormal accounts can be divided into multiple types such as credit card anomalies, identity theft, account fraud and so on.

[0074] Therefore, based on the performance of accounts on different features, a clustering algorithm can be used to group accounts with similar feature performance into the same cluster, so that different types of accounts can be represented based on different clusters.

[0075] In one embodiment of the present application, the clustering algorithm may be K-means (K-means clustering algorithm), a hierarchical clustering algorithm, or DBSCAN (Density-Based Spatial Clustering of Applications with Noise), etc.

[0076] Operation 103, labeling the unlabeled accounts in each cluster to obtain the labeling results of each cluster; labeling the unlabeled accounts in each cluster, including: calculating the similarity between each unlabeled account in each cluster and the positive sample account; labeling each unlabeled account in each cluster according to the similarity corresponding to each unlabeled account.

[0077] For each cluster, when labeling the unlabeled accounts in it, it can be done by calculating the similarity between the unlabeled accounts and the positive sample accounts. For example, unlabeled accounts whose similarity exceeds a threshold are labeled as updated positive sample accounts, and unlabeled accounts that do not exceed the threshold are labeled as updated negative sample accounts. The threshold can be configured according to actual needs, such as the accuracy requirements of anomaly detection, which will not be repeated here.

[0078] Operation 104 : training a sub-account detection model corresponding to each cluster based on the labeling results of each cluster and the positive sample accounts in each cluster.

[0079] In one embodiment of the present application, a sub-account detection model corresponding to each cluster is trained based on the labeling results of each cluster and the positive sample accounts within each cluster, including: for each cluster, the updated negative sample accounts and updated positive sample accounts shown in its labeling results are combined with the positive sample accounts within it to form a training account set to obtain a training account set corresponding to each cluster; based on the sample features of each account in the training account set corresponding to each cluster, a sub-account detection model for each cluster is trained.

[0080] The labeling results of each cluster show the account types of the unlabeled accounts in it. The account types include updated positive sample accounts and updated negative sample accounts. The labeling results of each cluster and its corresponding original positive sample accounts are used as a training account set.

[0081] The basic model is trained according to each training account set and the sample features of each account in the training account set to obtain the sub-account detection model corresponding to each cluster. The sample features can be understood as the account features of the corresponding account.

[0082] In one embodiment of the present application, the basic model may be a classifier, a deep learning model, etc. Preferably, the basic model is a classifier, and the classifier is XGBOOST (eXtreme Gradient Boosting, an integrated algorithm based on gradient boosting decision trees).

[0083] Operation 105 , based on the clustering result, aggregate the sub-account detection models corresponding to each cluster to obtain a target account detection model.

[0084] In one embodiment of the present application, the target account detection model can be expressed as:

[0085] in, represents the account to be detected and its sample features, m is the number of clusters in the clustering result, F is the clustering function based on the clustering result, which means selecting the cluster corresponding to the account to be detected from the clustering result and selecting the sub-account detection model of the corresponding cluster. The sub-account detection result of the sub-account detection model corresponding to the j-th cluster on the account to be detected. The account detection result of the account to be detected.

[0086] Therefore, the target account detection model can be constructed through the clustering results and the sub-account detection model corresponding to each cluster. When an account needs to be detected, the account and the corresponding account features are input into the target account detection model, and the target account detection model can obtain the cluster corresponding to the account to be detected in the clustering results, and detect the account to be detected according to the sub-account detection model of the corresponding cluster.

[0087] Therefore, the embodiment of the present application labels the unlabeled accounts in the sample data set by similarity calculation, so as to obtain more positive sample accounts, expand the positive sample accounts, and solve the problem of data imbalance. Therefore, when subsequently training the target account detection model, it helps the target account detection model to learn more positive sample accounts, that is, abnormal accounts, thereby improving the accuracy and efficiency of the target account detection model training, thereby improving the accuracy of account detection.

[0088] Figure 2 A schematic diagram of the implementation flow of the clustering operation of the account detection model training method provided in an embodiment of the present application is shown.

[0089] refer to Figure 2 In one embodiment of the present application, the above operation 102, performing clustering processing on each sample in the sample data set, includes:

[0090] Operation 201, obtaining sample features of each sample;

[0091] Operation 202, performing initial clustering on the positive sample accounts in the sample data set according to the sample characteristics of each positive sample account, to obtain an initial clustering result;

[0092] Operation 203, determining a cluster center according to the initial clustering result;

[0093] In operation 204 , each unlabeled account is clustered according to the cluster center and the sample features of each unlabeled account.

[0094] In one embodiment of the present application, obtaining sample features of each sample includes: obtaining transaction flow data corresponding to the sample data set; obtaining the transaction flow of each positive sample account and the transaction flow of each unlabeled account from the transaction flow data; extracting the feature data to be processed of the corresponding account from each transaction flow, and obtaining the feature data to be processed of each positive sample account and each unlabeled account; performing data feature preprocessing on each feature data to be processed, and obtaining a data feature preprocessing result, the data feature preprocessing includes outlier processing, missing value processing and binning processing; extracting sample features of each positive sample account and the sample features of each unlabeled account from the data feature preprocessing result, the sample features include user basic features, account basic features and account flow features.

[0095] First, the transaction flow data of all accounts in the sample data set are obtained. The transaction flow data includes the transaction flow of each account.

[0096] Positive sample accounts are usually blacklisted accounts. Since the blacklist needs to be updated at regular intervals, when account detection is required, the blacklist and whitelist may be from the past few months, while the transaction flow is from the recent months. Therefore, it is first necessary to match the positive sample accounts and unlabeled accounts with their corresponding transaction flows, that is, extract the transaction flows corresponding to the positive sample accounts and unlabeled accounts.

[0097] In one embodiment of the present application, after obtaining the transaction flow of each account, the transaction flow is also subjected to data cleaning and data inspection. Data cleaning mainly involves cleaning the flow without abnormalities, and data inspection mainly involves checking whether the transaction flow includes set fields, such as the timestamp field. If not, it needs to be obtained and supplemented. Data inspection also includes checking whether the flow is continuous. If it is not continuous, it needs to be supplemented by a data supplementation method. In this way, based on data cleaning and data inspection, the data quality can be improved, thereby improving the training and operation efficiency of subsequent models.

[0098] After obtaining the transaction flow of each account, it is necessary to extract useful features from the transaction flow, namely, the feature data to be processed. The feature data to be processed may include the user feature data of the corresponding account, the basic feature data of the account, and the account flow feature data.

[0099] In order to extract features that are more relevant to account detection, data feature preprocessing is also performed on the feature data to be processed.

[0100] In this embodiment of the present application, data feature preprocessing includes outlier processing, missing value processing and binning processing.

[0101] Among them, outlier processing includes: using a box plot to determine the value of the feature data with abnormal values ​​and removing them. Missing value processing includes: determining the missing value rate of the feature, removing abnormal data of the feature with a missing value rate within the first threshold range, and removing features with a missing value rate within the second threshold range. The first threshold range and the second threshold range can be configured according to actual needs, such as the requirements for data accuracy. The first threshold range can be configured to 70~80, and the second threshold range can be configured to 95~100.

[0102] The purpose of binning is to convert continuous features (or dense discrete features) into discrete features, so as to improve the robustness of the model when the target account detection model is trained later. The binning process includes rough binning and fine binning using the bad sample rate badRate. The specific binning process can refer to the conventional feature binning process, which will not be repeated here.

[0103] After preprocessing the feature data to be processed, the user basic features, account basic features and account flow features of each positive sample account and each unlabeled account are extracted from the data feature preprocessing results to obtain the sample features of each positive sample account and each unlabeled account. The sample features are account features, including user basic features, account basic features and account flow features.

[0104] After obtaining the sample features of each positive sample account and each unlabeled account, the positive sample accounts in the sample data set can be initially clustered using a conventional clustering algorithm, such as K-means, and multiple cluster centers can be determined based on the results of the initial clustering.

[0105] After obtaining the cluster centers, the same clustering algorithm as the initial clustering algorithm or a different clustering algorithm can be used to cluster the unlabeled accounts in the sample data set to each cluster center to obtain the cluster corresponding to each cluster center. Preferably, the clustering algorithm used for the unlabeled accounts in this application is different from the clustering algorithm used for the initial clustering, and the clustering algorithm can be Seeded-KMeans (Seeded K-Means Clustering Algorithm, seed-based K-means clustering algorithm).

[0106] Figure 3 A schematic diagram of the implementation flow of the similarity calculation operation of the account detection model training method provided in an embodiment of the present application is shown.

[0107] refer to Figure 3 In one embodiment of the present application, the similarity between each unlabeled account in each cluster and the positive sample account is calculated, including:

[0108] Operation 301, for each cluster, obtaining k samples adjacent to each unlabeled account;

[0109] Operation 302, determining the proportion of positive sample accounts in the k samples;

[0110] In operation 303 , the proportion of the positive sample accounts is used as the similarity between the unlabeled accounts and the positive sample accounts.

[0111] For the unlabeled accounts in each cluster, obtain the corresponding k nearest neighbor samples, calculate the proportion of positive sample accounts in the k nearest neighbor samples, and take the proportion of positive sample accounts in the k nearest neighbor samples of each unlabeled account in each cluster as its similarity.

[0112] In one embodiment of the present application, the proportion of positive sample accounts in the samples of the k nearest neighbors of each unlabeled account can be determined by the following formula:

[0113] in, is the proportion of positive sample accounts in the samples of the k nearest neighbors of the i-th unlabeled account, are the k nearest neighbors in the cluster to which the i-th unlabeled account belongs, j represents the j-th sample, represents the account label of the jth sample, which includes positive sample accounts and unlabeled accounts. Indicates that the jth sample is a positive sample account, 1 represents a positive sample account, is an indicator function. When the condition in the brackets is true (i.e. the condition is met), the value of the indicator function is 1. When the condition in the brackets is false (i.e. the condition is not met), the value of the indicator function is 0.

[0114] After obtaining the proportion of positive sample accounts among the k nearest neighbors of the unlabeled account i through the above formula, the proportion of positive sample accounts among the k nearest neighbors of the unlabeled account i is determined as the similarity corresponding to the account i.

[0115] Figure 4 A schematic diagram of the implementation process of the labeling operation of the account detection model training method provided in an embodiment of the present application is shown.

[0116] refer to Figure 4 In one embodiment of the present application, the labeling result includes the account label of each unlabeled account, and the account label includes an updated negative sample account and an updated positive sample account; accordingly, each unlabeled account in each cluster is labeled according to the similarity corresponding to each unlabeled account, including:

[0117] Operation 401, randomly generating a random number that follows a uniform distribution of [0, 1] for each unlabeled account;

[0118] Operation 402, comparing the similarity of each unlabeled account with the random number to obtain a comparison result of each unlabeled account;

[0119] Operation 403, when the comparison result shows that the similarity is greater than or equal to the random number, the corresponding unlabeled account is labeled as an updated positive sample account;

[0120] Operation 404 : When the comparison result shows that the similarity is less than the random number, the corresponding unlabeled account is labeled as an updated negative sample account.

[0121] For each unlabeled account, a random number that obeys the [0, 1] uniform distribution can be generated using a random number generation method. The [0, 1] uniform distribution means a uniform distribution within the range of [0, 1].

[0122] After obtaining the random number of each unlabeled account, compare the random number of each unlabeled account with its similarity. When the similarity is greater than or equal to the random number, the current unlabeled account is labeled as an updated positive sample account, that is, the account label of the current unlabeled account is determined to be an updated positive sample account; when the similarity is less than the random number, the current unlabeled account is labeled as an updated negative sample account, that is, the account label of the current unlabeled account is determined to be an updated negative sample account.

[0123] In one embodiment of the present application, the account label of the unlabeled account can be determined by the following formula:

[0124]

[0125] in, represents the actual account label of the i-th unlabeled account, It can be divided into 1 and -1, 1 means updating the positive sample account, -1 means updating the negative sample account. A random number representing the i-th unlabeled account.

[0126] In one embodiment of the present application, a trained target account detection model is also used to detect the account to be detected, specifically including: receiving the account to be detected and the transaction flow of the account to be detected; extracting the account characteristics of the account to be detected from the transaction flow of the account to be identified; using the target account detection model and the account characteristics to detect the account to be detected, and obtaining an account detection result; wherein, using the target account detection model and the account characteristics to detect the account to be detected, and obtaining an account detection result, includes: determining the cluster to which the account to be detected belongs based on the account characteristics of the account to be detected and the clustering results; inputting the account characteristics of the account to be detected into the sub-account detection model corresponding to the cluster to which the account to be detected belongs, and obtaining an account detection result.

[0127] After extracting the account features of the account to be tested from the transaction flow of the account to be tested, the account features can be input into the target account detection model . Then, the target account detection model determines which cluster the target account belongs to in the clustering result based on the account features, and detects the account features through the sub-account detection model of the cluster to which the target account belongs to obtain the account type of the target account.

[0128] In one embodiment of the present application, the above Figure 3 Specifically, we can obtain k sample data sets for k and configure the feasible range of k, and then use Figures 1 to 4The labeling method is to label the unlabeled accounts in the k sample data set corresponding to each k within the feasible range according to the numerical value, and finally calculate the number of positive sample accounts labeled under each k, and take the k whose number of labeled positive sample accounts meets the set conditions as the k of the embodiment of the present application. The set condition is that as k increases, the number of labeled positive sample accounts no longer increases.

[0129] For example, see Figure 5 , Figure 5 An example diagram showing the relationship between k and the number of positive sample account annotations in an embodiment of the present application. Figure 5 The inner vertical axis E_k represents the number of positive sample accounts marked, and the horizontal axis represents the value of k. Figure 5 It can be seen that as k increases, E_k increases continuously until E_k stabilizes when k is between 12 and 16, and then no longer increases. Therefore, k between 12 and 16 meets the set conditions, and the k value in the embodiment of the present application can be a value between 12 and 16. It should be noted that the selection of the above k value is only an exemplary description and is not specifically limited.

[0130] Figure 6 A schematic diagram of the composition structure of the training device for the account detection model provided in an embodiment of the present application is shown.

[0131] refer to Figure 6 Based on the above-mentioned account detection model training method, the embodiment of the present application also provides an account detection model training device, which includes: an acquisition module 601, used to acquire a sample data set, the samples of the sample data set include multiple unlabeled accounts and multiple positive sample accounts; a clustering module 602, used to cluster each sample in the sample data set to obtain a clustering result, and the clustering result includes multiple clusters; a labeling module 603, used to label the unlabeled accounts in each cluster to obtain the labeling result of each cluster; labeling the unlabeled accounts in each cluster, including: calculating the similarity between each unlabeled account in each cluster and the positive sample account; labeling each unlabeled account in each cluster according to the similarity corresponding to each unlabeled account; a training module 604, used to train the sub-account detection model corresponding to each cluster according to the labeling result of each cluster and the positive sample account in each cluster; an aggregation module 605, used to aggregate the sub-account detection models corresponding to each cluster according to the clustering result to obtain the target account detection model.

[0132] In one embodiment of the present application, the clustering module 602 includes: a first acquisition submodule, used to acquire sample characteristics of each sample; a primary clustering submodule, used to perform primary clustering of positive sample accounts in the sample data set according to the sample characteristics of each positive sample account, to obtain primary clustering results; a center determination submodule, used to determine the clustering center according to the primary clustering results; and an account clustering submodule, used to cluster each unlabeled account according to the clustering center and the sample characteristics of each unlabeled account.

[0133] In one embodiment of the present application, a first acquisition submodule includes: an acquisition unit, which is used to acquire transaction flow data corresponding to a sample data set; a first extraction unit, which is used to acquire the transaction flow of each positive sample account and the transaction flow of each unlabeled account from the transaction flow data; a second extraction unit, which is used to extract the feature data to be processed of the corresponding account from each transaction flow, and obtain the feature data to be processed of each positive sample account and each unlabeled account; a processing unit, which is used to perform data feature preprocessing on each feature data to be processed, and obtain a data feature preprocessing result, wherein the data feature preprocessing includes outlier processing, missing value processing and binning processing; a third extraction unit, which is used to extract sample features of each positive sample account and the sample features of each unlabeled account from the data feature preprocessing result, wherein the sample features include user basic features, account basic features and account flow features.

[0134] In one embodiment of the present application, the labeling module 603 includes: a labeling submodule, which is used to obtain k samples adjacent to each unlabeled account for each cluster; a first determination submodule, which is used to determine the proportion of positive sample accounts in the k samples; and an assignment submodule, which is used to use the proportion of positive sample accounts as the similarity between the unlabeled account and the positive sample account.

[0135] In one embodiment of the present application, the labeling result includes an account label for each unlabeled account, and the account label includes an updated negative sample account and an updated positive sample account; accordingly, the labeling module 603 includes: a generation submodule, which is used to randomly generate a random number that obeys a uniform distribution of [0, 1] for each unlabeled account; a comparison submodule, which is used to compare the similarity of each unlabeled account with the random number to obtain a comparison result for each unlabeled account; a first update submodule, which is used to label the corresponding unlabeled account as an updated positive sample account when the comparison result shows that the similarity is greater than or equal to the random number; and a second update submodule, which is used to label the corresponding unlabeled account as an updated negative sample account when the comparison result shows that the similarity is less than the random number.

[0136] In one embodiment of the present application, the training module 604 includes: a composition submodule, which is used to, for each cluster, combine the updated negative sample accounts and updated positive sample accounts shown in the labeling results with the positive sample accounts therein to form a training account set, so as to obtain a training account set corresponding to each cluster; a training submodule, which is used to train a sub-account detection model for each cluster based on the sample features of each account in the training account set corresponding to each cluster.

[0137] In one embodiment of the present application, the device also includes: a receiving module, which is used to receive the account to be detected and the transaction flow of the account to be detected; an extraction module, which is used to extract the account characteristics of the account to be detected from the transaction flow of the account to be detected; a detection module, which is used to detect the account to be detected using the target account detection model and the account characteristics to obtain an account detection result; wherein, the account to be detected using the target account detection model and the account characteristics to obtain the account detection result includes: determining the cluster to which the account to be detected belongs based on the account characteristics of the account to be detected and the clustering result; and inputting the account characteristics of the account to be detected into the sub-account detection model corresponding to the cluster to which the account to be detected belongs, to obtain the account detection result.

[0138] It should be noted that the description of the device in the embodiment of the present application is similar to the description of the method embodiment above, and has similar beneficial effects as the method embodiment, so it will not be repeated. Figures 1 to 5 The present invention can be understood by referring to the description of any one of the accompanying drawings.

[0139] According to an embodiment of the present application, the present application also provides an electronic device and a non-transitory computer-readable storage medium.

[0140] Figure 7 A schematic block diagram of an example electronic device 70 that can be used to implement an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0141] like Figure 7As shown, the electronic device 70 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 70 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0142] Multiple components in the electronic device 70 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 70 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0143] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the training method of the account detection model. For example, in some embodiments, the training method of the account detection model may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 70 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the training method of the account detection model described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the training method of the account detection model in any other appropriate manner (e.g., by means of firmware).

[0144] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0145] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, implements the functions / operations specified in the flow chart and / or block diagram. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0146] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0147] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0148] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0149] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0150] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this application can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution disclosed in this application can be achieved, and this document is not limited here.

[0151] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A training method for an account detection model, characterized in that: The method comprises: Obtain a sample data set, where samples in the sample data set include multiple unlabeled accounts and multiple positive sample accounts; Performing clustering processing on each sample in the sample data set to obtain a clustering result, wherein the clustering result includes a plurality of clusters; Labeling the unlabeled accounts in each cluster to obtain the labeling results of each cluster; labeling the unlabeled accounts in each cluster, including: calculating the similarity between each unlabeled account in each cluster and the positive sample account; labeling each unlabeled account in each cluster according to the similarity corresponding to each unlabeled account; Based on the annotation results of each cluster and the positive sample accounts in each cluster, train the sub-account detection model corresponding to each cluster; According to the clustering result, the sub-account detection models corresponding to each cluster are aggregated to obtain a target account detection model; The clustering process for each sample in the sample data set includes: obtaining sample features of each sample; performing initial clustering on the positive sample accounts in the sample data set according to the sample features of each positive sample account to obtain initial clustering results; determining a cluster center according to the initial clustering results; clustering each unlabeled account according to the cluster center and the sample features of each unlabeled account; The calculating of the similarity between each unlabeled account and the positive sample account in each cluster includes: for each cluster, obtaining k samples adjacent to each unlabeled account; determining the proportion of the positive sample accounts in the k samples; and using the proportion of the positive sample accounts as the similarity between the unlabeled account and the positive sample account; k is obtained by the following operations: obtaining a k-sample data set for k and configuring a feasible range for k; labeling the unlabeled accounts in the k-sample data set corresponding to each k within the feasible range according to the numerical value; calculating the number of positive sample accounts labeled under each k, and taking the k whose number of labeled positive sample accounts meets the set condition as the final k, where the set condition is that as k increases, the number of labeled positive sample accounts no longer increases; The obtaining of the sample feature of each sample includes: Obtaining transaction flow data corresponding to the sample data set; Obtaining the transaction flow of each positive sample account and the transaction flow of each unlabeled account from the transaction flow data; Extract the feature data to be processed of the corresponding account from each transaction flow, and obtain the feature data to be processed of each positive sample account and each unlabeled account; Performing data feature preprocessing on each feature data to be processed to obtain a data feature preprocessing result, wherein the data feature preprocessing includes outlier processing, missing value processing and binning processing; The sample features of each positive sample account and the sample features of each unlabeled account are extracted from the data feature preprocessing result, wherein the sample features include user basic features, account basic features and account flow features.

2. The method according to claim 1, characterized in that The labeling result includes an account label for each unlabeled account, and the account label includes an updated negative sample account and an updated positive sample account; accordingly, The step of labeling each unlabeled account in each cluster according to the similarity corresponding to each unlabeled account includes: Generate a random number uniformly distributed between [0, 1] for each unlabeled account; Compare the similarity of each unlabeled account with the random number to obtain the comparison result of each unlabeled account; If the comparison result shows that the similarity is greater than or equal to the random number, the corresponding unlabeled account is labeled as an updated positive sample account; When the comparison result shows that the similarity is less than the random number, the corresponding unlabeled account is labeled as an updated negative sample account.

3. The method according to claim 2, characterized in that The sub-account detection model corresponding to each cluster is trained according to the labeling results of each cluster and the positive sample accounts in each cluster, including: For each cluster, the updated negative sample accounts and updated positive sample accounts shown in the annotation results are combined with the positive sample accounts within the cluster to form a training account set, so as to obtain a training account set corresponding to each cluster; According to the sample features of each account in the training account set corresponding to each cluster, a sub-account detection model for each cluster is trained.

4. The method according to claim 1, characterized in that: The method further comprises: Receiving an account to be detected and a transaction flow of the account to be detected; Extracting account features of the account to be detected from the transaction flow of the account to be detected; Using the target account detection model and account features to detect the account to be detected, and obtain an account detection result; The target account detection model and account features are used to detect the account to be detected, and the account detection result is obtained, including: Determining the cluster to which the account to be detected belongs according to the account characteristics of the account to be detected and the clustering result; The account features of the account to be detected are input into the sub-account detection model corresponding to the cluster to which the account to be detected belongs, to obtain the account detection result.

5. A training device for an account detection model, characterized in that: The device comprises: An acquisition module, used to acquire a sample data set, wherein the samples of the sample data set include a plurality of unlabeled accounts and a plurality of positive sample accounts; A clustering module, used for performing clustering processing on each sample in the sample data set to obtain a clustering result, wherein the clustering result includes a plurality of clusters; The labeling module is used to label the unlabeled accounts in each cluster to obtain the labeling results of each cluster; labeling the unlabeled accounts in each cluster includes: calculating the similarity between each unlabeled account in each cluster and the positive sample account; labeling each unlabeled account in each cluster according to the similarity corresponding to each unlabeled account; The training module is used to train the sub-account detection model corresponding to each cluster based on the labeling results of each cluster and the positive sample accounts in each cluster; An aggregation module, configured to aggregate the sub-account detection models corresponding to each cluster according to the clustering result to obtain a target account detection model; The clustering module includes: a first acquisition submodule, used to acquire the sample features of each sample; a primary clustering submodule, used to perform primary clustering on the positive sample accounts in the sample data set according to the sample features of each positive sample account, and obtain the primary clustering results; a center determination submodule, used to determine the cluster center according to the primary clustering results; an account clustering submodule, used to cluster each unlabeled account according to the cluster center and the sample features of each unlabeled account; The labeling module includes: a labeling submodule, for obtaining, for each cluster, k samples adjacent to each unlabeled account; a first determination submodule, for determining the proportion of positive sample accounts in the k samples; an assignment submodule, for using the proportion of positive sample accounts as the similarity between the unlabeled account and the positive sample account; k is obtained by the following operations: obtaining a k-sample data set for k and configuring a feasible range for k; labeling the unlabeled accounts in the k-sample data set corresponding to each k within the feasible range according to the numerical value; calculating the number of positive sample accounts labeled under each k, and taking the k whose number of labeled positive sample accounts meets the set condition as the final k, where the set condition is that as k increases, the number of labeled positive sample accounts no longer increases; The first acquisition submodule includes: an acquisition unit, which is used to acquire transaction flow data corresponding to the sample data set; a first extraction unit, which is used to acquire the transaction flow of each positive sample account and the transaction flow of each unlabeled account from the transaction flow data; a second extraction unit, which is used to extract the to-be-processed feature data of the corresponding account from each transaction flow, and obtain the to-be-processed feature data of each positive sample account and each unlabeled account; a processing unit, which is used to perform data feature preprocessing on each to-be-processed feature data, and obtain the data feature preprocessing result, and the data feature preprocessing includes outlier processing, missing value processing and binning processing; a third extraction unit, which is used to extract sample features of each positive sample account and the sample features of each unlabeled account from the data feature preprocessing result, and the sample features include user basic features, account basic features and account flow features.

6. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 4.

7. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to make a computer execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Circuit breaker fault diagnosis method based on multi-layer DBN model

    CN110263837A

  • Telecommunication anti-fraud identification method and system based on pseudo tag, and storage medium

    CN116416445A

  • Abnormal account detection model training method and device, electronic equipment and storage medium

    CN118094444A