Abnormal account classification method and device, electronic equipment and storage medium
By optimizing the soft labels of unlabeled account data through semi-supervised machine learning and combining them with the standard labels of labeled data, a target classification model is constructed. This solves the problems of high manpower and material consumption and the influence of false labels in abnormal account identification, and achieves the effect of efficiently identifying account anomalies and specific anomaly types.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INDUSTRIAL AND COMMERCIAL BANK OF CHINA
- Filing Date
- 2023-03-07
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies for identifying abnormal accounts suffer from problems such as high consumption of manpower and resources, inaccurate labeling, and the impact of false labels on identification results. Furthermore, it is difficult to simultaneously identify whether an account is abnormal and the specific type of abnormality.
A semi-supervised machine learning approach is adopted. The soft labels of unlabeled account data are optimized through T rounds of training. Combined with the standard labels of labeled data, a target classification model is constructed to identify whether the account is abnormal and the specific type of abnormality.
It improves the accuracy and practicality of abnormal account identification, and can identify multiple abnormal types at the same time without the need for multiple models, thus reducing the consumption of manpower and material resources and the complexity of models.
Smart Images

Figure CN116304943B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of big data and information security technology, specifically to a method, apparatus, electronic device, and storage medium for classifying abnormal accounts. Background Technology
[0002] In the scenario of identifying abnormal accounts, supervised machine learning or semi-supervised machine learning is generally used to identify abnormal accounts.
[0003] Supervised machine learning typically trains models using pre-labeled anomaly account samples, which requires significant manpower and resources for labeling. Furthermore, because some hidden accounts fail to take action when anomalies occur, inaccurate labeling can negatively impact anomaly detection. Semi-supervised learning, on the other hand, usually converts the semi-supervised learning model into a supervised learning model by pre-assigning pseudo-labels to unlabeled samples, thus achieving anomaly account detection. However, incorrect pseudo-labels can cause the model to learn incorrect information, affecting the anomaly account detection performance. Summary of the Invention
[0004] In view of the above problems, this disclosure provides a method, apparatus, electronic device and storage medium for classifying abnormal accounts.
[0005] According to the first aspect of this disclosure, a method for classifying abnormal accounts is provided, comprising:
[0006] Obtain data for M accounts to be tested, including data used to characterize account transaction features;
[0007] Input M account data into a trained target classification model and output M classification results that match the M account data. The classification results are used to characterize whether the account data is abnormal and the type of abnormality.
[0008] The target classification model is trained on N sample data and N sample data labels through T rounds. The N sample data include M unlabeled account data and (NM) labeled data. The N sample data labels include (NM) standard labels of labeled data and M soft labels of account data. The soft labels are obtained by optimizing the standard labels through T rounds. The standard labels include C anomaly types, T≥1, C≥1, N≥M≥1.
[0009] According to embodiments of this disclosure, the process of determining the target classification model includes:
[0010] Based on N sample data and the labels of N sample data, the initial classification model is trained for T rounds, and the Tth classification model obtained from the Tth round of training is used as the target classification model;
[0011] In each training round, the soft labels of M account data are updated using the standard labels of (NM) labeled data. The soft labels include the probability that the account data belongs to C anomaly types.
[0012] According to embodiments of this disclosure, the initial classification model is trained for T rounds based on N sample data and the labels of the N sample data, and the Tth classification model obtained from the Tth round of training is used as the target classification model, including:
[0013] For the t-th training round, 2≤t≤T, obtain the (t-1)-th classification model obtained during the (t-1)-th training round, and the M (t-1)-th neighbor datasets corresponding to the M account data. Each (t-1)-th neighbor dataset includes the K (t-1)-th labeled data that are most similar to the account data, where K≥1.
[0014] Input M account data and (NM) labeled data into the (t-1)th classification model and output the t1th prediction result dataset. The t1th prediction result dataset includes M prediction results obtained from the M account data in the tth round and the 1st training, and (NM) prediction results obtained from the (NM) labeled data in the tth round and the 1st training.
[0015] Based on the (t-1)th neighbor dataset corresponding to each account data, calculate the t-th soft label corresponding to each account data, obtaining M t-th soft labels corresponding to M account data; and
[0016] Based on the t1 prediction result dataset, M t-th soft labels, and (NM) standard labels, optimize the (t-1)-th classification model until the loss function meets the preset conditions, thus obtaining the t-th classification model.
[0017] According to embodiments of this disclosure, optimizing the (t-1)th classification model based on the t1th prediction result dataset, M tth soft labels, and (NM) standard labels until the loss function satisfies a preset condition, yields the tth classification model, including:
[0018] Based on the t1-th prediction result dataset, M t-th soft labels, and (NM) standard labels, calculate the first loss function value of the (t-1)-th classification model during the t-th round, 1st training process; and
[0019] Optimize the (t-1)th classification model based on the loss function value to obtain the t1th classification model;
[0020] Input M account data and (NM) labeled data into the t1 classification model and output the t2 prediction result dataset. The t2 prediction result dataset includes M prediction results obtained from the M account data in the t round and the second training, and (NM) prediction results obtained from the (NM) labeled data in the t round and the second training.
[0021] The second loss function value is calculated based on the t2 prediction result dataset, M t-th soft labels, and (NM) standard labels. After multiple rounds of training until the loss function value is minimized, the t-th classification model is determined.
[0022] According to embodiments of this disclosure, after determining the t-th classification model, the method further includes:
[0023] Input M account data and (NM) labeled data into the t-th classification model, and output the t-th prediction result dataset;
[0024] Transform the t-th prediction result dataset into a result sequence;
[0025] Based on the result sequence, calculate the t-th label similarity matrix, which represents the similarity between M account data and (NM) labeled data; and
[0026] For each account data, based on the t-th label similarity matrix, select the K labeled data with the highest similarity to each account data from (NM) labeled data, and use the K labeled data as the K t-th labeled data to form the M t-th neighbor datasets corresponding to the M account data.
[0027] According to embodiments of this disclosure, the method further includes:
[0028] For the t-th training round, t=1, the K-nearest neighbor algorithm is used to calculate the K most similar labeled data to each account data, and the K labeled data are used as K initial labeled data to form M initial neighbor datasets corresponding to M account data.
[0029] Based on the initial neighbor dataset corresponding to each account data, calculate the first soft label corresponding to each account data to obtain M first soft labels corresponding to M accounts;
[0030] Train an initial classification model using M account data, M first soft labels, (NM) labeled data, and (NM) standard labels to obtain the first classification model;
[0031] Input M account data points and (NM) labeled data points into the first classification model, and output the first prediction result dataset; and
[0032] Based on the first prediction result dataset and (NM) labeled data, determine M first neighbor datasets.
[0033] According to embodiments of this disclosure, the account data includes at least one of the following: total number of transactions, transaction amount, account opening time, account opening location, device type, and number of transactions in multiple scenarios; tagged data includes account data with identified abnormal types.
[0034] According to embodiments of this disclosure, the loss function of the target classification model includes a labeled loss term, an unlabeled loss term, and a contrastive loss term. The labeled loss term represents the loss of (NM) labeled data, the unlabeled loss term represents the loss of M account data, and the contrastive loss term represents the loss of G account data after filtering, 1≤G≤M. The G account data are obtained by filtering M account data based on (NM) labeled data.
[0035] According to embodiments of this disclosure, the process of determining the target classification model further includes:
[0036] Based on (NM) labeled data, calculate the class center of the anomaly type to obtain C class centers corresponding to C anomaly types;
[0037] Calculate the class similarity between M account data and C class centers to obtain M*C class similarity values; and
[0038] In each of the T training rounds, M*C class similarities are selected based on the soft labels of M account data and the total threshold, and G account data are determined from the M account data in order to calculate the contrastive loss term in the loss function.
[0039] A second aspect of this disclosure provides a classification device for abnormal accounts, comprising:
[0040] The acquisition module is used to acquire data for M accounts to be detected. The account data includes data that characterizes the transaction features of the accounts.
[0041] The classification module is used to input M account data into the trained target classification model and output M classification results that match the M account data. The classification results are used to characterize whether the account data is abnormal and the type of abnormality.
[0042] The target classification model is trained on N sample data and N sample data labels through T rounds. The N sample data include M unlabeled account data and (NM) labeled data. The N sample data labels include (NM) standard labels of labeled data and M soft labels of account data. The soft labels are obtained by optimizing the standard labels through T rounds. The standard labels include C anomaly types, T≥1, C≥1, N≥M≥1.
[0043] A third aspect of this disclosure provides an electronic device comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the aforementioned method for classifying abnormal accounts.
[0044] A fourth aspect of this disclosure also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the aforementioned method for classifying abnormal accounts.
[0045] The fifth aspect of this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method for classifying abnormal accounts.
[0046] The embodiments of this disclosure improve the accuracy of soft labels in the semi-supervised model by optimizing the soft labels of the account data to be detected during the training process of the target classification model. This allows the target classification model to learn more accurate features based on the standard labels and soft labels, thereby improving the accuracy of the target classification model in identifying abnormal accounts. Furthermore, since the standard labels include C anomaly types, the target classification model can simultaneously learn features of multiple anomaly types. This allows it to identify not only whether account data is abnormal but also the specific anomaly type of the abnormal account, eliminating the need to determine multiple anomaly types through multiple models and improving the practicality of abnormal account identification. Attached Figure Description
[0047] The foregoing contents, as well as other objects, features, and advantages of this disclosure, will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0048] Figure 1 This illustration schematically depicts an application scenario of the method for classifying abnormal accounts according to embodiments of the present disclosure;
[0049] Figure 2 A flowchart illustrating a method for classifying abnormal accounts according to an embodiment of this disclosure is shown schematically.
[0050] Figure 3 A flowchart illustrating a method for training a target classification model according to an embodiment of the present disclosure is shown schematically.
[0051] Figure 4 A flowchart illustrating a method for determining a neighbor dataset according to an embodiment of the present disclosure is shown schematically;
[0052] Figure 5 A flowchart illustrating a method for determining a target classification model according to a specific embodiment of the present disclosure is shown schematically.
[0053] Figure 6 A schematic block diagram illustrating a classification apparatus for abnormal accounts according to embodiments of the present disclosure is shown; and
[0054] Figure 7 A block diagram of an electronic device suitable for a classification method of abnormal accounts according to an embodiment of the present disclosure is shown schematically. Detailed Implementation
[0055] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0056] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0057] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0058] When using expressions such as "at least one of A, B, and C", they should generally be interpreted in accordance with the meaning that is commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B, and C, etc.).
[0059] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of data (including but not limited to user personal information) comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and they do not violate public order and good morals.
[0060] In the banking industry, malicious deception against users can be detected through unusual bank card transfers and other suspicious activity. In these scenarios, relevant technologies typically use machine learning to identify abnormal accounts involved in malicious activity. For example, supervised or semi-supervised machine learning methods can be used to train models, which then identify abnormal accounts.
[0061] For supervised machine learning, various factors such as missing transaction information, information asymmetry, numerous transaction operations, and diverse transaction types lead to significant manpower and resources being required for labeling training samples, with poor labeling results. Furthermore, some abnormal accounts, when subjected to fraudulent activities, fail to report abnormal information to banks or other institutions, and may even be marked as legitimate accounts, thus affecting the model's recognition accuracy.
[0062] In semi-supervised machine learning, the cost of labeling operations is typically reduced by training the model with a small number of labeled samples and a large number of unlabeled samples. However, the accuracy of pseudo-labels in semi-supervised machine learning can affect the model's recognition accuracy. For example, even slight errors in pseudo-labels can be amplified by multiple iterations during model training, causing the final classification model to learn incorrect information. This not only affects recognition accuracy but also the model's generalization performance.
[0063] Furthermore, in the scenario of abnormal account identification, there is also the problem of machine learning tasks being too limited. For example, the trained model can only predict whether an account is abnormal, but cannot identify the specific type of abnormality. Alternatively, models can be built to identify specific abnormality types. However, as the number of abnormality types increases, the number of models also increases, leading to increased costs for abnormal account identification and affecting the practicality of the identification model.
[0064] The embodiments of this disclosure provide a method for classifying abnormal accounts, including: acquiring M account data to be detected, the account data including data used to characterize account transaction characteristics; inputting the M account data into a trained target classification model, and outputting M classification results matching the M account data, the classification results being used to characterize whether the account data is abnormal and the type of abnormality; wherein, the target classification model is obtained by training N sample data and the labels of N sample data for T rounds, the N sample data including M unlabeled account data and (NM) labeled data, the labels of the N sample data including (NM) standard labels of labeled data and soft labels of M account data, the soft labels being obtained by optimizing the standard labels for T rounds, the standard labels including C abnormal types, T≥1, C≥1, N≥M≥1.
[0065] Figure 1 The illustration depicts an application scenario of the method for classifying abnormal accounts according to embodiments of the present disclosure.
[0066] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0067] Users can interact with server 105 via network 104 using at least one of the first terminal device 101, second terminal device 102, and third terminal device 103 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, second terminal device 102, and third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0068] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0069] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0070] It should be noted that the abnormal account classification method provided in this embodiment can generally be executed by server 105. Correspondingly, the abnormal account classification device provided in this embodiment can generally be located in server 105. The abnormal account classification method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the abnormal account classification device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0071] For example, the server can obtain M account data to be detected from the first terminal device 101, the second terminal device 102, and the third terminal device 103. After obtaining the M account data, the server inputs the M account data into a trained target classification model and outputs M classification results to determine whether the M accounts matching the M classification results are abnormal and, if the accounts are abnormal, to determine the specific anomaly type. The target classification model is trained on N sample data and N sample data labels through T rounds. The N sample data includes M unlabeled account data and (NM) labeled data. The labels of the N sample data include (NM) standard labels for the labeled data and soft labels for the M account data. The soft labels are obtained by optimizing the standard labels, which include C anomaly types, where T≥1, C≥1, and N≥M≥1.
[0072] According to embodiments of this disclosure, the server can also perform a training process for a target classification model to obtain a trained target classification model. For example, after acquiring M account data to be detected, (NM) labeled data are acquired, and a target classification model is trained based on the M account data, (NM) labeled data, the soft labels of the M account data, and the standard labels of the (NM) labeled data to obtain a trained target classification model. During the training of the target classification model, the server can also optimize the soft labels of the M account data based on the standard labels of the (NM) labeled data.
[0073] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0074] For ease of understanding, the terminology used in this instruction manual will be explained uniformly below.
[0075] Soft labeling: The account data itself is unlabeled data. The labels calculated from the labeled data can be classified into C anomaly types. Therefore, the label value for each anomaly type in the account data is not 0 or 1, but a probability between 0 and 1, hence the term "soft labeling".
[0076] Standard Labels: The anomaly types of labeled data are predetermined, so the label value for each anomaly type in the labeled data is either 0 or 1.
[0077] Neighbor dataset: A collection of labeled data that are most similar to the account data.
[0078] Prediction Results Dataset: A collection of multiple prediction results output by the model.
[0079] Labeled data: Account data with known anomaly types and labeled according to the anomaly types, used to train classification models.
[0080] Figure 2 A flowchart illustrating a method for classifying abnormal accounts according to an embodiment of this disclosure is shown schematically.
[0081] like Figure 2 As shown, the method includes operations S210 to S220.
[0082] In operation S210, data on M accounts to be detected is obtained. The account data includes data used to characterize the transaction features of the accounts.
[0083] According to embodiments of this disclosure, account data includes data used to characterize account transaction features. For example, account data includes account identifiers, which serve as a carrier reflecting account transaction features, such as account number, account name, etc. In applications involving abnormal account identification, based on each piece of account data, it can be determined whether the account corresponding to that account data is abnormal.
[0084] According to embodiments of this disclosure, account data may include at least one of the following: total number of transactions, transaction amount, account opening time, account opening location, device type, and number of transactions in multiple scenarios; the tagged data includes account data with identified abnormal types.
[0085] According to embodiments of this disclosure, each account data may also include other data characterizing the account's transaction features, such as usage frequency, peak transaction periods, etc.
[0086] For example, account data is in the form of {"account identifier", "account opening time"}, where the account identifier can be a string or number of a preset length. The account data for user "Zhang X" is {"111110", "1020"}, which means that user "Zhang X" has an account identifier of 111110 and an account opening time of October 20th.
[0087] In operation S220, M account data are input into the trained target classification model, and M classification results matching the M account data are output. The classification results are used to characterize whether the account data is abnormal and the type of abnormality. The target classification model is trained on N sample data and N sample data labels through T rounds. The N sample data include M unlabeled account data and (NM) labeled data. The N sample data labels include (NM) standard labels of labeled data and M soft labels of account data. The soft labels are obtained by optimizing the standard labels through T rounds. The standard labels include C abnormality types, T≥1, C≥1, N≥M≥1.
[0088] According to embodiments of this disclosure, M account data are input into a trained target classification model. The target classification model analyzes each account data in turn to obtain a classification result corresponding to each account data. The target classification model can not only identify whether the account data is abnormal, but also identify the specific type of abnormality.
[0089] According to embodiments of this disclosure, the target classification model is trained based on a semi-supervised machine learning method. The N sample data used for training include M unlabeled account data to be detected, and (NM) labeled data.
[0090] According to embodiments of this disclosure, the labeled data includes C anomaly types. For example, anomalies in asset origin, asset destination, and transactions at different addresses simultaneously. Labeled data can characterize transaction features of multiple anomaly types, and the trained target classification model can learn these features. As the number of anomaly types increases, there is no need to change the structure or framework of the target classification model; simply adding labeled data is sufficient to ensure the target classification model can identify the increased anomaly types.
[0091] According to embodiments of this disclosure, during the training of the target classification model, since the sample data includes unlabeled account data, before training the target classification model using N sample data, soft labels for M account data can be obtained using (NM) labeled data including multiple anomaly types. Then, based on a semi-supervised machine learning method, the model is trained using the soft labels of the M account data and (NM) labeled data to obtain the trained target classification model. The trained target classification model can identify whether the input account data is abnormal and the type of abnormality based on the input account data.
[0092] The embodiments of this disclosure improve the accuracy of soft labels in the semi-supervised model by optimizing the soft labels of the account data to be detected during the training process of the target classification model. This allows the target classification model to learn more accurate features based on the standard labels and soft labels, thereby improving the accuracy of the target classification model in identifying abnormal accounts. Furthermore, since the standard labels include C anomaly types, the target classification model can simultaneously learn features of multiple anomaly types. This allows it to identify not only whether account data is abnormal but also the specific anomaly type of the abnormal account, eliminating the need to determine multiple anomaly types through multiple models and improving the practicality of abnormal account identification.
[0093] According to an embodiment of this disclosure, the process of determining the target classification model includes: training an initial classification model for T rounds based on N sample data and the labels of N sample data, and using the Tth classification model obtained from the Tth round of training as the target classification model; wherein, in each round of training, the soft labels of M account data are updated using the standard labels of (NM) labeled data, and the soft labels include the probability that the account data belongs to C anomaly types.
[0094] According to embodiments of this disclosure, in each round of training, it is necessary to optimize the classification model and the soft labels. Before optimizing the classification model, the soft labels for M account data are calculated. Then, based on the M account data and their corresponding soft labels, and (NM) labeled data and their corresponding standard labels, the optimal classification model for this round of training is obtained through multiple optimizations.
[0095] According to embodiments of this disclosure, during the T rounds of training the initial classification model to obtain the target classification model, each round of training optimizes the classification model and the soft labels of the M account data used to train the classification model. This ensures that the soft labels are optimized synchronously with the classification model, enabling the classification model to learn more accurate features from more accurate soft labels.
[0096] The embodiments disclosed herein can avoid the problem that the recognition accuracy of the classification model is affected due to the lack of synchronous optimization of soft tags.
[0097] The following will be based on Figure 1 The described scene, through Figures 3-5 The training method of the target classification model of the disclosed embodiments is described in detail.
[0098] According to embodiments of this disclosure, during the training of the target classification model, samples with a large amount of missing data can be deleted by observing the data missing situation to obtain N sample data.
[0099] Figure 3 A flowchart illustrating a training method for a target classification model according to an embodiment of the present disclosure is shown.
[0100] like Figure 3 As shown, the training method of the target classification model in this embodiment includes operations S310 to S340. Operations S310 to S340 can be used as a specific embodiment from the second round of training to the Tth round of training.
[0101] In operation S310, for the t-th round of training, 2≤t≤T, obtain the (t-1)-th classification model obtained in the (t-1)-th round of training and the M (t-1)-th neighbor datasets corresponding to the M account data. Each (t-1)-th neighbor dataset includes the K (t-1)-th labeled data that are most similar to the account data, K≥1.
[0102] According to embodiments of this disclosure, the (t-1)th classification model represents the optimal classification model during the (t-1)th training round. The (t-1)th neighbor dataset represents the neighbor dataset updated based on the (t-1)th classification model during the (t-1)th training round. In the next training round, a soft label corresponding to the account data can be calculated based on the (t-1)th neighbor dataset corresponding to the account data.
[0103] According to embodiments of this disclosure, during T rounds of training, the K labeled data in the neighbor dataset may include one or more anomaly types.
[0104] For example, during the t-th round of training, if the K most similar (t-1)-th labeled data points to account data B include multiple anomaly types, it indicates that account data B may belong to multiple anomaly types, and the classification model can be further optimized. If the K most similar (t-1)-th labeled data points to account data B include only one anomaly type C, it indicates that account data B belongs to anomaly type C, and the classification model can be further optimized, or optimization can be stopped.
[0105] As another specific implementation, for the Tth round of training, after obtaining the (T-1)th classification model and the M (T-1)th neighbor datasets corresponding to the M account data, it can be determined whether the K (T-1)th labeled data in the (T-1)th neighbor dataset belong to the same anomaly type.
[0106] If it is determined that the K (T-1)th labeled data belong to the same anomaly type, continue with the final round of training, i.e., the Tth round. If it is determined that the K (T-1)th labeled data do not belong to the same anomaly type, increase the number of training rounds by a preset number, for example, expand the number of training rounds from T rounds to T+T / 2 rounds.
[0107] In operation S320, M account data and (NM) labeled data are input into the (t-1)th classification model, and the t1th prediction result dataset is output.
[0108] In operation S330, based on the (t-1)th neighbor dataset corresponding to each account data, the tth soft label corresponding to each account data is calculated, resulting in M tth soft labels corresponding to M account data.
[0109] In operation S340, the (t-1)th classification model is optimized based on the t1th prediction result dataset, M tth soft labels, and (NM) standard labels until the loss function meets the preset conditions, thus obtaining the tth classification model.
[0110] According to an embodiment of this disclosure, after determining the optimal classification model obtained in the (t-1)th round of training, the (t-1)th classification model optimized in the (t-1)th round is used to re-predict M account data and (NM) labeled data to obtain a t1th prediction result dataset including N prediction results. The t1th prediction result dataset includes M prediction results obtained from the M account data in the tth round, 1st training, and (NM) prediction results obtained from the (NM) labeled data in the tth round, 1st training.
[0111] According to embodiments of this disclosure, the loss function of the (t-1)th classification model can be calculated based on the t-th prediction result dataset, the t-th soft label of M account data, and the standard label of (NM) labeled data. The (t-1)-th classification model is continuously optimized based on the loss function until the loss function meets preset conditions, resulting in the optimal classification model for the t-th round, i.e., the t-th classification model.
[0112] According to embodiments of this disclosure, before determining the loss function of the model, the soft labels of M account data are calculated based on the (t-1)th neighbor dataset updated in the (t-1)th round, thereby realizing the update of soft labels within the training round t.
[0113] According to embodiments of this disclosure, for tagged data, the standard label y of the i-th tagged data... i satisfy in, This represents the probability that the anomaly type of the standard label of the i-th labeled data belongs to the c-th class.
[0114] Soft labels are used to characterize the probability that account data belongs to C outlier types. The soft label for the i-th account data satisfies... in, This represents the probability that the anomaly type of the soft label of the i-th account data belongs to the c-th class.
[0115] According to embodiments of this disclosure, the neighbor dataset corresponding to the i-th account data is represented as follows: The soft label for the i-th account data can be calculated using the following formula:
[0116]
[0117] in, This represents the number of labeled data points in the neighbor dataset whose anomaly type belongs to the c-th class, and K represents the number of labeled data points in the neighbor dataset.
[0118] According to an embodiment of this disclosure, during t rounds of training, 2≤t≤T, the t-th soft label of M account data can be calculated according to the above formula (1).
[0119] The embodiments of this disclosure utilize the neighbor dataset obtained in the (t-1)th training round to calculate the t-th soft label of the account data, and then use the t-th soft label to optimize the (t-1)-th classification model multiple times to achieve model optimization and obtain the t-th classification model. By optimizing the soft label and the model in each round of training, the prediction accuracy of the classification model improves as the accuracy of the soft label increases.
[0120] According to embodiments of this disclosure, obtaining the t-th classification model includes: calculating the first loss function value of the (t-1)-th classification model during the t-th round of training based on the t1-th prediction result dataset, M t-th soft labels, and (NM) standard labels; optimizing the (t-1)-th classification model based on the loss function value to obtain the t1-th classification model; inputting M account data and (NM) labeled data into the t1-th classification model to output the t2-th prediction result dataset, which includes M prediction results obtained from the M account data during the t-th round of training and (NM) prediction results obtained from the (NM) labeled data during the t-th round of training; calculating the second loss function value based on the t2-th prediction result dataset, M t-th soft labels, and (NM) standard labels, and performing multiple rounds of training until the obtained loss function value is minimized to determine the t-th classification model.
[0121] According to embodiments of this disclosure, during each training round, the parameters of the classification model are adjusted multiple times using standard labels and soft labels updated in the current round to obtain the optimal classification model in the current training round.
[0122] According to embodiments of this disclosure, the preset conditions include obtaining the minimum loss function value.
[0123] During the t-th round of training, after obtaining the t-th soft label, (NM) standard labels, and the t1 prediction result dataset output by the (t-1)-th classification model for each of the M account data, the parameters of the (t-1)-th classification model are adjusted by calculating the first loss function value obtained in the t-th round of training, thus obtaining the t1-th classification model optimized in the t-th round of training.
[0124] Then, input M account data points and (NM) labeled data points into the t1-th classification model, outputting the t2 prediction result dataset. Similar to the t-th round of training, in the t-th round of training, based on the M t-th soft labels, (NM) standard labels, and the t2 prediction result dataset, calculate the second loss function value obtained in the t-th round of training, adjust the parameters of the t1-th classification model, and obtain the t2-th optimized classification model in the t-th round. Continue to execute the t-th round of training, further optimizing the t2-th classification model, until the calculated loss function value is minimized, obtaining the t-th classification model.
[0125] According to embodiments of this disclosure, an optimized classification model can be obtained by minimizing the loss function using a target algorithm. For example, the target algorithm includes stochastic gradient descent.
[0126] According to embodiments of this disclosure, the loss function value changes as the classification model is optimized. The calculation method for the loss function value is similar in each training round.
[0127] According to embodiments of this disclosure, the loss function of the target classification model includes a labeled loss term, an unlabeled loss term, and a contrastive loss term. The labeled loss term represents the loss of (NM) labeled data, the unlabeled loss term represents the loss of M account data, and the contrastive loss term represents the loss of G account data after filtering, 1≤G≤M. The G account data are obtained by filtering M account data based on (NM) labeled data.
[0128] According to embodiments of this disclosure, a labeled loss term can be obtained by calculating the loss between the standard labels of (NM) labeled data and the prediction results of the classification model.
[0129] According to embodiments of this disclosure, the process of calculating the labeled loss term for N sample data satisfies:
[0130]
[0131] in, Indicates a labeled loss term, x i Let x represent the i-th sample. i ∈X L This indicates that the i-th sample belongs to the labeled data, X. L This represents a labeled dataset, containing (NM) labeled data points. C indicates that there are C anomaly types. This represents the probability that the current classification model predicts the anomaly type to be the c-th class. This represents the probability that the anomaly type of the labeled data belongs to the c-th class.
[0132] For example, in the first optimization process of the t-th round, based on the standard labels of (NM) labeled data and the t1-th prediction result dataset, the labeled loss term in the first loss function value is obtained, as shown in formula (2). It can be the prediction result of labeled data in the t1 prediction result dataset; in the second optimization process of the t-th round, based on the standard labels of (NM) labeled data and the t2 prediction result dataset, the second loss function value is obtained, in formula (2) It can be the prediction result for the t2 prediction result dataset that contains labeled data.
[0133] According to embodiments of this disclosure, an unlabeled loss term can be obtained by calculating the loss between the soft labels of M account data and the prediction results of the current classification model.
[0134] According to embodiments of this disclosure, the process of calculating the unlabeled loss term satisfies:
[0135]
[0136] in, x represents the unlabeled loss term. i Let x represent the i-th sample. i ∈X U This indicates that the i-th sample belongs to account data, X U This represents an account dataset, containing data from M accounts. C indicates that there are C exception types. This represents the probability that the current classification model predicts the anomaly type to be the c-th class. This represents the probability that the anomaly type of the account data belongs to the c-th class.
[0137] For example, in the first optimization process of the t-th round, based on the t-th soft label and the t1-th prediction result dataset of M account data, the unlabeled loss term in the first loss function value is obtained, as shown in formula (3). It can be the prediction result of the account data in the t1 prediction result dataset; in the second optimization process of the t-th round, based on the t-th soft label of M account data and the t2 prediction result dataset, the unlabeled loss term in the second loss function value is obtained, in formula (3) It can be the prediction result of the account data in the t2 prediction result dataset.
[0138] According to embodiments of this disclosure, a comparative loss term is obtained by calculating the loss between the G account data and the center of each anomaly category based on the standard labels of the filtered G account data and (NM) labeled data.
[0139] According to embodiments of this disclosure, the process of calculating the comparison loss term satisfies:
[0140]
[0141] Among them, L CL x represents the comparison loss term. i This represents the i-th sample. This indicates that the i-th sample belongs to the filtered account data. This represents the filtered account dataset, including G filtered account data. C indicates that there are C anomaly types, c j and c c Let represent the class centers of the j-th class and the c-th class, respectively, in (NM) labeled data points. τ represents a hyperparameter, which can be determined during the training process.
[0142] According to embodiments of this disclosure, the loss function satisfies:
[0143]
[0144] Here, α and β are hyperparameters that adjust the weights of the unlabeled loss term and the contrastive loss term, and can be determined according to the actual situation.
[0145] For example, if the t-th classification model is obtained through P optimizations in each training round, then in the T training rounds, a total of P calculations are required. T The loss function value is p > 1. Each time the loss function value is calculated, the labeled loss term, the unlabeled loss term, and the contrastive loss term need to be calculated.
[0146] The embodiments of this disclosure employ a loss function that includes a labeled loss term and an unlabeled loss term to fully utilize the structural information of unlabeled account data and labeled data in the feature space to optimize the classification model. The loss function also constructs a contrastive loss term by selecting account data with high confidence levels, reducing the adverse effects of noise points on the classification model, improving the model's generalization ability, and enhancing the accuracy of the classification model in identifying abnormal accounts and determining anomaly types.
[0147] It should be noted that, in the embodiments of this disclosure, since the target classification model can identify multiple anomaly types, improving the generalization ability of the model can ensure the accuracy of identifying multiple anomaly types using the same target classification model.
[0148] According to embodiments of this disclosure, in determining the target classification model, a loss function needs to be calculated, and the loss function includes a contrastive loss term. Therefore, the training process further includes: calculating class centers for anomaly types based on (NM) labeled data, obtaining C class centers corresponding to C anomaly types; calculating class similarities between M account data and the C class centers, obtaining M*C class similarities; and in each of the T training rounds, filtering the M*C class similarities based on the soft labels and a total threshold of the M account data, determining G account data from the M account data for calculating the contrastive loss term in the loss function.
[0149] According to embodiments of this disclosure, before calculating the contrastive loss term within the loss function, class centers for each of the C anomaly types can be calculated based on (NM) labeled data.
[0150] When the exception type is the c-th class, the process of calculating the class center of the c-th class satisfies:
[0151]
[0152] Among them, c c This indicates the class center of the c-th class, x. i This represents the i-th sample. This indicates that the i-th sample is labeled data, and the anomaly type of this labeled data is the c-th class. This represents the number of labeled data in the c-th class.
[0153] According to embodiments of this disclosure, for C exception types, each class center can be calculated using formula (6).
[0154] According to embodiments of this disclosure, after determining the class center for each anomaly type, the class similarity between M account data and the class center for each anomaly type is calculated. Specifically, the class similarity between the account data and the class center for the j-th anomaly type satisfies:
[0155]
[0156] Where, x i Let x represent the i-th sample. i ∈X U This indicates that the i-th sample belongs to account data, X U This represents an account dataset, which includes data for M accounts. Let τ represent the class similarity between the class center of the i-th sample and the j-th anomaly type, and let τ represent the hyperparameter, which can be determined according to the training process.
[0157] According to embodiments of this disclosure, class similarity with C class centers is calculated for M account data, resulting in M*C class similarity scores.
[0158] According to an embodiment of this disclosure, during the t-th training process, if the class similarity between the account data and the j-th anomaly type is greater than the total threshold, and the t-th soft label of the account data indicates that the account data belongs to the j-th anomaly type, then the account data is selected as one of the G filtered account data.
[0159] According to embodiments of this disclosure, the total threshold is 1 / C, where C represents the number of abnormal types.
[0160] For example, if the class similarity between account data A and the class center of the j-th anomaly type is greater than 1 / C, and Therefore, account data A belongs to
[0161] According to embodiments of this disclosure, after selecting G account data, the account data can be paired with the determined class centers to form positive sample pairs for calculating the contrast loss term.
[0162] For example, for a certain account data x in the filtered G account data i , transfer account data x i The class center c of the j-th class j Construct positive sample pairs (x) i ,c j Then, construct the comparison loss term according to formula (4).
[0163] The embodiments of this disclosure construct contrastive learning terms by screening account data with high confidence through two parts: threshold comparison and soft labeling. This is to narrow the distance between the class centers of high-confidence account data and labeled data with the same label in the loss function, while widening the distance between the class centers of high-confidence account data and labeled data with different labels. This improves the feature learning accuracy of the classification model, thereby increasing the accuracy of identifying abnormal accounts and abnormal types.
[0164] Figure 4 A flowchart illustrating a method for determining a neighbor dataset according to an embodiment of this disclosure is shown schematically.
[0165] like Figure 4 As shown, the method for determining the neighbor dataset in this embodiment includes operations S410 to S440. For the t-th training round, operations S410 to S440 can occur after the t-th classification model has been determined.
[0166] In operation S410, M account data and (NM) labeled data are input into the t-th classification model, and the t-th prediction result dataset is output.
[0167] In operation S420, the t-th prediction result dataset is converted into a result sequence.
[0168] In operation S430, based on the result sequence, the t-th label similarity matrix is calculated. The t-th label similarity matrix represents the similarity between M account data and (NM) labeled data.
[0169] In operation S440, for each account data, based on the t-th label similarity matrix, select the K labeled data with the highest similarity to each account data from (NM) labeled data, and use the K labeled data as the K t-th labeled data to form the M t-th neighbor datasets corresponding to the M account data.
[0170] According to embodiments of this disclosure, after obtaining the optimal classification model for round t, M account data and (NM) labeled data are input into the classification model to obtain the prediction result dataset t, which includes N prediction results. After determining the prediction result dataset t, the prediction result dataset t is converted into a result sequence. This is so that a label similarity matrix can be calculated based on the resulting sequence.
[0171] According to embodiments of this disclosure, the process of calculating the label similarity matrix satisfies:
[0172]
[0173] in, This represents the label similarity matrix, where each element in the label similarity matrix is a label similarity matrix. This represents the probability that the i-th sample and the j-th sample belong to the same anomaly type, where the i-th sample and the j-th sample can be labeled data or account data. The sigmoid() function is used to map the probability to the range (0,1).
[0174] According to an embodiment of this disclosure, during the t-th round of training, the t-th label similarity matrix can be calculated according to formula (8).
[0175] According to embodiments of this disclosure, for each account data, during each round of training, K most similar labeled data are selected for each account data based on the label similarity matrix, and these are used as the updated neighbor dataset, so that the soft label of the account data can be determined through the updated neighbor dataset in the next round of training.
[0176] According to embodiments of this disclosure, in each training round of T rounds, after determining the optimal classification model for the current round, the neighbor dataset for the current round can be updated through operations similar to operations S410 to S440, so as to update the soft labels of the account data in the next training round.
[0177] The embodiments of this disclosure, after determining the optimal classification model for round t, optimize the neighbor dataset during each training round by calculating the label similarity matrix for round t, so as to improve the accuracy of the soft labels in the next round.
[0178] According to embodiments of this disclosure, for the t-th training round (t=1), the K-nearest neighbor algorithm is used to calculate the K most similar labeled data to each account data, and these K labeled data are used as K initial labeled data to form M initial neighbor datasets corresponding to M account data. Based on the initial neighbor datasets corresponding to each account data, the first soft label corresponding to each account data is calculated to obtain M first soft labels corresponding to the M accounts. An initial classification model is trained using the M account data, M first soft labels, (NM) labeled data, and (NM) standard labels to obtain a first classification model. The M account data and (NM) labeled data are input into the first classification model to output a first prediction result dataset. And based on the first prediction result dataset and (NM) labeled data, the M first neighbor datasets are determined.
[0179] According to embodiments of this disclosure, during the first round of training, before optimizing the initial classification model, soft labels for M account data are calculated based on (NM) labeled data. Specifically, the K nearest neighbors algorithm can be used to determine the K most similar labeled data to each account data from the (NM) labeled data, forming M initial neighbor datasets corresponding to the M account data.
[0180] In the first round of training, the first soft label for each account is calculated using the initial neighbor dataset, resulting in M first soft labels corresponding to M accounts. An initial classification model is trained using the M account data, M first soft labels, (NM) labeled data, and (NM) standard labels to obtain the first classification model. Then, after obtaining the first classification model, the first neighbor dataset is calculated. The method for obtaining the first classification model is similar to the method for obtaining the t-th classification model, and the method for obtaining the first neighbor dataset is similar to the method for obtaining the t-th neighbor dataset; therefore, it will not be repeated here.
[0181] Figure 5 A flowchart illustrating a method for determining a target classification model according to a specific embodiment of the present disclosure is shown.
[0182] like Figure 5As shown, the sample data includes labeled accounts 501 and unlabeled accounts 502. The number of labeled accounts 501 is far less than the number of unlabeled accounts 502. During the first round of training, the K-nearest neighbor algorithm is used to determine the neighbor dataset 503 corresponding to each unlabeled account 502 from the labeled accounts 501. Based on the neighbor dataset 503, soft labels corresponding to the unlabeled accounts 502 are calculated, resulting in soft-labeled accounts 504. On one hand, soft-labeled accounts 504 can be used to construct the unlabeled loss term 507; on the other hand, after processing, soft-labeled accounts 504 can be used to construct the contrastive loss term 508.
[0183] By calculating the class similarity between soft-labeled account 504 and the class centers of multiple anomaly types, high-confidence soft-labeled accounts are selected from soft-labeled accounts 504 to obtain the selected soft-labeled accounts 505. The selected soft-labeled accounts 505 are used to construct the contrastive loss term 508.
[0184] The loss function of the classification model includes a labeled loss term 506, an unlabeled loss term 507, and a contrastive loss term 508.
[0185] The loss function value of the classification model can be determined based on the labeled loss term 506, the unlabeled loss term 507, and the contrastive loss term 508. The classification model can be continuously optimized based on the calculated loss function value to obtain the optimal classification model 509 for this round of training. After determining the optimal classification model 509 for this round, labeled accounts 501 and unlabeled accounts 502 are input into the classification model 509 to obtain the prediction result dataset 510. Based on the prediction result dataset 510 for this round of training, the label similarity matrix 511 for this round can be calculated. The neighbor dataset for this round is then calculated based on the label similarity matrix 511.
[0186] In the next round of training, information such as soft-labeled accounts and filtered soft-labeled accounts are recalculated based on the neighbor dataset from the previous round to further optimize the classification model. After T rounds of training, a well-trained target classification model can be obtained.
[0187] The target classification model obtained according to the embodiments of this disclosure outperforms traditional machine learning algorithms in terms of accuracy, recall, and overall evaluation score for identifying abnormal accounts, and can more accurately predict whether an account is involved in anomalies and the specific type of anomaly.
[0188] Figure 6 A schematic block diagram of a classification device for abnormal accounts according to an embodiment of the present disclosure is shown.
[0189] like Figure 6 As shown, the abnormal account classification device 600 of this embodiment includes an acquisition module 610 and a classification module 620.
[0190] The acquisition module 610 is used to acquire data on M accounts to be detected, including data characterizing account transaction features. In one embodiment, the acquisition module 610 can be used to perform the operation S210 described above, which will not be repeated here.
[0191] The classification module 620 is used to input M account data points into a pre-trained target classification model and output M classification results that match the M account data points. These classification results characterize whether the account data is abnormal and the type of abnormality. The target classification model is trained on N sample data points and their labels over T rounds. The N sample data points include M unlabeled account data points and (NM) labeled data points. The labels of the N sample data points include standard labels for (NM) labeled data points and soft labels for the M account data points. The soft labels are obtained by optimizing the standard labels over T rounds. The standard labels include C abnormality types, where T≥1, C≥1, and N≥M≥1. In one embodiment, the classification module 620 can be used to perform the operation S220 described above, which will not be repeated here.
[0192] According to embodiments of this disclosure, the abnormal account classification device 600 further includes a training module.
[0193] The training module is used to determine the target classification model, including: training the initial classification model for T rounds based on N sample data and the labels of N sample data, and using the Tth classification model obtained from the Tth round of training as the target classification model; wherein, in each round of training, the soft labels of M account data are updated using the standard labels of (NM) labeled data, and the soft labels include the probability that the account data belongs to C anomaly types.
[0194] According to embodiments of this disclosure, the training module includes a first training unit, a second training unit, a third training unit, and a fourth training unit.
[0195] The first training unit is used for training in the t-th round, where 2≤t≤T, to obtain the (t-1)-th classification model obtained during the (t-1)-th training round, and the M (t-1)-th neighbor datasets corresponding to the M account data. Each (t-1)-th neighbor dataset includes the K (t-1)-th labeled data points most similar to the account data, where K≥1. In one embodiment, the first training unit can be used to perform the operation S310 described above, which will not be repeated here.
[0196] The second training unit is used to input M account data and (NM) labeled data into the (t-1)th classification model and output the t1th prediction result dataset. The t1th prediction result dataset includes M prediction results obtained from the M account data in the tth round, 1st training, and (NM) prediction results obtained from the (NM) labeled data in the tth round, 1st training. In one embodiment, the second training unit can be used to perform the operation S320 described above, which will not be repeated here.
[0197] The third training unit is used to calculate the t-th soft label corresponding to each account data based on the (t-1)-th neighbor dataset corresponding to each account data, thereby obtaining M t-th soft labels corresponding to M account data. In one embodiment, the third training unit can be used to perform the operation S330 described above, which will not be repeated here.
[0198] The fourth training unit is used to optimize the (t-1)th classification model based on the t1th prediction result dataset, M tth soft labels, and (NM) standard labels until the loss function meets the preset conditions, thus obtaining the tth classification model. In one embodiment, the fourth training unit can be used to perform the operation S340 described above, which will not be repeated here.
[0199] According to embodiments of this disclosure, the fourth training unit includes a first training subunit, a second training subunit, a third training subunit, and a fourth training subunit.
[0200] The first training subunit is used to calculate the first loss function value of the (t-1)th classification model in the t-th round of training, based on the t1-th prediction result dataset, M t-th soft labels, and (NM) standard labels.
[0201] The second training subunit is used to optimize the (t-1)th classification model based on the loss function value to obtain the t1th classification model.
[0202] The third training subunit is used to input M account data and (NM) labeled data into the t1 classification model and output the t2 prediction result dataset. The t2 prediction result dataset includes M prediction results obtained from the M account data in the t round and the second training, and (NM) prediction results obtained from the (NM) labeled data in the t round and the second training.
[0203] The fourth training subunit is used to calculate the second loss function value based on the t2 prediction result dataset, M t-th soft labels, and (NM) standard labels. After multiple rounds of training until the loss function value is minimized, the t-th classification model is determined.
[0204] According to embodiments of this disclosure, the training module further includes a fifth training unit, a sixth training unit, a seventh training unit, and an eighth training unit.
[0205] The fifth training unit is used to input M account data and (NM) labeled data into the t-th classification model and output the t-th prediction result dataset. In one embodiment, the fifth training unit can be used to perform the operation S410 described above, which will not be repeated here.
[0206] The sixth training unit is used to convert the t-th prediction result dataset into a result sequence. In one embodiment, the sixth training unit can be used to perform the operation S420 described above, which will not be repeated here.
[0207] The seventh training unit is used to calculate the t-th label similarity matrix based on the result sequence. The t-th label similarity matrix represents the similarity between M account data and (NM) labeled data. In one embodiment, the seventh training unit can be used to perform the operation S430 described above, which will not be repeated here.
[0208] The eighth training unit is used to select, for each account data, the K labeled data with the highest similarity to each account data from (NM) labeled data according to the t-th label similarity matrix, and use the K labeled data as the K t-th labeled data to form the M t-th neighbor datasets corresponding to the M account data. In one embodiment, the eighth training unit can be used to perform the operation S440 described above, which will not be repeated here.
[0209] According to embodiments of this disclosure, the training module further includes an initial training unit. The initial training unit includes a first initial training subunit, a second initial training subunit, a third initial training subunit, a fourth initial training subunit, and a fifth initial training subunit.
[0210] The first initial training subunit is used for training in round t, t=1. It uses the K nearest neighbor algorithm to calculate the K most similar labeled data to each account data, and uses the K labeled data as the K initial labeled data to form the M initial neighbor datasets corresponding to the M account data.
[0211] The second initial training subunit is used to calculate the first soft label corresponding to each account data based on the initial neighbor dataset corresponding to each account data, so as to obtain M first soft labels corresponding to M accounts.
[0212] The third initial training subunit is used to train an initial classification model using M account data, M first soft labels, (NM) labeled data, and (NM) standard labels to obtain the first classification model.
[0213] The fourth initial training subunit is used to input M account data and (NM) labeled data into the first classification model and output the first prediction result dataset.
[0214] The fifth initial training subunit is used to determine M first neighbor datasets based on the first prediction result dataset and (NM) labeled data.
[0215] According to embodiments of this disclosure, account data includes at least one of the following: total number of transactions, transaction amount, account opening time, account opening location, device type, and number of transactions in multiple scenarios; tagged data includes account data with identified abnormal types.
[0216] According to embodiments of this disclosure, the loss function of the target classification model includes a labeled loss term, an unlabeled loss term, and a contrastive loss term. The labeled loss term represents the loss of (NM) labeled data, the unlabeled loss term represents the loss of M account data, and the contrastive loss term represents the loss of G account data after filtering, 1≤G≤M. The G account data are obtained by filtering M account data based on (NM) labeled data.
[0217] According to embodiments of this disclosure, the training module further includes a filtering unit. The filtering unit includes a first filtering subunit, a second filtering subunit, and a third filtering subunit.
[0218] The first filtering subunit is used to calculate the class center of the anomaly type based on (NM) labeled data, and obtain C class centers corresponding to C anomaly types.
[0219] The second filtering subunit is used to calculate the class similarity between M account data and C class centers, resulting in M*C class similarity scores.
[0220] The third screening subunit is used to screen M*C class similarities based on the soft labels of M account data and the total threshold during each of the T rounds of training, and to determine G account data from the M account data in order to calculate the contrastive loss term in the loss function.
[0221] According to embodiments of this disclosure, any plurality of modules in the acquisition module 610 and the classification module 620 may be combined into one module, or any one of these modules may be split into multiple modules. Alternatively, at least a portion of the functionality of one or more of these modules may be combined with at least a portion of the functionality of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the acquisition module 610 and the classification module 620 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in any one of software, hardware, and firmware methods, or in a suitable combination of any of these. Alternatively, at least one of the acquisition module 610 and the classification module 620 may be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.
[0222] Figure 7 A block diagram of an electronic device suitable for a classification method of abnormal accounts according to an embodiment of the present disclosure is shown schematically.
[0223] like Figure 7 As shown, an electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage portion 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0224] RAM 703 stores various programs and data required for the operation of electronic device 700. Processor 701, ROM 702, and RAM 703 are interconnected via bus 704. Processor 701 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 702 and / or RAM 703. It should be noted that the programs may also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.
[0225] According to embodiments of this disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to a bus 704. The electronic device 700 may also include one or more of the following components connected to the I / O interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.
[0226] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0227] According to embodiments of this disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, such as, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703 described above.
[0228] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to enable the computer system to implement the abnormal account classification method provided in embodiments of this disclosure.
[0229] When the computer program is executed by the processor 701, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0230] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 709, and / or installed from a removable medium 711. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0231] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, it performs the functions defined in the system of this disclosure embodiment. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0232] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0233] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0234] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0235] The specific embodiments described above further illustrate the purpose, technical solutions, and beneficial effects of this disclosure. It should be understood that the above descriptions are merely specific embodiments of this disclosure and are not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.
Claims
1. A method for classifying abnormal accounts, comprising: Acquire data for M accounts to be detected, wherein the account data includes data characterizing the transaction features of the accounts; The M account data are input into the trained target classification model, and the model outputs M classification results that match the M account data. The classification results are used to characterize whether the account data is abnormal and the type of abnormality. The target classification model is obtained by training N sample data and the labels of the N sample data for T rounds. The Tth classification model obtained in the Tth round of training is the target classification model. The N sample data includes M unlabeled account data and (NM) labeled data. The labels of the N sample data include the standard labels of the (NM) labeled data and the soft labels of the M account data. In each round of training, the soft labels of the M account data are updated using the standard labels of the (NM) labeled data. The standard labels include C anomaly types, and the soft labels include the probability that the account data belongs to the C anomaly types, where T≥1, C≥1, N>M≥1. The process of determining the target classification model includes: For the t-th training round, 2≤t≤T, obtain the (t-1)-th classification model obtained during the (t-1)-th training round, and the M (t-1)-th neighbor datasets corresponding to the M account data, wherein each (t-1)-th neighbor dataset includes the K (t-1)-th labeled data that are most similar to the account data, K≥1; The M account data and the (NM) labeled data are input into the (t-1)th classification model, and the t1th prediction result dataset is output. The t1th prediction result dataset includes the M prediction results obtained from the M account data in the tth round and the 1st training, and the (NM) prediction results obtained from the (NM) labeled data in the tth round and the 1st training. Based on the (t-1)th neighbor dataset corresponding to each account data, calculate the t-th soft label corresponding to each account data, and obtain M t-th soft labels corresponding to the M account data; and The (t-1)th classification model is optimized based on the t1th prediction result dataset, the M tth soft labels, and (NM) standard labels until the loss function meets the preset conditions, thus obtaining the tth classification model.
2. The method according to claim 1, wherein, The step of optimizing the (t-1)th classification model based on the t1th prediction result dataset, the M tth soft labels, and (NM) standard labels until the loss function satisfies a preset condition, to obtain the tth classification model, includes: Based on the t1-th prediction result dataset, the M t-th soft labels, and (NM) standard labels, calculate the first loss function value of the (t-1)-th classification model during the t-th round, 1st training process; and The (t-1)th classification model is optimized based on the loss function value to obtain the t1th classification model; The M account data and the (NM) labeled data are input into the t1 classification model, and the t2 prediction result dataset is output. The t2 prediction result dataset includes the M prediction results obtained from the M account data in the t round and the second training, and the (NM) prediction results obtained from the (NM) labeled data in the t round and the second training. The second loss function value is calculated based on the t2 prediction result dataset, the M t-th soft labels, and (NM) standard labels. After multiple rounds of training until the loss function value is minimized, the t-th classification model is determined.
3. The method according to claim 1, further comprising, after determining the t-th classification model: Input the M account data and the (NM) labeled data into the t-th classification model, and output the t-th prediction result dataset; Convert the t-th prediction result dataset into a result sequence; Based on the result sequence, calculate the t-th label similarity matrix, which represents the similarity between the M account data and the (NM) labeled data; as well as For each account data, based on the t-th label similarity matrix, select the K labeled data with the highest similarity to each account data from the (NM) labeled data, and use the K labeled data as the K t-th labeled data to form the M t-th neighbor datasets corresponding to the M account data.
4. The method according to claim 1, further comprising: For the t-th training round, t=1, the K-nearest neighbor algorithm is used to calculate the K most similar labeled data to each account data, and the K labeled data are used as K initial labeled data to form M initial neighbor datasets corresponding to the M account data. Based on the initial neighbor dataset corresponding to each account data, calculate the first soft label corresponding to each account data to obtain M first soft labels corresponding to M accounts; An initial classification model is trained using the M account data, the M first soft labels, the (NM) labeled data, and the (NM) standard labels to obtain the first classification model; Input the M account data and the (NM) labeled data into the first classification model, and output the first prediction result dataset; as well as Based on the first prediction result dataset and the (NM) labeled data, determine M first neighbor datasets.
5. The method according to claim 1, wherein, The account data includes at least one of the following: total number of transactions, transaction amount, account opening time, account opening location, device type, and number of transactions in multiple scenarios; the tagged data includes account data with identified abnormal types.
6. The method according to claim 1, wherein the loss function of the target classification model includes a labeled loss term, an unlabeled loss term, and a contrastive loss term, wherein, The labeled loss term represents the loss of the (NM) labeled data, the unlabeled loss term represents the loss of the M account data, and the comparative loss term represents the loss of the filtered G account data, 1≤G≤M, wherein the G account data are obtained by filtering the M account data based on the (NM) labeled data.
7. The method according to claim 6, wherein, The process of determining the target classification model also includes: Based on the (NM) labeled data, calculate the class center of the anomaly type to obtain C class centers corresponding to the C anomaly types; Calculate the class similarity between the M account data and the C class centers to obtain M*C class similarity values; and In each of the T rounds of training, the M*C class similarities are filtered based on the soft labels and total threshold of the M account data, and the G account data are determined from the M account data in order to calculate the contrastive loss term in the loss function.
8. A classification device for abnormal accounts, comprising: The acquisition module is used to acquire data of M accounts to be detected, wherein the account data includes data used to characterize the transaction features of the accounts; The classification module is used to input the M account data into a trained target classification model and output M classification results that match the M account data. The classification results are used to characterize whether the account data is abnormal and the type of abnormality. The target classification model is obtained by training N sample data and the labels of the N sample data for T rounds. The Tth classification model obtained in the Tth round of training is the target classification model. The N sample data includes M unlabeled account data and (NM) labeled data. The labels of the N sample data include the standard labels of the (NM) labeled data and the soft labels of the M account data. In each round of training, the soft labels of the M account data are updated using the standard labels of the (NM) labeled data. The standard labels include C anomaly types, and the soft labels include the probability that the account data belongs to the C anomaly types, where T≥1, C≥1, N>M≥1. The process of determining the target classification model includes: For the t-th training round, 2≤t≤T, obtain the (t-1)-th classification model obtained during the (t-1)-th training round, and the M (t-1)-th neighbor datasets corresponding to the M account data, wherein each (t-1)-th neighbor dataset includes the K (t-1)-th labeled data that are most similar to the account data, K≥1; The M account data and the (NM) labeled data are input into the (t-1)th classification model, and the t1th prediction result dataset is output. The t1th prediction result dataset includes the M prediction results obtained from the M account data in the tth round and the 1st training, and the (NM) prediction results obtained from the (NM) labeled data in the tth round and the 1st training. Based on the (t-1)th neighbor dataset corresponding to each account data, calculate the t-th soft label corresponding to each account data, and obtain M t-th soft labels corresponding to the M account data; and The (t-1)th classification model is optimized based on the t1th prediction result dataset, the M tth soft labels, and (NM) standard labels until the loss function meets the preset conditions, thus obtaining the tth classification model.
9. An electronic device, comprising: One or more processors; Storage device for storing one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 7.
11. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 7.