Abnormal transaction identification model training method and abnormal transaction identification method

By building a deep forest model of multi-level taxonomies, training the decision tree step by step, and generating a second sample set based on the error samples of the previous decision tree for training, the problem of low recognition accuracy of credit card abnormal transactions in the existing technology is solved, and higher recognition accuracy and model stability are achieved.

CN114565040BActive Publication Date: 2025-05-02INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210190183.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-05-02
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

In the prior art, the identification accuracy of credit card abnormal transactions is low, mainly due to the unbalanced data set, the recognition effect of the random forest model is poor.

Method used

A deep forest model of multi-level taxonomies is adopted, and a decision tree in each taxonomy is trained step by step, and a second sample set is generated based on the error samples and training sample sets of the previous decision tree for training to improve the recognition accuracy of abnormal transactions.

Benefits of technology

Through this method, the identification accuracy of abnormal transactions can be significantly improved, the processing ability of unbalanced data sets can be improved, and the robustness and generalization ability of the model can be enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114565040B_ABST
    Figure CN114565040B_ABST
Patent Text Reader

Abstract

The present application relates to a method and device for training an abnormal transaction identification model, an abnormal transaction identification method and device, a computer device, a storage medium and a computer program product, and relates to the field of artificial intelligence technology. The method comprises: obtaining a training sample set corresponding to a first type classification unit; training the first decision tree in the first type classification unit by using the training sample set; for each decision tree except the first decision tree, obtaining a second sample set based on the sample of the abnormal transaction type misclassified by the previous decision tree and the training sample set, and training the decision tree; determining a new training sample set based on the classification result of the training sample set by the trained first type classification unit and the training sample set, and returning to execute the training of the first decision tree in the first type classification unit until the preset training end condition is reached. The abnormal transaction identification model obtained by the present method can improve the accuracy of identifying abnormal transactions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an abnormal transaction identification model training method and device, an abnormal transaction identification method and device, a computer device, a storage medium, and a computer program product. Background Art

[0002] Credit card fraud refers to stealing someone else's identity information to impersonate a credit card, thereby maliciously overdrawing, cashing out, and stealing. Credit card anti-fraud detection is to identify the risks of credit card transactions in order to prevent abnormal transactions such as credit card fraud in a timely manner. It is one of the keys to controlling risks in the financial industry.

[0003] In the related art, credit card anti-fraud detection is to classify credit card transaction data through a classification model to identify whether it is an abnormal transaction or a normal transaction, wherein the classification model is generally a trained random forest model. However, in real business scenarios, the sample size of abnormal credit card transactions is far less than that of normal transactions, which is an unbalanced data set, and the accuracy of identifying abnormal transactions using a random forest model is low. Summary of the invention

[0004] Based on this, it is necessary to provide an abnormal transaction identification model training method and device, abnormal transaction identification method and device, computer equipment, storage medium and computer program product that can improve the accuracy of identifying abnormal transactions in response to the above technical problems.

[0005] In a first aspect, the present application provides a method for training an abnormal transaction identification model. The abnormal transaction identification model is a deep forest model including multiple levels of classification units, each level of the classification unit includes a first type of classification unit, and the method includes:

[0006] Obtaining a training sample set corresponding to the first type classification unit to be trained; the training sample set includes samples of abnormal transaction types and samples of normal transaction types;

[0007] Training a first decision tree in the first type classification unit to be trained by using the training sample set;

[0008] For each decision tree other than the first decision tree in the first type classification unit to be trained, a second sample set corresponding to the decision tree is obtained according to samples of abnormal transaction types misclassified by a previous decision tree and the training sample set, and the decision tree is trained according to the second sample set to obtain a trained first type classification unit;

[0009] According to the classification result of the trained first-type classification unit on the training sample set and the training sample set, a new training sample set is determined, and the first-type classification unit included in the next-level classification unit is used as the first-type classification unit to be trained, and the training of the first decision tree in the first-type classification unit to be trained by the training sample set is returned to be executed until the preset training end condition is reached, so as to obtain the abnormal transaction recognition model.

[0010] In one embodiment, each level of the classification unit further includes a second type of classification unit, and the second type of classification unit is a random forest classification unit; the method further includes:

[0011] The second type classification unit to be trained is trained by using the training sample set to obtain the trained second type classification unit; the first type classification unit to be trained and the second type classification unit to be trained belong to the same level of classification units;

[0012] The step of determining a new training sample set based on the classification result of the training sample set by the trained first type classification unit and the training sample set comprises:

[0013] A new training sample set is determined according to the classification result of the trained first type classification unit on the training sample set, the classification result of the trained second type classification unit on the training sample set, and the training sample set.

[0014] In one embodiment, obtaining a second sample set corresponding to the decision tree based on samples of abnormal transaction types misclassified by a previous decision tree and the training sample set includes:

[0015] Determine the sample of abnormal transaction type misclassified by the previous decision tree as the first target sample, and use the K nearest neighbor algorithm to determine the nearest neighbor sample corresponding to the first target sample in the training sample set;

[0016] Synthesize the first target sample and the neighboring sample to obtain a second target sample;

[0017] A second sample set corresponding to the decision tree is obtained according to the first target sample, the second target sample, and the training sample set; the second sample set includes the first target sample and the second target sample.

[0018] In one embodiment, the method further comprises:

[0019] Determine an initial weight of each sample in the training sample set; wherein the initial weight of the sample of abnormal transaction type in the training sample set is greater than the initial weight of the sample of normal transaction type;

[0020] The step of training the first decision tree in the first type classification unit to be trained by using the training sample set includes:

[0021] Sampling is performed according to the initial weight of each sample in the training sample set, and the first decision tree in the first type classification unit to be trained is trained by using the samples obtained by the sampling.

[0022] In one embodiment, determining the initial weight of each sample in the training sample set includes:

[0023] Counting the number of samples of abnormal transaction types and the number of samples of normal transaction types in the training sample set;

[0024] The inverse of the number of samples of the abnormal transaction type is determined as the initial weight of the samples of the abnormal transaction type in the training sample set, and the inverse of the number of samples of the normal transaction type is determined as the initial weight of the samples of the normal transaction type in the training sample set.

[0025] In a second aspect, the present application also provides a method for identifying abnormal transactions. The method comprises:

[0026] Preprocessing the transaction data to be identified to obtain target feature data; the transaction data to be identified includes cardholder identity information and transaction behavior information;

[0027] The target feature data is input into an abnormal transaction identification model for decision classification to obtain a transaction type corresponding to the transaction data to be identified; the transaction type includes an abnormal transaction type and a normal transaction type; wherein the abnormal transaction identification model is trained by the abnormal transaction identification model training method described in the first aspect.

[0028] In one embodiment, the abnormal transaction identification model comprises multiple levels of classification units, each level of the classification units comprises a first type of classification unit and a second type of classification unit; the abnormal transaction identification model further comprises an output layer;

[0029] The step of inputting the target feature data into an abnormal transaction identification model for decision classification to obtain a transaction type corresponding to the transaction data to be identified includes:

[0030] Input the target feature data into an abnormal transaction recognition model to obtain a first-class vector and a second-class vector; wherein the first-class vector is a class vector output by a first-type classification unit included in the last-level classification unit in the multi-level classification unit, and the second-class vector is a class vector output by a second-type classification unit included in the last-level classification unit;

[0031] By using the output layer included in the abnormal transaction recognition model, the first category vector and the second category vector are averaged to obtain a target category vector;

[0032] The maximum value in the target class vector is determined, and the transaction type corresponding to the transaction data to be identified is determined according to the number of digits corresponding to the maximum value.

[0033] In a third aspect, the present application also provides an abnormal transaction identification model training device. The abnormal transaction identification model is a deep forest model including multiple levels of classification units, each level of the classification unit includes a first type of classification unit, and the device includes:

[0034] An acquisition module, used to acquire a training sample set corresponding to the first type classification unit to be trained; the training sample set includes samples of abnormal transaction types and samples of normal transaction types;

[0035] A first training module, used for training a first decision tree in the first type classification unit to be trained by using the training sample set;

[0036] A second training module is used for obtaining, for each decision tree except the first decision tree in the first type classification unit to be trained, a second sample set corresponding to the decision tree according to samples of abnormal transaction types misclassified by a previous decision tree and the training sample set, and training the decision tree according to the second sample set to obtain a trained first type classification unit;

[0037] The first determination module is used to determine a new training sample set based on the classification result of the trained first type classification unit on the training sample set and the training sample set, and use the first type classification unit included in the next level classification unit as the first type classification unit to be trained, and return to execute the training of the first decision tree in the first type classification unit to be trained through the training sample set until the preset training end condition is reached to obtain the abnormal transaction recognition model.

[0038] In one embodiment, each level of the classification unit further includes a second type of classification unit, and the second type of classification unit is a random forest classification unit; the device also includes a third training module for:

[0039] The second type classification unit to be trained is trained by using the training sample set to obtain the trained second type classification unit; the first type classification unit to be trained and the second type classification unit to be trained belong to the same level of classification units;

[0040] The first determining module is specifically configured to:

[0041] A new training sample set is determined according to the classification result of the trained first type classification unit on the training sample set, the classification result of the trained second type classification unit on the training sample set, and the training sample set.

[0042] In one embodiment, the second training module is specifically used for:

[0043] The sample of abnormal transaction type misclassified by the previous decision tree is determined as the first target sample, and the K nearest neighbor algorithm is used to determine the neighbor sample corresponding to the first target sample in the training sample set; the first target sample and the neighbor sample are synthesized to obtain the second target sample; according to the first target sample, the second target sample, and the training sample set, a second sample set corresponding to the decision tree is obtained; the second sample set includes the first target sample and the second target sample.

[0044] In one of the embodiments, the apparatus further includes a second determining module, configured to:

[0045] Determine an initial weight of each sample in the training sample set; wherein the initial weight of the sample of abnormal transaction type in the training sample set is greater than the initial weight of the sample of normal transaction type;

[0046] The first training module is specifically used for:

[0047] Sampling is performed according to the initial weight of each sample in the training sample set, and the first decision tree in the first type classification unit to be trained is trained by using the samples obtained by the sampling.

[0048] In one embodiment, the second determining module is specifically configured to:

[0049] Count the number of samples of abnormal transaction types and the number of samples of normal transaction types in the training sample set; determine the reciprocal of the number of samples of the abnormal transaction type as the initial weight of the samples of the abnormal transaction type in the training sample set, and determine the reciprocal of the number of samples of the normal transaction type as the initial weight of the samples of the normal transaction type in the training sample set.

[0050] In a fourth aspect, the present application also provides an abnormal transaction identification device. The device comprises:

[0051] A preprocessing module, used to preprocess the transaction data to be identified to obtain target feature data; the transaction data to be identified includes cardholder identity information and transaction behavior information;

[0052] A classification module is used to input the target feature data into an abnormal transaction identification model for decision classification to obtain a transaction type corresponding to the transaction data to be identified; the transaction type includes an abnormal transaction type and a normal transaction type; wherein the abnormal transaction identification model is trained by the abnormal transaction identification model training method described in the first aspect.

[0053] In one embodiment, the abnormal transaction identification model comprises multiple levels of classification units, each level of the classification units comprises a first type of classification unit and a second type of classification unit; the abnormal transaction identification model further comprises an output layer;

[0054] The classification module is specifically used for:

[0055] The target feature data is input into an abnormal transaction identification model to obtain a first-category vector and a second-category vector; wherein the first-category vector is a class vector output by a first-type classification unit included in the last-level classification unit in the multi-level classification unit, and the second-category vector is a class vector output by a second-type classification unit included in the last-level classification unit; through the output layer included in the abnormal transaction identification model, the first-category vector and the second-category vector are averaged to obtain a target class vector; the maximum value in the target class vector is determined, and the transaction type corresponding to the transaction data to be identified is determined according to the number of digits corresponding to the maximum value.

[0056] In a fifth aspect, the present application further provides a computer device, wherein the computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in the first aspect or the second aspect are implemented.

[0057] In a sixth aspect, the present application further provides a computer-readable storage medium, wherein a computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method described in the first aspect or the second aspect are implemented.

[0058] In a seventh aspect, the present application further provides a computer program product, wherein the computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect or the second aspect are implemented.

[0059] The above-mentioned abnormal transaction identification model training method and device, abnormal transaction identification method and device, computer equipment, storage medium and computer program product, by constructing a deep forest model of multi-level classification units, each level of classification units includes a first type of classification unit, for the current level classification unit to be trained, a training sample set corresponding to the first type of classification unit to be trained (that is, the first type of classification unit included in the current level classification unit) is obtained, and then the decision trees in the first type of classification unit are trained one by one through the training sample set, wherein, for each decision tree except the first decision tree, the second sample set corresponding to the decision tree is determined according to the samples of the abnormal transaction type misclassified by the previous decision tree and the training sample set, which is used to train the decision tree, thereby obtaining the trained first type of classification unit (included in the current level classification unit), and then according to the classification result of the trained first type of classification unit on the training sample set and the training sample set, a new training sample set is determined (that is, the classification result output by the current level classification unit and the training sample set are passed to the next level together), and then the first type of classification unit included in the next level classification unit is trained through the new training sample set until the preset training end condition is reached, that is, the abnormal transaction identification model is obtained.

[0060] In this method, since the second sample set used when training each decision tree except the first decision tree in the first type classification unit is obtained based on the samples of abnormal transaction types that are incorrectly classified by the previous decision tree and the training sample set, on the one hand, the proportion of abnormal transaction type samples in the second sample set is larger than that in the training sample set, that is, the distribution of the two types of samples in the sample set is relatively more balanced, which is conducive to the learning and identification of abnormal transaction type samples. On the other hand, the decision tree trained by the second sample set can make up for the recognition errors of some abnormal transaction type samples in the previous decision tree. The first type classification unit is obtained by training, and the abnormal transaction recognition model obtained thereby can improve the recognition accuracy of abnormal transactions. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 A schematic diagram of a flow chart of a method for training an abnormal transaction identification model in one embodiment;

[0062] Figure 2 A schematic diagram of a process of obtaining a second sample set corresponding to a decision tree in one embodiment;

[0063] Figure 3 A schematic diagram of a flow chart of an abnormal transaction identification method in one embodiment;

[0064] Figure 4 A schematic diagram of the structure of an abnormal transaction identification model in one embodiment;

[0065] Figure 5It is a structural block diagram of an abnormal transaction identification model training device in one embodiment;

[0066] Figure 6 is a structural block diagram of an abnormal transaction identification device in one embodiment;

[0067] Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0068] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0069] First, before specifically introducing the technical solution of the embodiment of the present application, the technical background or technical evolution context based on the embodiment of the present application is introduced. Credit card fraud refers to stealing the identity information of others to impersonate credit cards, thereby engaging in malicious overdrafts, cash withdrawals, and thefts. Credit card anti-fraud detection is to identify the risks of credit card transactions in order to promptly prevent abnormal transactions such as credit card fraud and reduce financial losses, which is one of the keys to controlling risks in the financial industry. In the related art, credit card anti-fraud detection is to classify credit card transaction data through a classification model to identify it as an abnormal transaction or a normal transaction, wherein the classification model is generally a trained random forest model. However, since in real business scenarios, the sample size of abnormal credit card transactions is far less than the sample size of normal transactions, it is an unbalanced data set, and thus the recognition accuracy of abnormal transactions using a random forest model is low. Based on this background, the applicant has proposed the abnormal transaction recognition model training method of the present application through long-term research and development and experimental verification, and the obtained abnormal transaction recognition model can improve the recognition accuracy of abnormal transactions. In addition, it should be noted that the applicant has made a lot of creative work in the discovery of the technical problem of the present application and the technical solution introduced in the following embodiment.

[0070] In one embodiment, Figure 1As shown, a method for training an abnormal transaction identification model is provided. This embodiment uses the method applied to a terminal as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal can be but is not limited to various personal computers, laptops, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The server 104 can be implemented as an independent server or a server cluster composed of multiple servers. In this embodiment, the abnormal transaction identification model is a deep forest model including multiple levels of classification units, and each level of classification unit includes a first type of classification unit. The abnormal transaction identification model training method includes the following steps:

[0071] Step 101: Obtain a training sample set corresponding to a first type classification unit to be trained.

[0072] The training sample set includes samples of abnormal transaction types and samples of normal transaction types.

[0073] In implementation, the terminal can construct a deep forest model containing multiple levels of classification units, and train it level by level, that is, first train the first-level classification units, and then train the second-level classification units, until the training end condition is reached (such as training a preset number of classification units, or when the classification effect of the current-level classification unit is no longer improved compared to the classification effect of the previous-level classification unit, the training is terminated and the current-level classification unit is used as the last-level classification unit). For the current-level classification unit to be trained, the first type of classification unit it contains is the first type of classification unit to be trained. For example, if the current-level classification unit to be trained is a first-level classification unit, the terminal can obtain the training sample set corresponding to the first-level classification unit as the training sample set corresponding to the first-type classification unit to be trained.

[0074] Among them, the training sample set corresponding to the first-level classification unit can be obtained based on historical transaction data. For example, historical credit card transaction data can be collected, including cardholder identity information (such as age, gender, name, city of residence, etc.), transaction behavior data (such as transaction time, transaction amount, transaction location, transaction form (online or offline), transaction object information (such as merchant information), etc.), and transaction type (abnormal transaction type or normal transaction type). Then, the collected credit card historical transaction data can be pre-processed by denoising, filling missing values, normalization, etc. to obtain the initial training sample set. The terminal can use the initial training sample set as the training sample set corresponding to the first-level classification unit. Each sample in the training sample set contains feature data and transaction type data. The feature data is obtained based on the cardholder's identity information and transaction behavior data in the historical transaction data, and the transaction type data is determined based on the transaction type of the historical transaction data. For example, a training sample set containing n samples (which can be denoted as S) can be expressed as:

[0075] S={(x i ,y i )|i=1,2…,n,y i ∈{1,0}}

[0076] Among them, x i is the characteristic data of the i-th sample, y i is the transaction type data of the i-th sample. The transaction type data corresponding to the abnormal transaction type can be set to 0, and the transaction type data corresponding to the normal transaction type can be set to 1.

[0077] For classification units other than the first-level classification units, the training sample set corresponding to each level of classification unit can be obtained according to the initial training sample set and the classification result of the training sample set by the previous level classification unit. The specific processing process is described in detail below.

[0078] Step 102: training the first decision tree in the first type classification unit to be trained by using the training sample set.

[0079] The first type classification unit includes multiple decision trees, and the number of decision trees included can be preset, such as 300. The training method for the first type classification unit is to train the decision trees it contains one by one, that is, first train the first decision tree, and then train the next decision tree, until the preset number of decision trees are trained.

[0080] In implementation, the terminal may train the first decision tree in the first type of classification unit to be trained by using the training sample set corresponding to the first type of classification unit to be trained obtained in step 101. For example, the terminal may extract a preset number of samples from the training sample set for training the first decision tree, wherein the sampling method may be sampling with replacement.

[0081] Step 103, for each decision tree except the first decision tree in the first type classification unit to be trained, a second sample set corresponding to the decision tree is obtained according to the samples of abnormal transaction types misclassified by the previous decision tree and the training sample set, and the decision tree is trained according to the second sample set to obtain the trained first type classification unit.

[0082] In implementation, for each decision tree except the first decision tree in the first type classification unit to be trained, the terminal can obtain the second sample set corresponding to the current decision tree to be trained based on the samples of the abnormal transaction type that are misclassified by the previous decision tree and the training sample set. For example, if the current decision tree to be trained is the second decision tree, the terminal can classify the training sample set according to the first decision tree after training, obtain the classification result corresponding to the first decision tree, and filter out the samples of the abnormal transaction type that are misclassified in the classification result. The classification error refers to the misclassification of the samples of the abnormal transaction type as the normal transaction type. Then, the terminal can obtain the second sample set corresponding to the second decision tree based on the samples of the abnormal transaction type that are misclassified and the training sample set. For example, the terminal can extract a preset number of samples from the training sample set, and then combine the extracted samples with the samples of the abnormal transaction type that are misclassified into a second sample set. The terminal can also synthesize new approximate samples based on the samples of the abnormal transaction type that are misclassified, and then combine the samples extracted from the training sample set, the synthesized new approximate samples, and the samples of the abnormal transaction type that are misclassified into the second sample set.

[0083] Afterwards, the terminal can train the decision tree according to the second sample set corresponding to the decision tree to be trained, until a preset number of decision trees are trained to obtain the trained first type classification unit. Optionally, the terminal can train a certain proportion of decision trees exceeding the preset number, and then select a preset number of decision trees with better classification effects as the decision trees finally included in the first type classification unit. The classification effect of each decision tree can be evaluated by the classification error rate or the classification accuracy rate.

[0084] Step 104, based on the classification results of the trained first-type classification unit on the training sample set and the training sample set, determine a new training sample set, and use the first-type classification unit included in the next-level classification unit as the first-type classification unit to be trained, return to execute the training of the first decision tree in the first-type classification unit to be trained through the training sample set, until the preset training end condition is reached, and obtain the abnormal transaction recognition model.

[0085] In implementation, after training the first type of classification unit included in the current level classification unit, the terminal can use the trained first type of classification unit to classify the training sample set to obtain a classification result. The classification result can be a class vector, and each digit of the class vector represents a transaction type. In this example, it is a two-dimensional class vector, that is, it contains two digits, representing the normal transaction type and the abnormal transaction type respectively. The value on each digit represents the probability value of the sample being the transaction type corresponding to the digit. Then, the terminal can determine a new training sample set based on the classification result and the training sample set. For example, for the training sample set corresponding to the second-level classification unit, the terminal can perform feature splicing on the classification result of each sample in the training sample set by the first-level classification unit and the initial training sample set to obtain a new training sample set as the training sample set corresponding to the second-level classification unit. It can be understood that for the training sample set corresponding to the third-level classification unit, the terminal can obtain the feature splicing of the classification result corresponding to the second-level classification unit and the initial training sample set, or it can replace the class vector value spliced ​​in the training sample set corresponding to the previous level (i.e., the second level) classification unit with the class vector value output by the previous level (i.e., the second level) classification unit. The same applies to other levels of classification units after the third level.

[0086] Then, the terminal can use the first type of classification unit contained in the next level classification unit as the first type of classification unit to be trained, and return to execute step 102, that is, train the first decision tree in the first type of classification unit to be trained through the training sample set (the new training sample set obtained after feature splicing or feature replacement) until the preset training end condition is reached. The preset training end condition can be that the classification units of a preset level are trained, or an adaptive method is used. When the classification effect of the current level classification unit under training is no longer improved compared with the classification effect of the previous level classification unit, the level increase is terminated and the current level classification unit is used as the last level classification unit. In addition, the trained model can also be tested using a test sample set. When the test accuracy reaches a preset value or the test error rate is lower than a preset value, the training end condition is met, and the trained model can be used as an abnormal transaction identification model.

[0087] In the above-mentioned abnormal transaction identification model training method, since the second sample set used in training each decision tree except the first decision tree in the first type classification unit is obtained based on the samples of abnormal transaction types incorrectly classified by the previous decision tree and the training sample set, on the one hand, the proportion of abnormal transaction type samples in the second sample set is larger than that in the training sample set, that is, the distribution of the two types of samples in the sample set is relatively more balanced, which is conducive to the learning and identification of abnormal transaction type samples. On the other hand, the decision tree trained by the second sample set can make up for the recognition errors of some abnormal transaction type samples in the previous decision tree. The first type classification unit is obtained by training, and the abnormal transaction identification model obtained in this way can improve the recognition accuracy of abnormal transactions.

[0088] In one embodiment, each level of classification unit further includes a second type of classification unit, and the second type of classification unit is a random forest classification unit. The method further includes the following steps: training the second type of classification unit to be trained by using a training sample set to obtain a trained second type of classification unit. The first type of classification unit to be trained and the second type of classification unit to be trained belong to the same level of classification unit.

[0089] In the implementation, in the deep forest model of the constructed multi-level classification unit, each level of the classification unit can also include a second type of classification unit, wherein the second type of classification unit is a random forest classification unit. Specifically, the terminal can use the training sample set obtained in step 101 as the training sample set for training the second type of classification unit to be trained, and the first type of classification unit to be trained and the second type of classification unit to be trained belong to the same level of classification unit. That is, the terminal can obtain the training sample set corresponding to the current level of classification unit as the training sample set corresponding to the first type of classification unit and the second type of classification unit contained in the current level of classification unit. The terminal can train the second type of classification unit to be trained through the training sample set to obtain a random forest classification unit. For details, please refer to the existing training method of the random forest model, which will not be repeated here.

[0090] Correspondingly, the process of determining a new training sample set in step 104 specifically includes the following steps: determining a new training sample set based on the classification results of the trained first type classification unit on the training sample set, the classification results of the trained second type classification unit on the training sample set, and the training sample set.

[0091] In implementation, after training the first type classification unit and the second type classification unit included in the current level classification unit, the terminal can classify the training sample set with the trained first type classification unit and the second type classification unit, respectively, to obtain the classification result (first type vector) corresponding to the first type classification unit and the classification result (second type vector) corresponding to the second type classification unit. Then, the terminal can obtain a new training sample set according to the classification result corresponding to the first type classification unit, the classification result corresponding to the second type classification unit, and the training sample set. For example, the first type vector, the second type vector, and the training sample set can be obtained by feature concatenation or feature replacement.

[0092] In this embodiment, in the deep forest model of multi-level classification units constructed, each level of classification units includes first-type classification units and second-type classification units, wherein the second-type classification units are random forest classification units. Thus, the trained abnormal transaction recognition model can take into account the high accuracy of the first-type classification units in identifying abnormal transactions and the recognition advantage of the second-type classification units (random forests) in identifying normal transaction types, thereby improving the robustness and generalization ability of the model, and improving the comprehensive recognition accuracy of transaction types.

[0093] In one embodiment, Figure 2 As shown, the process of obtaining the second sample set corresponding to the decision tree in step 103 specifically includes the following steps:

[0094] Step 201, determine the sample of abnormal transaction type misclassified by the previous decision tree as the first target sample, and use the K nearest neighbor algorithm to determine the nearest neighbor sample corresponding to the first target sample in the training sample set.

[0095] In implementation, for each decision tree except the first decision tree in the first type classification unit, the terminal may determine the sample of the abnormal transaction type that is misclassified by the previous decision tree as the first target sample, and then the terminal may use the K-Nearest Neighbor (KNN) algorithm to determine the k nearest (i.e., the nearest in the feature space) samples of the first target sample in the training sample set, and then the terminal may screen out a preset number of samples from the k nearest samples as the nearest neighbor samples corresponding to the first target sample. The preset number may be set according to the proportion of abnormal transaction type samples in the training sample set or experience, for example, the preset number may be set to 3 to 10.

[0096] Step 202: synthesize the first target sample and the neighboring sample to obtain a second target sample.

[0097] In implementation, the terminal may synthesize the first target sample and its corresponding neighboring sample to obtain a second target sample. In one example, the second target sample includes feature data and transaction type data, wherein the transaction type data is data corresponding to an abnormal transaction type (such as 0), and the feature data (which may be recorded as synic) may be calculated using the following formula:

[0098]

[0099] Among them, x′ i is the characteristic data of the first target sample, y′ i is the transaction type data of the first target sample, x′ j is the characteristic data of the neighboring samples, y′ j is the transaction type data of the neighbor sample, and r is the random seed.

[0100] Step 203: Obtain a second sample set corresponding to the decision tree according to the first target sample, the second target sample, and the training sample set.

[0101] The second sample set includes a first target sample and a second target sample.

[0102] In implementation, the terminal may obtain the second sample set corresponding to the decision tree to be trained according to the first target sample, the second target sample and the training sample set. For example, a certain number of samples may be extracted from the training sample set, and the extracted samples may be combined with the first target sample and the second target sample to form the second sample set.

[0103] In this embodiment, a certain number of new samples (i.e., second target samples) are synthesized based on samples of abnormal transaction types that are misclassified by the previous decision tree (i.e., first target samples) and their neighboring samples, that is, the samples of abnormal transaction types that are misclassified are expanded, and then the misclassified samples and the synthesized new samples are used as part of the second sample set for training the decision tree. This can reasonably increase the proportion of samples of abnormal transaction types in the second sample set for training the decision tree. The decision tree is trained based on the second sample set, and the abnormal transaction recognition model obtained by training can further improve the recognition accuracy of abnormal transactions.

[0104] In one embodiment, the abnormal transaction identification model method further includes the following step: determining an initial weight of each sample in the training sample set.

[0105] In implementation, after the terminal obtains the training sample set corresponding to the first type classification unit to be trained, the initial weight of each sample in the training sample set can be determined. Among them, the initial weight of the sample of abnormal transaction type in the training sample set is greater than the initial weight of the sample of normal transaction type. The initial weight can be preset or calculated according to other strategies.

[0106] Correspondingly, the process of training the first decision tree in step 102 specifically includes the following steps: sampling according to the initial weight of each sample in the training sample set, and training the first decision tree in the first type classification unit to be trained with the sampled samples.

[0107] In implementation, the terminal may perform sampling according to the initial weight of each sample in the training sample set, extract a preset number of samples, and then use the extracted samples to train the first decision tree in the first type classification unit.

[0108] Furthermore, the process of training each decision tree except the first decision tree may specifically include the following steps: updating the weight of each sample in the training sample set according to the classification result of the previous decision tree on the training sample set.

[0109] Specifically, the terminal can classify the training sample set through the previous decision tree to obtain a classification result, and then calculate the classification error rate corresponding to the previous decision tree based on the classification result. The classification error rate corresponding to the decision tree can be calculated using the following formula:

[0110]

[0111] Among them, e t is the classification error rate of the t-th decision tree, and the classification error rate of the first decision tree is e1; n is the number of samples in the training sample set, i is the i-th sample, y i is the transaction type data (actual data) of the i-th sample, h i is the transaction type data (prediction data) determined based on the classification result of the t-th decision tree for the i-th sample in the training sample set. For example, based on the classification result (class vector) obtained by the t-th decision tree for classifying the i-th sample, the digit with the largest value in the class vector can be determined, and then the transaction type data corresponding to the digit can be determined, that is, the predicted transaction type data h can be obtained. i .

[0112] Then, according to the classification error rate of the previous decision tree, the weight coefficient corresponding to the previous decision tree is calculated. The weight coefficient of the decision tree can be calculated using the following formula:

[0113]

[0114] Among them, α t is the weight coefficient of the tth decision tree, e t is the classification error rate of the tth decision tree.

[0115] Afterwards, the weight of each sample in the training sample set is updated according to the weight coefficient of the previous decision tree and the classification result of the previous decision tree on the training sample set. The new weight of each sample in the training sample set can be calculated using the following formula:

[0116]

[0117] in, is the weight of the i-th sample corresponding to the t-th decision tree, that is, the sample set used to train the t-th decision tree is based on the weight Sampling is performed to obtain; α t is the weight coefficient of the tth decision tree, y i is the transaction type data (actual data) of the i-th sample, h i is the transaction type data (prediction data) determined by the classification result of the t-th decision tree for the i-th sample in the training sample set, Z i is the normalization factor. For the t+1th decision tree, the weight corresponding to the tth (i.e., the previous) decision tree can be Update and determine the new weight As the weight corresponding to the t+1th decision tree, and according to the new weight Sampling is performed, and then based on the sampled samples and the samples of abnormal transaction types that were misclassified by the previous decision tree (i.e., the t-th decision tree), a second sample set corresponding to the t+1-th decision tree is obtained for training the t+1-th decision tree.

[0118] In this embodiment, by setting the initial weight of each sample in the training sample set, wherein the initial weight of the sample of the abnormal transaction type is greater than the initial weight of the sample of the normal transaction type, sampling according to the initial weight can increase the probability of extracting the abnormal transaction type sample, that is, in the sample set used to train the first decision tree, the distribution of the two transaction types is more balanced than the distribution in the initial training sample set, and thus the trained abnormal transaction recognition model can improve the recognition accuracy of abnormal transactions. In addition, the weights of the samples in the training sample set can be further updated to sample according to the updated weights for training the next decision tree, which can further improve the recognition accuracy of abnormal transactions by the abnormal transaction recognition model. The first type of classification unit in this embodiment is a new classification method for unbalanced data sets, which can be called an LC-AdaBoost classifier.

[0119] In one embodiment, the process of determining the initial weight specifically includes the following steps: counting the number of samples of abnormal transaction types and the number of samples of normal transaction types in the training sample set; determining the inverse of the number of samples of abnormal transaction types as the initial weight of the samples of abnormal transaction types in the training sample set, and determining the inverse of the number of samples of normal transaction types as the initial weight of the samples of normal transaction types in the training sample set.

[0120] In implementation, the terminal can count the number of samples of abnormal transaction types and the number of samples of normal transaction types in the training sample set, and determine the reciprocal of the number of samples of abnormal transaction types as the initial weight of the samples of abnormal transaction types in the training sample set, and determine the reciprocal of the number of samples of normal transaction types as the initial weight of the samples of normal transaction types in the training sample set. Since the proportion of samples of abnormal transaction types in the training sample set is relatively small, by setting the initial weight of each type of sample according to the number of samples, the sampling probability of samples of abnormal transaction types can be reasonably set, thereby ensuring the accuracy of the trained abnormal transaction recognition model in identifying abnormal transactions.

[0121] In one embodiment, Figure 3 As shown, a method for identifying abnormal transactions is also provided, the method comprising the following steps:

[0122] Step 301, pre-processing the transaction data to be identified to obtain target feature data.

[0123] Among them, the transaction data to be identified includes cardholder identity information and transaction behavior information.

[0124] In implementation, the terminal can obtain the transaction data to be identified, and pre-process the transaction data to be identified, such as filling missing values ​​and normalizing the data, to obtain target feature data. The pre-processing method can adopt the method of pre-processing historical transaction data in the abnormal transaction identification model training method.

[0125] Step 302: Input the target feature data into the abnormal transaction identification model for decision classification to obtain the transaction type corresponding to the transaction data to be identified.

[0126] The abnormal transaction identification model is trained by the abnormal transaction identification model training method mentioned above.

[0127] In implementation, the terminal can input the target feature data into a pre-trained abnormal transaction recognition model, and make a decision classification on the target feature data through the abnormal transaction recognition model to obtain the transaction type corresponding to the transaction data to be recognized. The transaction type includes abnormal transaction type and normal transaction type.

[0128] In this embodiment, the abnormal transaction identification model trained by the above-mentioned abnormal transaction identification model training method is used to identify the transaction type of the transaction data to be identified, which can improve the recognition accuracy of the transaction data to be identified, especially for the transaction data to be identified of abnormal transaction types, the recognition accuracy is relatively high.

[0129] In one embodiment, the abnormal transaction identification model includes a multi-level classification unit, each level of the classification unit includes a first type of classification unit and a second type of classification unit, and the abnormal transaction identification model also includes an output layer. Accordingly, the process of obtaining the transaction type corresponding to the transaction data to be identified in step 302 specifically includes the following steps: inputting the target feature data into the abnormal transaction identification model to obtain a first type vector and a second type vector; wherein the first type vector is a class vector output by the first type of classification unit included in the last level of the classification unit in the multi-level classification unit, and the second type vector is a class vector output by the second type of classification unit included in the last level of the classification unit; through the output layer included in the abnormal transaction identification model, the first type vector and the second type vector are averaged to obtain a target class vector; determine the maximum value in the target class vector, and determine the transaction type corresponding to the transaction data to be identified according to the number of digits corresponding to the maximum value.

[0130] In practice, the structure of the abnormal transaction identification model is as follows Figure 4As shown, the abnormal transaction identification model includes a multi-level classification unit and an output layer, wherein each level of the classification unit includes a first type classification unit (LC-AdaBoost classifier) ​​and a second type classification unit (Random Forest RF). The terminal can input the target feature data into the abnormal transaction identification model, and make a decision classification on the target feature data through the multi-level classification unit in the abnormal transaction identification model, and obtain the first type vector and the second type vector output by the last level classification unit in the multi-level classification unit. Specifically, the first type classification unit and the second type classification unit in the first level classification unit classify the target feature data respectively, and obtain the class vector corresponding to the first type classification unit and the class vector corresponding to the second type classification unit. Then, the two class vectors can be feature spliced ​​with the target feature data, and the spliced ​​data can be used as the input of the next level classification unit, so that the first type classification unit and the second type classification unit included in the next level classification unit can classify the spliced ​​data respectively, and then the obtained two class vectors are spliced ​​with the target feature data, and continue to be passed to the next level classification unit until the first type vector and the second type vector output by the last level classification unit are obtained. Then, the first-class vector and the second-class vector output by the last-level classification unit are used as the input of the output layer, and the mean of the first-class vector and the second-class vector is calculated through the output layer, that is, the mean of the values ​​with the same digits in the two class vectors is calculated to obtain the target class vector. Then, the terminal can determine the maximum value in the target class vector, and determine the transaction type corresponding to the transaction data to be identified based on the digits of the maximum value, that is, complete the transaction type identification of the transaction data to be identified. The terminal can generate and display the identification results of the transaction data to be identified for the staff to view.

[0131] In this embodiment, the mean calculation is performed based on the first-category vector and the second-category vector output by the last-level classification unit of the abnormal transaction recognition model to obtain a target class vector, and then the transaction type corresponding to the transaction data to be identified is determined based on the number of digits where the maximum value in the target class vector is located, which can take into account both recognition accuracy and recognition efficiency.

[0132] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0133] Based on the same inventive concept, the embodiment of the present application also provides an abnormal transaction identification model training device for implementing the abnormal transaction identification model training method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more abnormal transaction identification model training device embodiments provided below can refer to the limitations of the abnormal transaction identification model training method above, and will not be repeated here.

[0134] In one embodiment, Figure 5 As shown, an abnormal transaction identification model training device 500 is provided, the abnormal transaction identification model is a deep forest model including multiple levels of classification units, each level of classification unit includes a first type of classification unit, and the abnormal transaction identification model training device includes: an acquisition module 501, a first training module 502, a second training module 503 and a first determination module 504, wherein:

[0135] The acquisition module 501 is used to acquire a training sample set corresponding to the first type classification unit to be trained; the training sample set includes samples of abnormal transaction types and samples of normal transaction types.

[0136] The first training module 502 is used to train the first decision tree in the first type classification unit to be trained by using a training sample set.

[0137] The second training module 503 is used to obtain, for each decision tree except the first decision tree in the first type classification unit to be trained, a second sample set corresponding to the decision tree based on samples of abnormal transaction types misclassified by the previous decision tree and the training sample set, and train the decision tree based on the second sample set to obtain the trained first type classification unit.

[0138] The first determination module 504 is used to determine a new training sample set according to the classification results of the trained first-type classification unit on the training sample set and the training sample set, and use the first-type classification unit included in the next-level classification unit as the first-type classification unit to be trained, and return to execute the training of the first decision tree in the first-type classification unit to be trained through the training sample set until the preset training end condition is reached to obtain the abnormal transaction recognition model.

[0139] In one embodiment, each level of classification unit also includes a second type of classification unit, and the second type of classification unit is a random forest classification unit. The abnormal transaction identification model training device also includes a third training module, which is used to train the second type of classification unit to be trained through a training sample set to obtain a trained second type of classification unit; the first type of classification unit to be trained and the second type of classification unit to be trained belong to the same level of classification units.

[0140] Accordingly, the first determination module 504 is specifically used to determine a new training sample set according to the classification results of the trained first type classification unit on the training sample set, the classification results of the trained second type classification unit on the training sample set, and the training sample set.

[0141] In one embodiment, the second training module 503 is specifically used to: determine the sample of abnormal transaction type that is misclassified by the previous decision tree as the first target sample, and use the K nearest neighbor algorithm to determine the neighbor sample corresponding to the first target sample in the training sample set; synthesize the first target sample and the neighbor sample to obtain the second target sample; obtain the second sample set corresponding to the decision tree based on the first target sample, the second target sample, and the training sample set; the second sample set includes the first target sample and the second target sample.

[0142] In one embodiment, the abnormal transaction identification model training device also includes a second determination module, which is used to: determine the initial weight of each sample in the training sample set; wherein the initial weight of the sample of the abnormal transaction type in the training sample set is greater than the initial weight of the sample of the normal transaction type.

[0143] Correspondingly, the first training module 502 is specifically used to: perform sampling according to the initial weight of each sample in the training sample set, and train the first decision tree in the first type classification unit to be trained by using the sampled samples.

[0144] In one embodiment, the second determination module is specifically used to: count the number of samples of abnormal transaction types and the number of samples of normal transaction types in the training sample set; determine the inverse of the number of samples of abnormal transaction types as the initial weight of the samples of abnormal transaction types in the training sample set, and determine the inverse of the number of samples of normal transaction types as the initial weight of the samples of normal transaction types in the training sample set.

[0145] Based on the same inventive concept, the embodiment of the present application also provides an abnormal transaction identification device for implementing the abnormal transaction identification method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more abnormal transaction identification device embodiments provided below can refer to the limitations of the abnormal transaction identification method above, and will not be repeated here.

[0146] In one embodiment, Figure 6 As shown, an abnormal transaction identification device 600 is provided, and the abnormal transaction identification device includes: a pre-processing module 601 and a classification module 602, wherein:

[0147] The preprocessing module 601 is used to preprocess the transaction data to be identified to obtain target feature data; the transaction data to be identified includes the cardholder's identity information and transaction behavior information.

[0148] The classification module 602 is used to input the target feature data into the abnormal transaction identification model for decision classification to obtain the transaction type corresponding to the transaction data to be identified; the transaction type includes abnormal transaction type and normal transaction type; wherein the abnormal transaction identification model is trained by the abnormal transaction identification model training method involved above.

[0149] In one embodiment, the abnormal transaction identification model includes a multi-level classification unit, each level of the classification unit includes a first type of classification unit and a second type of classification unit, and the abnormal transaction identification model also includes an output layer. The classification module 602 is specifically used to: input the target feature data into the abnormal transaction identification model to obtain a first type vector and a second type vector; wherein the first type vector is a class vector output by the first type of classification unit included in the last level of the classification unit in the multi-level classification unit, and the second type vector is a class vector output by the second type of classification unit included in the last level of the classification unit; through the output layer included in the abnormal transaction identification model, the first type vector and the second type vector are averaged to obtain a target class vector; determine the maximum value in the target class vector, and determine the transaction type corresponding to the transaction data to be identified according to the number of digits corresponding to the maximum value.

[0150] Each module in the abnormal transaction identification model training device and the abnormal transaction identification device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0151] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 7As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, an abnormal transaction identification model training method or an abnormal transaction identification method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse, etc.

[0152] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0153] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0154] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0155] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0156] The abnormal transaction identification model training method and device, abnormal transaction identification method and device, computer equipment, storage medium and computer program product provided in this application relate to the field of artificial intelligence technology and can be used in the field of financial technology or other related fields. This application does not limit the application field of the abnormal transaction identification model training method and device, abnormal transaction identification method and device, computer equipment, storage medium and computer program product.

[0157] It should be noted that the user information (including but not limited to cardholder identity information, transaction behavior information) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0158] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0159] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0160] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for training an abnormal transaction identification model, characterized in that: The abnormal transaction identification model is a deep forest model including multiple levels of classification units, each level of the classification units including first type classification units, and the method includes: Obtaining a training sample set corresponding to the first type classification unit to be trained; the training sample set includes samples of abnormal transaction types and samples of normal transaction types; Training a first decision tree in the first type classification unit to be trained by using the training sample set; For each decision tree other than the first decision tree in the first type classification unit to be trained, a second sample set corresponding to the decision tree is obtained according to samples of abnormal transaction types misclassified by a previous decision tree and the training sample set, and the decision tree is trained according to the second sample set to obtain a trained first type classification unit; According to the classification result of the trained first-type classification unit on the training sample set and the training sample set, a new training sample set is determined, and the first-type classification unit included in the next-level classification unit is used as the first-type classification unit to be trained, and the training of the first decision tree in the first-type classification unit to be trained by the training sample set is returned to be executed until a preset training end condition is reached, thereby obtaining the abnormal transaction recognition model; The step of obtaining a second sample set corresponding to the decision tree based on the samples of abnormal transaction types misclassified by the previous decision tree and the training sample set includes: Determine the sample of abnormal transaction type misclassified by the previous decision tree as the first target sample, and use the K nearest neighbor algorithm to determine the nearest neighbor sample corresponding to the first target sample in the training sample set; Synthesize the first target sample and the neighboring sample to obtain a second target sample; A second sample set corresponding to the decision tree is obtained according to the first target sample, the second target sample, and the training sample set; the second sample set includes the first target sample and the second target sample.

2. The method according to claim 1, characterized in that Each level of the classification unit further includes a second type of classification unit, and the second type of classification unit is a random forest classification unit; the method further includes: The second type classification unit to be trained is trained by using the training sample set to obtain the trained second type classification unit; the first type classification unit to be trained and the second type classification unit to be trained belong to the same level of classification units; The step of determining a new training sample set based on the classification result of the training sample set by the trained first type classification unit and the training sample set comprises: A new training sample set is determined according to the classification result of the trained first type classification unit on the training sample set, the classification result of the trained second type classification unit on the training sample set, and the training sample set.

3. The method according to claim 1, characterized in that The method further comprises: Determine an initial weight of each sample in the training sample set; wherein the initial weight of the sample of abnormal transaction type in the training sample set is greater than the initial weight of the sample of normal transaction type; The step of training the first decision tree in the first type classification unit to be trained by using the training sample set includes: Sampling is performed according to the initial weight of each sample in the training sample set, and the first decision tree in the first type classification unit to be trained is trained by using the samples obtained by the sampling.

4. The method according to claim 3, characterized in that The determining of the initial weight of each sample in the training sample set comprises: Counting the number of samples of abnormal transaction types and the number of samples of normal transaction types in the training sample set; The inverse of the number of samples of the abnormal transaction type is determined as the initial weight of the samples of the abnormal transaction type in the training sample set, and the inverse of the number of samples of the normal transaction type is determined as the initial weight of the samples of the normal transaction type in the training sample set.

5. A method for identifying abnormal transactions, characterized in that: The method comprises: Preprocessing the transaction data to be identified to obtain target feature data; the transaction data to be identified includes cardholder identity information and transaction behavior information; The target feature data is input into an abnormal transaction identification model for decision classification to obtain a transaction type corresponding to the transaction data to be identified; the transaction type includes an abnormal transaction type and a normal transaction type; wherein the abnormal transaction identification model is trained by the abnormal transaction identification model training method described in any one of claims 1 to 4.

6. The method according to claim 5, characterized in that The abnormal transaction identification model comprises multiple levels of classification units, each level of the classification units comprises a first type of classification unit and a second type of classification unit; the abnormal transaction identification model further comprises an output layer; The step of inputting the target feature data into an abnormal transaction identification model for decision classification to obtain a transaction type corresponding to the transaction data to be identified includes: Input the target feature data into an abnormal transaction recognition model to obtain a first-class vector and a second-class vector; wherein the first-class vector is a class vector output by a first-type classification unit included in the last-level classification unit in the multi-level classification unit, and the second-class vector is a class vector output by a second-type classification unit included in the last-level classification unit; By using the output layer included in the abnormal transaction recognition model, the first category vector and the second category vector are averaged to obtain a target category vector; The maximum value in the target class vector is determined, and the transaction type corresponding to the transaction data to be identified is determined according to the number of digits corresponding to the maximum value.

7. An abnormal transaction identification model training device, characterized in that: The abnormal transaction identification model is a deep forest model including multiple levels of classification units, each level of the classification unit includes a first type of classification unit, and the device includes: An acquisition module, used to acquire a training sample set corresponding to the first type classification unit to be trained; the training sample set includes samples of abnormal transaction types and samples of normal transaction types; A first training module, used for training a first decision tree in the first type classification unit to be trained by using the training sample set; A second training module is used for obtaining, for each decision tree except the first decision tree in the first type classification unit to be trained, a second sample set corresponding to the decision tree according to samples of abnormal transaction types misclassified by a previous decision tree and the training sample set, and training the decision tree according to the second sample set to obtain a trained first type classification unit; A first determination module is used to determine a new training sample set according to the classification result of the trained first type classification unit on the training sample set and the training sample set, and use the first type classification unit included in the next level classification unit as the first type classification unit to be trained, and return to execute the training of the first decision tree in the first type classification unit to be trained by the training sample set until a preset training end condition is reached, so as to obtain the abnormal transaction recognition model; The step of obtaining a second sample set corresponding to the decision tree based on the samples of abnormal transaction types misclassified by the previous decision tree and the training sample set includes: Determine the sample of abnormal transaction type misclassified by the previous decision tree as the first target sample, and use the K nearest neighbor algorithm to determine the nearest neighbor sample corresponding to the first target sample in the training sample set; Synthesize the first target sample and the neighboring sample to obtain a second target sample; A second sample set corresponding to the decision tree is obtained according to the first target sample, the second target sample, and the training sample set; the second sample set includes the first target sample and the second target sample.

8. An abnormal transaction identification device, characterized in that: The device comprises: A preprocessing module, used to preprocess the transaction data to be identified to obtain target feature data; the transaction data to be identified includes cardholder identity information and transaction behavior information; A classification module is used to input the target feature data into an abnormal transaction identification model for decision classification to obtain a transaction type corresponding to the transaction data to be identified; the transaction type includes an abnormal transaction type and a normal transaction type; wherein the abnormal transaction identification model is trained by the abnormal transaction identification model training method described in any one of claims 1 to 4.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Risk transaction identification method and device, server and storage medium

    CN110309840A