Transaction anti-fraud user identification method based on fusion model

Through the transaction anti-fraud user identification method that integrates expert rules and machine learning models, the limitations of a single model in fraud detection are solved, and more efficient and accurate fraud user identification is achieved, and the interpretability and flexibility of the model are enhanced.

CN120297979APending Publication Date: 2025-07-11BANK OF NANJING CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510381495.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing anti-fraud systems rely on a single expert rule or machine learning model to have limitations in identifying fraudulent users, expert rules lack coverage and interpretability, machine learning models are difficult to explain decision-making processes and insufficient iteration speed, resulting in limited fraud detection efficiency and accuracy.

Method used

The transaction anti-fraud user identification method based on the fusion model is adopted, combined with expert rules and machine learning models, and the fraud risk probability value is output by generating a set of expert rules and a prediction model. The binning statistical technology is used to fusion to generate a risk identification management strategy, including comprehensive evaluation of card account data, user transaction flow data and black and gray sample data.

Benefits of technology

It improves the accuracy and interpretability of fraud detection, forms a multi-layer protection mechanism, flexibly responds to changes in the fraud environment, enhances the interpretability and accuracy of the model, and adapts to complex fraud scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297979A_ABST
    Figure CN120297979A_ABST
Patent Text Reader

Abstract

The invention discloses a transaction anti-fraud user identification method based on a fusion model, and relates to the field of transaction fraud identification, and the method comprises the steps: receiving user risk control data, carrying out the calculation and collection of the user risk control data, and generating a direction index layer which reflects the transaction business quantification based on a collection result; generating an expert rule set based on the direction index layer and the expert rule, and obtaining an account risk scoring result according to the expert rule set; and constructing a prediction model based on the direction index layer to output a fraud risk probability value, and fusing the fraud risk probability value and the account risk scoring result by using a binning statistical technology to obtain a risk identification management and control strategy. Through introduction of external data, black and gray samples and card account customer risk data are supplemented, and the bottleneck problems that a single structure cannot obtain blocking points of cross-mechanism cross-platform fraud of fraudulent personnel, and a single mechanism is insufficient in black and gray samples and cannot better, more and more accurately identify fraud modes are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of transaction fraud identification, and more specifically, to a method for identifying transaction anti-fraud users based on a fusion model. Background Art

[0002] In the context where funds are increasingly valued, financial institutions continuously improve their anti-fraud work systems, continuously update and iterate risk identification technologies, and enhance fraud risk monitoring capabilities, so as to achieve the goals of protecting the safety of bank funds and purifying the financial environment.

[0003] Due to the diversification of fraud methods, multiple complex transaction models have emerged at the transaction end. Existing anti-fraud systems mostly adopt single technical means, such as relying solely on expert rules or only using machine learning models. These two methods have certain limitations when identifying fraud users separately:

[0004] 1. Expert rules have subjectivity and biases caused by different personal experiences. Expert rules rely on manual experience, may not cover all fraud scenarios, are summarized based on individual cases and special situations, and are difficult to directly apply to large-scale transaction data, resulting in the efficiency and accuracy of fraud control being affected;

[0005] 2. Machine learning algorithms have advantages in dealing with complex problems and large amounts of data, but they also have strong dependencies on training samples and massive data. The quality and quantity of data limit the performance of the model. Some machine learning models, such as deep learning models, may be difficult to explain their decision-making processes. In fraud detection, the interpretability of the model is crucial for understanding fraud behavior and gaining the trust of regulatory agencies; the update speed of fraud methods challenges the iteration speed of the model. Compared with expert rules, machine learning models usually require more time to retrain and adjust, which may cause the detection system to be unable to effectively identify new types of fraud behavior for a period of time.

[0006] In response to the problems in the related art, no effective solutions have been proposed yet. Summary of the Invention

[0007] In response to the problems in the related art, the present invention proposes a method for identifying transaction anti-fraud users based on a fusion model to overcome the above-mentioned technical problems existing in the existing related art.

[0008] To this end, the specific technical solution adopted by the present invention is as follows:

[0009] A method for identifying transaction anti-fraud users based on a fusion model, comprising:

[0010] Receiving user risk control data, calculating and summarizing the user risk control data, and generating a direction index layer that reflects the quantification of transaction operations based on the summary result;

[0011] Generate an expert rule set based on the directional indicator layer and expert rules, and obtain the account risk score result according to the expert rule set;

[0012] Based on the directional indicator layer, a prediction model is built to output the fraud risk probability value, and the fraud risk probability value is integrated with the account risk scoring result using binning statistics technology to obtain risk identification and control strategies.

[0013] Preferably, the user risk control data includes card account data, user transaction flow data, user behavior data and black and gray sample data; the directional indicator layer includes static indicators, behavioral indicators, transaction indicators and related indicators.

[0014] Preferably, generating an expert rule set based on the direction indicator layer and the expert rules, and obtaining the account risk score result according to the expert rule set includes:

[0015] Based on historical transaction anti-fraud data, fraud accounts and reported accounts involved in the case are defined as negative samples, and accounts excluded after investigation by branches are defined as positive samples;

[0016] Customize expert rules based on the directional indicator layer and user trading scenarios, combine the expert rules with the directional indicator layer to generate an expert rule set, and determine the number of triggered accounts and hit accounts of the expert rule set;

[0017] The negative sample hit ratio of the expert rule set is calculated based on the number of triggered accounts and the number of hit accounts, and the lift index of each rule in the expert rule set is calculated using the negative sample hit ratio;

[0018] The score of each rule is calculated using the lift index, and the scores of each rule are added together to obtain the risk score result of the predicted account.

[0019] Preferably, a prediction model is constructed based on the directional indicator layer to output a fraud risk probability value, and the fraud risk probability value is integrated with the account risk score result using binning statistics technology to obtain a risk identification and control strategy including:

[0020] Based on the discretization technology, the directional indicator layer is grouped to obtain the indicator group, the evidence weight of the indicator group is calculated according to the grouping result, and the total information value of the indicator group is calculated using the evidence weight;

[0021] Combine the total information value with the gradient boosting decision tree to build a prediction model, and use the prediction model to output the predicted account fraud risk value;

[0022] Normalize the account risk scoring results to obtain the expert scoring results, and use binning statistics technology to integrate the expert scoring results and the predicted account fraud risk values ​​into grouping processes;

[0023] Generate the recognition risk grading result according to the fusion grouping result, and match the corresponding risk control strategy based on the recognition risk grading result.

[0024] Preferably, the calculation formula for the weight of evidence is:

[0025]

[0026] In the formula, WOE i represents the weight of evidence of the i-th index group, represents the proportion of responsive customers in the i-th index group in this index group, represents the proportion of non-responsive customers in the i-th index group in this index group, y i represents the amount of data of responsive customers in the i-th index group, y T represents the total amount of data of responsive customers in the i-th index group, n i represents the amount of data of non-responsive customers in the i-th index group, n T represents the total amount of data of non-responsive customers in the i-th index group.

[0027] Preferably, the calculation formula for the total value of information value is:

[0028]

[0029] In the formula, IV represents the total value of information value, IV i represents the information value of the i-th index group, n represents the number of variable groupings, represents the proportion of responsive customers in the i-th index group in this index group, represents the proportion of non-responsive customers in the i-th index group in this index group, WOE i represents the weight of evidence of the i-th index group.

[0030] Preferably, combine the total value of information value with the gradient boosting decision tree to construct a prediction model, and use the prediction model to output the predicted account fraud risk value, including:

[0031] Combine the total value of index information value with the index screening rule result, retain the index with the total value of index information value greater than the preset value, and divide the training set and the test set according to the retained index;

[0032] Combine the training set with the gradient boosting decision tree for model training to obtain an initial prediction model, and use the test set to debug the area under the curve value and the Lorenz curve value of the initial prediction model;

[0033] After the area under the curve value and the Lorenz curve value reach the target value, stop debugging to obtain the prediction model, use the prediction model to output the fraud risk probability value of the predicted account, and convert the fraud risk probability value into a fraud risk probability score in percentage system.

[0034] Preferably, the account risk scoring result is normalized to obtain an expert scoring result, and the binning statistical technique is used to perform a fusion grouping process on the expert scoring result and the predicted account fraud risk value, including:

[0035] Normalize the account risk scoring result, obtain the expert scoring result on a percentile scale based on the processing result, and sort the expert scoring result and the fraud risk probability score in ascending order as bins;

[0036] Calculate the Lorenz curve value of each bin, and select the bin corresponding to the maximum Lorenz curve value as the eigenvalue. Generate five groups of bins based on the eigenvalue and the chi-square binning statistical technique.

[0037] Preferably, generate an identified risk grading result according to the fusion grouping result, and match the corresponding degree risk control strategy based on the identified risk grading result, including:

[0038] Generate five bins of expert scores corresponding to the fraud risk probability according to the chi-square binning result, and combine the generated result with the two-dimensional binning result for comprehensive decision-making to divide the risk levels;

[0039] Use the risk level division result as the identified risk grading result, and match the corresponding gradient risk control strategy according to the grading result. The identified risk grading result includes the fifth level, the fourth level, the third level, the second level, and the first level.

[0040] Preferably, when the identified risk grading result is the fifth level, directly control the account; when the identified risk grading result is the fourth level, strengthen the verification of the account; when the identified risk grading result is the third level, conduct manual verification; when the identified risk grading result is the second level and the first level, no control is required.

[0041] The beneficial effects of the present invention are:

[0042] 1. By introducing external data, the present invention supplements black and gray samples and card account customer risk data, solves the problem of the single structure being unable to know the bottleneck of fraudsters' cross-institutional and cross-platform fraud, as well as the insufficient black and gray samples of a single institution, which cannot better, more, and more accurately identify fraud patterns.

[0043] 2. The present invention changes the problem that the traditional single expert rule only deals with the identification and prevention of fraud techniques in a certain fraud scenario. Based on the model scoring system, a comprehensive evaluation of more than 200 rule sets of the account is carried out to obtain the comprehensive risk score of the expert scoring, achieving the purpose of solving the low accuracy of a single rule. At the same time, the present invention adopts an index screening system based on the IV value to improve the interpretability and accuracy of the machine learning model.

[0044] 3. By comprehensively evaluating the expert scoring results and machine learning results, the present invention improves the accuracy and interpretability of the overall risk identification solution. Machine learning models, especially complex deep learning models, are often difficult to explain the decision-making process, while expert rules are highly interpretable. When used in combination, they can improve the interpretability of the model while maintaining its performance, making the results of fraud detection easier to understand and accept, fully leveraging expert knowledge and data-driven insights, and improving the accuracy of fraud identification.

[0045] 4. The present invention provides multi-layer protection. Expert rules and machine learning models can complement each other to form a multi-layer fraud detection system, enabling other levels of strategies to still function even if the detection strategy at a certain level fails, thus providing more comprehensive protection, while also being flexible and scalable. Expert rules can be adjusted according to changing scenario requirements and regulations, and machine learning models can expand their capabilities by adding new data or features. When used in combination, they can flexibly adapt to the changing fraud environment and business needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0047] Figure 1 is a flowchart of a method for identifying users in transaction anti-fraud based on a fusion model according to an embodiment of the present invention;

[0048] Figure 2 is a schematic diagram of risk level division in a method for identifying users in transaction anti-fraud based on a fusion model according to an embodiment of the present invention;

[0049] Figure 3 is an overall architecture diagram of a method for identifying users in transaction anti-fraud based on a fusion model according to an embodiment of the present invention;

[0050] Figure 4 is a flowchart of data preprocessing in a method for identifying users in transaction anti-fraud based on a fusion model according to an embodiment of the present invention;

[0051] Figure 5 is a flowchart of the operation of an expert rule engine in a method for identifying users in transaction anti-fraud based on a fusion model according to an embodiment of the present invention;

[0052] Figure 6It is a flowchart of machine learning model training and prediction in a method for identifying users in transaction anti-fraud based on a fusion model according to an embodiment of the present invention;

[0053] Figure 7 It is a logic diagram of the fusion decision-making step in a method for identifying users in transaction anti-fraud based on a fusion model according to an embodiment of the present invention. Specific embodiments

[0054] To further illustrate the embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. They are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention.

[0055] According to an embodiment of the present invention, a method for identifying users in transaction anti-fraud based on a fusion model is provided.

[0056] Now, the present invention will be further described in conjunction with the accompanying drawings and specific embodiments. As Figure 1 shown, the method for identifying users in transaction anti-fraud based on a fusion model according to an embodiment of the present invention includes:

[0057] Step S1, receiving user risk control data, calculating and summarizing the user risk control data, and generating a direction index layer reflecting the quantification of transaction operations based on the summary result;

[0058] Step S2, generating an expert rule set based on the direction index layer and expert rules, and obtaining an account risk score result according to the expert rule set;

[0059] Step S3, constructing a prediction model based on the direction index layer to output a fraud risk probability value, and fusing the fraud risk probability value and the account risk score result using binning statistical techniques to obtain a risk identification and control strategy.

[0060] In one embodiment, when generating an expert rule set based on the direction index layer and expert rules and obtaining an account risk score result according to the expert rule set, fraud accounts and reported accounts involved in cases can be defined as negative samples based on historical transaction anti-fraud data, and accounts excluded by branch inspections and confirmations can be defined as positive samples; expert rules can be customized according to the direction index layer and user transaction scenarios, and the expert rules and the direction index layer can be combined to generate an expert rule set, and the number of triggered accounts and the number of hit accounts in the expert rule set can be judged; the negative sample hit ratio value of the expert rule set can be calculated based on the number of triggered accounts and the number of hit accounts, and the lift index of each rule in the expert rule set can be calculated using the negative sample hit ratio value; the scores of each rule can be calculated using the lift index, and the scores of each rule can be added together to obtain the risk score result of the predicted account.

[0061] In one embodiment, when constructing a prediction model based on the direction index layer to output the fraud risk probability value, and using the binning statistical technique to fuse the fraud risk probability value with the account risk scoring result to obtain the risk identification and control strategy, the direction index layer can be grouped based on the discretization technique to obtain index groups, the evidence weight of the index groups can be calculated according to the grouping result, and the total information value of the index groups can be calculated using the evidence weight; the total information value is combined with the gradient boosting decision tree to construct a prediction model, and the prediction model is used to output the predicted account fraud risk value; the account risk scoring result is normalized to obtain the expert scoring result, and the binning statistical technique is used to perform a fusion grouping process on the expert scoring result and the predicted account fraud risk value; the identification risk grading result is generated according to the fusion grouping result, and the corresponding degree risk control strategy is matched based on the identification risk grading result.

[0062] In one embodiment, when combining the total information value with the gradient boosting decision tree to construct a prediction model and using the prediction model to output the predicted account fraud risk value, the total index information value and the index screening rule result can be used to retain the indexes with the total index information value greater than the preset value, and the training set and the test set are divided according to the retained indexes; the training set is combined with the gradient boosting decision tree for model training to obtain the initial prediction model, and the area under the curve value and the Lorenz curve value of the initial prediction model are debugged using the test set; after the area under the curve value and the Lorenz curve value reach the target value, the debugging is stopped to obtain the prediction model, the prediction model is used to output the fraud risk probability value of the predicted account, and the fraud risk probability value is converted into a fraud risk probability score in percentage system.

[0063] In one embodiment, when normalizing the account risk scoring result to obtain the expert scoring result and using the binning statistical technique to perform a fusion grouping process on the expert scoring result and the predicted account fraud risk value, the account risk scoring result can be normalized, and the expert scoring result in percentage system is obtained based on the processing result. The expert scoring result and the fraud risk probability score are sorted from low to high as bins; the Lorenz curve value of each bin is calculated, and the bin corresponding to the maximum Lorenz curve value is selected as the eigenvalue, and five bins are generated based on the eigenvalue and the chi-square binning statistical technique as the division basis.

[0064] In one embodiment, when generating the identification risk grading result according to the fusion grouping result and matching the corresponding degree risk control strategy based on the identification risk grading result, the expert scoring five bins and the prediction model scoring five bins corresponding to the fraud risk probability can be generated according to the chi-square binning result, and the generated result is combined with the two-dimensional binning result for comprehensive decision-making to divide the risk level; the risk level division result is used as the identification risk grading result, and the corresponding gradient risk control strategy is matched according to the grading result.

[0065] As Figure 3As shown in the figure, it should be noted that the purpose of this embodiment is to provide an anti-fraud user identification model that integrates expert experience and machine learning models when identifying fraud transactions. The aim is to combine the advantages of quickly identifying known fraud transaction methods through expert rules and the advantages of machine learning algorithms in improving fraud identification efficiency and identifying new frauds, so as to improve the accuracy and efficiency of fraud user identification. To facilitate the understanding of the above technical solution of the present invention, the working principle or operation method of the present invention in the actual process will be described in detail as follows:

[0066] Step 1: Data processing;

[0067] As Figure 4 shown, based on the capabilities of the data middle platform, in-house data and external data are processed into the index layer according to four categories: static indicators, transaction indicators, behavior indicators, and static indicators, and batch processing is performed daily to lay the foundation for subsequent processing.

[0068] (1) Data layer:

[0069] The data layer refers to the original data stored in the database or other storage systems. This data is unprocessed and contains a large amount of detailed information, mainly including but not limited to the following types of data:

[0070] Card-account-customer data: mainly includes the business attribute data of bank cards, accounts, and customers stored in-house. At the same time, external data is innovatively introduced to supplement the lack of risk portrait data of card-account-customers by a single institution. Taking bank card information as an example, the basic information such as the card type, account opening time, card-issuing institution, limit, and whether it is under control of the bank cards stored in-house is stored, but it cannot reflect key risk information such as whether the bank card has been traded on the dark web and whether the card has a close relationship with the risk cards of other banks. By introducing external data, the lack of in-house information is supplemented. The external data used in this embodiment includes but is not limited to customer risk portraits, bank risk portraits, IP risk portrait data, etc.;

[0071] Transaction flow data: such as core transactions, online payment transactions, mobile banking transaction flows, etc.;

[0072] Behavior data: such as behavior records of customers logging in to mobile phones and clicking on banks, logging in to and clicking on online banking, modifying passwords, modifying limits, etc.;

[0073] Black and gray sample data: such as black and gray samples (involved accounts, involved accounts of other banks or branch inspections, etc.), mainly referring to the known bank cards used for transaction anti-fraud. In order to identify more fraud methods and paths through black and gray samples, in addition to the feedback from branch inspections, data such as involved case notifications and cut-off card clue data are mainly introduced for further supplementation to enrich the sample data.

[0074] (2) Index layer:

[0075] The data at the index layer is sourced from the data layer and represents further processing and analysis of the data in the data layer. It is usually calculated and aggregated based on the data in the data layer to reflect the quantitative information of the business. In this embodiment, based on the basic data in the data layer, more than 200 indicators are processed according to 4 risk identification directions, as follows:

[0076] Static indicators: such as account type, mobile phone number, age, occupation, payroll label, etc. Taking the processing of the occupation indicator as an example, based on the information such as the customer's work unit and industry collected in the data layer, the occupation indicator is aggregated and processed, specifically including public institutions, government agencies, universities, military enterprises, etc.;

[0077] Behavior indicators: such as the number of mobile banking logins on the same day, the number of password modifications in the past 1 month, the modified amount in the past 1 month, the number of ATM cash withdrawals on the same day, the same-day over-the-counter cash withdrawals, the number of early morning cash withdrawals, the modification of the reserved mobile phone number (as well as frequent early morning cash withdrawals, modification of the reserved mobile phone or payroll, etc.);

[0078] Transaction indicators: such as the number of account trading counterparts, the number of account transactions, the account transaction amount, the number of over-the-counter transactions, the proportion of small transactions, the transaction time difference, the proportion of quick in and quick out (as well as the proportion of transactions in high-risk areas and night transactions), etc.;

[0079] Association indicators: such as association of the same registered mobile phone number, association of the same registered address, association of the same registered transaction IP address, and association of the same legal person, etc.

[0080] Step 2: Expert rule scoring;

[0081] (1) Selection of good and bad samples: In the expert scoring of this embodiment, the fraud accounts identified and verified by the branch and the reported involved accounts are used as negative samples (bad samples), and the accounts excluded by the branch after investigation and verification are used as positive samples (good samples).

[0082] (2) As Figure 5 shown, for various fraud scenarios, expert rules are customized based on the index layer to form an expert rule set. For example, for the risk monitoring of quick in and quick out of funds, for non-whitelisted customers, indicators such as the number of transactions in the past 1 hour, the transaction time difference, and the out-of-account amount / into-account amount are combined using logical operators "OR" and "AND" to form 1 expert rule for this scenario. And so on, finally forming a rule set of more than 200 rules.

[0083] (3) Calculate the bad_rate of the rule set: For the rule set of more than 200 rules, according to the number of triggered accounts acct_cnt, the number of "bad" accounts hit badacct_cnt, and the proportion of "bad" accounts hit bad_rate for each rule:

[0084]

[0085] In the formula, badacct_cnt k represents the number of bad samples where the k-th rule is triggered and hits, and acct_cnt k represents the total number of all accounts where the k-th rule triggers a warning, and badratet k represents the proportion of bad samples hit by the k-th rule.

[0086] (4) Based on badrate0 of the rule set, calculate the lift index (lift value) of a single rule:

[0087] lift k = badrate k / badrate0;

[0088] Among them, badrate0 represents the proportion of all bad samples to the total number of accounts triggering warnings.

[0089] (5) Assume that the single rule with the largest lift value is scored as score0, and calculate the score of a single rule as:

[0090]

[0091] In the formula, rule_score0 represents the score corresponding to the single rule with the largest lift value among all existing rules. In this embodiment, it is set to 10, and it can be adjusted appropriately according to the distribution of the lift values of the rule set. The principle is to make the scores more distinguishable. lift0 is the largest lift value among all existing rules, and lift k is the lift value corresponding to the k-th rule.

[0092] (6) Aggregate the account-level risk scores: For each predicted account, traverse all the expert rules hit by this account, add up the scores of the corresponding expert rules, and obtain the final score of this account. The higher the final score, the greater the identified risk. Specifically:

[0093]

[0094] In the formula, rule_score k represents the score of the k-th rule; m represents the total number of all rules, and acct_score j represents the expert rule score of the j-th account.

[0095] Step Three: Prediction by the machine learning model;

[0096] (1) Selection of good and bad samples: Similarly, take the fraud accounts identified through branch inspections and the reported accounts involved in cases as negative samples, and the accounts excluded through branch inspections as positive samples.

[0097] (2) Index processing and feature engineering: As Figure 6 shown (first obtain the source data, then perform data processing to obtain a rich feature set, clean the feature set to obtain clean data, and label the positive and negative samples for model training data prediction results), model using the more than 200 indicators generated in the data processing steps. In this embodiment, innovatively in this scenario, according to the IV value of each indicator, that is, the Information Value, perform the first round of indicator screening. The larger the IV value, the stronger the prediction ability of the variable. In this embodiment, indicators with an IV value above 0.1 are used to enter the model for modeling, specifically as follows:

[0098] Calculation of WOE: First, group the indicators (also called discretization, binning). After grouping, for the grouping result of the i-th group, the calculation formula of WOE is as follows:

[0099]

[0100] In the formula, WOE i represents the weight of evidence of the i-th indicator group, represents the proportion of responding customers in the i-th indicator group in this indicator group, represents the proportion of non-responding customers in the i-th indicator group in this indicator group, y i represents the data volume of responding customers in the i-th indicator group, y T represents the total data volume of responding customers in the i-th indicator group, n i represents the data volume of non-responding customers in the i-th indicator group, n T represents the total data volume of non-responding customers in the i-th indicator group.

[0101] Calculation of IV value: For a grouped variable, the IV value of the i-th group is calculated as follows:

[0102]

[0103] According to the IV values of the variable in each group, the IV value of the entire variable is:

[0104]

[0105] In the formula, IV represents the total information value, IV i represents the information value of the i-th indicator group, n represents the number of variable groups, represents the proportion of responding customers in the i-th indicator group in this indicator group, represents the proportion of non-responding customers in the i-th indicator group in this indicator group, WOE i represents the weight of evidence of the i-th indicator group.

[0106] Index screening is performed according to the IV value of each index. In this embodiment, only more than 50+ indexes with an IV value above 0.1 are retained for subsequent modeling, mainly including the cumulative transaction amount with a third-party payment institution as the debit counterparty in the account in the past 1 month (IV value 0.153), the cumulative number of ATM transactions at 23:01-01:00 in the account in the past 1 month (IV value 0.121), the cumulative number of transactions with a decimal transaction amount in the account in the past 1 month (IV value 0.132), and so on.

[0107] (3) Modeling: Randomly select 80% from the data set (the final data result of data processing) as the training set, and the remaining 20% as the test set. Based on the training set data, use the GBDT (Gradient Boosting Decision Tree) model for training, and continuously debug the parameter values of the area under the curve AUC (Area Under Curve) and the Kolmogorov-Smirnov curve KS that meet the expectations based on the test set, and use this model for final prediction. In this embodiment, the expected value of AUC is 0.8 or above, and the expected value of KS is 0.4 or above.

[0108] (4) Prediction: For each predicted account, the model outputs the fraud risk probability value, which is converted according to the percentage system into a risk score between 1 and 100.

[0109] Step Four: Fusion Decision-making;

[0110] Traditional anti-fraud identification usually uses expert rule scoring or machine learning results alone for control, which may have problems such as poor accuracy or poor model interpretability. In this embodiment, Best-KS binning is used to group the results of expert rule scoring and machine learning prediction into 5 groups respectively for the machine learning results, as Figure 7 shown, to form a comprehensive judgment of risk identification and further improve the reliability and interpretability of the fusion results, as follows:

[0111] (1) Normalization processing: Since the expert scoring module is not in the percentage system and there is a dimensional difference from the percentage scores of the machine learning prediction results, this embodiment first performs normalization processing on the expert scoring results and uniformly converts them into the percentage system:

[0112]

[0113] In the formula, max{acct_score1, acct_score2, …, acct_score n} represents the highest expert score of the existing accounts, and min{acct_score1, acct_score2, …, acct_scoren} represents the lowest score of the existing account expert scoring, acct_score_hun j represents the score of the j-th account.

[0114] (2) Fusion method: Use chi-square binning to perform 5-group binning on the expert scoring and machine learning scoring respectively. First, sort the expert scoring / machine learning scoring from small to large, take each value as a bin, and calculate the KS value of each bin:

[0115]

[0116] In the formula, goodacct_cnt l represents the number of good customers in the l-th bin, goodacct_cnt represents the number of all good customers, badacct_cnt l represents the number of bad customers in the l-th bin, badacct_cnt represents the number of all bad customers.

[0117] Find the bin corresponding to the maximum KS value, that is, the eigenvalue, and use this eigenvalue as the division basis to divide the data into two parts of data SET1 and SET2 (lower than this eigenvalue and higher than this eigenvalue), and repeat the recursive division of the left and right data sets until they are divided into 5 groups.

[0118] (3) Fusion model decision logic: According to the results of chi-square binning, obtain 5 bins of expert rule scoring (the risk is getting higher and higher) and 5 bins of machine learning model scoring (the risk is getting higher and higher) respectively. Based on the two-dimensional binning results, make a comprehensive decision, combined with the sample distribution after grading, as Figure 2 shown, divided into the following 5 risk levels:

[0119] First level: Expert level 1 and machine learning level 1, Expert level 2 and machine learning level 1, Expert level 1 and machine learning level 2;

[0120] Second level: Expert level 3 and machine learning level 1, Expert level 2 and machine learning level 2, Expert level 1 and machine learning level 3, Expert level 4 and machine learning level 1, Expert level 3 and machine learning level 2, Expert level 2 and machine learning level 3, Expert level 1 and machine learning level 4;

[0121] Third level: Expert level 5 and machine learning level 1, Expert level 4 and machine learning level 2, Expert level 3 and machine learning level 3, Expert level 2 and machine learning level 4, Expert level 1 and machine learning level 5;

[0122] Fourth level: Expert Level 5 and Machine Learning Level 2, Expert Level 5 and Machine Learning Level 3, Expert Level 4 and Machine Learning Level 3, Expert Level 4 and Machine Learning Level 4, Expert Level 3 and Machine Learning Level 4, Expert Level 3 and Machine Learning Level 5, Expert Level 2 and Machine Learning Level 5;

[0123] Fifth level: Expert Level 5 and Machine Learning Level 4, Expert Level 5 and Machine Learning Level 5, Expert Level 4 and Machine Learning Level 5.

[0124] (4) Risk control logics for different levels: According to the final risk grading results, corresponding gradient risk control strategies are matched. For example, accounts determined to be of high risk in Level Five are directly controlled; accounts in Level Four are subject to enhanced verification and are directly controlled if they fail; accounts in Level Three are subject to manual verification; accounts in Level Two and Level One are of low risk and do not require control.

[0125] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for identifying users in transaction anti-fraud based on a fusion model, characterized in that, Including: Receiving user risk control data, calculating and summarizing the user risk control data, and generating a direction index layer that reflects the quantification of transaction business based on the summary result; Generating an expert rule set based on the direction index layer and expert rules, and obtaining an account risk score result according to the expert rule set; Constructing a prediction model based on the direction index layer to output a fraud risk probability value, and using binning statistical techniques to fuse the fraud risk probability value and the account risk score result to obtain a risk identification and control strategy.

2. The method for identifying transaction anti-fraud users based on a fusion model according to claim 1, wherein The user risk control data includes card account customer data, user transaction flow data, user behavior data, and black and gray sample data; the direction index layer includes static indicators, behavior indicators, transaction indicators, and correlation indicators.

3. The method for identifying users in transaction anti-fraud based on a fusion model according to claim 2, wherein, The generating an expert rule set based on the direction index layer and expert rules, and obtaining an account risk score result according to the expert rule set includes: Defining fraud accounts and reported involved accounts as negative samples based on historical transaction anti-fraud data, and defining accounts excluded by branch inspections as positive samples; Customizing expert rules according to the direction index layer and user transaction scenarios, combining the expert rules with the direction index layer to generate an expert rule set, and judging the number of triggered accounts and the number of hit accounts in the expert rule set; Calculating the negative sample hit ratio value of the expert rule set based on the number of triggered accounts and the number of hit accounts, and calculating the lift index of each rule in the expert rule set using the negative sample hit ratio value; Calculating the scores of each rule using the lift index, and adding up the scores of each rule to obtain the risk score result of the predicted account.

4. A method for identifying users in transaction anti-fraud based on a fusion model according to claim 1, wherein, The constructing a prediction model based on the direction index layer to output a fraud risk probability value, and using binning statistical techniques to fuse the fraud risk probability value and the account risk score result to obtain a risk identification and control strategy includes: Performing grouping processing on the direction index layer based on discretization technology to obtain index groups, calculating the evidence weight of the index groups according to the grouping results, and calculating the total information value of the index groups using the evidence weight; Combining the total information value with a gradient boosting decision tree to construct a prediction model, and using the prediction model to output the fraud risk value of the predicted account; Performing normalization processing on the account risk score result to obtain an expert score result, and using binning statistical techniques to perform fusion grouping processing on the expert score result and the fraud risk value of the predicted account; Generating an identification risk grading result according to the fusion grouping result, and matching a corresponding degree risk control strategy based on the identification risk grading result.

5. The method for identifying users in transaction anti-fraud based on a fusion model according to claim 4, characterized in that, The calculation formula for the evidence weight is: Where, WOE i represents the weight of evidence for the i-th indicator group, represents the proportion of responding customers in the i-th indicator group within this indicator group, represents the proportion of non-responding customers in the i-th indicator group within this indicator group, y i represents the data volume of responding customers in the i-th indicator group, y T represents the total data volume of responding customers in the i-th indicator group, n i represents the data volume of non-responding customers in the i-th indicator group, n T represents the total data volume of non-responding customers in the i-th indicator group.

6. The method for identifying users in transaction anti-fraud based on a fusion model according to claim 5, wherein The calculation formula for the total information value is: Wherein, IV represents the total information value, IV i represents the information value of the i-th index group, n represents the number of variable groups, represents the proportion of responsive customers in the i-th index group in this index group, represents the proportion of non-responsive customers in the i-th index group in this index group, WOE i represents the weight of evidence of the i-th index group.

7. A method for identifying transaction anti-fraud users based on a fusion model according to claim 6, characterized in that, The combining the total information value with a gradient boosting decision tree to construct a prediction model, and using the prediction model to output the fraud risk value of the predicted account includes: Combining the index total information value with the index screening rule result, retaining the indexes with the index total information value greater than the preset value, and dividing the training set and the test set according to the retained indexes; Combining the training set with a gradient boosting decision tree to perform model training to obtain an initial prediction model, and debugging the area under the curve value and the Lorenz curve value of the initial prediction model using the test set; Stop debugging to obtain a prediction model after the area value under the curve and the Lorenz curve value reach the target value. Use the prediction model to output the fraud risk probability value of the predicted account, and convert the fraud risk probability value into a fraud risk probability score on a percentage scale.

8. The method for identifying users in transaction anti-fraud based on a fusion model according to claim 7, wherein Normalize the account risk scoring result to obtain an expert scoring result, and use the binning statistical technique to perform a fusion grouping process on the expert scoring result and the predicted account fraud risk value, including: Normalize the account risk scoring result, obtain an expert scoring result on a percentage scale based on the processing result, and sort the expert scoring result and the fraud risk probability score in ascending order as bins. Calculate the Lorenz curve value of each bin, and select the bin corresponding to the maximum Lorenz curve value as the eigenvalue. Generate five groups of bins based on the eigenvalue and the chi-square binning statistical technique as the division basis.

9. A method for identifying users in transaction anti-fraud based on a fusion model according to claim 8, characterized in that Generate an identification risk grading result according to the fusion grouping result, and match the corresponding degree of risk control strategy based on the identification risk grading result, including: Generate five bins of expert scores corresponding to the five bins of the fraud risk probability according to the chi-square binning result, and combine the generated result with the two-dimensional binning result for comprehensive decision-making to divide the risk level. Take the risk level division result as the identification risk grading result, and match the corresponding gradient risk control strategy according to the grading result. The identification risk grading result includes the fifth level, the fourth level, the third level, the second level, and the first level.

10. A method for identifying transaction anti-fraud users based on a fusion model according to claim 9, characterized in that, When the identification risk grading result is the fifth level, directly control the account. When the identification risk grading result is the fourth level, strengthen the verification of the account. When the identification risk grading result is the third level, manual verification is required. When the identification risk grading result is the second level and the first level, no control is required.

Citation Information

Cited By

  • Electric fraud transaction beforehand detection method, device and equipment and storage medium

    CN121190061A