Internet credit region fraud identification model construction method

By constructing an Internet credit regional fraud identification model, using technologies such as layered retention sampling, chi-square inspection and feature cross-section, the problem of traditional systems being difficult to integrate heterogeneous data and adapt to changes in fraud patterns is solved, and fraud identification with high predictability and interpretability is achieved.

CN120450853APending Publication Date: 2025-08-08重庆富民银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510683099.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Traditional Internet credit fraud identification systems are difficult to effectively integrate heterogeneous data and cannot adapt to changes in fraud patterns in real time, resulting in insufficient identification accuracy and comprehensiveness.

Method used

A model for identifying fraud in the Internet credit region is constructed, using hierarchical retention sampling, chi-square test, feature crossover and time window statistics derived variables, combined with random forest ensemble learning and Bayesian smoothing technology, to configure risk weights, monitor feature distribution in real time, and update the model dynamically.

Benefits of technology

It has achieved high predictive, strong interpretability and good adaptability of fraud identification, improved the accuracy and adaptability of fraud identification, met regulatory requirements, and reduced the cost of compliance review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120450853A_ABST
    Figure CN120450853A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to an Internet credit region fraud identification model construction method. Comprising the steps of constructing a risk tag system, and performing feature engineering preprocessing; dividing the preprocessed data set into a model development sample and an out-of-time test sample; fraud association features are mined, the correlation between variables and regional fraud tags is analyzed by applying chi-square test, and information values are calculated to screen out variables with significant correlation and strong predictive ability; deriving expert features, and generating feature crossover and time window statistic derived variables; performing variable screening and optimization; configuring risk weights to the screened variables; carrying out model training and optimization; carrying out model verification and calibration; deploying a production environment; and carrying out model full-life-cycle management, and regularly injecting new credit data and expert analysis experience. According to the technical scheme, a fraud identification model with high predictability, high interpretability and good adaptability can be constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and device for constructing an Internet credit regional fraud identification model. Background Art

[0002] In the context of internet lending, fraud is becoming increasingly complex and diverse, posing a significant challenge to fraud detection. Traditional fraud detection systems, based on single-dimensional rules, are gradually exposing their limitations in addressing this complex and volatile fraud landscape.

[0003] On the one hand, data heterogeneity poses a significant obstacle to traditional systems. Internet credit operations involve data from a wide range of sources, encompassing user basic information, credit records, transaction behavior, social network data, and geographic location information. This data exhibits significant differences in format, semantics, and structure. Traditional systems struggle to effectively integrate and utilize this heterogeneous data, unable to fully uncover the fraud patterns and correlations hidden within the data, thus limiting the accuracy and comprehensiveness of fraud detection.

[0004] On the other hand, the dynamic nature of fraud patterns poses a significant challenge to traditional static models. Fraudsters constantly adjust their methods to adapt to financial institutions' anti-fraud strategies, resulting in a continuous evolution of fraud patterns. Traditional fraud detection models, however, are mostly built based on historical data and are static in nature, making them difficult to adapt to changes in fraud patterns in real time. Over time, the model's features may become ineffective, performance degrading, and the model may be unable to accurately identify new fraudulent activities, exposing financial institutions to significant risks of fraud losses. Summary of the Invention

[0005] The purpose of the present invention is to propose a method and device for constructing an Internet credit regional fraud identification model, which can construct a fraud identification model with high predictability, strong interpretability and good adaptability.

[0006] To achieve the above objectives, in a first aspect, the present invention provides a method for constructing an Internet credit regional fraud identification model, comprising: Build a risk labeling system and perform feature engineering preprocessing; A stratified holdout sampling strategy is used to divide the preprocessed dataset into model development samples and out-of-time testing samples; Mining fraud-related features, using the chi-square test to analyze the correlation between variables and regional fraud labels, and calculating information values to screen out variables with significant correlation and strong predictive power; Derive expert features, generate feature cross- and time window statistics derived variables; Conduct variable screening and optimization, analyze PSI stability, test VIF collinearity, and use stepwise regression to select the optimal variable subset; Assign risk weights to the selected variables; Perform model training and tuning, using the random forest ensemble learning framework for model training and the grid search method for iterative optimization of hyperparameters; Perform model validation and calibration, perform out-of-sample testing using out-of-time test samples, and optimize the probability calibration curve using Bayesian smoothing techniques; Deploy the production environment, build the model API service layer, deploy the model using a grayscale release strategy, and establish a real-time feature computation pipeline; Carry out full life cycle management of the model and regularly inject new credit data and expert analysis experience.

[0007] Beneficial effects of the basic solution: Experience quantification and two-way validation. During the feature derivation phase, experts' empirical assumptions about regional fraud patterns are converted into verifiable statistical hypotheses. Through quantitative methods such as chi-square tests and information value calculations, business logic is embedded in feature engineering. This closed loop of "expert hypothesis → data verification → feature solidification" retains the interpretability of traditional expert scoring cards while avoiding the bias of subjective experience, ensuring that model features are both business-relevant and data-significant.

[0008] The two-tiered weighting system constructs a two-tiered system of "variable weight + feature cross-weight". In the variable screening stage, the basic weight is determined through stepwise regression, and additional weight is given to the composite features designed by experts. This realizes the deep integration of data-driven and business experience, ensuring that the model scoring results can not only reflect the objective data laws, but also conform to the risk control business logic.

[0009] A real-time feature pipeline and time series monitoring uses streaming computing technology to build a real-time feature computation pipeline, dynamically capturing temporal changes in user GPS trajectories and lending behavior. Combined with a rolling time window strategy, it continuously monitors changes in feature distribution. When abnormal fluctuations are detected, it automatically triggers local model updates, allowing rapid adaptation to the spread of new fraudulent tactics.

[0010] Drift Response and Model Iteration: The Population Stability Index (PSI) monitors feature distribution stability in real time. Once data drift is detected, the model calibration process is immediately initiated. Combined with expert analysis of new fraud patterns, feature engineering rules are dynamically adjusted to proactively combat the dynamic evolution of fraudulent tactics.

[0011] This hybrid architecture utilizes "feature screening (random forest + stepwise regression) → performance calibration (Bayesian smoothing) → expert scorecard decision-making." Machine learning algorithms are used at the feature level to uncover implicit patterns in the data. Bayesian smoothing is then used at the result level to optimize probability outputs. Finally, expert scorecard rules are used to manually review or adjust thresholds for high-risk cases. This ensures both model prediction accuracy and regulatory requirements for interpretability of risk decisions.

[0012] The Expert Scorecard module, which enhances compliance and manages risk, serves as a "safety valve" for final decision-making. It flexibly adjusts scoring rules based on regulatory policies (such as credit restrictions in specific regions) to ensure that model outputs meet regulatory requirements. Furthermore, by visualizing feature importance and weight distribution, it provides regulators with a clear basis for risk decision-making, reducing compliance review costs.

[0013] A weighted scoring system is constructed based on multi-dimensional data such as geographic location, lending behavior, and time-series data. Variable weights are determined through rigorous statistical tests (chi-square test, VIF collinearity analysis) and algorithmic screening (stepwise regression, information value ranking). This system prevents individual subjective intentions, experience, or ability from interfering with fraud detection results. This system ensures that model scoring relies solely on objective data features, significantly improving the consistency and reliability of fraud detection. It is particularly suitable for standardized risk assessment in large-scale internet lending scenarios.

[0014] As an implementable and preferred solution, a risk labeling system is constructed, specifically including the following contents: Extract credit application data covering different regions and credit products from the core business database of financial institutions Perform preliminary cleaning on the retrieved data to remove duplicate records, incorrectly formatted data, and data records with serious missing key information; Cross-check sample data based on multi-dimensional standards, including geographical dimensions, application behavior dimensions, and credit business dimensions; The sample data is divided into positive samples and negative samples. For positive samples, the Y variable is assigned a value of 1; for negative samples, the Y variable is assigned a value of 0.

[0015] As an implementable and preferred solution, feature engineering preprocessing is performed, which specifically includes the following: Conduct a comprehensive check on the sample data, count the number and proportion of missing values in each field, check whether the data format complies with the specifications, and compare the consistency of the same information in different data sources; For fields with missing values, multiple imputation is used to process them. Based on the correlation with other information of the applicant, similar sample data is screened out. Several groups of values are generated using statistical models. The missing values are imputed multiple times, and the most appropriate group is selected as the final imputation result. Numerical variables were processed using the Z-Score standardization method or the Min-Max normalization method.

[0016] As an implementable and optimal solution, mining fraud-related features specifically includes the following: The chi-square test method is used to analyze the correlation between variables in multiple dimensions and regional fraud labels. The formula is:

[0017] in, is the observation frequency, is the expected frequency, r is the number of rows, and c is the number of columns. By calculating the chi-square value, if the chi-square value is greater than the critical value at a certain significance level, it means that there is a significant correlation between the regional variable and the fraud label. Calculate the information value IV value to measure the predictive ability of the variable to the target variable. The formula is:

[0018]

[0019] Where n is the number of boxes, is the evidence weight of the ith box, is the proportion of fraud samples in the i-th box to the total fraud samples, is the proportion of normal samples in the i-th box to the total normal samples; the larger the IV value, the stronger the predictive ability of the variable for fraud identification.

[0020] As an implementable optimal solution, a derivative variable based on feature cross-pollination is generated, and feature cross-pollination operations are performed in combination with credit business logic; a derivative variable of time window statistics is generated, and time windows of different lengths are set to calculate changes in the equipment application region within the time window.

[0021] As an implementable preferred solution, risk weights are configured, specifically including the following: The selected variables are binned using weight of evidence conversion, and the ratio of fraudulent samples to normal samples in each bin is calculated, and then the WOE is calculated using the formula:

[0022] in, is the proportion of fraud samples in the box to the total fraud samples, is the ratio of normal samples in the box to the total normal samples; A two-tier weighting system is constructed. Based on the WOE binning results, risk scale values are assigned according to the statistical significance of the variables. The risk scale values are normalized to obtain the variable-level weights of different bins within each variable. Starting from different feature dimensions, the weight of each feature dimension is determined, and the variable-level weight is multiplied by the feature dimension weight to obtain the final comprehensive weight of each variable.

[0023] As an implementable and preferred solution, model validation and calibration is carried out, which specifically includes the following: Use out-of-sample testing of the optimized model using out-of-sample testing samples to evaluate the performance of the model on data that was not used in training; Perform cross-group validation, divide the out-of-time test samples into multiple groups by region, and calculate the KS statistic and AUC value for each group; The Bayesian smoothing technique is used to optimize the probability calibration curve, and the formula is:

[0024] in, is the number of samples in the box, is the actual fraud rate in the box, is the prior sample size, is the prior fraud rate; The predicted probability is divided into bins according to intervals. For the samples in each bin, the actual proportion of fraud is counted. The Bayesian formula is used to combine the prior probability and the likelihood function to calculate the posterior probability of each bin as the calibrated predicted probability.

[0025] As an implementable and preferred solution, deploying a production environment specifically includes the following: Use microservice architecture to build the model API service layer; A grayscale release strategy is used to deploy the model in the risk control decision engine. Credit applicants are divided into a grayscale testing group and an original model group according to certain rules. For application requests from the grayscale testing group, both the newly constructed fraud detection model and the original fraud detection system are used for prediction. The prediction results and business indicators of the two are compared. If the new model performs better than the original model and all indicators are stable and reliable, the grayscale testing group proportion is gradually expanded until the original model is completely replaced. If significant problems are found in the new model, the original model is rolled back and the new model is re-analyzed and optimized. Conduct A / B testing in a production environment, randomly dividing users into Group A and Group B, comparing key business metrics between the two groups, and quantitatively evaluating the effectiveness of the new model in improving fraud detection capabilities and ensuring customer experience. A real-time feature computing pipeline is established based on the Hive+Spark hybrid computing architecture. Various data from the credit application process are collected in real time through data collection tools and transmitted to Spark Streaming for processing. The real-time collected data is processed according to pre-defined feature engineering logic to generate the feature data required by the model. The feature data obtained by real-time calculation is stored in a distributed database for real-time query and call by the model API service layer.

[0026] As an implementable and preferred solution, the model's full life cycle management is carried out, specifically including the following: Build a monitoring dashboard to track the model's health indicators in real time, including regularly calculating the PSI value of each variable to monitor the stability of the variable distribution, regularly calculating the KS value of the model on out-of-time test samples or real-time production data to evaluate the model's discrimination ability, and dynamically plotting the model's AUC-ROC curve to evaluate the fluctuation of the model's overall performance; Establish a rolling time window mechanism to regularly inject newly generated credit data and experts' analysis experience on the latest regional fraud patterns into the model, collect new credit application data and fraud annotation data for processing to generate new sample data. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a logic diagram of a method for building an Internet credit regional fraud identification model.

[0028] Figure 2 FIG. 2 is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0029] In order to make the technical solution and advantages of the present application clearer, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings. It will be understood that the specific embodiments described herein are only partial embodiments of the present invention, which are only used to explain the present application, rather than to limit the present application. It should be noted that the technical features or combinations of technical features described in the following embodiments should not be considered to be isolated, and they can be combined with each other to achieve better technical effects. The same reference numerals appearing in the drawings of the following embodiments represent the same features or components, which can be applied to different embodiments.

[0030] In addition, unless otherwise defined, technical or scientific terms used in the description of the present invention should have the common meanings understood by those skilled in the art in the art to which the present invention belongs.

[0031] The present invention will be further described in detail below with reference to the accompanying drawings: Reference numerals: electronic device 500 , processor 501 , communication interface 502 , memory 503 , bus 504 .

[0032] Reference Figure 1 The present disclosure provides a method for constructing an Internet credit regional fraud identification model, including: Step S100: Building a risk labeling system, including: In step S101, credit application data covering various regions and loan product types from the past several years is retrieved from the financial institution's core business database. This data must include complete application information, such as applicant basic information (name, ID number, contact information, home address, etc.), requested credit amount, application date, approval result, repayment history, and location-related information (IP address, device GPS location information, permanent address, etc.). The retrieved data is initially cleaned to remove duplicate records, incorrectly formatted data, and data records with significant missing key information to ensure the accuracy and completeness of the sample data.

[0033] Step S102 cross-checks the sample data based on multi-dimensional criteria. From a regional perspective, the association between the applicant's geographic background and the likelihood of fraud is analyzed, taking into account factors such as the region's economic development level, credit business activity, and fraud history. From an application behavior perspective, the regularity of application timing, whether the application frequency is unusual, and the compatibility of the application amount with the applicant's financial situation are examined. From a credit business perspective, the risk characteristics of different credit products and the legitimacy of the applicant's application for each product are considered.

[0034] In step S103, the sample data is divided into positive samples (samples confirmed to contain regional fraud) and negative samples (normal samples without fraud). For positive samples, the Y variable is assigned a value of 1; for negative samples, the Y variable is assigned a value of 0. This clarifies the target variable for supervised learning and provides accurate labeled data for subsequent model training.

[0035] Step S104 stratifies the sample data based on key attributes such as region and credit product type. For example, by region, the data can be stratified by provincial administrative divisions. If the data volume is large enough, this can be further broken down to city or county levels. For credit product type, the data can be categorized into categories such as personal consumer credit and business operating credit.

[0036] For each stratum, samples are drawn according to a certain ratio to ensure that both positive and negative samples are representative of the business within each stratum. This stratified sampling not only avoids model training bias caused by excessive or insufficient data in certain strata, but also fully reflects the differences in fraud characteristics across different regions and credit product types, improving the model's generalization capabilities.

[0037] Step S200, performing feature engineering preprocessing, including: Step S201: Perform a comprehensive check on the sample data. The number and percentage of missing values for each field are counted. For data accuracy, the data format is checked for compliance, such as whether IP addresses conform to the standard dotted decimal format and whether ID card numbers comply with encoding rules. For data consistency, the consistency of the same information across different data sources is compared. For example, the applicant's permanent address recorded in the credit application system is compared with the address in the customer information management system.

[0038] Multiple imputation is used for fields with missing values. For example, if the applicant's permanent address field contains missing values, we first screen for similar sample data based on correlation with other applicant information (such as IP address and mobile phone number location). Based on the permanent address distribution of these similar samples, we apply statistical models (such as Bayesian estimation) to generate several sets of possible address values, and perform multiple imputations for the missing values. From these [M] sets of imputed values, the most appropriate set is selected as the final imputation result. This selection can be evaluated by calculating metrics such as the correlation between the imputed data and other relevant variables.

[0039] Step S202: For numerical variables, Z-Score standardization method is used, and the formula is:

[0040] in, is the original variable value, is the mean of the variable, is the standard deviation of the variable. This formula converts the variable values into standard normal distribution data with a mean of 0 and a standard deviation of 1, eliminating the impact of dimensional differences between different variables on model training.

[0041] For variables with specific value ranges, the Min-Max normalization method is used, and the formula is:

[0042] in, is the original variable value, is the minimum value of the variable, The maximum value of the variable. Map the variable value to the interval [0, 1] so that different variables can be compared on the same scale.

[0043] In step S300, a stratified retention sampling strategy is adopted to divide the original dataset after feature engineering preprocessing into model development samples (including training set and validation set) and out-of-time test samples according to a certain ratio (7:3 in this embodiment).

[0044] Stratification is performed based on attributes such as region and credit product type. For each stratum, 70% of the data is sampled as the model development sample, with the remaining 30% used as the out-of-time test sample. Within the model development sample, 70% of the data is further divided into a training set for model training, and 30% into a validation set to evaluate model performance during training and prevent overfitting. For example, within a specific region's business credit stratification, 700 data points are sampled as the model development sample, including 490 for the training set, 210 for the validation set, and 300 for the out-of-time test sample. This stratified, conservative sampling approach ensures representativeness of the model development sample across different regions and credit product types, while also enabling the out-of-time test sample to assess the model's adaptability to future data, effectively verifying the model's stability and generalization capabilities.

[0045] Step S400, mining fraud-related features, includes: In step S401, the chi-square test method is used to analyze the correlation between variables in multiple dimensions, such as application information, device fingerprints, and social graphs, and the regional fraud label (Y variable). The formula is:

[0046] in, is the observation frequency, is the expected frequency, r is the number of rows, and c is the number of columns. By calculating the chi-square value, if the chi-square value is greater than the critical value at a certain significance level, it means that there is a significant correlation between the regional variable and the fraud label.

[0047] Step S402: Information value (IV) is used to measure the predictive power of a variable on a target variable. For each variable to be analyzed, its value is binned. The formula is:

[0048]

[0049] Where n is the number of boxes, is the evidence weight of the ith box, is the proportion of fraud samples in the i-th box to the total fraud samples, is the proportion of normal samples in the i-th bin to the total normal samples. A larger IV value indicates a stronger predictive ability of the variable for fraud detection. Generally, an IV value greater than 0.5 indicates strong predictive ability, between 0.3 and 0.5 indicates moderate predictive ability, between 0.1 and 0.3 indicates weak predictive ability, and less than 0.1 indicates poor predictive ability. Using the chi-square test and IV value calculation, we screened for variables that were significantly correlated with regional fraud and possessed strong predictive ability. A risk factor correlation matrix was then constructed to visually demonstrate the correlation between each variable and fraud.

[0050] Step S500, deriving expert features, includes: Step S501 generates derived variables based on feature intersections, combining them with credit business logic to perform feature intersection operations. For example, the applicant's permanent address is intersected with the IP address at the time of application to generate a new variable, "Permanent Residence and Application Location Match." If the permanent address and IP address are consistent, a value of 1 is assigned; if they are inconsistent but belong to the same provincial administrative region, a value of 0.5 is assigned; if they are completely different regions, a value of 0 is assigned. For another example, the application time is intersected with regional holiday information. If the application time falls during a major holiday in a region that also experiences a high incidence of credit fraud during holidays, a new variable, "Holiday Application Flag," is generated and assigned a value of 1; otherwise, a value of 0 is assigned. This feature intersection approach uncovers regional fraud risk information hidden within different feature combinations, generating highly explanatory derived variables.

[0051] Step S502 generates variables derived from time window statistics. Using device fingerprint information as an example, time windows of varying lengths, such as 7 days and 30 days, are set. The number of credit applications for the same device within the time window is counted, generating variables such as "Number of device applications in the last 7 days" and "Number of device applications in the last 30 days." If the number of applications for a particular device in the last 7 days exceeds a threshold, there may be an abnormal risk. Simultaneously, the changes in the device application regions within the time window are calculated, such as the standard deviation of the device application regions. A large standard deviation indicates frequent changes in the device application regions, increasing the likelihood of fraud. This generates a variable called "Standard deviation of device application region changes." By using variables derived from time window statistics, we can capture changes in device behavior patterns within a certain timeframe, providing dynamic risk characteristics for regional fraud identification.

[0052] Step S600, variable screening and optimization, includes: Step S601: Analyze the univariate PSI stability. The univariate PSI (Population Stability Index) is used to measure the stability of a variable at different time points or in different data sets. The formula is as follows:

[0053] in, and The PSI is the proportion of the i-th interval of a variable in two different datasets (e.g., training set and validation set), respectively, where n is the number of intervals. It is generally believed that a PSI value less than 0.1 indicates good variable stability, values between 0.1 and 0.25 are acceptable, and values greater than 0.25 indicate poor variable stability and may require adjustment or removal.

[0054] Step S602: Detect multivariate VIF collinearity. Use variance inflation factor (VIF) to detect the degree of collinearity between multiple variables. The calculation formula of VIF is:

[0055] Where is the coefficient of determination obtained by linear regression using the jth variable as the dependent variable and the other variables as independent variables. It is generally believed that a VIF value greater than 10 indicates severe collinearity among the variables, which may affect the stability and accuracy of the model.

[0056] Step S603 uses stepwise regression to select the optimal variable subset. Stepwise regression is an iterative variable selection method that begins with an initial model (which can be a model containing only a constant term) and introduces or removes variables one at a time until the optimal model is achieved. When introducing variables, select those that contribute most to the model (e.g., increase the model's goodness of fit the most) and meet PSI and VIF requirements. When removing variables, select those that have the least impact on the model (e.g., decrease the model's goodness of fit the least) and do not meet stability or collinearity requirements. Stepwise regression ensures that the selected variables possess both good predictive power and business interpretability.

[0057] Step S700, configuring risk weights, includes: Step S701: Apply weight of evidence (WOE) transformation to bin the selected variables. Calculate the ratio of fraudulent samples to normal samples in each bin, and then calculate the WOE value. The formula is:

[0058] in, is the proportion of fraud samples in the box to the total fraud samples, is the ratio of normal samples in the box to the total normal samples.

[0059] Through WOE conversion, different values of variables are converted into numerical values with actual risk significance, which facilitates the subsequent assignment of risk scale values.

[0060] Step S702: constructing a two-tier empowerment system, including: In step S702-1, based on the WOE binning results, risk scale values are assigned according to the statistical significance of the variable. For example, bins with larger WOE values indicate a higher correlation between samples within those bins and fraud, and are therefore assigned higher risk scale values. Bins with smaller WOE values are assigned lower risk scale values. The risk scale values are normalized so that their sum equals 1, thus obtaining the variable-level weights for the different bins within each variable.

[0061] Step S702-2 comprehensively considers the importance of each dimension in identifying regional fraud, taking into account different feature dimensions such as application information, device fingerprints, and social graphs. The weight of each feature dimension is determined through expert experience evaluation and historical data analysis. The variable-level weight is multiplied by the feature dimension weight to obtain the final comprehensive weight for each variable. This constructs an expert scorecard framework with a two-tiered weighting system, enabling precise quantitative assessment of the risks of different variables and feature dimensions.

[0062] Step S800, model training and tuning, includes: Step S801 uses the Random Forest ensemble learning framework for model training. A Random Forest ensemble model consists of multiple decision trees. It generates multiple training subsets by randomly sampling the training data with replacement (bootstrapping), each used to train a decision tree. When constructing the decision tree, for each node, a subset of features is randomly selected to find the optimal split point, increasing the model's diversity and generalization capabilities.

[0063] Step S802: Visualize feature contributions using SHAP (SHapley Additive exPlanations) values to verify the importance of each feature within the ensemble learning framework. SHAP values, based on cooperative game theory, explain the contribution of each feature to the model's predictions.

[0064] Step S803: Determine the hyperparameters that need to be optimized and their value ranges based on the characteristics of the random forest model. Combine the random forest model with the hyperparameter search space and use a grid search method to iteratively optimize hyperparameters such as the decision threshold and weight coefficient. After the network search is completed, the optimal hyperparameters and model are obtained.

[0065] Step S900: Model validation and calibration. Use out-of-sample testing of the optimized model using out-of-time test samples to evaluate the performance of the model on data that has not participated in training. Perform cross-group validation, divide the out-of-time test samples into multiple groups by region (such as samples from different provinces), and calculate performance indicators such as KS statistics and AUC values for each group.

[0066] The Bayesian smoothing technique is used to optimize the probability calibration curve. The specific method is as follows: The predicted probability is binned into certain intervals (such as 0 - 0.1, 0.1 - 0.2, ..., 0.9 - 1).

[0067] For the samples in each box, the actual proportion of fraud that occurred (actual fraud rate) is calculated.

[0068] Using the Bayesian formula, we combine the prior probability (such as the fraud rate of the entire sample) and the likelihood function (the actual fraud rate of the sample in the box) to calculate the posterior probability of each box as the calibrated predicted probability. The formula for Bayesian smoothing is:

[0069] in, is the number of samples in the box, is the actual fraud rate in the box, is the prior sample size, is the prior fraud rate (e.g., the fraud rate of the entire sample). Bayesian smoothing can effectively alleviate the problem of large fluctuations in the actual fraud rate in small sample boxes, making the predicted probability closer to the true probability and improving the calibration and credibility of the model.

[0070] Step S1000, deploying the production environment, includes: Step S1001: Use the microservice architecture to build a model API service layer to enable convenient calling of the model in the production environment.

[0071] Step S1002: Deploy the model using a phased release strategy in the risk control decision engine to reduce the risk of launching the new model, including: Credit application users are divided into grayscale testing group and original model group according to certain rules.

[0072] For application requests from the grayscale testing group, the newly constructed fraud identification model and the original fraud identification system are called simultaneously for prediction, and the prediction results and business indicators of the two are compared.

[0073] If the new model performs better than the original model during the grayscale testing phase and all indicators are stable and reliable, the proportion of the grayscale testing group will be gradually expanded until the original model is completely replaced. If significant problems are found in the new model, roll back to the original model and re-analyze and optimize the new model.

[0074] Step S1003: Conduct an A / B test in the production environment. Randomly divide users into Group A (using the new model) and Group B (using the original model). Compare the key business indicators of the two groups over the same time period to quantitatively evaluate the actual effect of the new model in improving fraud detection capabilities and ensuring customer experience. Step S1004: Based on the Hive+Spark hybrid computing architecture, a real-time feature computing pipeline is established to ensure that the model can obtain the latest feature data in real time in the production environment. The specific process is as follows: Through data collection tools such as Flume, various data in the credit application process (such as IP address, device GPS location information, application time, etc.) are collected in real time and transmitted to Spark Streaming for processing.

[0075] In Spark Streaming, real-time data is processed based on predefined feature engineering logic (such as standardization, normalization, feature cross-pollination, and time window statistics) to generate the feature data required by the model. For example, features such as the number of applications for a device in the past five minutes, and the difference between the current application region and the historical application region can be calculated in real time.

[0076] The feature data calculated in real time is stored in a distributed database such as HBase, making it available for real-time query and access by the model API service layer. Through this real-time feature calculation pipeline, the model can promptly capture real-time behavioral changes and regional characteristics of applicants, improving its response to emerging regional fraud patterns.

[0077] Step S1100: Model full declaration cycle management, including: In step S1101, a monitoring dashboard is built using open-source tools such as Grafana. Visual charts (such as line charts, bar charts, and dashboards) are used to visually display the changes in various indicators. A model monitoring dashboard is developed to track the health indicators of the model in real time, including: The PSI value of each variable is calculated regularly to monitor the stability of the variable distribution. If the threshold is exceeded, an early warning mechanism is triggered, prompting technical personnel to analyze and investigate the variable.

[0078] Regularly calculate the KS value of the model on out-of-time test samples or real-time production data. If the KS value drops by more than the threshold compared to the initial value, it indicates that the model's discrimination ability has decreased and the model retraining process needs to be initiated.

[0079] Dynamically draw the AUC-ROC curve of the model, observe the changing trend of the area under the curve, and evaluate the fluctuation of the overall performance of the model.

[0080] Step S1102: Establish a rolling time window mechanism to regularly (e.g., monthly) inject new credit data and expert analysis of the latest regional fraud patterns into the model. Specific steps are as follows: Collect new credit application data and fraud labeling data within a past time window (such as one month), process them according to the risk label system construction, feature engineering preprocessing and other processes to generate a new sample data set.

[0081] An expert review meeting will be held to analyze recent trends and new methods of regional fraud, such as the emergence of new regional camouflage techniques and the characteristics of fraud rings in specific regions. Based on the analysis results, adjustments will be made to feature engineering logic (such as adding features related to new fraud patterns), variable binning strategies, or risk weighting configurations.

[0082] New sample data was merged with historical sample data, and the sample space was re-divided using a stratified sampling strategy. The model was then trained and tuned using updated feature engineering logic and expert experience. Automated scripts were used to partially implement the model training, validation, and deployment processes, while retaining expert review of key parameters and model structure. This enabled semi-automated iterative optimization of the model, ensuring it could continuously adapt to changes in regional fraud patterns.

[0083] Those skilled in the art will understand that all or part of the processes in a method for building a model for identifying regional fraud of Internet credit can be implemented by instructing related hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of various embodiments of a method for building a model for identifying regional fraud of Internet credit. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0084] This embodiment of the present application also provides an electronic device 500 that utilizes the aforementioned system for constructing a model for identifying regional fraud in internet credit. The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the aforementioned method for constructing a model for identifying regional fraud in internet credit are implemented. In this embodiment of the present application, the processor serves as the control center of the computer method and can be a processor of a physical machine or a processor of a virtual machine.

[0085] Reference Figure 2 The electronic device 500 includes: at least one processor 501, at least one communication interface 502, at least one memory 503, and at least one bus 504. Bus 504 is used to enable communication between these components, communication interface 502 is used to communicate signaling or data with other node devices, and memory 503 stores machine-readable instructions executable by processor 501. When the electronic device 500 is running, processor 501 communicates with memory 503 via bus 504. When the machine-readable instructions are called by processor 501, the steps of the aforementioned method for constructing an Internet credit regional fraud identification model are executed.

[0086] The above contents are merely embodiments of the present invention. Common knowledge such as the known specific structures and characteristics in the scheme is not described in detail here. A person of ordinary skill in the art is aware of all common technical knowledge in the technical field to which the invention belongs before the filing date or priority date, is able to obtain all existing technologies in the field, and has the ability to apply conventional experimental means before that date. A person of ordinary skill in the art can, under the guidance of this application, improve and implement this scheme in combination with his or her own abilities. Some typical known structures or known methods should not become an obstacle for a person of ordinary skill in the art to implement this application. It should be pointed out that for a person of ordinary skill in the art, several variations and improvements can be made without departing from the structure of the present invention, which should also be regarded as the scope of protection of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection claimed in this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. A method for constructing an Internet credit regional fraud identification model, characterized in that: include: Build a risk labeling system and perform feature engineering preprocessing; A stratified holdout sampling strategy is used to divide the preprocessed dataset into model development samples and out-of-time testing samples; Mining fraud-related features, using the chi-square test to analyze the correlation between variables and regional fraud labels, and calculating information values to screen out variables with significant correlation and strong predictive power; Derive expert features, generate feature cross- and time window statistics derived variables; Conduct variable screening and optimization, analyze PSI stability, test VIF collinearity, and use stepwise regression to select the optimal variable subset; Assign risk weights to the selected variables; Perform model training and tuning, using the random forest ensemble learning framework for model training and the grid search method for iterative optimization of hyperparameters; Perform model validation and calibration, perform out-of-sample testing using out-of-time test samples, and optimize the probability calibration curve using Bayesian smoothing techniques; Deploy the production environment, build the model API service layer, deploy the model using a grayscale release strategy, and establish a real-time feature computation pipeline; Carry out full life cycle management of the model and regularly inject new credit data and expert analysis experience.

2. The method for constructing an Internet credit regional fraud identification model according to claim 1, characterized in that: Build a risk labeling system, specifically including the following: Extract credit application data covering different regions and credit products from the core business database of financial institutions Perform preliminary cleaning on the retrieved data to remove duplicate records, incorrectly formatted data, and data records with serious missing key information; Cross-check sample data based on multi-dimensional standards, including geographical dimensions, application behavior dimensions, and credit business dimensions; The sample data is divided into positive samples and negative samples. For positive samples, the Y variable is assigned a value of 1; for negative samples, the Y variable is assigned a value of 0.

3. The method for constructing an Internet credit regional fraud identification model according to claim 2, characterized in that: Perform feature engineering preprocessing, which specifically includes the following: Conduct a comprehensive check on the sample data, count the number and proportion of missing values in each field, check whether the data format complies with the specifications, and compare the consistency of the same information in different data sources; For fields with missing values, multiple imputation is used to process them. Based on the correlation with other information of the applicant, similar sample data is screened out. Several groups of values are generated using statistical models. The missing values are imputed multiple times, and the most appropriate group is selected as the final imputation result. Numerical variables were processed using the Z-Score standardization method or the Min-Max normalization method.

4. The method for constructing an Internet credit regional fraud identification model according to claim 2, characterized in that: Mining fraud-related features, including the following: The chi-square test method is used to analyze the correlation between variables in multiple dimensions and regional fraud labels. The formula is: in, is the observation frequency, is the expected frequency, r is the number of rows, and c is the number of columns. By calculating the chi-square value, if the chi-square value is greater than the critical value at a certain significance level, it means that there is a significant correlation between the regional variable and the fraud label. Calculate the information value IV value to measure the predictive ability of the variable to the target variable. The formula is: Where n is the number of boxes, is the evidence weight of the ith box, is the proportion of fraud samples in the i-th box to the total fraud samples, is the proportion of normal samples in the i-th box to the total normal samples; the larger the IV value, the stronger the predictive ability of the variable for fraud identification.

5. The method for constructing an Internet credit regional fraud identification model according to claim 1, characterized in that: Generate derived variables based on feature crossover and perform feature crossover operations in combination with credit business logic; Generate time window statistical derivative variables, set time windows of different lengths, and calculate the changes in the equipment application region within the time window.

6. The method for constructing an Internet credit regional fraud identification model according to claim 1, characterized in that: Configure risk weights, specifically including the following: The selected variables are binned using weight of evidence conversion, and the ratio of fraudulent samples to normal samples in each bin is calculated, and then the WOE is calculated using the formula: in, is the proportion of fraud samples in the box to the total fraud samples, is the ratio of normal samples in the box to the total normal samples; A two-tier weighting system is constructed. Based on the WOE binning results, risk scale values are assigned according to the statistical significance of the variables. The risk scale values are normalized to obtain the variable-level weights of different bins within each variable. Starting from different feature dimensions, the weight of each feature dimension is determined, and the variable-level weight is multiplied by the feature dimension weight to obtain the final comprehensive weight of each variable.

7. The method for constructing an Internet credit regional fraud identification model according to claim 1, characterized in that: Conduct model validation and calibration, including the following: Use out-of-sample testing of the optimized model using out-of-sample testing samples to evaluate the performance of the model on data that was not used in training; Perform cross-group validation, divide the out-of-time test samples into multiple groups by region, and calculate the KS statistic and AUC value for each group; The Bayesian smoothing technique is used to optimize the probability calibration curve, and the formula is: in, is the number of samples in the box, is the actual fraud rate in the box, is the prior sample size, is the prior fraud rate; The predicted probability is divided into bins according to intervals. For the samples in each bin, the actual proportion of fraud is counted. The Bayesian formula is used to combine the prior probability and the likelihood function to calculate the posterior probability of each bin as the calibrated predicted probability.

8. The method for constructing an Internet credit regional fraud identification model according to claim 1, characterized in that: Deploy the production environment, including the following: Use microservice architecture to build the model API service layer; A grayscale release strategy is used to deploy the model in the risk control decision engine. Credit applicants are divided into a grayscale testing group and an original model group according to certain rules. For application requests from the grayscale testing group, both the newly constructed fraud detection model and the original fraud detection system are used for prediction. The prediction results and business indicators of the two are compared. If the new model performs better than the original model and all indicators are stable and reliable, the grayscale testing group proportion is gradually expanded until the original model is completely replaced. If significant problems are found in the new model, the original model is rolled back and the new model is re-analyzed and optimized. Conduct A / B testing in a production environment, randomly dividing users into Group A and Group B, comparing key business metrics between the two groups, and quantitatively evaluating the effectiveness of the new model in improving fraud detection capabilities and ensuring customer experience. A real-time feature computing pipeline is established based on the Hive+Spark hybrid computing architecture. Various data from the credit application process are collected in real time through data collection tools and transmitted to Spark Streaming for processing. The real-time collected data is processed according to pre-defined feature engineering logic to generate the feature data required by the model. The feature data obtained by real-time calculation is stored in a distributed database for real-time query and call by the model API service layer.

9. The method for constructing an Internet credit regional fraud identification model according to claim 1, characterized in that: Carry out model full life cycle management, including the following: Build a monitoring dashboard to track the model's health indicators in real time, including regularly calculating the PSI value of each variable to monitor the stability of the variable distribution, regularly calculating the KS value of the model on out-of-time test samples or real-time production data to evaluate the model's discrimination ability, and dynamically plotting the model's AUC-ROC curve to evaluate the fluctuation of the model's overall performance; Establish a rolling time window mechanism to regularly inject newly generated credit data and experts' analysis experience on the latest regional fraud patterns into the model, collect new credit application data and fraud annotation data for processing to generate new sample data.