Intelligent financial risk control system based on mobile device monitoring
By introducing an intelligent risk control system based on mobile device monitoring in the financial risk control system, collecting and analyzing user's personal and device behavior data, using automatic binning and multiple machine learning models to perform credit scores and credit limit prediction, and adjusting the model in real time through the dynamic risk control adjustment module, the problems of incomplete data, poor generalization capabilities of model and difficulty in dynamic adjustment in the existing system are solved, and efficient and accurate credit risk management is achieved.
Patent Information
- Application Number
- CN202510484609.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
AI Technical Summary
The existing financial risk control system has shortcomings in data sources, feature engineering, model architecture and dynamic adjustments, which leads to insufficient comprehensive credit assessment and poor real-time performance, and it is difficult to effectively evaluate users who lack traditional credit records.
The intelligent financial risk control system based on mobile device monitoring is adopted, through the collaborative work of user terminals, management terminals and servers, users' personal information and device behavior information are collected, data preprocessing and feature encoding are performed, and credit scores and quota prediction is used to use automatic binning modules and various machine learning models (such as logistic regression, encoder-decoder structure and gradient improvement decision tree), and the model is adjusted in real time through dynamic risk control adjustment modules.
It realizes accurate prediction of user credit risks and available quotas, improves the risk management efficiency and user service quality of financial institutions, and solves the problems of incomplete data, poor generalization capabilities of model and difficulty in dynamic adjustment in traditional systems.
Smart Images

Figure CN120013663A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent risk control technology, and in particular to an intelligent financial risk control system based on mobile device monitoring. Background Art
[0002] With the rapid development of financial technology and the high popularity of mobile devices, financial services are accelerating their evolution towards digitalization, intelligence, and personalization. As a core link in financial services, risk control is also constantly innovating and upgrading its data collection, modeling, prediction, and monitoring technologies. However, the existing financial risk control system still faces many deficiencies and challenges in terms of data sources, feature engineering, model architecture, and dynamic adjustment.
[0003] The traditional financial risk control system has a single data collection method, which mainly relies on the user's historical financial behavior data, such as bank statements, credit card bills, and credit records. This credit assessment method based on static data has obvious limitations. First, the data type is single, lacking in-depth analysis of user device behavior and timing characteristics. Second, information updates lag, and the frequency of financial data updates is low, making it difficult to reflect the user's credit status and financial behavior changes in real time. At the same time, for users who lack traditional credit records, it is difficult for the system to conduct effective assessments.
[0004] At present, many financial risk control systems still use traditional statistical models (such as logistic regression, decision trees, etc.) or single machine learning algorithms (such as random forests, neural networks) for credit scoring and credit limit prediction. However, these methods have certain limitations. Although traditional statistical methods have good interpretability, they perform poorly when dealing with complex nonlinear data and are difficult to accurately characterize user credit risks. In addition, a single machine learning model may face problems such as overfitting and lack of robustness when dealing with high-dimensional data, resulting in large fluctuations in prediction results between different user groups or different time periods.
[0005] The scorecard model is one of the main credit assessment tools in the current financial risk control field. Its core lies in feature engineering, especially the processing method of feature binning. However, traditional scorecard models usually rely on manually set feature binning methods, such as equidistant binning or equifrequency binning, but there are many challenges. The manually set binning method may not be able to adapt to the data distribution of different user groups, resulting in poor generalization of the model; fixed binning strategies are difficult to adapt to changes in the market environment and user behavior. After long-term operation, feature drift problems may occur, which reduces the predictive ability of the model. In addition, traditional scorecard models usually lack automatic adjustment capabilities. When the data distribution changes, manual intervention is required to readjust the binning rules, which increases maintenance costs and adjustment difficulties, affecting the long-term stability of the financial risk control system. Summary of the invention
[0006] The present application discloses an intelligent financial risk control system based on mobile device monitoring. Based on the optimization processing of data comprehensiveness, data quality, model performance and user-friendliness, it provides financial institutions with an efficient, intelligent and stable risk control system.
[0007] The technical solution adopted in this application is:
[0008] An intelligent financial risk control system based on mobile device monitoring, which includes a user terminal, a management terminal and a server;
[0009] The user terminal is used for user identity authentication, collecting user information, and for the user to initiate a financial service application to the management terminal; the user information includes user personal information and user device behavior information of the user terminal; the user personal information is actively collected by the user terminal; the user device information is collected based on the monitoring instruction initiated by the management terminal;
[0010] The management terminal is used to receive and process user information collected by the user terminal, and includes a data preprocessing module, a security storage module and a financial business management module; wherein the data preprocessing module is used to perform data preprocessing on the received user information and send it to the security storage module; the security storage module de-identifies the user information after data preprocessing based on the user ID mapping strategy, and then uploads the processed user information to the server to store it in a designated storage location; the financial business management module is used to receive the financial business application initiated by the user terminal, access the server based on the role-based access control mechanism, and call the corresponding business processing module on the server based on the mapped user ID and the financial business application content, and generate a business processing result in response to the financial business application initiated by the user based on the output of the currently called business processing module and return it to the user terminal; wherein the business processing module of the server includes a credit scoring module and a credit limit prediction module;
[0011] The server includes an automatic box sorting module, a credit scoring module and a credit limit prediction module; wherein the automatic box sorting module is used to encode the information features of the pre-processed user personal information uploaded by the management terminal, and send the feature codes of the user personal information to the credit limit prediction module and the credit scoring module; the credit scoring module predicts the user's credit score based on the feature codes of the user's personal information, and the credit limit prediction module predicts the user's loan limit based on the user's device behavior time series features and the feature codes of the user's personal information, wherein the user's device behavior time series features are obtained based on the user's device behavior information after data processing.
[0012] Furthermore, the credit scoring module of the present application can first predict the default probability of the current user through a logistic regression model, and then convert the default probability into a credit score (also called a credit point).
[0013] Furthermore, the credit scoring module of the server calculates the user's default probability based on the characteristics of personal information, and converts it into a credit score through a dynamic scoring strategy; the loan limit prediction module includes two prediction models; one of the prediction models (i.e., the first prediction model) adopts an encoder-decoder structure, and predicts the user's first loan prediction limit based on the time series characteristics of the user's device behavior; the other prediction model (i.e., the second prediction model) is a neural network structure based on gradient boosted decision trees (GBDT), which predicts the user's second loan prediction limit based on the feature encoding of the user's personal information; the final prediction limit of the loan limit prediction module is obtained based on the weighted sum of the first and second loan prediction limits.
[0014] Furthermore, the server also includes a dynamic risk control adjustment module, which periodically calculates the population stability index (Population Stability Index, PSI) of the information feature based on a set first detection period, and triggers the bin reconstruction processing of the automatic binning module when the population stability index is greater than or equal to a specified first threshold; the dynamic risk control adjustment module periodically calculates the statistics (Kolmogorov Smirnov, KS) of the credit scoring module based on a set second detection period, and re-performs the feature screening of the user information and adjusts the binning strategy when the fluctuations of the statistics in the last two times are greater than or equal to a specified second threshold, wherein the first detection period is less than the second detection period.
[0015] Furthermore, user personal information specifically includes: name, telephone number, gender, annual income, ID number, living status, education level, loan amount, loan term, loan interest rate, loan grade, loan purpose category, debt-to-income ratio, installment amount, employment title, years of employment and other basic identity information.
[0016] Furthermore, user device behavior information includes: installed application catalog, device usage behavior, IP address, number of daily logins, single usage duration, frequency of geographic location changes, and percentage of usage time of risky target applications.
[0017] Furthermore, the prediction model using the encoder-decoder structure is set to a network structure based on a long short-term memory network (LSTM) and an attention mechanism, wherein the encoder is a structure based on a long short-term memory network, the decoder is the inverse operation of the encoder, and the encoding features output by the encoder are processed by the attention mechanism network before being sent to the decoder, and then the output of the decoder is mapped to the first loan prediction amount based on the fully connected layer, and the model input of the prediction model is the user device behavior time series feature;
[0018] For example, the time series characteristics of user device behavior can be specifically set as: number of daily logins, single usage duration, frequency of geographic location changes, and usage duration of risk target applications. Then based on the specified time step , using past continuous The time series features of historical user device behavior for days are used as the input of the prediction model to predict the future The first loan forecast amount of the day, of which, The specified number of forecast days, for example, set it to 7.
[0019] Furthermore, the second prediction model (the prediction model based on the neural network structure of the gradient boosting decision tree) is specifically set as:
[0020] The feature encoding of user personal information is input into the gradient boosting decision tree model (such as the lightweight gradient boosting machine model (LightGBM) model). The output of the model is the default probability with a value range of 0 to 1, which is recorded as ;
[0021] After the characteristic coding of user personal information is standardized, The enhanced feature matrix is obtained by concatenation;
[0022] A neural network is constructed to predict the second loan prediction amount, and the input data of the neural network is the enhanced feature matrix.
[0023] Furthermore, the loss function used in the training of the second prediction model is set as:
[0024]
[0025] in, represents the loss function, N represents the number of samples, is the sample number, is the true label, Representation sample The probability of default (i.e., the probability of default) ), To set the contribution factor, real-time adjustments are made based on the value of information obtained by the automated binning module. The adjustment method is: ,in represents the information value of feature f, It indicates determining the maximum information value of all features.
[0026] Furthermore, the data preprocessing of the data preprocessing module includes: drawing histograms, calculating descriptive statistical information (such as maximum, minimum, mean and quantile, etc.), processing outliers and missing values, data cleaning, and converting categorical variables in user information using one-hot encoding, and eliminating features with variance inflation factors greater than the specified inflation factor threshold, and deleting invalid features with a degree of uniformity of values greater than or equal to the frequency threshold.
[0027] Furthermore, the user ID mapping strategy is as follows:
[0028] For the set target sensitive information (such as name, ID number, telephone number), a hash algorithm is used to generate a unique identifier, which is used as the data index of the server to obtain de-identified user information, and it is transmitted to the designated database on the server, thereby ensuring that the real identity information is not directly stored or transmitted during the transmission process.
[0029] Furthermore, the automatic binning module performs information feature extraction including: for continuous variables, the conditional entropy minimization algorithm is used for segmentation to ensure the internal consistency of the data in each bin; for discrete variables, the chi-square test is used to merge adjacent categories, and the encoding of a single feature dimension (such as WOE (Weight of Evidence Encodin) encoding) and information value (Infommation Value, IV) are calculated to verify the binning effect and further improve the interpretability of the model.
[0030] The technical solution provided by this application brings at least the following beneficial effects:
[0031] The financial risk control system proposed in this application can accurately predict credit risk (credit score output by the credit scoring module) and available credit limit (final predicted credit limit of the credit limit prediction module) by deeply analyzing the user's real-time device data and multi-dimensional static features, realize dynamic risk control strategy adjustment, and improve the risk management efficiency and user service quality of financial institutions. Compared with traditional risk control systems, the technical solution requested for protection in this application has achieved optimization in terms of data comprehensiveness, data quality, model performance, user-friendliness, etc., providing financial institutions with an efficient, intelligent and stable risk control solution. The system proposed in this application can be applied to the risk control needs of banks, consumer finance and other financial institutions. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0033] Figure 1 A schematic diagram of the structure of an intelligent financial risk control system based on mobile device monitoring provided in an embodiment of the present application.
[0034] Figure 2 This is a schematic diagram of the data interaction process of some functional modules of the intelligent financial risk control system based on mobile device monitoring.
[0035] Figure 3 This is a schematic diagram of the processing process of the credit limit prediction module of the intelligent financial risk control system based on mobile device monitoring. DETAILED DESCRIPTION
[0036] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions of the embodiments of the present application will be described in detail and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the embodiments described with reference to the drawings are exemplary and are intended to be used to explain the present application, and cannot be understood as limiting the present application.
[0037] The present application embodiment proposes an intelligent financial risk control system based on mobile device monitoring, see Figure 1, the system includes a user terminal, a management terminal and a server; wherein the user terminal is used for user identity authentication, user information collection, and for users to initiate financial service applications to the management terminal (such as loan amount prediction, credit score prediction, loan application, repayment application, etc.); wherein the user information includes user personal information and user device behavior information of the user terminal; wherein the user personal information is actively collected by the user terminal; the user device information is collected based on the monitoring instruction initiated by the management terminal; that is, in this application, the user terminal includes a user identity authentication module, a data collection module and a financial service processing module, wherein the user identity authentication module is used for user identity authentication, and a multi-factor identity authentication method can be used, including but not limited to: SMS verification code, face recognition, ID card recognition, etc., to ensure the authenticity of the user identity. The data collection module is used to actively collect user personal information on the one hand, and on the other hand, based on the monitoring instruction initiated by the management terminal, the user device behavior information of the user terminal is collected, and the collected information can be transmitted to the management terminal in an end-to-end encrypted manner. The financial service processing module is mainly used to provide users with a human-computer interaction interface for business operation and processing, and generate a financial service application for the corresponding financial service based on the user's operation input, and then send it to the management terminal. That is, the financial business processing module is used for users to initiate financial business applications to the management terminal. The management terminal is used to receive and process user information collected by the user terminal, and includes a data preprocessing module, a security storage module and a financial business management module; wherein the data preprocessing module is used to perform data preprocessing on the received user information and send it to the security storage module; the security storage module de-identifies the user information after data preprocessing based on the user ID mapping strategy, and then uploads the processed user information to the server to store it in a designated storage location; the financial business management module is used to receive financial business applications initiated by user terminals, access the server based on the role-based access control mechanism, and call the corresponding business processing module on the server based on the mapped user ID and the financial business application content, and generate a response based on the output of the currently called business processing module The business processing result of the financial business application initiated by the user is returned to the user terminal; wherein the business processing module of the server includes a credit scoring module and a credit limit prediction module; the server includes an automatic box sorting module, a credit scoring module and a credit limit prediction module; wherein the automatic box sorting module is used to encode the information feature of the pre-processed user personal information uploaded by the management terminal, obtain its WOE code, and send the WOE code to the credit limit prediction module and the credit scoring module; the credit scoring module predicts the user's credit score based on the feature coding of the user's personal information, and the credit limit prediction module predicts the user's loan limit based on the user's device behavior time series features and the user's personal information feature coding, wherein the user's device behavior time series features are obtained based on the user's device behavior information after data processing. In this application, both the user terminal and the management terminal can be implemented based on mobile devices.
[0038] In this application, for example, when a user initiates a financial service application for loan amount prediction through a financial service processing module, the financial service processing module will trigger the data collection module of the user terminal to actively collect user personal information. The data collection module uses the currently collected user personal information and the financial service application initiated by the user as the data payload of the financial service application message, and transmits it to the management terminal through end-to-end encryption; the management terminal determines whether the user terminal needs to perform the collection of user device information based on the current financial service type. If the current financial service type requires the use of user device information, it will initiate a monitoring instruction to the user terminal to collect user device information; the management terminal then performs data preprocessing on the collected information based on its data preprocessing module, and calls the corresponding business processing module on the server based on the specific financial service application to obtain its business processing result, that is, the user's predicted loan amount is then fed back to the corresponding user terminal. If the user initiates a loan / repayment application, the management terminal will call the loan / repayment processing business program pre-set on the server, and complete the loan / repayment business processing based on the information interaction of the loan / repayment business securely transmitted by the user terminal, etc.
[0039] Compared with the traditional risk control system, the intelligent financial risk control system based on mobile device monitoring provided by the embodiment of the present application can effectively utilize the user's real-time device data and improve the comprehensiveness of credit assessment. And based on the credit scoring module and credit limit prediction module set in this application, efficient, accurate and stable user credit assessment and risk management are realized. This system enhances the comprehensiveness and timeliness of credit assessment by actively collecting the user's personal static information (i.e., user personal information) and background collection of device behavior data (i.e., user device behavior information). In addition, in terms of feature engineering, the embodiment of the present application adopts automated binning technology and WOE encoding, combined with WOE and IV value screening for binning verification, to further improve the interpretability and generalization ability of the prediction model of the credit limit prediction module. In one embodiment, the model layer of the prediction model of the credit limit prediction module adopts a multi-level stacked learning framework, integrating LSTM (Long Short-Term Memory) and XGBoost (optimized distributed gradient enhancement library), LightGBM (Light Gradient Boosting Machine) models, optimizing prediction accuracy, improving efficiency, and at the same time, through dynamic adjustment strategies, PSI index monitoring and KS value control are performed to ensure model stability. The system also takes security into consideration, integrating user ID mapping with role-based access control mechanisms to ensure data privacy and security.
[0040] In one embodiment, the server also includes a dynamic risk control adjustment module, which regularly calculates the PSI of information features based on a set first detection cycle (for example, it can be set to monthly detection), and when the PSI is greater than or equal to a specified threshold (for example, PSI ≥ 8%), triggers the bin reconstruction processing of the automatic binning module; the dynamic risk control adjustment module regularly calculates the statistic KS of the credit scoring module based on a set second detection cycle (for example, it can be set to quarterly detection), and when the fluctuation of the statistics for the last two times is greater than or equal to a specified second threshold (for example, the fluctuation is greater than or equal to 2%), re-performs feature screening of user information and adjusts the binning strategy.
[0041] In one embodiment, the data collection module of the user terminal, the data preprocessing module and the security storage module of the management terminal, and the automatic box sorting module, the credit scoring module, the credit limit prediction module and the dynamic risk control adjustment module of the server interact as follows: Figure 2 As shown, the data collection module will send the collected user information (such as user personal information, user device behavior information, etc.) to the data preprocessing module; the data preprocessing module will send the corresponding data preprocessing results to the automatic binning module for information feature extraction; and send the extracted results to the credit scoring module and the credit limit prediction module respectively.
[0042] In one embodiment, the data collection module of the user terminal of the present application can adopt multi-data source collection technology to ensure the comprehensiveness and accuracy of the user credit assessment. The data sources include data actively input by the user (i.e., user personal information actively collected by the data collection module) and background device monitoring data, and format verification is performed during the data collection process to ensure the standardization and validity of the data. Among them, the data actively input by the user includes but is not limited to: name, telephone number, gender, annual income, ID number, living situation, education, loan amount, loan term, loan interest rate, loan grade, loan purpose category, debt-to-income ratio, installment amount, employment title, employment years and other basic identity information. The background monitoring data (user device behavior information passively collected by the data collection module) covers user device information, such as installed application directory, device usage behavior, IP address, and time series features include daily login times, single use duration (minutes), GPS (Global Positioning System) location change frequency (times / day), and high-risk application APP usage duration ratio (%). These data must meet strict format verification rules when collected. For example, the ID number must be 18 digits, the phone number must be 11 digits long, the IP address must be resolved to a specific geographic area, and the geographic location must be resolved to a specific area code, etc., to ensure the integrity and accuracy of the data. For the proportion of high-risk APP usage time, the high-risk APP defined in this embodiment comes from the published high-risk APP. The background collects its usage time and then compares it with the time used by all APPs on the device.
[0043] In one embodiment, the data preprocessing module of the management terminal can be used to clean and process the original data (user information collected by the data acquisition module), including filling missing values, removing outliers, etc., to ensure data quality. Specifically, the data preprocessing performed by the data preprocessing module includes: drawing histograms, calculating descriptive statistical information such as maximum, minimum, mean and quantile, processing outliers and missing values, and removing the feature when the missing rate of continuous variables exceeds 30%. The remaining features are interpolated using the K nearest neighbor algorithm to clean the data. In addition, for numerical variables (such as annual income and loan amount), outliers can be processed based on quantiles; and for categorical variables (such as cities and occupations), outliers can be processed based on frequency thresholds, and categories with a frequency of less than 1% can be removed.
[0044] For user information processed by the data preprocessing module, through feature engineering, categorical variables are encoded with One-Hot to generate binary features, and highly correlated or redundant features are removed, thereby improving the validity of the data and the training efficiency of the model. For example, one-hot encoding can be used to transform categorical variables, and then preliminary feature screening can be performed on the transformed features, such as eliminating features with variance inflation factors greater than 10, and deleting invalid features with a single value degree greater than 95%.
[0045] In order to ensure the security of data, the security storage module of the management terminal of this application adopts user ID mapping technology to de-identify user personal information, and uses a unique user ID instead of directly storing sensitive information such as user name, ID number, etc.; for example, a hash algorithm is used to generate a unique identifier for the user's sensitive information (such as name, ID number, phone number, etc.), and the mapping relationship is stored in a secure mapping table, and then the de-identified data is transmitted to the database and server to ensure that the real identity information is not directly stored or transmitted during the transmission process. At the same time, the system can also integrate a role-based access control mechanism to set the data access scope according to the permission level of different (management) users. The management terminal implements strict permission management in all data access, modification and storage operations, and records data access through a log audit system to improve data security and traceability.
[0046] In addition, in order to ensure the authenticity of the user's identity and the security of the data, this application system can also adopt a multi-level identity authentication and privacy protection mechanism. User identity authentication adopts a three-factor authentication method of SMS verification code verification, ID card information verification and face recognition to prevent identity fraud.
[0047] The processing of the automatic binning module of the server of this application includes: for continuous variables, the conditional entropy minimization algorithm is used for segmentation to ensure the internal consistency of the data in each bin; for discrete variables, the chi-square test is used to merge adjacent categories, and the WOE and IV values are calculated to verify the binning effect, further improving the interpretability of the model. The credit scoring module is a credit scoring model based on WOE coding. It calculates the probability of default of each user through methods such as logistic regression, and converts it into a credit score through a dynamic scoring formula. This process not only provides an intuitive credit score, but also can be used to evaluate performance through indicators such as AUC (AreaUnder Curve), KS, and PSI values to ensure the accuracy and robustness of the scoring model of the credit scoring module.
[0048] Specifically, the fully automatic binning optimization processing based on the automatic binning module can be set as follows:
[0049] Continuous variable binning method: Use the conditional entropy minimization criterion to find the optimal split point; iteratively merge adjacent bins until the sample proportion of each bin is ≥5%, each bin contains both normal and default samples, and the final number of bins ∈ [4,6];
[0050] The binning method is used for discrete variables: the correlation of variables is determined by calculating the chi-square value, and when the test statistic is less than a specified value (such as 3.841), adjacent enumeration values are merged;
[0051] Calculate IV value for the merged bins and keep .02 features;
[0052] Verification after binning: Perform a monotonicity test on non-age variables, requiring that the difference in WOE values between adjacent bins be ≥ 0.1; generate a binning quality report, including the sample distribution of each bin, WOE trend chart, and IV contribution, etc.
[0053] In one embodiment, see Figure 3 The feature extraction part of the credit limit prediction module can use a triple stacking model (LSTM, XGBoost, LightGBM), respectively inputting time series features (i.e., user device behavior time series features) and structural features (i.e., WOE encoding features). The default probability prediction values output by the three base models (LSTM model, XGBoost model, and LightGBM model) are: , whose dimensions are ,in Represents the number of samples; then the joint WOE code features (WOE code values) are horizontally spliced to form a feature matrix ; Finally, the final predicted amount can be output based on the fully connected layer (fully connected neural network), that is Figure 3 In .
[0054] In addition, in this application, the credit limit prediction module can also use the LSTM model based on the encoder-decoder framework to process time series features, and lightXGB to process structured features, and finally weight the outputs of the two models to obtain the final predicted credit limit. This ensures that both short-term trends and long-term risk factors are effectively modeled, improving the accuracy and stability of the prediction.
[0055] The dynamic risk control adjustment module of the server of this application calculates the PSI index every month to detect the stability of the model in different time periods. Once the PSI value exceeds the threshold, the system will trigger the bin reconstruction process to adjust the model structure or update the data in time to adapt to changes in user behavior or market. In addition, by controlling the fluctuation of the KS value, the effectiveness of the model in long-term use is guaranteed.
[0056] In summary, the intelligent financial risk control system based on mobile device monitoring proposed in the embodiment of the present application realizes efficient, accurate and stable user credit assessment and risk management through multi-device data collection, automated feature engineering, triple stacking model (LSTM, XGBoost, LightGBM) and dynamic risk control adjustment mechanism. The system enhances the comprehensiveness and timeliness of credit assessment by actively collecting users' personal static information (such as annual income, phone number) and collecting device behavior data (such as IP location, application usage, etc.) in the background. In terms of feature engineering, automated binning technology and WOE encoding are used, and binning verification is performed in combination with WOE and IV value screening to improve the interpretability and generalization ability of the model. The model layer adopts a triple stacking model to integrate time series features and structured features, optimize prediction accuracy, and improve efficiency. At the same time, through dynamic adjustment strategies, PSI index monitoring and KS value control are performed to ensure model stability. The system also considers security, integrates user ID mapping and role-based access control mechanism to ensure data privacy and security.
[0057] In one embodiment, a specific implementation of an intelligent financial risk control system based on mobile device monitoring proposed in an embodiment of the present application may include:
[0058] (1) Automatically bin continuous variables and set features is a continuous variable, its subscript j is the feature index, and the target variable y (i.e., the true label) is default or normal. Taking age as an example, first sort all the feature values, that is, traverse the values of all the split points of the current feature , such as 20 years old, 25 years old, 30 years old, etc., calculate the "conditional entropy" at each segmentation point: and , and then calculate the value of each split point The joint entropy of , select joint entropy The smallest eigenvalue is used as the optimal split point to divide the data into two parts, set the stop condition, and repeat the above steps. represents the sorting function, represents the conditional probability, n represents the total number of samples in the currently processed data subset, Represents the features in the current data subset Less than or equal to The number of samples, Represents the features in the current data subset Greater than The number of samples, It is an indicator value, which is rounded to 0 or 1. 0 indicates normal and 1 indicates default.
[0059] In this embodiment, the stopping conditions during the binning process are as follows: , j represents the jth continuous variable, m represents the number of bin states finally calculated for each continuous variable, that is, the number of bins, They represent the 1st to mth bin states of the jth continuous variable, and the proportion of samples in each bin must be greater than or equal to 5%. , each bin must contain normal samples and default samples, and except for age, other variables are kept as monotonic as possible.
[0060] (2) Automatically bin discrete variables and set features For discrete variables, taking education level as an example, it can be classified into junior high school, high school, undergraduate, master, and doctoral degrees. The initial state is that each type is a separate box. First, a chi-square test is performed, and then adjacent boxes are grouped and bound in pairs. The chi-square test is performed on "each group value and the target classification" to calculate the test statistic. Values: ,in represents the observation frequency, i.e. The number of samples detected at the intersection of the row and the category of the vth column, represents the expected frequency, Respectively represent the frequencies of the four cells in the 2x2 contingency table, such as ,parameter ,Right now represents the total frequency. Then merge The smaller the value, the closer the actual frequency is to the theoretical frequency, which means that the chi-square test is valid, proving that grouping and classification are not related. Traverse all adjacent bin pairs and merge the adjacent bins with the smallest chi-square value. Finally, the chi-square value is significant and the number of bins is less than 5. Finally, we can get the value of each discrete variable. Get the bin status set ,in Represents the number of bin states finally calculated for discrete variables, Respectively represent The first ~ The status of the sub-box.
[0061] (3) For each feature and The final binning result and Calculate the WOE code value and IV value:
[0062] , ;
[0063] in, represents the WOE code of the mth bin of the jth feature, represents the mth bin of the jth feature, Indicates The true label of samples (indicating default or normal), Representation characteristics The number of default samples in the mth bin, represents the number of default samples in the entire data, and N represents the number of samples. .
[0064] (4) Then the user characteristics and The WOE code values after automated binning are converted into linear combinations and input into the logistic regression model. The formula for predicting the probability of default is:
[0065]
[0066]
[0067] in, represents the predicted probability of default, Represents the natural base; represents the initial logistic regression coefficient; represents the logistic regression coefficient of the jth feature; Represents the WOE encoding value of the jth feature, which is determined by the bin in which the sample is located. decide; and if ,but ; J represents the total number of features input into the logistic regression model. For any logistic regression coefficient , which is used to characterize the influence weight of the feature on the probability of default, This indicates that this feature increases the risk of user default, such as frequent changes of IP addresses. Indicates that the feature reduces the risk of default, such as high income. The model can then output the probability of default p.
[0068] The above formula for predicting the probability of default is based on the characteristic For example, the feature Similarly, the corresponding predicted default probability can be obtained.
[0069] (5) Convert the default probability p into a credit score. The smaller the default probability p, the higher the credit score. Assuming that the probability of a customer defaulting is p, the normal probability is 1 − p. From this, we can get the default probability: , , the traditional scorecard is a "linear function that converts the logarithm of the default probability Odds into a score", and its corresponding expression is:
[0070]
[0071] Parameters A and B are determined during construction, but this results in the score not being dynamically adjusted as user characteristics and time change. Therefore, the present application embodiment introduces dynamic parameter adjustment to allow the credit score to change dynamically:
[0072]
[0073] in, The risk characteristics of the group are reflected by calculating the average WOE value of each variable in time t. At the same time, the values of A and B are calculated by setting two conditions. When the probability of default Time corresponding credit score , It is usually set to 500, 600, 700, etc. When the default probability doubles, that is, When the credit score is reduced by PDO, PDO is usually set to 50 or 100. The two conditions can be linked to obtain the values of A and B.
[0074] (6) Verify the model by calculating AUC, KS value, and PSI. The classification performance of the model is evaluated by AUC, and the value range of AUC is between 0 and 1. The closer to 1, the better the model performance. Calculate the true positive rate and false positive rate, then draw the ROC (Receiver Operating Characteristic Curve) curve and calculate the area under the line to get the AUC value. The AUC value is required to be > 0.85; the model's discrimination ability is verified by the KS statistic, sort the predicted values of the test set samples in descending order, and calculate the cumulative proportion of positive and negative samples for each deck position k: ,in, , are the number of positive and negative samples, respectively. Represents the indicator function. Then calculate the maximum absolute difference KS of the cumulative proportion of positive and negative samples, requiring the KS index>0.3; the model stability is evaluated by the population stability index (PSI), and each feature is divided into equal frequency bins based on the training set, and the proportion of samples in each bin of the test set is calculated, and the PSI index value is calculated:
[0075]
[0076] In this embodiment, the calculated PSI is required to be less than 8%, where: , Respectively represent the sample proportions of the test set and the training set in each bin, is the bin index, and B is the number of bins.
[0077] (7) Amount forecast:
[0078] At level 1, we design an encoder-decoder combined with the LIST model to process the temporal characteristics of user device behavior, and integrate the self-attention module to focus on global information. We use the data from the past period of time to discover hidden potential high-risk users, so that we can take measures in advance and better manage liabilities. Input data includes the number of logins per day. , Single use duration (minutes) , GPS location change frequency (times / day) , High-risk APP usage time ratio (%) , the time step of each feature is Days, using data from the past 30 days to predict the future The daily quota is calculated using the sliding window technique. This implementation formats the data as ,in Indicates the batch size, represents the time step, i.e. the number of days that have passed, Represents the feature dimension of the input.
[0079] First, design the encoder part to convert the input time series data , where each is a vector containing four features (corresponding to the number of daily logins, single usage duration, frequency of geographic location changes, and proportion of high-risk APP usage duration), which specifically represents the features of the tth day and is converted into a fixed-length context vector Specifically, it is implemented through the LSTM layer. When X is input into the network, LSTM will update the values of the input gate, forget gate, and output gate at each time step t, and calculate the hidden state of each time step. and cell status , and finally LSTM at the last time step t= , output the final hidden state , denoted as , the specific formula is as follows:
[0080]
[0081] in, is the input at time t, and are the hidden state and cell state at the previous time step, is the output of the LSTM network, is the hidden state at the current time step.
[0082] In order to allow the model to better handle the volatility, non-stationarity, periodicity, nonlinearity and other features that often appear in financial time series, and pay more attention to the part of the input sequence that is most relevant to the output, a self-attention mechanism module is added after the encoder to help the model automatically select important time steps, thereby weighting each part of the input sequence. For features that have a great impact on the output, a larger weight is used. The calculation formula is: ,in, represents the feature dimension, Q, K, and V are the context vectors The linear transformations with different weight matrices are performed to represent the query, key, and value respectively. The specific calculation is: , , ,in , , They represent the weight matrices corresponding to Q, K, and V learned during the training process, and finally the weighted hidden state can be obtained .
[0083] Then design the decoder part to generate the prediction sequence based on the context information output by the encoder. As the input of the decoder, the core of the decoder is still an LSTM network, which generates a prediction value at each time step. The decoder will generate a loan amount prediction value at each time step. The specific calculation is as follows ,in is the decoder output at the previous time step, is the hidden state of the decoder at the previous time step, is the weighted input hidden state, and finally connected to the fully connected layer: , mapping the decoder output to the predicted amount ,in is the weight matrix of the fully connected layer, is the bias term.
[0084] Level 2: To improve the training speed and deployment efficiency of the model, LightGBM is used instead of XGBoost. While maintaining the prediction performance, it has faster training speed and lower memory usage. This model is used to process structured features. The input data is the WOE encoded features (20 dimensions) after binning, including annual income binning, living conditions, gender, education, loan amount, loan term, loan interest rate, loan grade, loan purpose category, debt-to-income ratio, installment amount, employment title, employment years WOE value, device static features (such as installed application directory, device usage behavior, IP address), etc. LightGBM adopts a leaf-based splitting method, limiting the number of leaves to 31, controlling the minimum number of leaf nodes to 20, and randomly retaining 70% of the features. The gradient boosting decision tree is used for model training, and the default probability is finally output. The above number of leaves and the minimum number of samples of leaf nodes can be adjusted and set according to the actual application scenario.
[0085] Next, we adjust the dimensions of the output of the LightGBM model. , while the original WOE feature matrix (i.e., the characteristic encoding of user personal information) may be distributed in , so Z-score standardization is adopted: ,in represents the mean of the WOE matrix, Represents the standard deviation of the WOE matrix, and then performs feature splicing. After horizontal splicing, the following enhanced feature matrix is formed: .
[0086] Build a neural network architecture and directly input 21-dimensional Feature matrix, hidden layer 1 sets 64 neurons, forward propagation calculates output, uses ReLU activation function, combined with He normal distribution for initialization, and then uses dropout layer to discard 30% of neurons to prevent overfitting. The process can be expressed as follows;
[0087]
[0088]
[0089] in, represents the output features of hidden layer 1, , denote the weight matrix and bias of hidden layer 1 respectively, represents the output of the ReLU activation function, Represents the ReLU activation function.
[0090] Hidden layer 2 is set with 32 neurons, and the output is calculated by forward propagation. The same initialization strategy is set as hidden layer 1, and L2 regularization is added to the loss function during training: ,in, represents the dynamic cross entropy loss function, represents each weight matrix of hidden layer 2, Represents the weight matrix The matrix elements of the matrix, that is, the subscripts i and j are weight matrices The element index of represents the weight of L2 regularization, and the weight Can be set to ; At the same time, set the penalty term to effectively constrain the size of the weight and prevent the model from overfitting the training data. Finally, connect a linear activation function to output the predicted amount That is, for the enhanced feature matrix obtained , output the predicted amount through the constructed prediction network module When , the prediction network module can be set to: at least two fully connected layers, and finally output the amount through a linear activation function .
[0091] It is worth noting that during the training process of the above prediction network module, the optimized dynamic binary cross entropy loss function can be used to update the corresponding network parameters:
[0092]
[0093] in, is the true label, is the default probability predicted by the model (i.e. the output of the LightGBM model), represents the number of samples, is the sample number, To emphasize the contribution of certain key features to the loss, the specific design is to dynamically update the weights based on the IV values calculated by the previous automatic binning. For each feature f, the weight can be designed according to the calculated IV value. The higher the IV value, the greater the contribution of the feature to the predicted probability: ,in, represents the information value of feature f, represents the maximum value function, Used to obtain the maximum information value of all features. This embodiment improves the model's understanding of key variables by prioritizing the impact of important features.
[0094] In this embodiment, the Adam (Adaptive Moment Estimation) optimizer is used for gradient optimization, and momentum and adaptive learning rate are combined to accelerate convergence and avoid local optimality. The training batch is 64, and the early stopping strategy is adopted. Cross-validation uses 5-fold cross-validation to generate stacked features.
[0095] (8) Take the weighted average of the first prediction amount output by LSTM and the second prediction amount output by the neural network to obtain the final prediction amount: ,in, Indicates the weight of the first predicted amount. K-fold cross validation can be used to determine the weight value.
[0096] In one embodiment, the automatic binning module can be dynamically regulated based on the dynamic risk control adjustment module of the server. Specifically, dynamic regulation is performed through PSI monitoring and KS monitoring. The characteristic PSI index is calculated every month to measure the change of the input data distribution of the model (realizing automatic binning) over time. When the data distribution changes significantly, the prediction ability of the model will decrease, and the model needs to adapt to the latest user characteristics. The specific setting is that when the index is greater than 8%, the binning reconstruction process is triggered. The PSI calculation is as follows: ,in, Indicates the proportion of samples in the i-th bin in the current time period, Indicates the proportion of samples in the i-th bin in the benchmark time period, is the number of bins.
[0097] Update the model data once a quarter to ensure that the model adapts to the new market situation. Use the KS indicator to measure the fluctuation of model performance. When the KS value fluctuates by more than 2%, re-screen the features and adjust the binning strategy. Use the AUC index to evaluate the prediction ability of the new model and compare it with the old model to ensure that the performance does not decrease. KS is calculated as follows: , is the cumulative distribution function of all actual defaulting users, is the cumulative distribution function of all normal users.
[0098] The embodiment of the present application is based on the LSTM model of the encoder-decoder, and the hybrid model architecture based on the LightGBM model to predict the user's loan amount, wherein the LSTM model is combined with the attention mechanism to capture the long-term dependence of the time series features, and the LightGBM model can efficiently parse the structured features. The combination of the two can simultaneously characterize dynamic behavior data and static credit factors, improve the comprehensive performance of the risk control model, and finally obtain the predicted amount through the weighted fusion decision-making mechanism. This implementation method can not only help banks and financial institutions reduce the risk of default, but also improve user experience and personalization of financial services. For example, by predicting the credit limit demand and user credit changes in the future period of time, the system can promptly identify high-risk customers and take corresponding measures (such as freezing accounts, adjusting credit limits, pushing repayment reminders, etc.). In addition, the system can also provide customers with personalized credit limit adjustment suggestions based on the user's historical behavior patterns and future credit assessments, thereby realizing intelligent and sustainable credit management.
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An intelligent financial risk control system based on mobile device monitoring, including a user terminal, a management terminal and a server, characterized in that: The user terminal is used for user identity authentication, collecting user information, and for the user to initiate a financial service application to the management terminal; wherein the user information includes user personal information and user device behavior information of the user terminal; wherein the user personal information is actively collected by the user terminal; the user device information is collected based on the monitoring instruction initiated by the management terminal; The management terminal is used to receive and process user information collected by the user terminal, and includes a data preprocessing module, a security storage module and a financial business management module; wherein the data preprocessing module is used to perform data preprocessing on the received user information and send it to the security storage module; the security storage module de-identifies the user information after data preprocessing based on the user ID mapping strategy, and then uploads the processed user information to the server to store it in a designated storage location; the financial business management module is used to receive the financial business application initiated by the user terminal, access the server based on the role-based access control mechanism, and call the corresponding business processing module on the server based on the mapped user ID and the financial business application content, and generate a business processing result in response to the financial business application initiated by the user based on the output of the currently called business processing module and return it to the user terminal; wherein the business processing module of the server includes a credit scoring module and a credit limit prediction module; The server includes an automatic box sorting module, a credit scoring module and a credit limit prediction module; wherein the automatic box sorting module is used to encode the information features of the pre-processed user personal information uploaded by the management terminal, and send the feature codes of the user personal information to the credit limit prediction module and the credit scoring module; the credit scoring module predicts the user's credit score based on the feature codes of the user's personal information, and the credit limit prediction module predicts the user's loan limit based on the user's device behavior time series features and the feature codes of the user's personal information, wherein the user's device behavior time series features are obtained based on the user's device behavior information after data processing.
2. The intelligent financial risk control system based on mobile device monitoring as claimed in claim 1, characterized in that: The credit scoring module of the server calculates the user's default probability based on the personal information characteristics and converts it into a credit score through a dynamic scoring strategy; the credit limit prediction module includes two prediction models; one prediction model adopts an encoder-decoder structure and predicts the user's first loan prediction limit based on the time series characteristics of the user's device behavior; the other prediction model is a neural network structure based on a gradient boosting decision tree, which predicts the user's second loan prediction limit based on the feature encoding of the user's personal information; The final predicted amount of the amount prediction module is obtained based on the weighted sum of the first and second loan prediction amounts.
3. The intelligent financial risk control system based on mobile device monitoring as claimed in claim 1, characterized in that: The server also includes a dynamic risk control adjustment module, which regularly calculates the group stability index of information features based on a set first detection period, and triggers bin reconstruction processing of the automatic binning module when the group stability index is greater than or equal to a specified first threshold; the dynamic risk control adjustment module regularly calculates the statistics of the credit scoring module based on a set second detection period, and re-performs feature screening of user information and adjusts the binning strategy when the fluctuations of the statistics for the last two times are greater than or equal to a specified second threshold, wherein the first detection period is less than the second detection period.
4. The intelligent financial risk control system based on mobile device monitoring as claimed in claim 1, characterized in that: User personal information specifically includes: name, telephone number, gender, annual income, ID number, residential status, education level, loan amount, loan term, loan interest rate, loan grade, loan purpose category, debt-to-income ratio, installment amount, employment title and years of employment.
5. The intelligent financial risk control system based on mobile device monitoring as claimed in claim 1, characterized in that: User device behavior information includes: installed application catalog, device usage behavior, IP address, number of daily logins, single usage duration, frequency of geographic location changes, and usage time percentage of risky target applications.
6. The intelligent financial risk control system based on mobile device monitoring as claimed in claim 2, characterized in that: The prediction model using an encoder-decoder structure is set as a network structure based on a long short-term memory network and an attention mechanism, wherein the encoder is a structure based on a long short-term memory network, the decoder is the inverse operation of the encoder, and the encoded features output by the encoder are processed by the attention mechanism network before being sent to the decoder, and then the output of the decoder is mapped to the first loan prediction amount based on the fully connected layer. The model input of the prediction model specifically includes: number of daily logins, single usage duration, frequency of geographic location changes, and proportion of usage time of risk target applications.
7. The intelligent financial risk control system based on mobile device monitoring as claimed in claim 2, characterized in that: The prediction model of the neural network structure based on the gradient boosting decision tree is specifically set as: The feature encoding of user personal information is input into the gradient boosting decision tree model. The output of the model is the default probability with a value range of 0 to 1, which is recorded as ; After the characteristic coding of user personal information is standardized, The enhanced feature matrix is obtained by concatenation; A neural network is constructed to predict the second loan prediction amount, and the input data of the neural network is the enhanced feature matrix.
8. The intelligent financial risk control system based on mobile device monitoring as claimed in claim 7, characterized in that: The loss function used in the training of the prediction model based on the neural network structure of the gradient boosting decision tree is set as: ; in, represents the loss function, N represents the number of samples, is the sample number, is the true label, Representation sample The probability of default, To set the contribution factor, real-time adjustments are made based on the value of information obtained by the automated binning module. The adjustment method is: ,in represents the information value of feature f, It indicates determining the maximum information value of all features.
9. The intelligent financial risk control system based on mobile device monitoring as claimed in claim 1, characterized in that: The data preprocessing of the data preprocessing module includes: drawing histograms, calculating descriptive statistics, handling outliers and missing values, data cleaning, and converting categorical variables in user information using one-hot encoding.
10. The intelligent financial risk control system based on mobile device monitoring according to claim 1, characterized in that: The user ID mapping strategy is as follows: For the set target sensitive information, a hash algorithm is used to generate a unique identifier, and the unique identifier is used as the data index of the server to obtain the de-identified user information; And transmit the de-identified user information to the server.
Citation Information
Patent Citations
Microloan system based on big data intelligent risk control and microloan method thereof
CN107330785A
Credit line management and control method and device based on artificial intelligence, equipment and medium
CN114971866A
Financial data prediction method and device, electronic equipment and storage medium
CN116308735A
Risk control model creation method and device, electronic equipment and storage medium
CN116542511A