Method and system for predicting debt escaping behavior after expiration
By collecting and processing multi-dimensional behavioral data, using two-way long and short-term memory networks and deep reinforcement learning for timing modeling, and combining multi-objective decision-making and fuzzy logic for risk assessment, the problem of real-time data and model update lag in the existing technology is solved, and real-time prediction and efficient risk prevention and control of debt evasion after overdue debt is achieved.
Patent Information
- Application Number
- CN202510303020.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-14
AI Technical Summary
In the prediction of debt evasion after overdue, the existing technology has problems such as insufficient real-time data, lagging model updates, strong dependence on manual labeling, single feature extraction, and limited time-series dynamic change capture, which makes it difficult for the timeliness and accuracy of risk prediction results to meet financial risk control requirements.
By collecting user's multi-dimensional behavioral data, data cleaning and standardization are performed, features are extracted and time sequence data are constructed. The two-way long and short-term memory network and deep reinforcement learning technology are used to perform timing dynamic modeling and parameter adjustment, and risk assessment and decision intervention are combined with multi-objective decision algorithms and fuzzy logic, and adaptive updates of the model are achieved through incremental learning and Bayesian optimization.
Real-time prediction and early warning of debt evasion after overdue debt is achieved, providing financial institutions with efficient risk prevention and control tools, significantly improving the accuracy of risk assessment and the robustness of prediction models.
Smart Images

Figure CN120219083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and financial risk control, and particularly relates to a method and system for predicting the behavior of defaulting on debts after overdue. Background Art
[0002] The technology for predicting the behavior of defaulting on debts after overdue aims to identify potential debt default risks in advance by deeply analyzing the behavior of users after overdue, and provide a basis for financial institutions to prevent risks in advance. The existing technology mainly relies on social network data and determines the debt default group by constructing a classification model based on XGBoost. Although it realizes risk identification to a certain extent, there are defects such as insufficient data timeliness, lag in model update, strong dependence on manual annotation, and single feature extraction. In addition, the traditional method has limited capture of the temporal dynamic changes of user behavior, resulting in the timeliness and accuracy of risk prediction results being difficult to meet the requirements of financial risk control. Summary of the Invention
[0003] In view of the above-mentioned many problems existing in the prior art, the present invention provides a method and system for predicting the behavior of defaulting on debts after overdue. The present invention collects multi-dimensional behavior data of users, extracts features after data cleaning and standardization processing, and fuses them into time series data by a fixed window method; then uses a bidirectional long short-term memory network to model the time series data, captures the bidirectional time series dependence of user behavior, and combines deep reinforcement learning technology to dynamically adjust the model parameters to reduce prediction errors; finally, conducts comprehensive risk scoring and numerical conversion on the prediction results through a multi-objective decision-making algorithm and fuzzy logic, and uses an auxiliary statistical method for cross-validation. Finally, the model is adaptively updated by incremental learning and Bayesian optimization to generate accurate final prediction results. The present invention realizes real-time prediction and early warning of the behavior of defaulting on debts after overdue, and provides an efficient risk prevention and control tool for financial institutions.
[0004] A method for predicting the behavior of defaulting on debts after overdue includes the following steps:
[0005] Collect multi-dimensional behavior data of users, perform data cleaning and data standardization processing on the behavior data, extract features from the standardized data, and generate a fused feature dataset;
[0006] Construct time series data based on the fused feature dataset, train with a bidirectional long short-term memory network to generate time series risk prediction data for each user, and use deep reinforcement learning to adjust the parameters of the bidirectional long short-term memory network to form a risk prediction model and output risk prediction data;
[0007] Based on the risk prediction data, a multi-objective decision-making algorithm is used to calculate a comprehensive risk score, and fuzzy logic is used to numerically adjust the comprehensive risk score to form decision intervention data. At the same time, an auxiliary statistical model is used to cross-validate the risk prediction data, and risk assessment data verified by statistics is output;
[0008] Using the risk assessment data verified by statistics as feedback, the incremental learning method is used to update the parameters of the risk prediction model, and the hyperparameters of the risk prediction model are adjusted by an adaptive optimization algorithm, so as to generate the final prediction result.
[0009] Preferably, the multi-dimensional behavior data is collected from the user account transaction records, user social interaction records and user device usage records according to a preset data format, and the data outside the statistical interval is removed by an outlier detection algorithm. At the same time, a missing data filling algorithm is used to compensate for the missing items in the data, and the collected data is normalized by the min-max normalization method, so as to generate the standardized original data with a unified numerical range.
[0010] Preferably, the features in the standardized original data are calculated by a feature extraction algorithm based on a fixed window. In the calculation process, the account changes, social interaction frequencies and device usage patterns of the user are quantified respectively, and the quantified results are accumulated according to a preset formula to form a fused feature dataset.
[0011] Preferably, the time series data constructed based on the fused feature dataset is trained by a bidirectional long short-term memory network. The bidirectional long short-term memory network sets a forward layer and a backward layer in the network structure, and uses the backpropagation algorithm to update the network parameters to generate the time series risk prediction data of each user.
[0012] Preferably, the deep reinforcement learning adopts the proximal policy optimization algorithm to dynamically adjust the weights of the bidirectional long short-term memory network according to a preset reward function, where the reward function uses the risk prediction error as the feedback input, so as to form a risk prediction model and output risk prediction data.
[0013] Preferably, the multi-objective decision-making algorithm calculates the respective risk scores of the financial behavior risk, social behavior risk and device behavior risk according to preset weighting coefficients, and generates comprehensive risk score data through cumulative operation; the comprehensive risk score data is converted by a preset membership function to discretize the continuous risk score into decision-making parameters, forming decision intervention data.
[0014] Preferably, the auxiliary statistical model adopts a five-fold cross-validation method, divides the risk prediction data into a training set and a validation set, and compares the consistency between the risk prediction data and the decision intervention data by calculating statistical indicators, so as to output the risk assessment data verified by statistics.
[0015] Preferably, the incremental learning method adopts the online gradient descent algorithm. When the newly collected user behavior data reaches a preset threshold, the new data is incorporated into the risk prediction model to update the model parameters in real time.
[0016] Preferably, the adaptive optimization algorithm adopts Bayesian optimization technology to automatically adjust the hyperparameters of the risk prediction model within a preset hyperparameter search range, and generates a final prediction result based on the statistically verified risk assessment data as feedback.
[0017] A system for predicting the behavior of evading debts after overdue, which is used to implement the method for predicting the behavior of evading debts after overdue, and the system includes:
[0018] A collection module, which is used to collect multi-dimensional behavior data of user account transaction records, user social interaction records, and user device usage records according to a preset data collection scheme;
[0019] A data processing module, which is used to perform data cleaning, abnormal data elimination, missing data filling, and min-max standardization processing on the multi-dimensional behavior data, and extract features from the standardized data to generate a fused feature dataset;
[0020] A time series modeling module, which is used to construct time series data based on the fused feature dataset, and use a bidirectional long short-term memory network to train the time series data to generate time series risk prediction data for each user. At the same time, use deep reinforcement learning to adjust the parameters of the weights of the bidirectional long short-term memory network to form a risk prediction model and output risk prediction data;
[0021] A decision-making module, which is used to calculate the risk scores of each risk dimension by using a multi-objective decision-making algorithm based on the risk prediction data, calculate the comprehensive risk score according to a preset weighting coefficient, and perform numerical conversion on the comprehensive risk score through a preset membership function to form decision intervention data. At the same time, use an auxiliary statistical module to perform cross-validation on the risk prediction data and output statistically verified risk assessment data;
[0022] A model update module, which is used to update the parameters of the risk prediction model by using an online gradient descent incremental learning method according to the feedback of the risk assessment data, and automatically adjust the hyperparameters of the risk prediction model within a preset search range by using Bayesian optimization technology, so as to generate a final prediction result.
[0023] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:
[0024] Through the multi-modal data fusion technology, the present invention realizes the comprehensive collection and standardized processing of multi-dimensional behavioral data of users, thereby generating a high-quality fused feature dataset;
[0025] Through the bidirectional long short-term memory network and deep reinforcement learning technology, the present invention realizes the temporal dynamic modeling and real-time parameter adjustment of users' behaviors after overdue, thereby outputting stable risk prediction data;
[0026] Through the multi-objective decision-making algorithm and fuzzy logic technology, the present invention realizes the quantitative integration of different risk dimensions and the generation of decision intervention data;
[0027] Through the cross-validation of the auxiliary statistical model and the online gradient descent and Bayesian optimization technology, the present invention realizes the real-time update and adaptive optimization of the risk prediction model, thereby generating the final prediction result.
[0028] The comprehensive application of these technical means not only solves the problems of poor data real-time performance and untimely model update in the existing methods, but also significantly improves the accuracy of risk assessment and the robustness of the prediction model, providing scientific and automated technical support for financial risk control. Brief Description of the Drawings
[0029] Figure 1 It is a schematic flow chart of the method of the present invention;
[0030] Figure 2 It is a schematic diagram of deep reinforcement learning in the present invention;
[0031] Figure 3 It is a structural block diagram of the system of the present invention. Detailed Embodiments
[0032] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.
[0033] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0034] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0035] As Figure 1 shown, a method for predicting the behavior of defaulting on debts after overdue includes the following steps:
[0036] Collect multi-dimensional behavioral data of users, perform data cleaning and data standardization processing on the behavioral data, extract features from the standardized data, and generate a fused feature dataset;
[0037] For the prediction of the behavior of defaulting on debts after overdue, the present invention first constructs a basic dataset reflecting the behavioral characteristics of users through data collection and preprocessing. The overall process includes collecting multi-dimensional behavioral data of users, cleaning and standardizing the collected data, and extracting features from the standardized data and generating a fused feature dataset.
[0038] In this step, in the data collection stage, through a preset data collection scheme, behavioral data such as user account transaction records, user social interaction data, and user device usage records are obtained from multiple information sources; subsequently, through data cleaning techniques, incorrect data is removed and outliers are detected from the original data, and a missing data filling algorithm is applied to compensate for the missing items in the data; then, a numerical standardization technique is used to uniformly map data with different dimensions and value ranges into a fixed interval, thereby generating standardized original data. The standardized data is used as the input for subsequent feature extraction. Through segmented statistics and calculations, key information such as changes in user accounts, social interaction frequencies, and device usage patterns is extracted, and each index is integrated according to preset rules to form a comprehensive fused feature dataset. This dataset not only has the characteristics of unity and standardization in data structure, but also can fully reflect the temporal dynamics of user behavior in terms of content, providing a high-quality input basis for the training and parameter optimization of subsequent risk prediction models. Through the above data collection and preprocessing process, the present invention comprehensively and systematically records and processes user behavior at the data level, ensuring the accuracy and stability of the subsequent prediction process, and at the same time providing sufficient data support and a reliable theoretical basis for risk assessment.
[0039] Preferably, the multi-dimensional behavioral data is collected according to a preset data format from user account transaction records, user social interaction records, and user device usage records, and data outside the statistical interval is removed through an outlier detection algorithm. At the same time, a missing data filling algorithm is used to compensate for the missing items in the data, and the collected data is normalized using the minimum-maximum normalization method to generate standardized original data with a unified value range.
[0040] In the present invention, the multi-dimensional behavior data is specifically composed of user account transaction records, user social interaction records, and user device usage records, and its collection process is automatically completed according to a preset data format. First, the system extracts the user's account transaction records, social interaction records, and device usage records from the financial trading platform, social media, and mobile device logs respectively according to a fixed data template. Each type of data includes a clear timestamp, numerical indicators, and relevant identifiers.
[0041] After collection, various types of data are initially screened through an outlier detection algorithm. The algorithm uses a statistical distribution model (such as a normal distribution or exponential distribution model) to calculate the mean and standard deviation of each piece of data, and eliminates the values that exceed the set statistical interval. At the same time, a missing data filling algorithm is adopted to compensate for the missing items in the data according to the mean or median of the same user's historical data to ensure data integrity. Then, the minimum-maximum normalization method is used for the data processed above to map each piece of data into a fixed numerical interval (such as [0,1]). The data after normalization is the original data after standardization. This process realizes automated data processing through programming to ensure the consistency of the data formats and numerical ranges of data from various sources, providing unified and reliable input data for subsequent feature extraction.
[0042] Preferably, the features in the original data after standardization are calculated using a feature extraction algorithm based on a fixed window. In the calculation process, the user's account changes, social interaction frequency, and device usage patterns are quantified respectively, and the quantified results are accumulated according to a preset formula to form a fused feature dataset.
[0043] After generating the original data after standardization, the present invention further uses a feature extraction algorithm based on a fixed window to quantify the key information in the data to construct a fused feature dataset.
[0044] In specific operations, the system first sets a fixed time window, for example, one hour, one day, or one week as a window, and divides the original data after standardization into multiple sub-datasets according to the window. Within each window, statistical calculations are respectively performed on the user account transaction records, social interaction records, and device usage records. For the user account transaction records, the system calculates the average value, volatility, and transaction frequency of the transaction amount within each window; for the user social interaction records, it statistically calculates the number of interactions, information transmission frequency, and interaction time distribution within each window; for the user device usage records, it quantifies the device activation duration and usage frequency. Each statistical indicator is calculated using a preset formula. For example, the formula for calculating the window mean of a certain indicator is set as:
[0045]
[0046] Among them, μ w is the mean value within the window, N represents the total number of data within the window, and x i is the i-th data value. Subsequently, the system accumulates various statistical indicators weighted according to preset rules to form the final fused feature dataset, where each weighting coefficient is determined in advance based on historical data performance and expert experience. This weighted accumulation process ensures that data in different dimensions can reflect their respective importance during fusion, and through the accumulation formula:
[0047]
[0048] Among them, F represents the fused feature value, α j is the weighting coefficient of the j-th feature, and f j is the statistic of the j-th feature within the window. Finally, the fused feature values generated within all windows are integrated into continuous time-series data to provide data support for subsequent risk prediction model training. This fixed-window feature extraction method has the advantages of simple operation, clear parameters, and strong repeatability, and can stably reflect key behavioral characteristics among different user groups.
[0049] Based on the fused feature dataset, construct time-series data, use a bidirectional long short-term memory network for training to generate time-series risk prediction data for each user, and use deep reinforcement learning to adjust the parameters of the bidirectional long short-term memory network to form a risk prediction model and output risk prediction data;
[0050] The present invention aims at the problem of predicting the behavior of defaulting on debts after overdue, uses the time-series data constructed from the fused feature dataset for risk prediction, and dynamically adjusts the parameters of the time-series model through deep reinforcement learning, thereby constructing a risk prediction model that can output risk prediction data. Specifically, the present invention first forms standardized data after preprocessing multi-dimensional behavioral data, then extracts key features from it to generate a fused feature dataset; based on this fused feature dataset, the system constructs continuous time-series data and uses a bidirectional long short-term memory network for training to generate risk prediction data for each user in future time periods. Subsequently, the system introduces deep reinforcement learning technology to dynamically adjust the weights of each layer in the bidirectional long short-term memory network, aiming to reduce the prediction error and improve the robustness of the model. Finally, a risk prediction model optimized by reinforcement learning is formed, and the risk prediction data output by this model can be directly used for subsequent risk management and decision-making.
[0051] In specific implementation, the present invention first collects various types of user behavior data, including financial transaction records, social interaction records, and device usage records. The data cleaning technology is used to eliminate abnormal data, and the standardized processing is adopted to map the data to a unified numerical interval, ensuring that each data source can be compared and integrated on the same scale. Subsequently, through the feature extraction algorithm based on a fixed window, the standardized data is statistically analyzed hour by hour or day by day, and key indicators reflecting user account changes, social activity levels, and device usage patterns are extracted. Each indicator is accumulated according to the preset weight to generate a fused feature dataset.
[0052] Next, the fused feature dataset is used as input to construct continuous time series data, which retains the information on the behavior changes of users within different time windows. To capture this time series dynamics, a bidirectional long short-term memory network (Bi-LSTM) is used for model training. By setting a forward layer and a backward layer, Bi-LSTM can consider the influence of both historical information and future information on the current state simultaneously, thus more accurately predicting the behavior risks of users. During the training process of the model, the backpropagation algorithm is used to adjust the network parameters, and a loss function (such as mean squared error) is used to measure the deviation between the predicted output and the real data, so as to continuously iterate and optimize.
[0053] After initially obtaining the time series risk prediction data, the present invention further introduces the deep reinforcement learning technology to dynamically adjust the parameters of the bidirectional long short-term memory network. By adopting the proximal policy optimization (PPO) algorithm, the system dynamically corrects the network weights according to the preset reward function, with the risk prediction error as the feedback input, reducing the model prediction error. This process not only enables the model to adapt to the dynamic changes of user behavior, but also ensures that the model still has strong prediction ability and stability under continuously updated data. Finally, the risk prediction model formed through this optimization process can output risk prediction data, serving as the basic data for subsequent risk management, decision-making, and intervention measures. This overall solution constitutes a complete data processing and prediction link for predicting the behavior of defaulting on debts after overdue from data collection, feature fusion, time series modeling, reinforcement learning optimization to model output, providing technical support and decision-making basis for risk prevention and control.
[0054] Preferably, the time series data constructed based on the fused feature dataset is trained using a bidirectional long short-term memory network. The bidirectional long short-term memory network sets a forward layer and a backward layer in the network structure, and uses the backpropagation algorithm to update the network parameters to generate the time series risk prediction data for each user.
[0055] The time series data constructed based on the fused feature dataset is trained using a bidirectional long short-term memory network. Its specific implementation includes explicitly setting a forward layer and a backward layer in the network structure. The forward layer is used to process the chronological information of the input data, while the backward layer is used to capture the impact of future states on the current decision in reverse. In the specific operation, first, according to the preset time window, the fused feature dataset is divided into several sub-datasets in continuous chronological order, and each sub-dataset corresponds to a fixed time interval, such as every hour or every day. The data points in each sub-dataset contain the feature values obtained after preprocessing and fusion of the user account transaction records, social interaction records, and device usage records, and these feature values reflect the user's behavior state within the time window.
[0056] After constructing the time series data, this data is passed as input into the bidirectional long short-term memory network. The network structure adopts a standard bidirectional structure design. When the forward layer processes the time series data, it gradually transmits information backward from the starting time and captures the long-term dependencies of the early data through memory units; while the backward layer transmits information backward from the end of the sequence to capture the impact of the later data on the current state. The bidirectional long short-term memory network uses activation functions (such as the tanh and sigmoid functions) and structures such as forget gates, input gates, and output gates for selective memory and forgetting of information, so as to automatically retain the key information beneficial to prediction during the information transmission process.
[0057] During the training process, the network uses the backpropagation algorithm to calculate the loss function value (commonly such as the mean squared error, formula:
[0058]
[0059] where, y i is the actual risk state, is the predicted risk value, and N is the number of samples) to determine the error between the predicted output and the actual output, and this error is propagated backward to each layer of the network through the gradient descent method to update the network parameters. During the parameter update process, the system sets a fixed learning rate and can use the momentum method to accelerate convergence, so as to ensure the stability and efficiency of the training process.
[0060] For example, in practical applications, for the behavior data of a certain user within 24 consecutive hours, the fused feature data extracted through a fixed window can generate input vectors at 24 moments, and each vector consists of multiple dimensions. After inputting these vectors into the bidirectional long short-term memory network, the network outputs the hidden states corresponding to each time point respectively, and these hidden states form the final risk prediction data after being processed by the fully connected layer. This process makes full use of the time bidirectional dependency information captured by the bidirectional structure to ensure that the prediction result reaches a balance between considering historical trends and future dynamics.
[0061] In addition, during the model training process, the system regularly saves the model weights and evaluates the model's performance through the validation set to ensure the stability and repeatability of the training results. The cross-validation method is used to verify the model's prediction ability, and the training parameters are adjusted according to the verification results, further improving the robustness of data extraction and model training. The entire process realizes automated data processing through programming, ensuring that the model training steps can be reproduced in different data collection cycles, providing solid data support and algorithm guarantee for the prediction of debt evasion behaviors after overdue.
[0062] Preferably, as Figure 2 shown, the deep reinforcement learning adopts the Proximal Policy Optimization algorithm to dynamically adjust the weights of the bidirectional long short-term memory network according to a preset reward function, where the reward function uses the risk prediction error as the feedback input, thereby forming a risk prediction model and outputting risk prediction data.
[0063] The deep reinforcement learning adopts the Proximal Policy Optimization algorithm to dynamically adjust the weights of the bidirectional long short-term memory network through a preset reward function. Its specific implementation involves using the risk prediction error as the feedback input to realize real-time correction of the model parameters, thereby forming a risk prediction model and outputting risk prediction data. In this technology, deep reinforcement learning, as an adaptive adjustment mechanism, its core principle is to continuously interact with the environment and update the policy parameters according to the reward signal to continuously reduce the prediction error. Specifically, during implementation, first compare the sequential risk prediction data output by the bidirectional long short-term memory network with the actual risk data to calculate the risk prediction error, which is usually measured using the mean square error or the absolute error. Let the risk prediction error be where y is the actual risk state, is the predicted risk value. The preset reward function determines the reward signal according to the error value. For example, a positive reward is given when the error |e| is less than a predetermined threshold, otherwise a negative reward is given.
[0064] Under the framework of the Proximal Policy Optimization (PPO) algorithm, the policy network uses the current weights of the bidirectional long short-term memory network as the initial policy, and calculates the policy update objective function using the sampled risk prediction data and the corresponding reward values. The update objective function can be expressed as:
[0065]
[0066] where represents the ratio of the new and old policies, is the advantage function, and ∈ is the clipping parameter. This formula ensures that the policy update is carried out within a certain range to prevent the update amplitude from being too large and destroying the stability of the original model.
[0067] In actual operation, the system first generates several batches of samples through a simulation environment. Each sample contains the current state, the actions taken, the rewards obtained, and the information of the next state. The advantage function is calculated based on the sample data, and the weights of the bidirectional long short-term memory network are adjusted through the PPO algorithm. This process uses the batch gradient descent method to calculate the gradient of the loss function, and the gradient calculation process uses an automatic differentiation tool to achieve accurate calculation. The adjusted network weights are used as a new policy in the next sample collection, thus realizing continuous policy iteration and update.
[0068] For example, for the behavior data of a specific user within multiple consecutive time windows, the system compares the predicted risk with the actual risk, calculates a series of error values, then generates corresponding reward signals according to a preset reward function, and uses the proximal policy optimization algorithm to update the current policy. After multiple rounds of iteration, when the model converges, the prediction error is significantly reduced, thus forming a risk prediction model with relatively high robustness. The risk prediction data output by this model has relatively high credibility in risk assessment and decision-making.
[0069] In addition, to ensure the stability of the deep reinforcement learning process during parameter adjustment, the system monitors the range of policy changes after each round of policy update and sets hyperparameters to control the policy change rate, thus preventing the model from falling into a local optimum due to too fast convergence. The entire deep reinforcement learning optimization process is implemented through modular programming in the program and is repeatedly trained and verified in combination with actual risk data to ensure its operability and stability in the prediction of defaulting on debts after overdue. Through real-time feedback and dynamic adjustment, the present invention enables the bidirectional long short-term memory network to more accurately capture the subtle changes in users' risk behaviors, providing strong support for the output of the final risk prediction data.
[0070] Based on the risk prediction data, a multi-objective decision-making algorithm is used to calculate a comprehensive risk score, and fuzzy logic is used to numerically adjust the comprehensive risk score to form decision intervention data. At the same time, an auxiliary statistical model is used to cross-validate the risk prediction data, and risk assessment data verified by statistics is output;
[0071] In the present invention, based on the risk prediction data, a multi-objective decision-making algorithm is used to calculate a comprehensive risk score, and fuzzy logic is utilized to adjust the score numerically so as to form decision intervention data, ultimately providing a quantitative risk assessment result for the prediction of debt evasion behavior after overdue. The core of this process lies in integrating risk assessment indicators from multiple behavioral dimensions and forming a unified risk score through a comprehensive decision-making mechanism, which is an important basis for subsequent decision intervention. Specifically, the comprehensive risk score is obtained by calculating the risk performance of the user in behaviors such as finance, social interaction, and device usage in a weighted cumulative manner to obtain an overall risk score. Subsequently, fuzzy logic is used to convert the score from a continuous value into a discrete decision parameter, thereby forming decision intervention data to guide subsequent risk control and management measures. In addition, the system cross-verifies the risk prediction data and decision intervention data through an auxiliary statistical model to ensure the accuracy and stability of the scoring model. Through five-fold cross-validation, the performance of the model is verified to ensure the effectiveness of the risk assessment data, thus providing a solid theoretical basis and data support for the final prediction of overdue debt evasion.
[0072] Preferably, the multi-objective decision-making algorithm calculates the respective risk scores for financial behavior risk, social behavior risk, and device behavior risk according to preset weighting coefficients, and generates comprehensive risk score data through cumulative operation; the comprehensive risk score data is converted via a preset membership function to discretize the continuous risk score into decision parameters, forming decision intervention data.
[0073] In the application of the multi-objective decision-making algorithm, the present invention first takes financial behavior risk, social behavior risk, and device behavior risk as the main evaluation dimensions. The risk score of each dimension is calculated through preset weighting coefficients and the comprehensive risk score data is obtained through cumulative operation. During the calculation of the risk score of each dimension, normalization processing is first performed based on historical data to eliminate the dimensional differences between different dimensions, making the scores of each dimension comparable. For financial behavior risk, the system calculates by evaluating indicators such as the user's credit history, arrears record, and overdue situation. The specific methods include changes in credit limits, fluctuations in account balances, etc.; for social behavior risk, the system calculates the risk brought by the user's social activity by analyzing data such as the interaction frequency, content posting frequency, and social circle activity on the social platform; for device behavior risk, the system scores based on factors such as the usage pattern of the user's device, device switching frequency, and the geographical location of the logged-in device.
[0074] During the risk score calculation process, preset weighted coefficients are set. For example, the weight of financial behavior risk may be 0.4, the weight of social behavior risk is 0.3, and the weight of device behavior risk is 0.3. These weighted coefficients are adjusted based on the experience of risk management experts and the performance of historical data, and the relative importance of each dimension in the comprehensive score is determined through data analysis. The specific calculation process is as follows:
[0075] R total = α1 × R finance + α2 × R social + α3 × R device
[0076] Wherein, R total is the final comprehensive risk score, R finance , R social and R device are the scores of financial, social, and device behavior risks respectively, and α1, α2, and α3 are the corresponding weighted coefficients. Through such weighted calculation, the system can comprehensively consider the risk situations of each dimension, thereby generating an accurate comprehensive risk score.
[0077] After obtaining the comprehensive risk score data, the present invention uses fuzzy logic to adjust the score and convert the continuous risk score value into a decision parameter. The role of fuzzy logic in this process is to map the originally continuous risk score to a discrete decision interval through the membership function for subsequent decision intervention. The membership function is a fuzzy mathematics method that can map a numerical input to a membership value within an interval, and this value represents the degree of membership of the numerical value in a specific category. For example, during the adjustment of the risk score, assuming the range of the comprehensive risk score is from 0 to 100, the membership function can map different score values to categories such as "low risk", "medium risk", and "high risk" according to preset rules.
[0078] For example, the set membership function can be defined in the following way:
[0079] When the risk score R total is between [0, 30], the membership value is the value in the "low risk" region;
[0080] When R total is between [30, 70], the membership value is the value in the "medium risk" region;
[0081] When R total is between [70, 100], the membership value is the value in the "high risk" region.
[0082] The specific implementation of this adjustment process is carried out through a fuzzy rule engine, in which the input risk scores are classified by rules, and finally decision intervention data is generated. These decision data can be used for subsequent risk intervention measures, such as automatically generating warning signals, triggering further reviews, notifying relevant departments to take actions, etc. The core of the present invention lies in the fine adjustment of risk scores through fuzzy logic, enabling the system to make flexible and appropriate decisions for different risk levels.
[0083] Preferably, the auxiliary statistical model adopts a five-fold cross-validation method, divides the risk prediction data into a training set and a validation set, and compares the consistency between the risk prediction data and the decision intervention data by calculating statistical indicators, so as to output statistically verified risk assessment data.
[0084] To ensure the stability of the risk prediction model and the accuracy of the prediction results, the present invention uses an auxiliary statistical model for cross-validation to further verify the consistency between the risk prediction data and the decision intervention data. Cross-validation is a commonly used verification method. It evaluates the prediction ability of the model by dividing the data set into multiple subsets and alternately using each subset as the validation set and the remaining part as the training set. In the present invention, the five-fold cross-validation method is adopted for cross-validation, that is, the data set is divided into 5 subsets, and each time one of the subsets is selected as the validation set and the rest as the training set. Through multiple cycles of the training and validation processes, the stability and generalization ability of the model are finally calculated.
[0085] In the process of cross-validation, first, the risk prediction model is trained using the training set to generate a set of preliminary risk prediction data; then, these prediction data are compared with the actual data, the prediction error is calculated, and the model parameters are adjusted according to the error. After each round of cross-validation, the system will calculate the average error of each fold to ensure the performance consistency of the model on different data sets.
[0086] Cross-validation not only verifies the accuracy of the risk prediction data, but also ensures a high consistency of the risk assessment data in the decision-making process. For example, when cross-validating the financial risk scores, the system may find that there are obvious differences in the financial behaviors of certain user groups, and these differences may affect the risk assessment results. Through cross-validation, the model can identify and correct these potential biases, so that the finally output risk assessment data can effectively reflect the real behavior patterns of users. Finally, the verified data will be used as the basis for decision intervention, providing scientific support for subsequent risk management and overdue debt handling.
[0087] Using the statistically verified risk assessment data as feedback, the risk prediction model is updated with parameters by an incremental learning method, and the hyperparameters of the risk prediction model are adjusted by an adaptive optimization algorithm to generate a final prediction result.
[0088] In the present invention, based on the statistically verified risk assessment data, the risk prediction model is updated with parameters by an incremental learning method, and the hyperparameters of the model are adjusted by an adaptive optimization algorithm to generate a final prediction result. The combination of the incremental learning method and the adaptive optimization algorithm enables the risk prediction model to adapt to changes in newly collected user behavior data in real time, improving the flexibility and prediction accuracy of the model. In a specific implementation, when the new user behavior data reaches a preset threshold, the system incorporates the new data into the existing model in real time by the incremental learning method, updating the model parameters to ensure that the model can always reflect the latest user behavior characteristics. At the same time, the system automatically adjusts the hyperparameters of the model by using Bayesian optimization technology through the adaptive optimization algorithm, further improving the generalization ability and prediction accuracy of the model in different data environments. The whole process is automatically executed without manual intervention, can continuously optimize the prediction result, and provides a highly reliable decision-making basis for the prediction of debt evasion behavior after overdue.
[0089] Preferably, the incremental learning method adopts an online gradient descent algorithm. When the newly collected user behavior data reaches a preset threshold, the new data is incorporated into the risk prediction model, and the model parameters are updated in real time, so that the model can reflect the latest user behavior characteristics.
[0090] In the implementation process of the incremental learning method, the present invention adopts an online gradient descent algorithm to ensure that when the newly collected user behavior data reaches a preset threshold, the new data can be quickly and real-time incorporated into the risk prediction model to dynamically adjust the model parameters. The online gradient descent algorithm is a commonly used incremental learning method suitable for processing streaming data. Its core principle is that when new data arrives each time, only the new data is used to fine-tune the model parameters instead of training the model from scratch each time, thus improving the computational efficiency. In the present invention, the incremental learning process is realized through the following steps:
[0091] When new user behavior data is collected, the data is first processed by data cleaning and standardization to ensure the quality and consistency of the data.
[0092] Use the current model to predict the new data and calculate the prediction error. Usually, the mean squared error (MSE) is used as the loss function:
[0093]
[0094] where y i is the true value, is the model prediction value, and N is the number of data samples.
[0095] The gradient of the loss function with respect to the model parameters is calculated using the gradient descent method, and the parameters are adjusted according to the learning rate. The formula is as follows:
[0096]
[0097] Among them, θ t is the current parameter of the model, η is the learning rate, is the gradient of the loss function with respect to the parameter.
[0098] In this way, every time new data arrives, the system makes appropriate adjustments to the existing model, ensuring the timeliness and accuracy of the model, while avoiding the high computational cost of training from scratch.
[0099] For example, when new financial transaction data or social interaction behaviors of users change, the system can quickly adapt through this method and update the model parameters in real time to ensure that the model can reflect the most accurate risk prediction.
[0100] Preferably, the adaptive optimization algorithm adopts Bayesian optimization technology to automatically adjust the hyperparameters of the risk prediction model within a preset hyperparameter search range, and generates a final prediction result based on statistically verified risk assessment data as feedback.
[0101] In the application of the adaptive optimization algorithm, the present invention uses Bayesian optimization technology to automatically adjust the hyperparameters of the risk prediction model. Bayesian optimization is an efficient global optimization method, especially suitable for high-dimensional and complex optimization problems, and can find the optimal solution with limited computational resources. The core principle of Bayesian optimization is to establish a probability model of the objective function (usually using Gaussian process regression), and based on this model, predict the hyperparameter configuration that is most likely to lead to the optimal result, so as to select the hyperparameters that are most likely to improve the model performance for testing in each iteration.
[0102] Specifically, in the present invention, the Bayesian optimization process can be divided into the following steps:
[0103] 1. Construct a surrogate model: By sampling a part of the points in the current hyperparameter space and using the performance of these points as observations, a surrogate model is constructed. Usually, Bayesian optimization uses a Gaussian process regression model (Gaussian Process Regression, GPR) to fit the objective function, representing the relationship between hyperparameters and model performance.
[0104] 2. Acquisition Function: Based on the surrogate model, Bayesian optimization uses an acquisition function (such as the Expected Improvement (EI) function) to determine the next hyperparameter point that is most likely to improve the model performance. The Expected Improvement function makes a selection by calculating the expected improvement of the current point relative to the known optimal solution:
[0105]
[0106] where f(θ) is the objective function (such as the model performance metric), θ * is the current known optimal hyperparameter configuration, and θ is the hyperparameter configuration to be tested.
[0107] 3. Evaluation and Update: Train the model with the selected hyperparameters and evaluate its performance. After obtaining new results, update the surrogate model and continue to use the model to search the hyperparameter space until the stopping conditions (such as the number of iterations, error threshold, etc.) are met.
[0108] Through this method, Bayesian optimization can quickly find the optimal or approximately optimal configuration in a large range of hyperparameter spaces, while avoiding redundant calculations in traditional grid search or random search methods, thus greatly improving the optimization efficiency. In the present invention, Bayesian optimization is mainly applied to adjust the hyperparameters (such as the learning rate, the number of hidden layer units, etc.) of the Long Short-Term Memory Network (LSTM) to improve the accuracy and robustness of the risk prediction model. Through adaptive adjustment, the Bayesian optimization technique ensures that the model can always maintain good generalization ability in different data environments, and as the data changes, it can automatically optimize the model parameters and improve the prediction accuracy.
[0109] For example, if the model obtained through training with the initial hyperparameter configuration performs poorly, Bayesian optimization can dynamically adjust parameters such as the learning rate and the number of LSTM layers, repeatedly test and optimize, and finally generate a model with higher prediction accuracy.
[0110] By combining incremental learning and Bayesian optimization, the present invention can achieve continuous optimization of the risk prediction model and use statistically verified risk assessment data as feedback to ensure that each model update can improve the prediction performance. During the model optimization process, incremental learning continuously absorbs new data to keep the model up-to-date, while Bayesian optimization focuses on improving the overall performance of the model by intelligently searching the hyperparameter space. Finally, after multiple rounds of incremental learning and hyperparameter adjustment, the model can maintain stable and efficient prediction ability in an environment of real-time data update, providing scientific and reliable support for the prediction of post-delinquency debt evasion behavior. Through the cyclic iteration of this process, the present invention effectively improves the accuracy and stability of the risk prediction model, ensuring the timeliness and accuracy of decision-making.
[0111] Such as Figure 3As shown in the figure, a system for predicting the behavior of evading debts after overdue is used to implement the method for predicting the behavior of evading debts after overdue. The system includes:
[0112] An acquisition module, which is used to acquire multi-dimensional behavior data of user account transaction records, user social interaction records, and user device usage records according to a preset data acquisition scheme; the acquisition module uses a dedicated data acquisition board card, which is composed of a microcontroller and a high-speed communication interface, integrates dedicated sensors and network interface circuits, reads multi-dimensional behavior data in real time from the user account transaction system, social platform server, and device log interface through a preset protocol, and uses a buffer memory to temporarily cache the data to ensure that the data is not lost during the transmission process. The acquisition module transmits a complete data packet to the next processing unit.
[0113] A data processing module, which is used to perform data cleaning, abnormal data elimination, missing data filling, and min-max standardization processing on the multi-dimensional behavior data, and extract features from the standardized data to generate a fused feature dataset; the data processing module is based on a hardware platform composed of a field programmable gate array and a digital signal processor to implement data cleaning, abnormal data elimination, missing data filling, and min-max standardization processing. This module embeds a dedicated algorithm core, performs high-speed filtering and statistical analysis on the input data through hardware logic, and at the same time uses the built-in memory to store standardization parameters to uniformly transform the data; on this basis, the system also uses a hardware accelerator to calculate the features extracted from the original data to achieve fast feature fusion, generate a high-quality fused feature dataset, and transmit the processing result to the time series modeling module through a high-speed interface.
[0114] A time series modeling module, which is used to construct time series data based on the fused feature dataset, and use a bidirectional long short-term memory network to train the time series data to generate time series risk prediction data for each user. At the same time, use deep reinforcement learning to adjust the parameters of the weights of the bidirectional long short-term memory network to form a risk prediction model and output risk prediction data; the time series modeling module mainly relies on a graphics processing unit or a dedicated deep learning acceleration card to implement the hardware implementation of the bidirectional long short-term memory network (Bi-LSTM). This module organizes the fused feature dataset output by the data processing module into continuous time series data, stores and schedules the data in chronological order through a dedicated memory, and uses a pre-programmed neural network accelerator to perform forward propagation and backward propagation calculations to generate time series risk prediction data for each user in real time. At the same time, the embedded deep reinforcement learning engine uses the proximal policy optimization algorithm to adjust the network parameters in real time to reduce the prediction error, form a risk prediction model optimized by hardware, and transmit the prediction data to the decision-making module through a high-speed data bus.
[0115] A decision-making module, which is used to calculate the risk scores of each risk dimension by using a multi-objective decision-making algorithm based on the risk prediction data, calculate the comprehensive risk score according to a preset weighting coefficient, and perform numerical conversion on the comprehensive risk score through a preset membership function to form decision intervention data. At the same time, it uses an auxiliary statistical module to perform cross-validation on the risk prediction data and outputs statistically verified risk assessment data. The decision-making module is integrated within a system-on-chip and is jointly composed of an embedded processor and a custom logic circuit, and is used to implement the multi-objective decision-making algorithm. This module performs weighted calculations on the risk prediction data output by the timing modeling module according to the financial, social, and device risk dimensions respectively, and generates comprehensive risk score data through the preset weighting coefficient and cumulative operation. Subsequently, a preset membership function implemented by hardware is used to convert the continuous score into a discrete decision parameter to form decision intervention data. In addition, the decision-making module has a built-in auxiliary statistical unit, which uses a cross-validation algorithm to group and compare the risk prediction data and calculate statistical indicators to ensure that the output risk assessment data has been verified at the hardware level and meets the system decision requirements.
[0116] A model update module, which is used to update the parameters of the risk prediction model by using an online gradient descent incremental learning method according to the feedback of the risk assessment data, and automatically adjust the hyperparameters of the risk prediction model within a preset search range by using Bayesian optimization technology, so as to generate a final prediction result. The model update module uses an online gradient descent accelerator and a Bayesian optimization engine integrated in the system to update the risk prediction model in real time. This module receives the risk assessment feedback data returned by the decision-making module through a high-speed interface. When the newly collected data reaches the preset quantization threshold, the updater immediately triggers the online gradient descent algorithm to fine-tune the current model parameters. At the same time, the best configuration is automatically searched within the preset hyperparameter search range by using Bayesian optimization technology to adjust the model hyperparameters and achieve adaptive optimization. The updated model parameters are fed back to the timing modeling module in real time through memory sharing and high-speed communication, and finally an optimized final prediction result is generated to ensure that the system always maintains the prediction accuracy and response speed in practical applications.
[0117] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.
[0118] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for predicting debt evasion behavior after overdue payment, characterized in that: The following steps are involved: Collect multi-dimensional behavior data of users, perform data cleaning and data standardization on the behavior data, extract features from the standardized data, and generate a fused feature data set; Based on the fused feature data set, time series data is constructed, a bidirectional long short-term memory network is used for training to generate time series risk prediction data for each user, and deep reinforcement learning is used to adjust parameters of the bidirectional long short-term memory network to form a risk prediction model and output risk prediction data; Based on the risk prediction data, a multi-objective decision-making algorithm is used to calculate a comprehensive risk score, and fuzzy logic is used to numerically adjust the comprehensive risk score to form decision intervention data, and an auxiliary statistical model is used to cross-validate the risk prediction data, and statistically validated risk assessment data is output; Based on the statistically verified risk assessment data as feedback, the risk prediction model is updated with parameters using an incremental learning method, and the hyperparameters of the risk prediction model are adjusted using an adaptive optimization algorithm to generate a final prediction result.
2. The method according to claim 1, characterized in that The multi-dimensional behavioral data is collected from user account transaction records, user social interaction records and user device usage records in a preset data format, and the data beyond the statistical range is eliminated through an outlier detection algorithm. At the same time, a missing data filling algorithm is used to compensate for missing items in the data, and the collected data is normalized using the minimum and maximum standardization method to generate standardized original data with a unified numerical range.
3. The method according to claim 2, characterized in that The features in the standardized raw data are calculated using a fixed window-based feature extraction algorithm. The calculation process quantifies the user's account changes, social interaction frequency, and device usage patterns, and accumulates the quantified results according to a preset formula to form a fused feature data set.
4. The method according to claim 1, characterized in that The time series data constructed based on the fused feature data set is trained using a bidirectional long short-term memory network. The bidirectional long short-term memory network sets a forward layer and a backward layer in the network structure, and uses a back propagation algorithm to update the network parameters to generate time series risk prediction data for each user.
5. The method according to claim 1, characterized in that: The deep reinforcement learning adopts a proximal strategy optimization algorithm to dynamically adjust the weights of the bidirectional long short-term memory network according to a preset reward function, wherein the reward function uses the risk prediction error as feedback input to form a risk prediction model and output risk prediction data.
6. The method according to claim 1, characterized in that The multi-objective decision algorithm calculates the respective risk scores for financial behavior risk, social behavior risk and device behavior risk according to preset weighting coefficients, and generates comprehensive risk score data through cumulative calculation; The comprehensive risk score data is converted through a preset membership function to discretize the continuous risk score into decision parameters to form decision intervention data.
7. The method according to claim 1, characterized in that The auxiliary statistical model adopts a five-fold cross-validation method to divide the risk prediction data into a training set and a validation set, and compares the consistency between the risk prediction data and the decision intervention data by calculating statistical indicators, thereby outputting statistically verified risk assessment data.
8. The method according to claim 1, characterized in that The incremental learning method adopts an online gradient descent algorithm. When the newly collected user behavior data reaches a preset threshold, the new data is integrated into the risk prediction model and the model parameters are updated in real time.
9. The method according to claim 1, characterized in that: The adaptive optimization algorithm adopts Bayesian optimization technology to automatically adjust the hyperparameters of the risk prediction model within a preset hyperparameter search range, and generates a final prediction result based on statistically verified risk assessment data as feedback.
10. A system for predicting overdue debt evasion behavior, used to implement the method for predicting overdue debt evasion behavior as claimed in any one of claims 1 to 9, characterized in that: The system includes: A collection module, used to collect multi-dimensional behavioral data of user account transaction records, user social interaction records, and user device usage records according to a preset data collection plan; A data processing module is used to perform data cleaning, abnormal data elimination, missing data filling and minimum and maximum standardization processing on the multi-dimensional behavior data, and extract features from the standardized data to generate a fused feature data set; A time series modeling module is used to construct time series data based on the fused feature data set, and train the time series data using a bidirectional long short-term memory network to generate time series risk prediction data for each user, and at the same time use deep reinforcement learning to adjust the weights of the bidirectional long short-term memory network to form a risk prediction model and output risk prediction data; A decision-making module is used to calculate the risk score of each risk dimension based on the risk prediction data using a multi-objective decision-making algorithm, calculate the comprehensive risk score according to the preset weighting coefficient, and perform numerical conversion on the comprehensive risk score through a preset membership function to form decision intervention data, and at the same time use an auxiliary statistical module to cross-validate the risk prediction data and output statistically validated risk assessment data; The model updating module is used to update the parameters of the risk prediction model based on the risk assessment data feedback, using the online gradient descent incremental learning method, and automatically adjust the hyperparameters of the risk prediction model within a preset search interval using the Bayesian optimization technology, so as to generate the final prediction result.
Citation Information
Patent Citations
An internet financial user loan overdue prediction method based on big data
CN109255506A
Enterprise debt evasion risk early warning system and construction method
CN111401798A
Evasion debt block chain predicting system
CN112200340A
Credit scoring method and device based on deep learning
CN113011966A
Information processing method and device, equipment and storage medium
CN113222732A