Credit card overdue risk prediction method and device, equipment, medium and product

By combining the dual deep Q network and the competitive deep Q network to form a credit card overdue risk prediction model, the accuracy and real-time problems of risk identification in traditional methods are solved, more accurate risk prediction and early warning are achieved, risk response strategies are provided, and the effectiveness of credit card overdue risk management is improved.

CN120725786APending Publication Date: 2025-09-30INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510869365.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Traditional credit card overdue risk identification methods are unable to identify risks in real time and dynamically and perform automatic adjustments, and their prediction accuracy is limited.

Method used

An overdue risk prediction model based on a combination of dual deep Q-network and competitive deep Q-network is adopted. By obtaining static attribute data and dynamic behavior data of credit card users, the loss function value of action benefit is trained to predict the overdue probability, and an early warning is initiated when the probability is greater than the threshold.

Benefits of technology

It improves the accuracy of credit card overdue risk prediction, reduces calculation complexity, provides intuitive risk warnings and response strategies, and improves the accuracy of risk identification and business security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725786A_ABST
    Figure CN120725786A_ABST
Patent Text Reader

Abstract

The invention discloses a credit card overdue risk prediction method, device and equipment, a medium and a product. The credit card overdue risk prediction method comprises the following steps: acquiring user state data of a credit card user; the user state data comprises user static attribute data and dynamic behavior data associated with credit card operation; inputting the user state data into an overdue risk prediction model, and obtaining an overdue probability output by the overdue risk prediction model; the overdue risk prediction model is obtained by training a loss function value of action earnings obtained by combining a dual-depth Q network and a competition depth Q network; and when the overdue probability is greater than a probability threshold, initiating overdue early warning for the credit card user. According to the technical scheme of the embodiment of the invention, the accuracy of the overdue risk prediction of the credit card can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of financial technology, and in particular to a credit card overdue risk prediction method, device, equipment, medium and product. Background Art

[0002] Credit card delinquency risk is one of the important risks faced by financial institutions, and risk identification is very important for controlling the risk losses of financial institutions.

[0003] Traditional overdue risk identification methods are primarily based on statistical models and rule engines. These methods often fail to identify risks dynamically and automatically in real time, and their prediction accuracy is limited. Therefore, improving the accuracy of credit card overdue risk identification has attracted widespread attention. Summary of the Invention

[0004] The present invention provides a credit card overdue risk prediction method, device, equipment, medium and product to solve the problem that the accuracy of overdue risk identification in the existing technology cannot meet the requirements.

[0005] According to one aspect of the present invention, a method for predicting credit card overdue risk is provided, comprising:

[0006] Obtaining user status data of a credit card user; the user status data includes user static attribute data and dynamic behavior data associated with credit card operations;

[0007] Inputting the user status data into an overdue risk prediction model to obtain an overdue probability output by the overdue risk prediction model; the overdue risk prediction model is trained based on a loss function value of an action benefit obtained by combining a dual deep Q network and a competitive deep Q network;

[0008] When the overdue probability is greater than a probability threshold, an overdue warning is initiated for the credit card user.

[0009] According to another aspect of the present invention, a credit card overdue risk prediction device is provided, comprising:

[0010] A status data acquisition module is used to acquire user status data of a credit card user; the user status data includes user static attribute data and dynamic behavior data associated with credit card operations;

[0011] An overdue probability acquisition module, configured to input the user status data into an overdue risk prediction model to obtain the overdue probability output by the overdue risk prediction model; the overdue risk prediction model is trained based on the loss function value of the action benefit obtained by combining a dual deep Q network and a competitive deep Q network;

[0012] The overdue warning module is used to initiate an overdue warning for the credit card user when the overdue probability is greater than a probability threshold.

[0013] According to another aspect of the present invention, an electronic device is provided, comprising:

[0014] at least one processor; and

[0015] a memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the credit card overdue risk prediction method described in any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the credit card overdue risk prediction method described in any embodiment of the present invention when executed.

[0018] According to another aspect of the present invention, a computer program product is provided, including a computer program, which, when executed by a processor, implements the credit card overdue risk prediction method of any embodiment of the present disclosure.

[0019] The technical solution of the embodiment of the present invention obtains the user status data of the credit card user, and then inputs the user status data into the overdue risk prediction model to obtain the overdue probability output by the overdue risk prediction model, wherein the overdue risk prediction model is trained based on the loss function value of the action benefit obtained by combining the dual deep Q network and the competitive deep Q network. When the overdue probability is greater than the probability threshold, an overdue warning is initiated for the credit card user. On the one hand, based on the dual deep Q, the action selection and action benefit evaluation can be decoupled to avoid over-estimation and improve the prediction accuracy. On the other hand, based on the competitive deep Q network, the cumulative benefit of the action is decomposed into state value and action advantage, which reduces the computational complexity, makes more accurate action benefit prediction, and further improves the accuracy of overdue risk prediction.

[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 This is a flowchart of a credit card overdue risk prediction method provided according to the first embodiment of the present invention;

[0023] Figure 2a This is a flowchart of a credit card overdue risk prediction method provided in accordance with the second embodiment of the present invention;

[0024] Figure 2b This is a flowchart of the overdue risk prediction model training method provided in accordance with the second embodiment of the present invention;

[0025] Figure 3 This is a schematic structural diagram of a credit card overdue risk prediction device provided according to a third embodiment of the present invention;

[0026] Figure 4 It is a structural diagram of an electronic device for implementing the credit card overdue risk prediction method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0029] Example 1

[0030] Figure 1 A flowchart of a credit card overdue risk prediction method is provided for the first embodiment of the present invention. This embodiment is applicable to the case where a risk prediction model is trained based on a deep Q network in a deep reinforcement learning algorithm. The method can be executed by a credit card overdue risk prediction device, which can be implemented in the form of hardware and / or software. The credit card overdue risk prediction device can be configured in various general computing devices, for example, the credit card overdue risk prediction device is a server. Figure 1 As shown, the method includes:

[0031] S110: Obtain user status data of the credit card user.

[0032] User status data includes static user attribute data and dynamic behavior data associated with credit card operations. For example, static user data includes age, occupation, income level, and historical credit history; dynamic behavior data includes credit card spending data, repayment data, credit score, interest rate changes, and unemployment rate within the current collection period, for example, within the past month. All of this user status data is collected through compliant collection methods with user authorization, and necessary confidentiality measures are implemented.

[0033] In this embodiment of the present invention, initial user status data representing the credit card user's status is first collected. Because this initial user status data may contain some data lacking usage information, it is necessary to preprocess the initial user status data. For example, this involves filling missing values ​​and filtering outliers. The preprocessed data is then used as the user status data.

[0034] S120: Input the user status data into the overdue risk prediction model to obtain the overdue probability output by the overdue risk prediction model.

[0035] The overdue risk prediction model is trained based on the action-reward loss function obtained by combining a dual deep Q-network and a competitive deep Q-network. Specifically, the overdue risk prediction model uses a deep neural network derived from the dual deep Q-network and the competitive deep Q-network as its base network and is trained using sample data pre-stored in the experience replay component. Each piece of sample data includes the credit card user's current user status, current action, immediate reward, and next user status.

[0036] The network structure, resulting from the combination of a dual deep Q network and a competitive deep Q network, consists of a main network and a target network. The main network is used to predict the cumulative reward of executing the current action in the current user state as the predicted reward, and also predicts the cumulative reward of executing each executable action in the next user state. The target action with the largest cumulative reward is provided to the target network. The target network is used to predict the cumulative reward of executing the target action in the next user state and determine the expected reward based on the cumulative reward and the immediate reward. The loss function value is determined based on the predicted reward and the expected reward.

[0037] After training the network structure (including the main network and target network) derived from the combination of the Dual Deep Q Network and the Competitive Deep Q Network, the trained main network is used as the overdue risk prediction model. After training, the main network is used to predict the cumulative benefits of each action performed under the input user state data, and outputs the overdue probability based on the cumulative benefits of the actions.

[0038] In an embodiment of the present invention, user status data is input into an overdue risk prediction model. The overdue risk prediction layer model predicts the cumulative benefits of each action under the user status data based on the user status data. Then, based on the cumulative benefits of each action, the overdue probability is calculated. For example, the softmax function is first used to normalize the cumulative benefits of each action, and the probability of each action being selected is determined based on the normalized result. The probability of each action being selected is weighted and summed to obtain the overdue probability. The weight of each action in the weighted sum is pre-set, and the weight of the action is positively correlated with the degree of intervention of the action on the user.

[0039] S130: When the overdue probability is greater than the probability threshold, initiate an overdue warning for the credit card user.

[0040] In this embodiment of the present invention, if the probability of overdue payment exceeds a probability threshold, an overdue warning is issued for the credit card user. The overdue warning includes the overdue probability. Based on the overdue probability in the overdue warning, the user can adopt a corresponding response strategy for the current credit card user to reduce losses.

[0041] The technical solution of the embodiment of the present invention obtains the user status data of the credit card user, and then inputs the user status data into the overdue risk prediction model to obtain the overdue probability output by the overdue risk prediction model, wherein the overdue risk prediction model is trained based on the loss function value of the action benefit obtained by combining the dual deep Q network and the competitive deep Q network. When the overdue probability is greater than the probability threshold, an overdue warning is initiated for the credit card user. On the one hand, based on the dual deep Q, the action selection and action benefit evaluation can be decoupled to avoid over-estimation and improve the prediction accuracy. On the other hand, based on the competitive deep Q network, the cumulative benefit of the action is decomposed into state value and action advantage, which reduces the computational complexity, makes more accurate action benefit prediction, and further improves the accuracy of overdue risk prediction.

[0042] Example 2

[0043] Figure 2a This is a flowchart of a credit card overdue risk prediction method provided by the second embodiment of the present invention. This embodiment further refines the above embodiment and provides specific steps for inputting user status data into the overdue risk prediction model, outputting the overdue probability from the overdue risk prediction model, and initiating an overdue warning for the credit card user when the overdue probability is greater than the probability threshold. Figure 2a As shown, the method includes:

[0044] S210. Obtain user status data of the credit card user; the user status data includes user static attribute data and dynamic behavior data associated with credit card operations.

[0045] S220: Input the user status data into the overdue risk prediction model, and the overdue risk prediction model predicts the cumulative benefit of performing each executable action under the user status data.

[0046] S230: Determine the overdue probability based on the accumulated benefits of each executable action using the overdue risk prediction model.

[0047] In an embodiment of the present invention, after the overdue risk prediction model predicts the cumulative action benefit of each executable action, the overdue probability is calculated based on the cumulative action benefit of each executable action. Specifically, the cumulative action benefit value of each executable action can be normalized first, and then the normalized results can be weighted and summed, wherein the weight of each executable action is positively correlated with the degree of intervention of the executable action on the user. By converting the cumulative action benefit output by the overdue risk prediction model into an expected probability, the deep Q network concept is applied to the risk probability prediction scenario, providing users with intuitive and reliable risk prompts.

[0048] In addition, all executable actions can be pre-classified into high-risk actions (i.e., actions that require a high degree of user intervention, such as reducing a credit limit or freezing an account) and low-risk actions (i.e., actions that require a low degree of user intervention, such as taking no action or sending a reminder). The cumulative benefits of high-risk actions are extracted and normalized. The normalized results are summed to form the overdue probability.

[0049] Optionally, the overdue risk prediction model determines the overdue probability based on the cumulative benefits of each executable action, including:

[0050] The overdue risk prediction model determines the probability of each executable action being selected based on the cumulative benefits of each executable action;

[0051] The selection probability of each executable action is weighted and summed to obtain the overdue probability; the weight of each executable action in the weighted sum is positively correlated with the degree of intervention of the executable action on the user.

[0052] In this optional embodiment, a specific method is provided for determining the overdue probability based on the cumulative benefit of each executable action by the overdue risk prediction model: first, the overdue risk prediction model determines the probability of selection of each executable action based on the cumulative benefit of each executable action. Specifically, the softmax function can be used to normalize the cumulative benefit of each executable action. Based on the normalization result, the probability of selection of each executable action is determined. The probability of selection of each executable action is weighted and summed to obtain the overdue probability. Among them, the weight of each executable action in the weighted sum is positively correlated with the degree of intervention of the executable action on the user. In the above manner, the cumulative benefit value of the action output by the deep Q network is converted into an expected probability, providing the user with an intuitive risk reference value.

[0053] Exemplarily, the executable actions include sending a reminder, adjusting the credit limit, and freezing an account, wherein the weight of sending a reminder is smaller than the weight of adjusting the credit limit, and the weight of adjusting the credit limit is smaller than the weight of freezing the account.

[0054] S240: When the overdue probability is greater than the probability threshold, based on the cumulative benefit of each executable action, determine at least one of the executable actions as a risk response strategy.

[0055] In an embodiment of the present invention, if the overdue probability output by the overdue risk prediction model is greater than a pre-set probability threshold, it indicates that a warning should be issued to the user. In this case, based on the cumulative action benefit of each executable action obtained during the overdue probability calculation process, at least one executable action can be determined as a risk response strategy. Specifically, multiple executable actions can be sorted in descending order of the cumulative action benefit of each executable action, and a predetermined number, for example, the first two executable actions, are ultimately selected as risk response strategies.

[0056] S250. Based on the risk response strategy and overdue probability, initiate overdue warning for credit card users.

[0057] In this embodiment of the present invention, risk response strategies and overdue probability are combined as early warning information to initiate overdue warnings for credit card users. This not only provides an intuitive overdue probability, but also provides users with at least one risk response strategy. Users can use this risk response strategy as a reference to address current overdue risks and reduce losses caused by credit card overdue payments.

[0058] Optionally, the overdue risk prediction model is trained in the following way:

[0059] Randomly extract a set number of sample data from the experience playback component; each sample data includes the credit card user's current user status, current action, immediate revenue, and next user status;

[0060] The main network composed of the dual deep Q network and the competitive deep Q network predicts the cumulative reward of performing the current action under the current user state as the predicted reward;

[0061] The main network predicts the cumulative benefit of executing each executable action in the next user state, and provides the executable action with the largest cumulative benefit as the target action to the target network composed of the dual deep Q network and the competitive deep Q network;

[0062] The target network predicts the cumulative benefit of executing the target action in the next user state, and determines the expected benefit based on the cumulative benefit and the immediate benefit;

[0063] The loss function value of the action benefit is determined based on the predicted benefit and the expected benefit, and the parameters of the main network and the target network are adjusted based on the loss function value until the training end condition is met to obtain a risk prediction model.

[0064] In this optional embodiment, a training method for an overdue risk prediction model is provided, such as Figure 2bAs shown in the figure, this embodiment of the present invention uses a dual deep Q-network and a competitive deep Q-network from a deep reinforcement learning algorithm as a foundational model to train a credit card overdue risk prediction model. The resulting model structure mainly includes a main network (i.e., a Q-network), a target network, and an experience replay component.

[0065] Among them, the main network and the target network are neural networks with the same structure, including a feature extraction network (for example, a convolutional neural network or a multi-layer fully connected network, etc.) and a value stream branch network and an advantage stream scoring network (for example, a fully connected network) respectively connected to the feature extraction network.

[0066] The experience playback component stores multiple sample data determined based on historical user data, each sample data includes (s t ,a t ,r t ,s t+1 ). Among them, s t is the current user status, for example, the user status data of the tth month; a t is the current action, for example, the intervention measures taken by the financial institution in response to the current user status; r t is the immediate income (calculated based on whether it is overdue in month t+1); t+1 The next user status, for example, the user status data of the t+1th month. In addition, the sample data may also include a flag indicating whether to terminate the training, which is used to end the training process.

[0067] The combination of a dual deep Q network and a competitive deep Q network results in a network structure that incorporates state, action, and reward. Specifically, in the credit card overdue risk prediction scenario, the state includes static user attribute data and dynamic behavior data associated with credit card operations. For example, static user attribute data includes age, occupation, income level, and historical credit history, while dynamic behavior data includes monthly spending and repayment data. The current action is the intervention measure taken by the bank based on the current user state, such as adjusting the credit limit, sending payment reminders, and freezing the account. The immediate reward is a feedback signal designed based on whether the user will subsequently default and the severity of the default. For example, a reward of +1 is given for repaying on time, and a reward of -5 is given for overdue payments.

[0068] In this embodiment of the present invention, a weight initialization operation is performed on the network structure obtained by combining the dual deep Q network and the competitive deep Q network, and then a set number of sample data are extracted from the experience replay component for model training. The set number of sample data is randomly extracted from the experience replay component each time to break the temporal correlation between the data.

[0069] Optionally, before randomly extracting a set number of sample data from the experience replay component, the following steps are also included:

[0070] Obtain historical source data of users using credit cards;

[0071] Based on historical source data, a sliding window is constructed in the time dimension to obtain the current user status;

[0072] Based on each sliding window, the sample data consists of the current user state, current action, immediate benefit, and next user state.

[0073] In this optional embodiment, historical source data of the user's credit card usage is obtained, and then based on the historical source data, a sliding window is constructed in the time dimension to obtain the current user status. Based on each sliding window, sample data is composed of the current user status, current action, immediate benefit and next user status.

[0074] Specifically, we obtain time series features from multiple dimensions and collect raw data related to credit card users, including user static attribute data, dynamic behavior data, and executable actions, as follows:

[0075] (1) User static attribute data may include:

[0076] Basic information of the user: age, occupation, income level, education level, etc.;

[0077] Historical credit data: historical overdue times, bandwidth balance, number of credit cards, etc.

[0078] (2) Dynamic behavior data may include:

[0079] Consumption data: monthly spending amount, frequency, and type of spending, such as large purchases or overseas purchases;

[0080] Repayment data: monthly repayment amount, repayment time, and minimum repayment ratio;

[0081] Credit score: Credit score updated monthly;

[0082] External economic environment: macroeconomic indicators such as changes in interest rates and unemployment rates.

[0083] (3) Executable actions may include:

[0084] Historical action data: historical actions taken by financial institutions, such as whether to adjust credit limits, send reminders, or freeze accounts;

[0085] Overdue label: whether the user has overdue payment and the number of overdue days, etc.

[0086] The source data may contain some data that does not have usage information, so it is necessary to preprocess the source data, for example, to process missing values ​​and outliers in the original data.

[0087] Exemplarily, missing value filling can be done by using the median / mode to fill in static attribute data of users, and using time series interpolation (such as linear interpolation) for dynamic behavior data. For example: Occupation / income missing: If the missing rate is lower than the set threshold (such as <5%), use the mode to fill in (such as the most common occupation category); if the missing rate is higher than the set threshold, mark it as an "Unknown" category to avoid introducing bias. Income level missing: Use the median income of the same occupation group to fill in (for example, the median income of the engineer occupation group); if no association is possible, use the median income of all users.

[0088] Outlier filtering can be used to filter out data with obviously unreasonable consumption amounts, repayment amounts, or transaction times. For example, by defining the rule "Consumption amount > credit limit × the percentage of excess credit allowed by the financial institution," we can filter out abnormal consumption amounts; by defining the rule "Repayment amount < 0 or Repayment amount > current bill amount × 2," we can filter out abnormal repayment amounts; by detecting transaction dates later than the current date or earlier than the account opening date, we can detect abnormal transaction times. Transactions conducted at abnormal times will be deemed invalid.

[0089] Furthermore, the pre-processed historical source data is used to construct a sliding window in the time dimension to obtain the current user status. For example, the data of the first three months to the current moment is used as the current user status s t . In addition, it is also necessary to perform feature standardization processing and time series feature extraction on the historical source data. For example, the continuous features (for example, the amount of consumption) are standardized by Z-score, and the categorical features (for example, occupation) are one-hot encoded. Time series feature extraction can be to calculate the month-on-month growth rate of the monthly consumption amount, the standard deviation of the repayment time in multiple repayment cycles, the frequency of large-amount consumption, and the proportion of nighttime consumption, etc., which are time series features that can characterize the user's consumer financial behavior habits.

[0090] Finally, based on each sliding window, sample data is composed of the current user state (preprocessed and standardized state data), current action, immediate benefit, and next user state, and is stored in the experience replay component.

[0091] Furthermore, the main network composed of the dual deep Q network and the competitive deep Q network predicts the cumulative benefit Q of performing the current action in the current user state. main(s, a) (i.e., the predicted Q value) is used as the predicted profit. The main network includes a feature extraction network (e.g., a convolutional neural network or a multi-layer fully connected network), a value stream branch network, and an advantage stream scoring network (e.g., a fully connected network) connected to the feature extraction network.

[0092] Furthermore, the main network predicts the cumulative benefit of executing each executable action in the next user state, and provides the executable action with the largest cumulative benefit as the target action to the target network composed of the dual deep Q network and the competitive deep Q network.

[0093] Optionally, the main network predicts the cumulative benefit of performing each action that can be performed in the next user state, including:

[0094] The feature extraction network of the main network extracts the feature data of the next user state; the value stream branch network connected to the feature extraction network in the main network outputs the state value of the next user state based on the feature data; the advantage stream branch network connected to the feature extraction network in the main network outputs the action advantage value of each executable action in the next user state based on the feature data; based on the state value and action advantage value, the cumulative action benefit of each executable action is predicted.

[0095] In this optional embodiment, the value stream branch network connected to the feature extraction network in the main network outputs the state value V(s) of the next user state based on the feature data. t+1 At the same time, the advantage flow branch network in the main network, connected to the feature extraction network, outputs the action advantage value A(a) for each action that can be performed in the next user state based on the feature data. Based on the state value and action advantage value, the cumulative action benefit of each action is predicted.

[0096] In a specific example, the current state is s, and after executing action a, it transitions to the next state s'. By inputting the next state s' into the main network, the cumulative action benefit Q of each action output by the main network can be obtained. main (s',a n ), for example, Q main (s',a1),Q main (s',a2),Q main (s',a3),….

[0097] Among them, Q main (s',a i )The specific calculation formula is as follows:

[0098]

[0099] Among them, V(s') is the state value of the next user state, A(s',a i ) is an executable action a i The action advantage value, |A| is the number of executable actions, ∑ a A(s',a) is the sum of the action advantage values ​​of all executable actions in the next user state s'.

[0100] Furthermore, the main network takes the executable action with the largest cumulative benefit as the target action a*=argmax a Q main (s',a i ) is provided to the target network, which is a combination of a dual-deep Q network and a competitive deep Q network. For example, in state s', the user is five days overdue and their spending has increased by 50% in the past week. The main network predicts the cumulative benefits of different possible actions taken by the financial institution. For example, sending a reminder: Q = 15, reducing the credit limit by 20%: Q = 25, and freezing the account: Q = 22. In summary, the action with the highest cumulative benefit is selected as the target action, i.e., a* is a 20% credit limit reduction.

[0101] Furthermore, the target network predicts the cumulative benefit Q of executing the target action in the next user state target (s',a*), and based on the cumulative benefit Q of the action target (s', a*) and the immediate return r, determine the expected return. The specific calculation formula is as follows:

[0102] Q target =r+γQ target (s',a*)

[0103] Where r is the immediate return, γ is the discount factor, and Q target (s',a*) is the action cumulative benefit of the target action output by the target network.

[0104] For example, the target network evaluates the cumulative benefit of the target action a* (derated by 20%) to 18, which is lower than the 25 output by the main network. The target network reduces the deviation caused by the overestimation of the value by the main network by delaying the update of parameters.

[0105] Finally, the loss function value of the action benefit is determined based on the predicted benefit and the expected benefit. Specifically, the mean square error of the predicted benefit and the expected benefit is used as the loss function, and the calculation formula is as follows:

[0106]

[0107] Among them, N is the number of samples extracted in this round of training, Q main (s,a) is the predicted return output by the main network, Q targetis the expected return of the target network output.

[0108] Based on the loss function value, the gradient descent method is used to adjust the parameters of the main network and the target network until the training end conditions are met to obtain the risk prediction model.

[0109] Optionally, a loss function value of the action benefit is determined based on the predicted benefit and the expected benefit, and parameters of the main network and the target network are adjusted based on the loss function value until the training end condition is met, thereby obtaining a risk prediction model, including:

[0110] Determining a loss function value of the action benefit based on the predicted benefit and the expected benefit, and adjusting parameters of the main network based on the loss function value;

[0111] At set time steps, the parameters of the main network are synchronized to the target network until the training end conditions are met to obtain a risk prediction model.

[0112] In this optional embodiment, a loss function value for determining the action benefit based on the predicted benefit and the expected benefit is provided, and based on the loss function value, the parameters of the main network and the target network are adjusted until the training end condition is met, thereby obtaining a specific method for the risk prediction model: the loss function value for determining the action benefit based on the predicted benefit and the expected benefit, and based on the loss function value, the parameters of the main network are adjusted by the gradient descent method. At every set time step, the parameters of the main network are synchronized to the target network until the training end condition is met, thereby obtaining a risk prediction model. By delaying the update of the target network, the target network remains stable within the set time step, avoiding over-estimation, alleviating training shocks, and reducing false triggering of high-risk actions.

[0113] The end-of-training adjustment can be calculated during the training process, when the loss function converges, when sample data containing the end-of-training flag is read, or when the number of training rounds reaches a certain number.

[0114] The technical solution of the embodiment of the present invention obtains user status data of the credit card user, inputs the user status data into the overdue risk prediction model, and the overdue risk prediction model predicts the cumulative benefit of executing each executable action under the user status data, and the overdue risk prediction model determines the overdue probability based on the cumulative benefit of each executable action. When the overdue probability is greater than the probability threshold, based on the cumulative benefit of each executable action, at least one item is determined as a risk response strategy among the executable actions. Finally, based on the risk response strategy and the overdue probability, an overdue warning is initiated for the credit card user. The cumulative benefit of the action obtained by the overdue risk prediction model in the process of calculating the overdue probability can be used to add a risk response strategy to the overdue warning, thereby improving the accuracy of risk identification and providing users with a risk warning strategy to enhance business security.

[0115] Example 3

[0116] Figure 3 This is a schematic diagram of the structure of a credit card overdue risk prediction device provided by the third embodiment of the present invention. Figure 3 As shown, the device includes:

[0117] Status data acquisition module 310, for acquiring user status data of a credit card user; the user status data includes user static attribute data and dynamic behavior data associated with credit card operations;

[0118] An overdue probability acquisition module 320 is configured to input the user status data into an overdue risk prediction model to obtain the overdue probability output by the overdue risk prediction model; the overdue risk prediction model is trained based on the loss function value of the action benefit obtained by combining a dual deep Q network and a competitive deep Q network;

[0119] The overdue warning module 330 is configured to initiate an overdue warning for the credit card user when the overdue probability is greater than a probability threshold.

[0120] The technical solution of the embodiment of the present invention obtains the user status data of the credit card user, and then inputs the user status data into the overdue risk prediction model to obtain the overdue probability output by the overdue risk prediction model, wherein the overdue risk prediction model is trained based on the loss function value of the action benefit obtained by combining the dual deep Q network and the competitive deep Q network. When the overdue probability is greater than the probability threshold, an overdue warning is initiated for the credit card user. On the one hand, based on the dual deep Q, the action selection and action benefit evaluation can be decoupled to avoid over-estimation and improve the prediction accuracy. On the other hand, based on the competitive deep Q network, the cumulative benefit of the action is decomposed into state value and action advantage, which reduces the computational complexity, makes more accurate action benefit prediction, and further improves the accuracy of overdue risk prediction.

[0121] Optionally, the overdue probability acquisition module 320 includes:

[0122] a benefit calculation unit, configured to input the user status data into an overdue risk prediction model, and have the overdue risk prediction model predict the cumulative benefit of executing each executable action under the user status data;

[0123] The overdue probability acquisition unit is used to determine the overdue probability based on the cumulative action benefit of each executable action using the overdue risk prediction model.

[0124] Optionally, the overdue warning module 330 is specifically configured to:

[0125] When the overdue probability is greater than a probability threshold, determining at least one of the executable actions as a risk response strategy based on the cumulative benefits of each executable action;

[0126] Based on the risk response strategy and the overdue probability, an overdue warning is initiated for the credit card user.

[0127] Optional, overdue probability acquisition unit, specifically used for:

[0128] Determining the selection probability of each executable action based on the cumulative benefits of each executable action by the overdue risk prediction model;

[0129] The selection probability of each executable action is weighted and summed to obtain the overdue probability; the weight of each executable action in the weighted sum is positively correlated with the degree of intervention of the executable action on the user.

[0130] Optionally, the credit card overdue risk prediction device further includes a model training module for training the overdue risk prediction model, including:

[0131] The sample extraction unit is used to randomly extract a set number of sample data from the experience playback component; each sample data includes the credit card user's current user status, current action, immediate benefits, and next user status;

[0132] a predicted benefit determining unit, configured to predict, by a main network combining a dual deep Q network and a competitive deep Q network, an action cumulative benefit of performing the current action in the current user state as the predicted benefit;

[0133] A cumulative benefit calculation unit is configured to predict, by the main network, the cumulative benefit of executing each executable action in the next user state, and provide the executable action with the largest cumulative benefit as the target action to the target network composed of the dual deep Q network and the competitive deep Q network;

[0134] an expected benefit determining unit, configured to predict, by the target network, the cumulative benefit of performing the target action in the next user state, and determine the expected benefit based on the cumulative benefit and the immediate benefit;

[0135] A model determination unit is used to determine the loss function value of the action benefit based on the predicted benefit and the expected benefit, and based on the loss function value, adjust the parameters of the main network and the target network until the training end condition is met to obtain a risk prediction model.

[0136] Optional, cumulative income calculation unit, specifically used to:

[0137] Extracting feature data in the next user state by a feature extraction network of the main network;

[0138] The value stream branch network in the main network connected to the feature extraction network outputs the state value of the next user state based on the feature data;

[0139] The advantage flow branch network in the main network connected to the feature extraction network outputs an action advantage value for each executable action in the next user state based on the feature data;

[0140] Based on the state value and action advantage value, the cumulative action benefit of each executable action is predicted.

[0141] Optionally, the model determination unit is specifically configured to:

[0142] Determining a loss function value of the action benefit based on the predicted benefit and the expected benefit, and adjusting parameters of the main network based on the loss function value;

[0143] At set time steps, the parameters of the main network are synchronized to the target network until the training end conditions are met to obtain a risk prediction model.

[0144] The credit card overdue risk prediction device provided by the embodiment of the present invention can execute the risk prediction model training method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0145] In the technical solution of the present invention, the information collected is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0146] Example 4

[0147] According to an embodiment of the present invention, the present invention further provides an electronic device, a readable storage medium and a computer program product.

[0148] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, appliances, blade appliances, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0149] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0150] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0151] Processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any other suitable processors, controllers, microcontrollers, etc. Processor 11 executes the various methods and processes described above, such as the credit card overdue risk prediction method.

[0152] In some embodiments, the credit card overdue risk prediction method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the credit card overdue risk prediction method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the credit card overdue risk prediction method in any other appropriate manner (for example, by means of firmware).

[0153] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0154] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or application.

[0155] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0156] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0157] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data applier), or a computing system that includes middleware components (e.g., an application applier), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0158] A computing system may include clients and applications. These clients and applications are typically remote from each other and typically interact via a communication network. This client-application relationship is established by computer programs running on the respective computers, establishing a client-application relationship. The application can be a cloud application, also known as a cloud computing application or cloud host. This is a host product within a cloud computing application ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS applications.

[0159] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0160] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A credit card overdue risk prediction method, characterized in that: include: Obtaining user status data of a credit card user; the user status data includes user static attribute data and dynamic behavior data associated with credit card operations; Inputting the user status data into an overdue risk prediction model to obtain an overdue probability output by the overdue risk prediction model; the overdue risk prediction model is trained based on a loss function value of an action benefit obtained by combining a dual deep Q network and a competitive deep Q network; When the overdue probability is greater than a probability threshold, an overdue warning is initiated for the credit card user.

2. The method according to claim 1, characterized in that Inputting the user status data into an overdue risk prediction model to obtain the overdue probability output by the overdue risk prediction model includes: Inputting the user status data into an overdue risk prediction model, and having the overdue risk prediction model predict the cumulative benefit of executing each executable action under the user status data; The overdue risk prediction model determines the overdue probability based on the cumulative benefit of each executable action.

3. The method according to claim 2, characterized in that When the overdue probability is greater than the probability threshold, initiating an overdue warning for the credit card user, including: When the overdue probability is greater than a probability threshold, determining at least one of the executable actions as a risk response strategy based on the cumulative benefits of each executable action; Based on the risk response strategy and the overdue probability, an overdue warning is initiated for the credit card user.

4. The method according to claim 2, characterized in that The overdue risk prediction model determines the overdue probability based on the cumulative benefit of each executable action, including: Determining the selection probability of each executable action based on the cumulative benefits of each executable action by the overdue risk prediction model; The selection probability of each executable action is weighted and summed to obtain the overdue probability; the weight of each executable action in the weighted sum is positively correlated with the degree of intervention of the executable action on the user.

5. The method according to claim 1, wherein The overdue risk prediction model is trained in the following way: Randomly extract a set number of sample data from the experience playback component; each sample data includes the credit card user's current user status, current action, immediate revenue, and next user status; The main network composed of the dual deep Q network and the competitive deep Q network predicts the cumulative reward of performing the current action under the current user state as the predicted reward; The main network predicts the cumulative benefit of executing each executable action in the next user state, and provides the executable action with the largest cumulative benefit as the target action to the target network composed of the dual deep Q network and the competitive deep Q network; The target network predicts the cumulative benefit of executing the target action in the next user state, and determines the expected benefit based on the cumulative benefit and the immediate benefit; The loss function value of the action benefit is determined based on the predicted benefit and the expected benefit, and the parameters of the main network and the target network are adjusted based on the loss function value until the training end condition is met to obtain a risk prediction model.

6. The method according to claim 5, characterized in that The main network predicts the cumulative benefit of executing each executable action in the next user state, including: Extracting feature data in the next user state by a feature extraction network of the main network; The value stream branch network in the main network connected to the feature extraction network outputs the state value of the next user state based on the feature data; The advantage flow branch network in the main network connected to the feature extraction network outputs an action advantage value for each executable action in the next user state based on the feature data; Based on the state value and action advantage value, the cumulative action benefit of each executable action is predicted.

7. The method according to claim 5, characterized in that Determining a loss function value of the action benefit based on the predicted benefit and the expected benefit, and adjusting parameters of the main network and the target network based on the loss function value until a training end condition is met, thereby obtaining a risk prediction model, including: Determining a loss function value of the action benefit based on the predicted benefit and the expected benefit, and adjusting parameters of the main network based on the loss function value; At set time steps, the parameters of the main network are synchronized to the target network until the training end conditions are met to obtain a risk prediction model.

8. A credit card overdue risk prediction device, characterized in that: include: A status data acquisition module is used to acquire user status data of a credit card user; the user status data includes user static attribute data and dynamic behavior data associated with credit card operations; An overdue probability acquisition module, configured to input the user status data into an overdue risk prediction model to obtain the overdue probability output by the overdue risk prediction model; the overdue risk prediction model is trained based on the loss function value of the action benefit obtained by combining a dual deep Q network and a competitive deep Q network; The overdue warning module is used to initiate an overdue warning for the credit card user when the overdue probability is greater than a probability threshold.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the credit card overdue risk prediction method described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the credit card overdue risk prediction method according to any one of claims 1 to 7 when executed.

11. A computer program product, characterized in that The computer program product includes a computer program, which, when executed by a processor, implements the credit card overdue risk prediction method according to any one of claims 1 to 7.