Method and device for adjusting emergency strategy of lost customer based on reinforcement learning, and medium

By dynamically optimizing the emergency strategy for lost customers based on reinforcement learning methods, the problem of insufficient adaptability of traditional strategies is solved, and efficient fund recovery and risk management are achieved.

CN120655115APending Publication Date: 2025-09-16GUANGDONG MECHANICAL & ELECTRICAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510433647.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

When financial institutions face the problem of lost contact with loan customers, their traditional response strategies cannot effectively adapt to changes in customer status, resulting in inefficient collection and difficulty in risk management. In particular, the loan default rate of lost customers is high and the collection costs increase.

Method used

A reinforcement learning-based method is used to select lost customer data through the training set, generate actions and execute them, use the reward function to adjust the Q-value table, dynamically optimize the emergency strategy, and select the action with the largest Q-value as the target emergency strategy.

Benefits of technology

It improves the collection efficiency and fund recovery success rate in complex market environments, reduces operating costs, and improves risk management and customer service levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655115A_ABST
    Figure CN120655115A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning-based missed customer emergency strategy adjustment method and device and a medium, and relates to the technical field of financial intelligent prediction, and the method comprises the steps: selecting a group of missed customer data in a preset training set, and extracting a first state of the missed customer data; generating a first action in a preset action space, and executing the first action to obtain a second state; calculating a reward value for executing the first action in the first state according to a preset reward function; adjusting a Q value corresponding to a mapping pair of the first state and the first action in a preset Q value table according to the reward value and the second state; if a preset termination condition is met and the Q value is converged, stopping adjusting the Q value table; and inputting the to-be-processed lost customer data into the Q value table to obtain Q values corresponding to the first actions, and determining the first action with the maximum Q value as a target emergency strategy. According to the method and the device, the technical problem of risk management and control caused by loss-of-communication scene heterogeneity and response lag of a financial institution is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of financial intelligent prediction technology, and in particular to a method, device, and medium for adjusting emergency strategies for lost customers based on reinforcement learning. Background Art

[0002] In the digital transformation of financial lending, the frequent occurrence of lost contact with loan customers has become a key bottleneck hindering the risk management effectiveness of lending institutions such as microfinance companies, banks, and financial platforms. Statistics show that the loan default rate for lost contact customers is significantly higher than that for regular customers, and the collection costs also increase significantly, severely eroding the profit margins and liquidity of financial institutions. The root cause is that lost contact behavior is often accompanied by a deterioration in the customer's repayment ability or a subtle shift in their willingness to repay. The dynamic nature and heterogeneity of these behaviors make traditional response strategies difficult to effectively adapt to the complex and ever-changing real-world scenarios. Currently, lending institutions' traditional strategies for dealing with lost contact with loan customers rely primarily on static rules and manual experience. A common practice is to mechanically select collection, litigation, or asset preservation measures based on characteristics such as the number of days overdue and the duration of loss of contact. For example, some financial institutions stipulate that litigation proceedings will be initiated if the number of days overdue exceeds a threshold, but this approach has numerous drawbacks in practice. Traditional strategies often fail to effectively adapt to different customer statuses, resulting in low collection efficiency and difficulties in achieving effective risk management. Summary of the Invention

[0003] The main purpose of this application is to provide a method for adjusting emergency strategies for lost customers based on reinforcement learning, aiming to solve the technical problems of risk management caused by the heterogeneity of loss of contact scenarios and delayed response of financial institutions.

[0004] To achieve the above objectives, this application proposes a method for adjusting the emergency strategy for lost customers based on reinforcement learning, including:

[0005] Selecting a set of lost customer data from a preset training set and extracting the first state of the lost customer data;

[0006] Generate a first action in a preset action space, and execute the first action to obtain a second state;

[0007] Calculating a reward value for performing a first action in a first state according to a preset reward function;

[0008] Adjust the Q value corresponding to the mapping pair of the first state and the first action in the preset Q value table according to the reward value and the second state;

[0009] If the preset termination condition is met and the Q value converges, stop adjusting the Q value table;

[0010] The lost customer data to be processed is input into the Q value table to obtain the Q value corresponding to each first action, and the first action with the largest Q value is determined as the target emergency strategy.

[0011] In one embodiment, the step of generating a first action in a preset action space includes:

[0012] Get the current exploration rate, where the current exploration rate decays as the number of training times increases;

[0013] Generate the current random number corresponding to the current training round;

[0014] If the current random number is less than the current exploration rate, any action in the action space is determined as the first action;

[0015] If the current random number is greater than or equal to the current exploration rate, the action with the highest Q value corresponding to the first state in the Q value table is determined as the first action.

[0016] In one embodiment, the step of extracting the first state of the lost customer data includes:

[0017] Generate corresponding loss pattern identification variables based on customer characteristics of lost customer data;

[0018] Calculate the ratio of the number of successful customer contacts within a preset time period to the total number of contacted customers in the lost customer data to obtain the customer response rate corresponding to the lost customer data;

[0019] Calculate the asset loss rate corresponding to the lost customer data, where the asset loss rate is used to monitor customer asset transfer behavior;

[0020] The loss pattern identification variable, the customer response rate, and the asset loss rate are determined as the first state corresponding to the lost customer data.

[0021] In one embodiment, the action space includes collection actions and litigation actions, the collection actions include a collection speech severity coefficient and a collection frequency, and the litigation actions include litigation intensity. The step of calculating a reward value for executing a first action in a first state based on a preset reward function includes:

[0022] The success reward is calculated based on the amount recovered after executing the first action in the first state and the response time of executing the first action;

[0023] Calculate cost penalties based on the frequency of collection and litigation intensity of the first action;

[0024] Maintenance rewards are calculated based on the customer response rate and the severity coefficient of the first action's collection speech;

[0025] According to the loss mode identification variable, the weights corresponding to the success reward, cost penalty, and maintenance reward are determined respectively. The success reward, cost penalty, and maintenance reward are weightedly calculated according to the respective weights to obtain the reward value for executing the first action in the first state.

[0026] In one embodiment, after the step of determining the weights corresponding to the success reward, the cost penalty, and the maintenance reward according to the loss mode identification variable, the following steps are performed:

[0027] Update the weight of the success reward based on the asset loss rate, where the asset loss rate is positively correlated with the weight of the success reward;

[0028] updating the weight of the cost penalty according to the customer response rate, wherein the customer response rate is negatively correlated with the weight of the cost penalty;

[0029] Update the weight of the maintenance reward based on the loss time in the lost customer data, where the loss time is positively correlated with the weight of the maintenance reward;

[0030] The success reward, cost penalty, and maintenance reward are weightedly calculated according to the updated weights to obtain the reward value of the first action in the first state.

[0031] In one embodiment, the step of adjusting the Q value corresponding to the mapping pair of the first state and the first action in the preset Q value table according to the reward value and the second state includes:

[0032] In the Q value table, find all new mapping pairs corresponding to the second state;

[0033] Get the new Q values ​​of all new mapping pairs and determine the maximum new Q value from all the new Q values;

[0034] Calculate the Q value increment based on the reward value and the maximum new Q value;

[0035] Add the Q value increment to the Q value corresponding to the mapping pair to obtain the updated Q value;

[0036] Use the updated Q value to update the corresponding Q value of the mapping pair.

[0037] In one embodiment, after the step of adjusting the Q value corresponding to the mapping pair of the first state and the first action in the preset Q value table according to the reward value and the second state, the step includes:

[0038] If the preset termination condition is not met, the process returns to the step of generating the first action in the preset action space.

[0039] In one embodiment, after the step of adjusting the Q value corresponding to the mapping pair of the first state and the first action in the preset Q value table according to the reward value and the second state, the step includes:

[0040] If the Q value does not converge, the process returns to the step of selecting a set of lost customer data from the preset training set.

[0041] In addition, to achieve the above-mentioned purpose, the present application also proposes a device for adjusting an emergency strategy for a lost customer based on reinforcement learning. The device for adjusting an emergency strategy for a lost customer based on reinforcement learning includes:

[0042] A state acquisition module is used to select a set of lost customer data from a preset training set and extract a first state of the lost customer data;

[0043] An action execution module, configured to generate a first action in a preset action space, and execute the first action to obtain a second state;

[0044] A reward calculation module, configured to calculate a reward value for performing a first action in a first state according to a preset reward function;

[0045] A Q-value adjustment module is used to adjust the Q-value corresponding to the mapping pair of the first state and the first action in the preset Q-value table according to the reward value and the second state;

[0046] A convergence judgment module is used to stop adjusting the Q value table if a preset termination condition is met and the Q value converges;

[0047] The emergency strategy module is used to input the lost customer data to be processed into the Q value table, obtain the Q value corresponding to each first action, and determine the first action with the largest Q value as the target emergency strategy.

[0048] In addition, to achieve the above-mentioned purpose, the present application also proposes a device for adjusting the emergency strategy for lost customers based on reinforcement learning. The device includes: a memory, a processor, and a computer program stored in the memory and runnable on the processor. The computer program is configured to implement the steps of the method for adjusting the emergency strategy for lost customers based on reinforcement learning as described above.

[0049] In addition, to achieve the above-mentioned purpose, the present application also proposes a medium, which is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, the steps of the lost customer emergency strategy adjustment method based on reinforcement learning as described above are implemented.

[0050] One or more technical solutions proposed in this application have at least the following technical effects:

[0051] In the training phase of the present application, a group of lost customers are first selected from a preset training set, and the first state of the lost customers is extracted; a first action is generated in a preset action space according to a preset strategy, and the first action is executed; the reward value for executing the first action in the first state is calculated based on a preset reward function; the Q value of the mapping pair between the first state and the first action in the preset Q value table is adjusted according to the reward value; if the preset termination condition is met and the Q value converges, the adjustment of the Q value table is stopped; in the application phase, the lost customer data to be processed is input into the trained Q value table to obtain the Q value corresponding to each first action, and the first action with the largest Q value is determined as the target emergency strategy. Through precise state recognition and dynamic adjustment mechanisms, the present application can not only adapt to different types of lost connection situations, but also continuously optimize strategies to improve collection efficiency and effectiveness. This enables institutions to respond quickly in a complex and changing market environment and take the most appropriate collection actions, thereby significantly improving the success rate of fund recovery and reducing operating costs. In addition, this reinforcement learning-based method provides financial institutions with an intelligent tool that helps improve their overall level of risk management and customer service. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0053] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0054] Figure 1 This is a flow chart of the first embodiment of the method for adjusting the emergency strategy for lost customers based on reinforcement learning of this application;

[0055] Figure 2 This is the Q-value convergence curve of the reinforcement learning-based lost customer emergency strategy adjustment method in this application;

[0056] Figure 3 This is a diagram of the adjustment mechanism of the lost customer emergency strategy adjustment method based on reinforcement learning in this application;

[0057] Figure 4 This is a flowchart of the implementation method of the lost customer emergency strategy adjustment method based on reinforcement learning in an embodiment of the present application;

[0058] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the lost customer emergency strategy adjustment method based on reinforcement learning in an embodiment of the present application.

[0059] The purpose, features and advantages of this application will be further explained with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION

[0060] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0061] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0062] It should be noted that the method for adjusting the emergency strategy for lost customers based on reinforcement learning is specifically a method for adjusting the emergency strategy for lost customers based on reinforcement learning, a device, and a computer-readable storage medium.

[0063] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device or terminal system capable of implementing the above functions. The following uses the system as an example to illustrate this embodiment and the following embodiments.

[0064] Based on this, this embodiment provides a method for adjusting the emergency strategy of lost customers based on reinforcement learning. Figure 1 , Figure 1 This is a flow chart of the method for adjusting the emergency strategy for lost customers based on reinforcement learning in this application. The method for adjusting the emergency strategy for lost customers based on reinforcement learning includes steps S10 to S60:

[0065] Step S10, selecting a set of lost customer data from a preset training set and extracting a first state of the lost customer data;

[0066] Step S20, generating a first action in a preset action space, and executing the first action to obtain a second state;

[0067] Step S30, calculating a reward value for performing the first action in the first state according to a preset reward function;

[0068] Step S40, adjusting the Q value corresponding to the mapping pair of the first state and the first action in the preset Q value table according to the reward value and the second state;

[0069] Step S50: If the preset termination condition is met and the Q value converges, stop adjusting the Q value table;

[0070] Step S60: input the lost customer data to be processed into a Q value table, obtain the Q value corresponding to each first action, and determine the first action with the largest Q value as the target emergency strategy.

[0071] In a feasible implementation manner, step S50 further includes steps T10-T20:

[0072] Step T10: If the preset termination condition is not met, the process returns to the step of generating the first action in the preset action space.

[0073] Step T20: If the Q value does not converge, the process returns to the step of selecting a set of lost customer data from the preset training set.

[0074] It should be noted that in this embodiment, lost customer data refers to multi-dimensional information recorded by lending institutions regarding lost customers, including but not limited to mobile phone number status and number of days overdue. The training set contains multiple sets of lost customer data. The first state is a set of information extracted from this data describing the customer's current status, such as whether the customer is in hide-and-seek mode, absconding with the funds, or faking disappearance. The first state is extracted based on a defined state space, which aims to comprehensively and accurately describe the characteristic states of loan customers under different lost contact modes, providing a basis for subsequent policy adjustments. Corresponding to the state space is the action space, which defines the set of all possible actions that can be taken, such as adjusting collection frequency or initiating legal proceedings. The first action is the specific operation selected from the action space based on the policy. The reward value is calculated using a preset reward function and is used to evaluate the outcome of executing a particular action. The Q-value table is a data structure that records the expected long-term benefits of state-action pairs and is used to make optimal decisions. The termination condition refers to a pre-defined criterion for stopping training, such as customer repayment, the duration of lost contact exceeding a threshold, or reaching the maximum number of iterations. Q-value convergence refers to the process of continuously updating the state-action pairs in the Q-value table until these values ​​become stable, indicating that the optimal strategy has been learned.

[0075] In this embodiment, steps S10-S50 are the training phase of the Q-value table, and step S60 is the application phase of the Q-value table. First, in step S10, a group of data on lost customers is selected from the training set (a historical data set of lost connections of financial institutions), and a first state describing the current situation of the customer is extracted from it, which includes identifying dynamic indicators such as loss of contact patterns, response rates, and asset loss rates. Next, in step S20, a specific action, namely the first action, is generated based on the defined action space, and then the action is executed to change the customer's state to the second state. Step S30 involves using a preset reward function to quantify the effect after executing the first action. This reward function comprehensively considers factors such as the success rate of debt collection, cost control, and customer relationship maintenance. Subsequently, in step S40, the corresponding entries in the Q-value table are updated based on the reward value and the new state. This is achieved through the core update rule of the Q-learning algorithm (a reinforcement learning algorithm) to optimize future decisions.

[0076] If the preset termination condition is not met, the process returns to the step that generated the first action in the preset action space. This means that even if the current round does not achieve the expected goal, such as if the customer does not repay the loan or the disconnection period has not expired, it is necessary to continue exploring and using known information to adjust its behavior strategy. Specifically, it is necessary to reselect and execute a new action based on the second state, calculate the reward value for executing the new action in the second state, and update the Q value table based on the reward value until the preset termination condition is met.

[0077] If the pre-set termination condition is met, the Q-value is determined to have converged. If the Q-value has not yet converged, meaning the model has not yet learned the optimal decision path, the process returns to the step of selecting a new set of lost customer data from the financial institution's historical lost customer data set, restarting the learning process from the beginning. If the termination condition is met and the Q-value has stabilized, the Q-value table update stops (step S50).

[0078] Finally, in step S60, the lost customer data to be processed is input into the trained Q-value table, and the action with the highest Q-value is selected as the target emergency strategy.

[0079] In the training phase of this embodiment, a group of lost customers are first selected from a preset training set, and the first state of the lost customers is extracted; a first action is generated in a preset action space according to a preset strategy, and the first action is executed; the reward value for executing the first action in the first state is calculated based on a preset reward function; the Q value of the mapping pair of the first state and the first action in the preset Q value table is adjusted according to the reward value; if the preset termination condition is met and the Q value converges, the adjustment of the Q value table is stopped; in the application phase, the lost customer data to be processed is input into the trained Q value table to obtain the Q value corresponding to each first action, and the first action with the largest Q value is determined as the target emergency strategy. This embodiment can not only adapt to different types of lost connection situations through precise state recognition and dynamic adjustment mechanisms, but also continuously optimize strategies to improve collection efficiency and effectiveness. This enables institutions to respond quickly in a complex and changing market environment and take the most appropriate collection actions, thereby significantly improving the success rate of fund recovery and reducing operating costs. In addition, this reinforcement learning-based method provides financial institutions with an intelligent tool that helps improve their overall level of risk management and customer service.

[0080] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above introduction, and no further details will be given later. On this basis, the steps of step S10 also include steps A10 to A40:

[0081] Step A10: generating a corresponding lost contact mode identification variable according to the customer characteristics of the lost contact customer data;

[0082] Step A20: Calculate the ratio of the number of successfully contacted customers in the lost customer data within a preset period to the total number of contacted customers to obtain the customer response rate corresponding to the lost customer data;

[0083] Step A30: Calculate the asset loss rate corresponding to the lost customer data, where the asset loss rate is used to monitor customer asset transfer behavior;

[0084] Step A40: Determine the lost connection mode identification variable, the customer response rate, and the asset loss rate as the first state corresponding to the lost connection customer data.

[0085] It should be noted that the loss pattern identification variable refers to the distinction between different types of loss of contact behaviors based on various customer characteristics (such as mobile phone number status, number of days overdue, etc.). The customer response rate measures the customer's responsiveness to collection actions by calculating the ratio of the number of successful customer contact attempts to the total number of contact attempts within a specific time period. The asset churn rate is used to monitor the transfer of assets of customers over a period of time, specifically the change in the value of collateral or liquid assets. The first state is a state description composed of the loss pattern identification variable, the customer response rate, and the asset churn rate, which is used to reflect the customer's current specific situation.

[0086] In this embodiment, step A10 uses key information in the lost customer data to generate a lost connection mode identification variable. This process involves analyzing the customer's mobile phone number status, permanent address, overdue amount and other characteristics, and judging which lost connection mode the customer belongs to based on predefined standards. Next, in step A20, the customer response rate is calculated. This requires counting the number of successful customer contact within a set time period and dividing it by the total number of contact attempts to obtain a proportional value, which reflects the customer's sensitivity to the collection strategy. Then, in step A30, the asset loss rate is determined by comparing the value changes of customer assets within a certain time interval. This indicator helps to understand whether the customer is transferring assets and how fast the transfer is taking place. Finally, in step A40, the lost connection mode identification variable, customer response rate and asset loss rate obtained above are combined to form a first state that describes the current state of the lost customer, providing a basis for subsequent action selection.

[0087] For example, if this embodiment is applied to a specific loan collection field, a state space is defined, and the first state in this embodiment is extracted based on the state space.

[0088] Based on the actual situation, we categorize loss patterns into three main types: hide-and-seek, absconding with funds, and feigned disappearance. Each of these patterns possesses unique characteristics and behavioral manifestations. In designing the state space, we fully consider the heterogeneous nature of loan customer loss behaviors, specifically the dynamic classification based on loss patterns, ensuring that the state space fully reflects the complexity and diversity of loan customer loss patterns.

[0089] (1) Lost connection mode identification variable

[0090] This study models the state space of loan customer loss contingency strategies as a multidimensional dynamic indicator set, with the core dimension being the loss pattern identifier variable M. Based on the heterogeneity of loan customer behavioral characteristics, M is defined as a discrete categorical variable with a value set of {0, 1, 2}, representing the following three loss patterns:

[0091] 1) Hide-and-seek mode (M=0)

[0092] The core characteristic of the "hide-and-seek" model is that customers use temporary means to avoid contact. For example, they reserve a temporary address to conceal their true residence, while keeping their mobile phone number online but frequently refusing collection calls. Furthermore, these customers often apply for credit loans and have weak social connections. This type of customer often uses temporary addresses and frequent phone rejections to avoid collections. Furthermore, the lack of collateral further increases the risk of loss of contact. Weak social connections also reduce the possibility of tracing through connected individuals.

[0093] Its status criterion can be formalized as:

[0094]

[0095] Among them, S phone Indicates the online status of the mobile phone number. The value range is {0,1,2,3,4,5}, which respectively represent normal but rejected, downtime, online but unavailable, empty number, not activated (or cancelled), and abnormal status; n reject represents the cumulative number of times a lending institution receives a busy tone or a powered-off signal when calling a loan customer's mobile phone number; τ1 represents the threshold for the number of times a call to a loan customer's mobile phone number is rejected; A address Indicates the permanent address feature, 0 indicates a temporary address, and 1 indicates a household registration address or business address; L type Indicates the loan type, with a value range of {0,1,2,3}, representing secured loan, mortgage loan, credit loan and discount loan respectively; L relation Indicates the relationship between the loan customer and the valid contact person, 0 means no immediate family member, 1 means there is an immediate family member;

[0096] τ1 can be determined by analyzing historical data showing the number of rejections when a customer's phone number is called. K-means clustering (a clustering algorithm) (k = 2) is then used to divide the group into groups of normal and lost customers. Based on the distribution of the groups, a threshold is selected that clearly distinguishes normal from lost customers. For example, if cluster analysis shows that most normal customers experience rejections between 0 and 5 times, while lost customers experience rejections of more than 10 times, τ1 can be set to 10 times.

[0097] 2) Run away with the money (M=1)

[0098] The "abscond with funds" model is characterized by clients disappearing with large loans. Significant indicators include significant overdue amounts, unusual mobile phone numbers (e.g., suspended, unused, unactivated, or unusual), and complex but ineffective social connections (i.e., numerous valid contacts who provide no valid information). Clients in this pattern typically have large overdue amounts and unusual mobile phone numbers, indicating a determination to sever ties. Complex but ineffective social connections may obscure their true whereabouts and require further verification through asset loss rates. The status criterion can be formalized as follows:

[0099]

[0100] Among them, L overdue represents the overdue loan amount, that is, the amount of debt owed by the customer that has not been repaid on the date agreed in the contract; τ2 represents the threshold value of the overdue loan amount; n valid is the number of valid contacts, indicating the number of non-empty and real mobile phone numbers in the loan customer's mobile phone address book; τ3 is the threshold value of the number of valid contacts.

[0101] τ2 can be determined by analyzing the overdue loan amounts of customers in historical data and using cluster analysis to categorize them into different groups. Based on the distribution characteristics of these groups, a threshold for overdue loan amounts is selected that clearly distinguishes between healthy customers and lost customers. For example, if cluster analysis shows that 90% of customers in historical data have overdue amounts below 500,000 yuan, then τ2 is set to 500,000 yuan.

[0102] τ3 can be determined by analyzing the number of valid contacts of customers in historical data and using cluster analysis to categorize them into different groups. Based on the distribution characteristics of the groups, a threshold for the number of valid contacts that clearly distinguishes between healthy customers and lost customers can be selected. For example, if cluster analysis shows that most healthy customers have between 50 and 100 valid contacts, while lost customers have fewer than 10 valid contacts, τ3 can be set to 10.

[0103] 3) False missing mode (M=2)

[0104] The fake disappearance pattern manifests itself as the client severing social contact channels. Identification criteria include canceling their mobile phone number, missing contact time exceeding a threshold, and a sudden drop in active contacts. Clients in this pattern often evade debt collection by canceling their mobile phone number and being out of touch for extended periods. A sudden drop in active contacts indicates they have actively severed social ties, and the degree of their concealment needs to be assessed in conjunction with the rate of asset loss. The status criteria can be formalized as follows:

[0105]

[0106] Among them, T lost is the number of days without contact, indicating the duration of the loan customer’s loss of contact; τ4 is the threshold for the number of days without contact, and τ5 is the threshold for the number of valid contacts.

[0107] τ4 can be determined using survival analysis to identify the number of days after a customer loses contact when the probability of repayment drops to 10%, using this as the initial threshold. For example, if the probability of repayment after 60 days is less than 5%, set τ4 to 60 days and update it regularly based on collection results.

[0108] τ5 can be determined by analyzing the difference in the number of valid contacts before and after a customer goes missing using a t-test, setting a critical value corresponding to the significance level (ρ < 0.01). For example, if the number of valid contacts for a pseudo-missing customer drops sharply to ≤ 5 (the normal mean is 8), then τ5 is set to 5.

[0109] (2) Dynamic response characteristics

[0110] In order to enhance the real-time performance of the state space, the customer response rate R is introduced response To quantify the reach effect. This indicator is defined as the ratio of successful reach times to total reach times within time period t:

[0111]

[0112] Among them, N contact is the number of times the lending institution reaches the loan customer through relevant channels (phone, email or other channels), and the kth successful reach I k The value of is 1, otherwise it is 0. The customer response rate can dynamically reflect the customer's feedback on the collection contact strategy. For example, in the hide-and-seek mode, R response It may fluctuate briefly, but in a pseudo-disappearance mode it approaches zero.

[0113] (3) Asset loss rate

[0114] Define the asset loss rate V asset To monitor customer asset transfer behavior:

[0115]

[0116] Among them, ΔC is the change in the value of the collateral or liquid assets during the period, and Δt is the time interval calculated in days. asset It is usually negative and has a large absolute value, reflecting the rapid transfer of assets; in the pseudo-disappearance mode, asset loss may be accompanied by long-term concealment.

[0117] Finally, the state space S can be expressed as:

[0118] S={M,R response ,V asset}

[0119] Therefore, for each lost customer, observe its dynamic features and extract the first state:

[0120]

[0121] Among them, M t The variable identified by the lost contact pattern of the lost customer, It is calculated by the percentage of successful touches within a period of time; It is updated based on the time rate of change of the value of the collateral or liquid assets.

[0122] In addition, with respect to the action space, this embodiment defines the action space as follows in combination with the actual business operations of the lending institution:

[0123] (1) Collection actions

[0124] Collection actions mainly focus on the frequency and method of communication with customers, specifically including the following two aspects:

[0125] 1) Collection frequency: The collection frequency is defined as the number of times collection is performed on a customer within a unit of time t (e.g., one month), denoted as F. The range of collection frequency is:

[0126] F∈{F min ,F min +1,…,F max}

[0127] Among them, F min and F max The specific value of is determined based on industry experience and actual conditions. Under different loss of contact modes, the adjustment of collection frequency needs to take into account the customer response rate R response and repayment willingness. For example, for customers in the hide-and-seek mode, if the customer response rate R responseAn upward trend over a certain period of time indicates that the customer is responsive to collection efforts. In this case, the collection frequency F can be appropriately reduced to avoid excessive collection efforts that could cause customer resentment. If the response rate remains low, F can be increased within a reasonable range, perhaps to a value between [10, 15] times / month. For customers who are in a "fake disappearance" pattern, since they actively cut off contact, excessive collection frequencies will have little effect, so F may be adjusted to between [3, 5] times / month.

[0128] 2) Severity coefficient of debt collection tactics

[0129] The severity coefficient W of debt collection words is introduced to quantify the severity of debt collection words. The value range of W is

[0130] W∈[0,1]

[0131] 0 represents gentle and friendly communication, while 1 represents intimidating and forceful communication. Similarly, the severity coefficient of the communication needs to be correlated with the customer's response rate and disconnection pattern. In the "hide-and-seek" mode, when the customer's response rate increases, the severity coefficient W can be appropriately lowered to a value between 0.3 and 0.5, using gentle reminders to maintain communication with the customer. If the customer is more resistant to debt collection, has a low response rate, and is unclear about repayment intentions, W can be set to 0.6-0.8, using forceful language such as emphasizing legal consequences.

[0132] (2) Legal Action

[0133] Legal actions mainly involve the timing of legal proceedings and the intensity of legal deterrence, specifically including the following two aspects:

[0134] 1) Legal action initiation time: The legal action initiation time is defined as the number of days of loan overdue D and the asset loss rate V. asset Function, denoted as T ι , the specific form is:

[0135]

[0136] Among them, τ6 is the preset overdue days threshold for initiating legal proceedings, and τ7 is the asset loss rate threshold. Both thresholds are determined by analyzing the effects of legal proceedings under different overdue days and asset loss conditions in historical data. For example, after data analysis, it was found that when the overdue period exceeds 90 days and the asset loss rate exceeds a certain proportion (such as a monthly asset reduction of more than 10% of total assets), the recovery rate of legal proceedings is higher. In this case, τ6 can be set to 90 days, and τ7 can be set to a monthly asset reduction rate of 10%. For customers who run away with the money, due to the high risk of asset transfer, if the asset loss rate reaches τ7, even if the overdue days do not reach τ6, legal proceedings should be considered.

[0137] 2) Legal deterrence strength

[0138] Legal deterrence intensity L is used to measure the degree of pressure exerted on clients during legal proceedings, and its value range is:

[0139] L∈[0,1]

[0140] 0 indicates only mild deterrent measures, such as sending a lawyer's letter, while 1 indicates full legal enforcement measures. The intensity of legal deterrence should be commensurate with the severity of the client's disappearance and their willingness to repay. For clients who abscond with funds, as their disappearance is more severe, a higher value can be used to maximize the deterrent effect and encourage them to repay or cooperate with the investigation. For clients who are feigning disappearance, if they have the potential to repay, a moderate value can be used initially, and adjustments can be made after observing their response.

[0141] (3) Debt restructuring actions

[0142] Debt restructuring actions mainly focus on the flexibility of the restructuring plan and the adjustment of the repayment plan, specifically including the following two aspects:

[0143] 1) Restructuring plan flexibility coefficient

[0144] Define the flexibility coefficient R of the restructuring plan, and the value range of R is

[0145] R∈[0,1]

[0146] A value of 0 indicates full compliance with the original loan agreement, with no restructuring; a value of 1 indicates comprehensive and in-depth adjustments to the loan agreement, such as extending the repayment period, lowering the interest rate, or adjusting the repayment method. When determining the R value, the client's repayment ability, asset status, and pattern of loss of contact must be comprehensively considered. For clients who have lost contact due to temporary cash flow difficulties but have significant repayment potential, a debt restructuring plan with a higher R value may be considered, such as extending the repayment period by 50% and reducing the interest rate by 20%. This will help alleviate repayment pressure and increase the bank's likelihood of recovering the loan.

[0147] 2) Adjustment of repayment plan

[0148] The repayment plan adjustment range ΔP is used to quantify the degree of change in the repayment plan. If the original repayment plan is a monthly repayment amount of P, the adjusted repayment amount is P ′ ,but

[0149]

[0150] When ΔP>0, it indicates an increase in the repayment amount, and when ΔP<0, it indicates a decrease in the repayment amount. The adjustment range of the repayment plan is correlated with the flexibility coefficient R of the restructuring plan, which can generally be expressed as:

[0151] ΔP=R·P

[0152] When the R value is high, in order to ensure that the customer has the repayment ability under the new repayment arrangement, P may be adjusted to a negative value accordingly, that is, reducing the repayment amount and extending the repayment period to ensure the feasibility and effectiveness of the debt restructuring plan.

[0153] In summary, the action space A can be expressed as:

[0154] A={F,W,T ι ,L,R,ΔP}

[0155] The action space encompasses three broad categories: collection actions, legal actions, and debt restructuring actions. The specific actions within each category are highly quantifiable and can be dynamically adjusted and optimized based on factors such as the client's loss of contact patterns, repayment willingness, and asset status. In practice, these three types of actions will work together to achieve optimal risk management and capital recovery.

[0156] This embodiment not only accurately models the problem of lost customer contact, but also enhances the model's adaptability and responsiveness to complex loss scenarios. This mechanism helps financial institutions take the most appropriate response measures in a timely manner, improve fund recovery efficiency, reduce collection costs, and enhance risk management capabilities.

[0157] In a feasible implementation, the steps in step S20 may include steps B10 to B40:

[0158] Step B10, obtaining the current exploration rate, wherein the current exploration rate decays as the number of training times increases;

[0159] Step B20, generating a current random number corresponding to the current training round;

[0160] Step B30: If the current random number is less than the current exploration rate, any action in the action space is determined as the first action;

[0161] Step B40: If the current random number is greater than or equal to the current exploration rate, the action with the highest Q value corresponding to the first state in the Q value table is determined as the first action.

[0162] It should be noted that the current exploration rate is a value that gradually decreases with the number of training cycles. It determines the probability of selecting a random action. The current random number is a value between 0 and 1 generated at the beginning of each training cycle and is compared with the current exploration rate to determine the strategy to adopt.

[0163] In step B10 of this embodiment, the exploration rate of the current training round is calculated. This value will gradually decrease as the number of training times increases. The purpose is to balance the ratio of exploring new actions and utilizing existing knowledge. Then, in step B20, a random number is generated for the current training round. Step B30 checks if this random number is less than the current exploration rate, which means that it should be inclined to explore unknown areas, so an action is randomly selected from the action space as the first action. On the contrary, if the random number is greater than or equal to the current exploration rate (step B40), it indicates that it is more inclined to utilize known information. At this time, the action with the highest Q value in the first state in the Q value table will be selected as the first action. This strategy ensures that the intelligent agent can fully explore the environment and effectively use the best strategy learned to optimize decision-making.

[0164] For example, the strategy for selecting the first action can refer to the following formula: randomly select an action with a probability of ε, select the action with the largest Q value in the current state with a probability of 1-ε, and balance exploring new actions and utilizing existing experience.

[0165]

[0166] Where ∈ is the exploration rate, which decays with the number of training rounds (e.g. ∈=∈0·e -kt ),α t It is the first action.

[0167] Through the above steps, this embodiment enables the agent to both broadly explore various possible action combinations during the learning process and effectively utilize learned knowledge to make optimal decisions, thereby improving the model's convergence speed and ultimate performance. This mechanism also helps improve financial institutions' ability to respond to the issue of lost loan customers, enhancing risk control and capital recovery efficiency.

[0168] In a feasible implementation manner, step S30 further includes steps C10 to C40:

[0169] Step C10, calculating a success reward based on the amount recovered after executing the first action in the first state and the response time for executing the first action;

[0170] Step C20, calculating the cost penalty based on the collection frequency and litigation intensity of the first action;

[0171] Step C30: Calculate the maintenance reward based on the customer response rate and the severity coefficient of the first action's collection speech;

[0172] Step C40, based on the loss mode identification variable, determines the weights corresponding to the success reward, cost penalty, and maintenance reward respectively, and performs weighted calculation on the success reward, cost penalty, and maintenance reward according to each weight to obtain the reward value for executing the first action in the first state.

[0173] It should be noted that in this embodiment, the success reward is a reward value calculated based on the amount recovered after executing the first action in the first state and the response time, which is intended to encourage fast and efficient collection actions. The cost penalty is a cost burden determined based on the frequency of collection and the intensity of litigation involved in executing the first action, with the purpose of controlling unnecessary expenses. The maintenance reward is evaluated by analyzing the customer response rate and the severity coefficient of the collection speech of the first action, and is used to measure the efforts to maintain good customer relationships. The weight is dynamically adjusted based on the loss of contact mode identification variable to ensure a balance between different goals. The final reward value is a comprehensive evaluation indicator obtained by weighting the success reward, cost penalty and maintenance reward.

[0174] First, in step C10, the calculation of the success reward takes into account the actual amount recovered and the time interval between action and recovery. This step emphasizes timeliness and effectiveness: recovering more funds faster will result in higher rewards. Next, in step C20, the calculation of the cost penalty takes into account the frequency of collection and the intensity of legal action. A higher frequency of collection or stronger legal deterrence typically means higher costs and, therefore, corresponding penalties. Then, in step C30, the maintenance reward is determined based on the customer response rate and the severity coefficient of the collection tactics. Gentle tactics and positive customer responses help maintain long-term customer relationships, thus earning positive rewards. Finally, in step C40, the weights of each reward item are determined based on the loss pattern identifier variable. The weighted sum of the success reward, cost penalty, and maintenance reward is then taken to determine the overall reward value for executing the first action in the first state. This process adjusts the weights to accommodate different loss patterns, ensuring that the strategy achieves both effective collection results and balanced cost control and customer relationship maintenance.

[0175] For example, this embodiment combines the actual business operations of lending institutions, comprehensively considers the three core goals of collection success rate, cost control, and customer relationship maintenance, and constructs a composite reward function, which is mathematically formalized as follows:

[0176] R total =α·R collect +β·R cost +γ·R relation

[0177] Among them, R collect R is the reward item for successful collection. cost is the cost penalty term, R relation is the relationship maintenance reward item; α, β, γ are weight coefficients, satisfying α+β+γ=1. The specific value is determined through experimental tuning. The definitions of each sub-item are as follows:

[0178] (1) Rewards for successful collection

[0179] This reward aims to maximize the amount of loan recovery while taking into account the timeliness of collection actions under different loss of contact modes. The recovery amount ratio is defined as:

[0180]

[0181] Among them, C actual is the actual recovery amount, C expected The expected recoverable amount (estimated based on the customer's asset value and overdue amount).

[0182] We further introduce a time penalty factor λ(t) to encourage the agent to respond quickly:

[0183] λ(t)=e -θ·t

[0184] Where t is the time interval from the loss of contact to the time of taking action, and θ is the decay coefficient (θ>0). The final collection reward is:

[0185] R collect =r collect ·λ(t)

[0186] (2) Cost penalty

[0187] Cost items are used to constrain the economic cost of debt collection actions, including labor costs, legal fees, and opportunity costs:

[0188] R cost =-(ω1·F+ω2·L+ω3·I legal )

[0189] Among them, F is the frequency of debt collection, L is the legal deterrent strength, I legal Initiation of legal proceedings legal =1 triggers litigation costs); ω1, ω2, and ω3 are unit cost coefficients, determined through historical data regression analysis. For example, ω1 could represent the labor cost of a single debt collection, and ω3 could be the fixed cost of legal action.

[0190] (3) Relationship maintenance rewards

[0191] This reward is intended to avoid excessive collection leading to customer loss or reputation loss, combined with the customer response rate R relation Designed with the severity of the words:

[0192] R relation =η·R response -μ·W 2

[0193] Among them, η is the response rate gain coefficient, μ is the severity penalty coefficient, η and μ are determined by adjusting the parameters of the multi-objective optimization algorithm to balance the collection efficiency and customer relationship. 2 Strengthen the suppression of high-pressure rhetoric to avoid sacrificing long-term customer value for short-term gains.

[0194] Therefore, the selected action A is executed t , triggering the corresponding collection, legal or restructuring operations. The reward value for executing the first action in the first state is:

[0195]

[0196] This implementation not only effectively manages the issue of lost loan customers but also significantly improves risk management and fund recovery efficiency. This mechanism enables financial institutions to formulate more scientific collection strategies, balance short-term profits with long-term customer value, and enhance their ability to navigate complex market environments. This approach also helps reduce unnecessary collection costs and improve overall operational efficiency.

[0197] In a feasible implementation manner, step C40 further includes steps C401 to C404:

[0198] Step C401: updating the weight of the success reward according to the asset loss rate, wherein the asset loss rate is positively correlated with the weight of the success reward;

[0199] Step C402: updating the weight of the cost penalty according to the customer response rate, wherein the customer response rate is negatively correlated with the weight of the cost penalty;

[0200] Step C403: updating the weight of the maintenance reward according to the loss time in the lost customer data, wherein the loss time is positively correlated with the weight of the maintenance reward;

[0201] Step C404 , performing weighted calculation on the success reward, cost penalty, and maintenance reward according to the updated weights to obtain a reward value for the first action in the first state.

[0202] It should be noted that the weight of the success reward is updated according to the asset loss rate. This means that if the customer's asset loss rate accelerates, the weight of the reward obtained by successfully recovering the loan will also increase accordingly to encourage faster action. The weight of the cost penalty is adjusted according to the customer response rate. A higher customer response rate will result in a lower weight of the cost penalty, because this indicates that the customer has a positive response to the collection, reducing unnecessary high-cost actions. The weight of the maintenance reward is updated by the length of time of loss of contact. The longer the loss of contact, the higher the weight of the maintenance reward, emphasizing the importance of maintaining customer relationships in the event of a long period of loss of contact. Finally, the weighted values ​​of the success reward, cost penalty, and maintenance reward are recalculated based on these updated weights to obtain the comprehensive reward value for performing the first action in the first state.

[0203] In this embodiment, step C401 adjusts the weight of the success reward based on the asset churn rate. Specifically, this involves monitoring the speed of customer asset transfers and dynamically adjusting the reward weight for successfully recovering a loan accordingly. A faster asset churn rate means quicker action is needed to mitigate losses, so increasing the weight of the success reward can incentivize swift action. Next, in step C402, the weight of the cost penalty is adjusted based on the customer's response rate. When a customer demonstrates positive response, such as answering a call or making a partial repayment, the weight of the cost penalty is appropriately reduced, reflecting the need to over-rely on costly collection methods in this situation. Then, in step C403, the weight of the maintenance reward is updated using the duration of disconnection. For customers who have been disconnected for an extended period, maintaining the relationship becomes more important, so the weight of the maintenance reward increases accordingly as the disconnection time increases. Finally, in step C404, the updated weights are used to perform a weighted sum of the success reward, cost penalty, and maintenance reward to obtain the combined reward value for executing the first action in the first state, ensuring a balance and optimization between different objectives.

[0204] For example, to adapt to the differentiated goals of different loss modes in the loan loss customer scenario, the weight coefficients α, β, and γ are dynamically adjusted according to the mode M:

[0205] Absence model (M=1): Focuses on asset preservation and legal deterrence, increasing α and β weights;

[0206] Hide-and-seek mode (M=0): Balance collection efficiency and customer relationships, and increase γ appropriately;

[0207] False disappearance mode (M=2): Focus on cost control and long-term tracking, and increase β.

[0208] The specific adjustment rules are:

[0209]

[0210] Among them, the weight coefficients α0, β0, and γ0 are initialized by expert experience; k is the asset loss sensitivity coefficient, which is optimized by gradient descent method to maximize the long-term cumulative reward; T lost is the duration of loss of contact, and τ4 is the threshold of loss of contact days.

[0211] Through the above steps, this embodiment not only accurately reflects the impact of different factors on reward values, but also enhances the model's ability to adapt to complex and changing situations. This mechanism helps financial institutions more flexibly respond to various loss of contact scenarios, improves fund recovery efficiency, reduces collection fees, and enhances risk management capabilities.

[0212] Based on the first or second embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the first or second embodiment can be referred to the above introduction and will not be described in detail. Steps D10 to D50 are also included after step S40:

[0213] Step D10, searching the Q value table for all new mapping pairs corresponding to the second state;

[0214] Step D20, obtaining new Q values ​​of all new mapping pairs, and determining the maximum new Q value from all new Q values;

[0215] Step D30, calculating the Q value increment based on the reward value and the maximum new Q value;

[0216] Step D40, adding the Q value increment to the Q value corresponding to the mapping pair to obtain an updated Q value;

[0217] Step D50: Use the updated Q value to update the Q value corresponding to the mapping pair.

[0218] It should be noted that in this embodiment, the second state refers to the new state of the loan customer after performing the first action. The new mapping pairs refer to all possible actions and their expected returns associated with the second state in the Q-value table. The maximum new Q-value is the highest expected return found among these new mapping pairs. The Q-value increment is calculated based on the current reward value and the maximum new Q-value, used to update the original Q-value. The updated Q-value is the sum of the original Q-value and the Q-value increment, reflecting the new learning results.

[0219] First, in step D10 of this embodiment, all new mapping pairs related to the new state (i.e., the second state) after executing the first action are searched from the Q-value table. This step prepares for subsequent updates by identifying all possible actions and their expected benefits. Then, in step D20, the new Q values ​​of all new mapping pairs are obtained, and the maximum new Q value is found from them. This step helps determine the maximum expected benefit that can be obtained by taking the best action in the current state. Then, in step D30, the current reward value and the maximum new Q value are used to calculate the Q-value increment. This increment represents the correction of the original knowledge based on the latest results. The specific formula usually adopts the update rule in the Q-learning algorithm, such as the Q-value increment is equal to the reward value plus the discount factor multiplied by the maximum new Q value minus the original Q value. Subsequently, in step D40, the calculated Q-value increment is added to the Q value corresponding to the original mapping pair to obtain the updated Q value. This process reflects the ability of the intelligent agent to gradually optimize its strategy through continuous learning. Finally, in step D50, the updated Q-values ​​are used to replace the old values ​​in the original Q-value table, ensuring that the agent can use the latest and better knowledge in future decisions.

[0220] For example, based on the reward obtained from performing the action and the new state after performing the action, the Q value of the corresponding state-action pair in the Q value table is adjusted:

[0221]

[0222] Among them, η is the learning rate, δ is the discount factor, S t+1 The new state after the action is executed.

[0223] This embodiment enables institutions to quickly respond and take optimal measures in complex and changing environments, improving fund recovery efficiency, reducing collection costs, and enhancing risk management capabilities. Furthermore, this mechanism helps agents continuously improve their strategies, enhancing overall operational efficiency and market competitiveness.

[0224] Furthermore, after adjusting the Q value of the corresponding state-action pair in the Q value table, if any of the following conditions is met, the current training round is terminated:

[0225] (1) Customer repayment or debt settlement;

[0226] (2) The duration of the loss of contact exceeds the threshold τ4;

[0227] (3) Reach the maximum number of iterations T max .

[0228] Determine whether the Q value converges. The Q value convergence curve is as follows: Figure 2As shown in the figure, the Q-value convergence curve rises rapidly at the beginning of training, then drops steeply, and then rises steadily again. This is because the purpose of building a dynamic adjustment mechanism for emergency loan customer loss strategies based on Q-learning is to improve collection success rates, control costs, or maintain customer relationships. The Q-value represents the expected long-term reward for taking a certain action under a certain state. As the curve rises, it means that the agent is able to learn the optimal decision-making strategy under different loss modes and environmental conditions, gradually discovering better action options, and the Q-value continues to increase, converging towards the optimal strategy.

[0229] If it does not converge, continue with state extraction, action selection and other operations. If it converges, output the optimal strategy π * :

[0230]

[0231] π * Embedded in the collection system, it enables dynamic adjustment of emergency strategies for loan customers who lose contact.

[0232] Combining all the embodiments, we can see that the adjustment mechanism diagram of this application is as follows: Figure 3 As shown, various customer information is first collected, including mobile phone number online status, permanent address, loan type, relationship with active contacts, overdue loan amount, number of active contacts, number of days lost contact, value of collateralized liquid assets, and number of rejected calls to the mobile phone number. Based on this data, a state space is defined: Based on the collected data, the customer's current state is identified, including different loss of contact patterns (hide-and-seek, absconding with funds, and fake disappearance), as well as dynamic response rates and asset churn rates. An action space is defined: Possible collection actions are defined, including frequency, severity of collection rhetoric, whether to initiate legal action, legal deterrence intensity, debt restructuring plans, and repayment schedules. A reward function is defined: Multiple objectives, including loan recovery amount, response speed, cost control, and customer relationship maintenance, are set to evaluate the effectiveness of different strategies. A dynamic adjustment mechanism, through dynamic adjustment of weight coefficients, refined mapping, and dynamic adaptation, ensures that strategies adapt to different loss of contact patterns and balances collection efficiency and customer relationships. Finally, the optimal action combination is output: the specific collection strategy combination is output, including the frequency of collection, the severity of the words, legal proceedings, the intensity of legal deterrence, debt restructuring plan and the adjustment range of the repayment plan.

[0233] The implementation flow chart of this application is as follows Figure 4As shown, a group of lost customers is first selected from the training set, and an initial state is defined. A Q-value table is initialized to store the expected reward for each state-action pair. Real-time customer data is collected and analyzed to determine the current state. Based on the ε-greedy strategy (a behavior selection strategy), a decision is made whether to randomly explore new actions or select the action with the highest current Q-value. Instantaneous rewards are calculated based on the executed actions. Reward weights are adjusted based on the immediate rewards to better reflect the effectiveness of the current strategy. The corresponding entries in the Q-value table are updated based on the new reward information. A determination is made as to whether the preset termination conditions (such as the number of training runs, convergence criteria, and maximum number of iterations) have been met. If not, training is continued; if the termination conditions are met, the current training round is terminated. The Q-value table is then determined to determine whether the values ​​have converged, i.e., whether the changes are sufficiently small. If not, training is continued; if so, the current Q-value table is output. The lost customer data to be processed is entered into the Q-value table to obtain the Q-value corresponding to each first action. The first action with the highest Q-value is selected as the optimal strategy, which is then deployed to the collection system for dynamic adjustment.

[0234] It should be noted that the above examples are only used to understand this application and do not constitute a limitation on the method for adjusting the emergency strategy for lost customers based on reinforcement learning in this application. More simple transformations based on this technical concept are all within the scope of protection of this application.

[0235] The present application also provides a device for adjusting a lost customer emergency strategy based on reinforcement learning, and the device for adjusting a lost customer emergency strategy based on reinforcement learning includes:

[0236] The state acquisition module 10 is used to select a set of lost customer data from a preset training set and extract the first state of the lost customer data;

[0237] An action execution module 20 is configured to generate a first action in a preset action space, and execute the first action to obtain a second state;

[0238] A reward calculation module 30 is used to calculate a reward value for performing a first action in a first state according to a preset reward function;

[0239] A Q value adjustment module 40 is configured to adjust the Q value corresponding to the mapping pair of the first state and the first action in a preset Q value table according to the reward value and the second state;

[0240] A convergence judging module 50 is configured to stop adjusting the Q value table if a preset termination condition is met and the Q value converges;

[0241] The emergency strategy module 60 is configured to input the lost customer data to be processed into a Q value table, obtain the Q value corresponding to each first action, and determine the first action with the largest Q value as the target emergency strategy.

[0242] The reinforcement learning-based emergency strategy adjustment device for lost customers provided in this application adopts the reinforcement learning-based emergency strategy adjustment method for lost customers in the above-mentioned embodiment, which can solve the technical problems of risk management of financial institutions caused by the heterogeneity of loss of connection scenarios and response lag. Compared with the existing technology, the beneficial effects of the reinforcement learning-based emergency strategy adjustment device for lost customers provided in this application are the same as the beneficial effects of the reinforcement learning-based emergency strategy adjustment method for lost customers provided in the above-mentioned embodiment, and the other technical features of the reinforcement learning-based emergency strategy adjustment device for lost customers are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.

[0243] The present application provides a device for adjusting an emergency strategy for a lost customer based on reinforcement learning. The device for adjusting an emergency strategy for a lost customer based on reinforcement learning includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for adjusting an emergency strategy for a lost customer based on reinforcement learning in the above-mentioned embodiment one.

[0244] Reference below Figure 5 , which shows a schematic structural diagram of a device for adjusting emergency strategies for lost customers based on reinforcement learning suitable for implementing an embodiment of the present application. The device for adjusting emergency strategies for lost customers based on reinforcement learning in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The reinforcement learning-based lost customer emergency strategy adjustment device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.

[0245] like Figure 5As shown, the device for adjusting the emergency strategy for lost customers based on reinforcement learning may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1002 or the program loaded from the storage device 1003 into the random access memory 1004. Various programs and data required for the operation of the device for adjusting the emergency strategy for lost customers based on reinforcement learning are also stored in the random access memory 1004. The processing device 1001, the read-only memory 1002 and the random access memory 1004 are connected to each other via a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems can be connected to the input / output interface 1006 which is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the reinforcement learning-based emergency strategy adjustment device for lost customers to communicate wirelessly or wired with other devices to exchange data. Although the figure shows a reinforcement learning-based emergency strategy adjustment device for lost customers with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have instead.

[0246] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.

[0247] The reinforcement learning-based emergency strategy adjustment device for lost customers provided in this application adopts the reinforcement learning-based emergency strategy adjustment method for lost customers in the above-mentioned embodiment, which can solve the technical problems of risk management of financial institutions caused by the heterogeneity of loss of connection scenarios and delayed response. Compared with the existing technology, the beneficial effects of the reinforcement learning-based emergency strategy adjustment device for lost customers provided in this application are the same as the beneficial effects of the reinforcement learning-based emergency strategy adjustment method for lost customers provided in the above-mentioned embodiment, and the other technical features of the reinforcement learning-based emergency strategy adjustment device for lost customers are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0248] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0249] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0250] The present application provides a medium, which is a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, and the computer-readable program instructions are used to execute the reinforcement learning-based lost customer emergency strategy adjustment method in the above-mentioned embodiment.

[0251] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0252] The above-mentioned computer-readable storage medium can be included in the lost customer emergency strategy adjustment device based on reinforcement learning; or it can exist independently without being assembled into the lost customer emergency strategy adjustment device based on reinforcement learning.

[0253] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the lost customer emergency strategy adjustment device based on reinforcement learning, the lost customer emergency strategy adjustment device based on reinforcement learning:

[0254] Selecting a set of lost customer data from a preset training set and extracting the first state of the lost customer data;

[0255] Generate a first action in a preset action space, and execute the first action to obtain a second state;

[0256] Calculating a reward value for performing a first action in a first state according to a preset reward function;

[0257] Adjust the Q value corresponding to the mapping pair of the first state and the first action in the preset Q value table according to the reward value and the second state;

[0258] If the preset termination condition is met and the Q value converges, stop adjusting the Q value table;

[0259] The lost customer data to be processed is input into the Q value table to obtain the Q value corresponding to each first action, and the first action with the largest Q value is determined as the target emergency strategy.

[0260] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0261] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0262] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0263] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned method for adjusting the emergency strategy for lost customers based on reinforcement learning, and can solve the technical problems of risk management and control caused by the heterogeneity of loss of connection scenarios and delayed response of financial institutions. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the method for adjusting the emergency strategy for lost customers based on reinforcement learning provided in the above-mentioned embodiment, and will not be elaborated here.

[0264] The present application also provides a product, which is a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the above-mentioned method for adjusting the emergency strategy for lost customers based on reinforcement learning.

[0265] The computer program product provided in this application can address the technical challenges of risk management for financial institutions caused by heterogeneity in disconnection scenarios and delayed responses. Compared to existing technologies, the beneficial effects of the computer program product provided in this application are similar to those of the reinforcement learning-based emergency strategy adjustment method for disconnected customers provided in the aforementioned embodiments, and are not further elaborated here.

[0266] The above are only some embodiments of the present application and are not intended to limit the patent scope of the present application. All equivalent structural transformations made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for adjusting emergency strategies for lost customers based on reinforcement learning, characterized in that: The reinforcement learning-based method for adjusting the lost customer emergency strategy includes: Selecting a set of lost customer data from a preset training set and extracting a first state of the lost customer data; Generate a first action in a preset action space, and execute the first action to obtain a second state; Calculating a reward value for performing the first action in the first state according to a preset reward function; Adjusting the Q value corresponding to the mapping pair of the first state and the first action in a preset Q value table according to the reward value and the second state; If the preset termination condition is met and the Q value converges, then stop adjusting the Q value table; The lost customer data to be processed is input into the Q value table to obtain the Q values ​​corresponding to the first actions, and the first action with the largest Q value is determined as the target emergency strategy.

2. The method for adjusting the emergency strategy for lost customers based on reinforcement learning according to claim 1, characterized in that: The step of generating a first action in a preset action space includes: Obtaining a current exploration rate, wherein the current exploration rate decays as the number of training times increases; Generate the current random number corresponding to the current training round; If the current random number is less than the current exploration rate, determining any action in the action space as the first action; If the current random number is greater than or equal to the current exploration rate, the action with the highest Q value corresponding to the first state in the Q value table is determined as the first action.

3. The method for adjusting the lost customer emergency strategy based on reinforcement learning according to claim 1, characterized in that: The step of extracting the first state of the lost customer data includes: Generate a corresponding lost contact mode identification variable according to the customer characteristics of the lost contact customer data; Calculating the ratio of the number of successfully contacted customers to the total number of contacted customers in the lost customer data within a preset time period to obtain the customer response rate corresponding to the lost customer data; Calculating the asset loss rate corresponding to the lost customer data, wherein the asset loss rate is used to monitor customer asset transfer behavior; The loss mode identification variable, the customer response rate, and the asset loss rate are determined as a first state corresponding to the lost customer data.

4. The method for adjusting the emergency strategy for lost customers based on reinforcement learning according to claim 3, characterized in that: The action space includes collection actions and litigation actions, the collection actions include a collection speech severity coefficient and a collection frequency, and the litigation actions include litigation intensity. The step of calculating a reward value for executing the first action in the first state according to a preset reward function includes: Calculate the success reward based on the amount recovered after executing the first action in the first state and the response time of executing the first action; Calculate cost penalties based on the collection frequency and litigation intensity of the first action; Calculate the maintenance reward based on the customer response rate and the severity coefficient of the collection speech of the first action; According to the loss mode identification variable, the weights corresponding to the success reward, the cost penalty, and the maintenance reward are determined respectively, and the success reward, the cost penalty, and the maintenance reward are weightedly calculated according to each weight to obtain the reward value for performing the first action in the first state.

5. The method for adjusting the emergency strategy for lost customers based on reinforcement learning according to claim 4, characterized in that: After the step of determining the weights corresponding to the success reward, the cost penalty, and the maintenance reward according to the loss mode identification variable, the following steps are included: updating the weight of the success reward according to the asset loss rate, wherein the asset loss rate is positively correlated with the weight of the success reward; updating the weight of the cost penalty according to the customer response rate, wherein the customer response rate is negatively correlated with the weight of the cost penalty; Updating the weight of the maintenance reward according to the loss of contact duration in the lost customer data, wherein the loss of contact duration is positively correlated with the weight of the maintenance reward; The success reward, the cost penalty, and the maintenance reward are weightedly calculated according to the updated weights to obtain a reward value for the first action in the first state.

6. The method for adjusting the emergency strategy for lost customers based on reinforcement learning according to claim 1, characterized in that: The step of adjusting the Q value corresponding to the mapping pair of the first state and the first action in a preset Q value table according to the reward value and the second state includes: Searching the Q value table for all new mapping pairs corresponding to the second state; Get the new Q values ​​of all new mapping pairs and determine the maximum new Q value from all the new Q values; Calculate a Q value increment according to the reward value and the maximum new Q value; Adding the Q value increment to the Q value corresponding to the mapping pair to obtain an updated Q value; Using the updated Q value, the Q value corresponding to the mapping pair is updated.

7. The method for adjusting the emergency strategy for lost customers based on reinforcement learning according to claim 1, characterized in that: After the step of adjusting the Q value corresponding to the mapping pair of the first state and the first action in the preset Q value table according to the reward value and the second state, the following steps are included: If the preset termination condition is not met, the process returns to the step of generating the first action in the preset action space.

8. The method for adjusting the lost customer emergency strategy based on reinforcement learning according to claim 1, characterized in that: After the step of adjusting the Q value corresponding to the mapping pair of the first state and the first action in the preset Q value table according to the reward value and the second state, the following steps are included: If the Q value does not converge, the process returns to the step of selecting a set of lost customer data from the preset training set.

9. A device for adjusting emergency strategies for lost customers based on reinforcement learning, characterized in that: The reinforcement learning-based lost customer emergency strategy adjustment device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the reinforcement learning-based lost customer emergency strategy adjustment method as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the method for adjusting the emergency strategy for lost customers based on reinforcement learning as described in any one of claims 1 to 8 are implemented.