Information processing device, method, and program
The system corrects causal scores using a combination of high-bias, low-variance and high-bias, high-variance models to improve accuracy and reduce operational costs in debt collection.
Patent Information
- Application Number
- JP2022099545
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-06-21
- Publication Date
- 2025-08-21
- Estimated Expiration
- 2042-06-21
AI Technical Summary
Existing machine learning models struggle to accurately predict the effect of operations on users due to high bias and low variance, leading to inefficiencies in debt collection efforts.
A system that corrects causal scores using a high-bias, low-variance model with a high-bias, high-variance model, employing a correction function to minimize error and improve accuracy.
Enhances the accuracy of scoring operations' effects on users, reducing operational costs without compromising debt collection rates.
Smart Images

Figure 0007727595000001 
Figure 0007727595000002 
Figure 0007727595000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to techniques for scoring the effect that an operation has on a user. [Background technology]
[0002] Conventionally, call center operators make phone calls to customers to urge them to pay their credit card bills, i.e., to collect debts (see Patent Document 1). Also known is a urging support technology that changes the tone of the sample phrases used by the operator when responding to calls depending on the time of day the urging work is being carried out (see Patent Document 2). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2001-282994 [Patent Document 2] Japanese Patent Application Laid-Open No. 2010-224617 Summary of the Invention [Problem to be solved by the invention]
[0004] It is conceivable to use a machine learning model to score the effect of a given operation on a user. However, if the machine learning model is trained by relying heavily on past user behavioral tendencies and / or results, the model may not be able to accurately predict future scores.
[0005] In view of the above-mentioned problems, an object of the present disclosure is to improve the accuracy of scoring the effect that an operation has on a user. [Means for solving the problem]
[0006] An example of the present disclosure is an information processing device that includes a first effect estimation means that acquires a first causal score that indicates the effect of a specified operation on a user by inputting the user's attributes into a first model; a second effect estimation means that acquires a second causal score that indicates the effect by inputting the user's attributes into a second model that has a smaller bias and a larger variance than the first model; a correction function determination means that determines a correction function for correcting the first causal score based on the first causal score and the second causal score calculated for each of a plurality of users; and a third effect estimation means that determines a third causal score that indicates the effect on the target user by applying the first causal score calculated for the target user to the correction function.
[0007] The present disclosure can be understood as an information processing device, a system, a method executed by a computer, or a program executed by a computer. The present disclosure can also be understood as such a program recorded on a recording medium readable by a computer or other device, machine, etc. Here, a recording medium readable by a computer, etc. refers to a recording medium that stores information such as data and programs by electrical, magnetic, optical, mechanical, or chemical action and can be read by a computer, etc. [Effects of the Invention]
[0008] According to the present disclosure, it is possible to improve the accuracy of scoring the effect that an operation has on a user. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a schematic diagram illustrating a configuration of an information processing system according to an embodiment. [Figure 2] FIG. 1 is a diagram illustrating an outline of a functional configuration of an information processing apparatus according to an embodiment. [Figure 3]FIG. 10 is a diagram illustrating an outline of a configuration in an embodiment in which a causal score output from a model with a relatively high bias and low variance is corrected by referring to the output from a model with a relatively low bias and high variance. [Figure 4] FIG. 10 is a diagram illustrating an overview of a process for determining a correction function using a first causal score and a second causal score in an embodiment. [Figure 5] FIG. 1 is a diagram illustrating a simplified concept of a decision tree of a machine learning model employed in an embodiment. [Figure 6] FIG. 10 is a diagram showing the relationship between estimated effects and risks and operation conditions in an embodiment. [Figure 7] 1 is a flowchart illustrating a flow of machine learning processing according to an embodiment. [Figure 8] 10 is a flowchart illustrating a flow of a correction function determination process according to the embodiment. [Figure 9] 10 is a flowchart showing the flow of a third causality score estimation process according to the embodiment. [Figure 10] 10 is a flowchart showing the flow of an operation condition output process according to the embodiment. [Figure 11] 10 is a flowchart showing a flow of variation (1) of the operation condition output process according to the embodiment. [Figure 12] 10 is a flowchart showing the flow of variation (2) of the operation condition output process according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of an information processing device, method, and program according to the present disclosure will be described with reference to the drawings. However, the embodiments described below are merely examples, and the information processing device, method, and program according to the present disclosure are not limited to the specific configurations described below. In implementing the present disclosure, a specific configuration according to the embodiment may be appropriately adopted, and various improvements and modifications may be made.
[0011] In this embodiment, an embodiment will be described in which the technology according to the present disclosure is implemented in an operation center management system for collecting receivables by urging payment of overdue credit card amounts. However, systems to which the technology according to the present disclosure can be applied are not limited to operation center management systems for urging payment of credit card amounts. The technology according to the present disclosure can be widely used in technologies for scoring the effects that operations have on users, and the application of the present disclosure is not limited to the examples shown in the embodiments. Furthermore, the types of operations to which the technology according to the present disclosure can be applied and the types of effects that the technology has on users are not limited to the examples shown in the embodiments.
[0012] Normally, credit card payments are made by debiting the user's account on the monthly withdrawal date or by the user depositing the amount by a specified date, but there are cases where the payment of the credit card amount is not completed by the specified date due to reasons such as insufficient funds in the user's account or the user not depositing the amount by the specified date. For this reason, operations such as calling customers or sending messages from an operation center (call center) have traditionally been carried out to urge users to pay their credit card amounts and collect receivables.
[0013] Generally, operations directed at users are effective in avoiding default (non-performance of debt), and the more operations are performed, the higher the debt collection rate. However, the more operations are performed, the higher the costs, such as labor costs for operations, system usage fees, and system maintenance costs, increase. Therefore, in consideration of the above-mentioned problems, the system disclosed herein employs technology for suppressing operation costs without reducing the debt collection rate. Note that, in this embodiment, an example will be described in which the predetermined operation is primarily a call to the user, but the content of the predetermined operation is not limited, and may be various operations for prompting the user to take a predetermined action.
[0014] <System configuration> FIG. 1 is a schematic diagram showing the configuration of an information processing system according to this embodiment. In the system according to this embodiment, an information processing device 1, an operation center management system 3, and a credit card management system 5 are connected to each other so that they can communicate with each other. An operation terminal (not shown) is installed in the operation center to perform operations according to instructions from the operation center management system 3, and an operator operates the operation terminal to perform operations for a user. The user is a credit card user who deposits the credit card usage amount via a financial institution or the like, and payment history data for the credit card usage amount is notified to the operation center management system 3 via the credit card management system 5.
[0015] The information processing device 1 is an information processing device for outputting data for controlling operations by the operation center management system 3. The information processing device 1 is a computer including a central processing unit (CPU) 11, a read-only memory (ROM) 12, a random access memory (RAM) 13, a storage device 14 such as an electrically erasable and programmable read-only memory (EEPROM) or a hard disk drive (HDD), a communication unit 15 such as a network interface card (NIC), etc. However, the specific hardware configuration of the information processing device 1 can be omitted, replaced, or added as appropriate depending on the embodiment. Furthermore, the information processing device 1 is not limited to a device consisting of a single housing. The information processing device 1 may be realized by multiple devices using so-called cloud or distributed computing technology, etc.
[0016] The operation center management system 3, the credit card management system 5, and the operation terminal are all computers equipped with a CPU, ROM, RAM, a storage device, a communication unit, an input device, an output device, etc. (not shown). Furthermore, these systems and terminals are not limited to devices consisting of a single housing. These systems and terminals may be realized by multiple devices using so-called cloud or distributed computing technology, etc.
[0017] 2 is a diagram showing an outline of the functional configuration of the information processing device 1 according to this embodiment. The information processing device 1 functions as an information processing device including an effect estimation unit 21, a correction function determination unit 22, a risk estimation unit 23, a machine learning unit 24, and a condition output unit 25, by reading a program recorded in a storage device 14 into a RAM 13 and executing it by a CPU 11, thereby controlling each piece of hardware included in the information processing device 1. Note that in this embodiment and other embodiments described below, each function included in the information processing device 1 is executed by the CPU 11, which is a general-purpose processor, but some or all of these functions may be executed by one or more dedicated processors.
[0018] The effect estimation unit 21 estimates the effect of a predetermined operation on a user to prompt the user to perform a predetermined action on whether or not the user will perform the action. In this embodiment, the predetermined action is payment of a credit card amount that is overdue. Note that the specific payment method is not limited, and may be a transfer to a designated account, payment at a designated counter, or the like. In this embodiment, the predetermined operation is a call to remind the user to pay the credit card amount that is overdue. Note that the call to the user may be an automated call using a recorded or mechanical voice, or a call in which an operator (human) speaks to the user. In this embodiment, the effect estimation unit 21 estimates the effect of the operation using a machine learning model that outputs a causality score that indicates the effect of the operation on the user in response to input of one or more user attributes related to the target user.
[0019] Here, to generate a machine learning model that outputs a causal score, for example, a machine learning framework based on ensemble learning and a machine learning framework based on gradient boosting decision trees can be used. However, although depending on the machine learning framework, training data, and other conditions, bias and variance in a machine learning model generally have a trade-off relationship, and when calculating a causal score, the output of a machine learning model tends to have high bias and low variance.
[0020] Therefore, the system according to this embodiment employs a configuration in which the causal score output from a relatively high-bias, low-variance model is corrected by referring to the output from a relatively low-bias, high-variance model that determines the causal score according to the actual recovery rate or default rate, thereby enabling more accurate scoring of the causal score.
[0021] 3 is a diagram showing an outline of a configuration in this embodiment in which the causal score output from a model with a relatively high bias and low variance is corrected by referring to the output from a model with a relatively low bias and high variance. (1) The first causal score S output from the relatively high-bias, low-variance model is corrected to a more accurate (relatively low-bias, low-variance) third causal score E; and (2) It can be seen that this correction uses a correction function that reduces (minimizes, if possible) the difference (e.g., L2 loss) between the second causal score Eh output from a relatively low-bias, high-variance model (base estimator) generated based on the label Y assigned to each user indicating actual collection status, and the first causal score S, which has a relatively high bias and low variance. Note that, in this embodiment, an example is described in which L2 loss is used to evaluate the error between the predicted value and the correct value, but the method of evaluating the error is not limited, and other evaluation methods may be adopted.
[0022] In order to perform the above-mentioned processing and obtain the third causal score E as the final output by the effect estimation unit 21, in this embodiment, the effect estimation unit 21 has a first effect estimation unit 21A that obtains a first causal score S with a relatively high bias and low variance, a second effect estimation unit 21B that obtains a second causal score Eh with a relatively low bias and high variance, and a third effect estimation unit 21C that corrects the first causal score to obtain a third causal score E with a relatively low bias (see Figure 2).
[0023] The first effect estimation unit 21A inputs the attributes of a user into the first model to obtain a first causal score S that indicates the effect of a predetermined operation on the user. As described above, the first model according to this embodiment has a relatively large bias, and therefore the first causal score S output from the first model is corrected by the third effect estimation unit 21C, which will be described later.
[0024] The second effect estimation unit 21B acquires a second causal score Eh indicating the effect by inputting the user's attributes into a second model that has a smaller bias and a larger variance than the first model. Here, a model that can be adopted as the second model may be any model that has a smaller bias and a larger variance than the first model, and the configuration of the model that can be adopted as the second model is not limited.
[0025] In this embodiment, an example will be described in which a model (histogram window estimator) that estimates a causal score using a histogram is adopted as the second model (base estimator). In this embodiment, the second model is generated by a method in which multiple users are assigned to multiple bins (here, n bins) in a histogram according to the attributes of each user, and a causal score corresponding to each user group assigned to the bin is calculated for each bin. The second effect estimation unit 21B then identifies a bin in the second model (histogram) corresponding to a target user based on the attributes of the target user whose causal score is to be estimated, and estimates the causal score calculated for the identified bin as the second causal score Eh for the target user. Therefore, unlike the correction function described below, the second model is neither a smooth function nor a monotonic function. However, the second model can obtain a low-bias second causal score Eh.
[0026] Here, the causal score corresponding to the user group assigned to each bin is generated by a method of calculating the causal score based on a statistical quantity relating to the rate of action execution by users who received an operation among multiple users having attributes corresponding to the bin, and a statistical quantity relating to the rate of action execution by users who did not receive the operation among these multiple users, based on a label Y that indicates whether or not a collection was actually performed, assigned to each user. In this embodiment, a default rate P1 is used as the statistical quantity relating to the rate of action execution by users who received an operation, and a default rate P0 is used as the statistical quantity relating to the rate of action execution by users who did not receive the operation, and the difference between these is taken as the second causal score Eh (Eh=P0-P1).
[0027] The third effect estimation unit 21C determines a third causal score E (E=F(S)) indicating the effect on the target user by applying the first causal score S calculated for the target user to a correction function F(x). Note that in the present disclosure, the processing content of the function used by the third effect estimation unit 21C is described using the expression calibration, but the processing content of the function may include regularization, normalization, modification, or the like.
[0028] The correction function determination unit 22 determines a correction function for correcting the first causality score S based on the first causality score S and the second causality score Eh calculated for each of the multiple users.
[0029] 4 is a diagram illustrating an outline of a process for determining a correction function using the first causality score and the second causality score in this embodiment. The correction function determination unit 22 determines the correction function and its parameters so that the difference (e.g., L2 loss) between the third causality score E and the second causality score Eh obtained when the first causality score S is applied to the correction function becomes smaller (minimized if possible) (L2 loss = min_{F}(E - Eh)^2), in other words, so that the third causality score E becomes closer to the second causality score Eh after correction.
[0030] Taking the linear function "F(x) = a*x + b" as an example, the correction function determination unit 22 determines the parameters a and b so that the difference (e.g., L2 loss) between the third causality score E and the second causality score Eh obtained when the first causality score S is applied to the correction function is smaller (minimized if possible) (L2 loss = min_{a, b}(aS + b - Eh)^2). Note that although the linear function "F(x) = a*x + b" is used here as an example of the correction function, the type of correction function and the number of parameters are not limited. For example, the correction function determination unit 22 may store multiple types of function formats in advance, identify an approximating function format based on the difference between the third causality score E and the second causality score Eh, and then search for parameters (coefficients) in the function format so that the loss is smaller. Furthermore, for example, the correction function determination unit 22 may refer to an approximation function of the first causal score S generated based on the first causal score S and an approximation function of the second causal score Eh generated based on the second causal score Eh, find a function that can convert the approximation function of the first causal score S into an approximation function of the second causal score Eh, and determine the function as the correction function.
[0031] In this embodiment, the correction function may be a smooth and monotonic function. Because the correction function is a smooth and monotonic function, the user's rank (rank based on the causal rank) determined based on the first causal score S can be maintained before and after application of the correction function, and the user's rank can be prevented from being reversed before and after correction. The smooth function may, for example, refer to a differentiable function, a continuously differentiable function, a k-th-order continuously differentiable function (k=1, 2, . . .), a function whose derivative is a continuous function, a piecewise differentiable function, or a continuously differentiable function with respect to a predetermined n. The monotonic function may, for example, refer to a monotonically increasing or monotonically decreasing function whose function value always increases or decreases as a variable increases or decreases, or a function whose function value increases or decreases as a variable increases or decreases within a predetermined range.
[0032] The risk estimation unit 23 estimates risk based on the probability that a user will not perform the predetermined action described above. The method of expressing risk is not limited, and various indicators may be employed. For example, risk can be expressed using the probability that a receivable will default without being collected. In this embodiment, the risk estimation unit 23 estimates risk based on the probability that a user will not perform an action using a machine learning model that outputs a risk indicator indicating the probability that a receivable will default without being collected for one or more user attributes related to a target user. However, risk may also be estimated according to, for example, a predetermined rule without using a machine learning model. For example, risk may be acquired by storing a value corresponding to each user attribute or combination of user attributes in advance and reading out this value. Furthermore, indicators other than the probability of default may be employed as indicators of risk. For example, risk may be classified (ranked), and the class (rank) of the risk may be used as an indicator.
[0033] The machine learning unit 24 generates and / or updates a machine learning model used for effect estimation by the effect estimation unit 21 and a machine learning model used for risk estimation by the risk estimation unit 23. The machine learning model for effect estimation is a machine learning model that, when data on one or more user attributes related to a target user is input, outputs a causal score indicating the effectiveness of an operation on the user. Furthermore, the machine learning model for risk estimation is a machine learning model that, when data on one or more user attributes related to a target user is input, outputs a risk index indicating the degree of risk based on the likelihood that the user will not perform an action. The user attributes input to these machine learning models may include, for example, demographic attributes, behavioral attributes, or psychographic attributes. Here, demographic attributes include, for example, the user's gender, family structure, age, etc.; behavioral attributes include, for example, whether or not a cash advance has been used, whether or not a revolving payment has been used, deposit and withdrawal history for a specified account, and commercial transaction history for some product including gambling or lotteries (which may include online transaction history in an online marketplace, etc.); and psychographic attributes include, for example, preferences for gambling or lotteries. However, usable user attributes are not limited to the examples given in this embodiment. For example, "time required for an operation (such as a phone call)" and "amount used by credit card" may also be used as attributes.
[0034] When generating and / or updating a machine learning model (a first model in this embodiment) for effect estimation, the machine learning unit 24 creates a machine learning model based on training data (machine learning data) in which, for each user attribute, a score based on a statistic related to the action execution rate (debt collection rate) of a user who received an operation among multiple users having a predetermined attribute and a statistic related to the action execution rate of a user who did not receive an operation among multiple users is defined as a causal score indicating the effect of the operation on the user having the attribute. In this embodiment, for example, if the content of the operation is a call to a user, a causal score is calculated based on the difference between the statistic values using the formula "(debt collection rate when the user made a call) - (debt collection rate when the user did not make a call)." The calculated causal score is combined with the attribute data of the corresponding user and input to the machine learning unit 24 as training data. Note that this embodiment describes an example in which an average value is used as the statistic. However, a statistical indicator such as a mode or a median may also be used as the statistic. Here, the statistic related to the action execution rate may be based on the past debt collection rate of each user within a predetermined period (e.g., a predetermined month). Furthermore, in this embodiment, a causal score is calculated based on the difference between each statistical quantity. However, if no statistically significant difference is found in the difference, the causal score may be set to zero or approximately zero. Here, existing statistical methods may be used to determine whether or not there is a significant difference. For example, standard error and confidence intervals may be considered for a set of average user response rates for each month, and the statistical significance of changes in response rates due to the presence or absence of calls may be considered. In this way, it is possible to calculate a causal score taking into account variations in average response rates among users within the same group.
[0035] When creating training data, a user who has been called once cannot become a user who has not been called thereafter. Therefore, for a user, it is possible to obtain only either the recovery rate when the user is called or the recovery rate when the user is not called. Therefore, training data on the effectiveness of calls is created for each user group having a common attribute. That is, a causal score indicating the effectiveness of calls on a user group consisting of users with a common attribute is obtained by, for example, dividing multiple users with the common attribute into a first sub-user group that is called and a second sub-user group that is not called, calculating the average debt recovery rate from the first sub-user group that is called and the average debt recovery rate from the second sub-user group that is not called, and calculating the difference between these average debt recovery rates based on the above-mentioned formula. For example, if the average debt recovery rate from the first sub-user group that is called is 80% and the average debt recovery rate from the second sub-user group that is not called is 70%, the causal score indicating the effectiveness of calls on that user group is "10."
[0036] A machine learning model generation framework that can be employed to implement the technology of the present disclosure is based on, for example, an ensemble learning algorithm. For example, a machine learning framework (e.g., LightGBM) based on a Gradient Boosting Decision Tree (GBDT) may be employed as the framework. In other words, the framework may be a machine learning framework based on a decision tree model that inherits the error between the correct answer and the predicted value between previous and next weak learners (weak classifiers). The predicted value here refers to, for example, a predicted value of a causality score or a risk index. Note that the framework may employ boosting methods such as LightGBM, XGBoost, or CatBoost. A framework using a decision tree can generate a machine learning model with relatively high performance with less parameter adjustment effort than a framework using a neural network. However, the machine learning model generation framework that can be employed to implement the technology of the present disclosure is not limited to the example described in this embodiment. For example, instead of a gradient boosting decision tree, other learning devices such as a random forest may be used as the learning device, or a learning device that is not a so-called weak learning device such as a neural network may be used. In particular, when a learning device that is not a so-called weak learning device such as a neural network is used, ensemble learning does not need to be used.
[0037] FIG. 5 is a simplified diagram illustrating the concept of a decision tree in a machine learning model employed in this embodiment. When employing a gradient boosting machine learning framework based on a decision tree algorithm, the branching conditions of each node in the decision tree are optimized. Specifically, in the gradient boosting machine learning framework based on a decision tree algorithm, a causal score indicating the effect of an operation is calculated for each user group having attributes indicated by two child nodes branching from a single parent node, and the branching condition of the parent node is optimized so that the difference between these causal scores is large (e.g., so that the difference is maximized or exceeds a predetermined threshold), i.e., so that the two child nodes branch cleanly. For example, if the attribute indicated as the branching condition of a node is age, the age set as the branching threshold may be changed, or the branching condition may be changed to an attribute other than age. In this way, by recursively optimizing the branching conditions of all nodes in the decision tree, the accuracy of estimating the effect of an operation can be improved.
[0038] When generating and / or updating a machine learning model for risk estimation, the machine learning unit 24 creates a machine learning model based on training data in which a statistical quantity (in this embodiment, an average value; however, a statistical indicator such as a mode or a median may also be used) related to the default incidence rate of multiple users having a predetermined attribute is defined as a risk index indicating the degree of risk of a user having the attribute. The calculated risk index is combined with the attribute data of the corresponding user and input to the machine learning unit 24 as training data. Furthermore, in generating or updating a machine learning model for risk estimation, there are no limitations on the machine learning model generation framework that can be adopted, but a gradient boosting machine learning framework based on a decision tree algorithm may be adopted, as in the case of generating and / or updating a machine learning model for effect estimation described above.
[0039] The condition output unit 25 determines and outputs conditions (hereinafter referred to as "operation conditions") for operations to be performed on users at the operation center based on the estimated effects and risks (in this embodiment, the causal scores and risk indices). In this embodiment, the operation conditions include at least one of whether the operation needs to be performed, the number of times the operation is to be performed, the order in which the operations are to be performed, the means of contacting the user in the operation, and the content of contacting the user in the operation. Examples of contact means include making a phone call or sending a message, and examples of the content of contact include the content to be conveyed to the user in the call and the content of the message.
[0040] 6 is a diagram showing the relationship between the estimated effects and risks and the operation conditions in this embodiment. Basically, the condition output unit 25 outputs operation conditions such that a higher priority is given to operations for users with higher estimated effects and at least some of the operations for users with higher estimated risks. Also, the condition output unit 25 outputs operation conditions such that a lower priority is given to operations for users with lower estimated effects and at least some of the operations for users with lower estimated risks.
[0041] Here, priority is a measure set for a user or an operation for that user, indicating the degree to which the operation is performed with priority. A user with a higher priority is more likely or frequently to receive the operation, while a user with a lower priority is less likely or frequently to receive the operation. Specifically, the condition output unit 25 outputs operation conditions, as conditions for assigning a higher priority, such as setting the operation to be performed as required, increasing the number of times the operation is performed, moving the order of execution of the operation earlier, or making the means or content of contact with the user for the operation more cost-effective or effective. Furthermore, the condition output unit 25 outputs operation conditions, as conditions for assigning a lower priority, such as setting the operation to be not performed, reducing the number of times the operation is performed, moving the order of execution of the operation later, or making the means or content of contact with the user for the operation less cost-effective or effective.
[0042] In this embodiment, the condition output unit 25 compares the estimated causal score and risk index with predetermined thresholds for the causal score and the risk index, respectively, and determines and outputs operation conditions according to the comparison results. Specifically, in the example shown in FIG. 6, a threshold C1 and a threshold C2 greater than the threshold C1 are set for the causal score, and a threshold R1 (first threshold) and a threshold R2 (second threshold) greater than the threshold R1 are set for the risk index. Furthermore, for the risk index, a threshold R3 is set to determine whether or not to set an operation condition according to the causal score, and a threshold R4 is set to determine a user to whom a high-priority operation condition is set regardless of the causal score. Note that in this embodiment, an example will be described in which the same value is used for the threshold R1 and the threshold R3, and the same value is used for the threshold R2 and the threshold R4 (see FIG. 6), but different values may be used for these thresholds.
[0043] Here, a case where an operation condition with a high priority is output will be described. Case (1): The condition output unit 25 determines and outputs high-priority operation conditions for at least some users whose causal scores are equal to or greater than the threshold C2 because the operation is highly effective. Case (2): Furthermore, the condition output unit 25 determines and outputs high-priority operation conditions for at least some of the users whose risk index is equal to or greater than the threshold R2, since the risk is high. Case (3): In particular, the condition output unit 25 may determine and output the highest priority operation condition for a user whose causal score is equal to or greater than the threshold C2 and whose risk index is equal to or greater than the threshold R2 (see the area UR indicated by the dashed line in Figure 6). Case (4): However, for users whose estimated risk index is lower than the threshold R3, the condition output unit 25 does not need to output operation conditions that are given high priority because the user is unlikely to become a default in the first place. To enhance the cost containment effect, it is preferable to determine and output low-priority or medium-priority operation conditions for users whose risk index is lower than the threshold R3, regardless of the causal score (the region UL indicated by the dashed line in FIG. 6 has a causal score of C2 or higher but a risk index lower than R3, and therefore does not become a high-priority operation condition). Therefore, the effect estimation unit 21 may estimate the effect of an operation for users whose risk index is estimated to be equal to or higher than the threshold R3 (third threshold), and may not estimate the effect of an operation for users whose risk index is estimated to be lower than the threshold R3 (third threshold) (omit the estimation process).
[0044] A case where an operation condition with a low priority is output will also be described. Case (5): For at least some users whose causal scores are less than the threshold C1, the effect of the operation is low, so the condition output unit 25 determines and outputs an operation condition with a low priority. Case (6): Furthermore, the condition output unit 25 determines and outputs low-priority operation conditions for at least some of the users whose risk index is less than the threshold R1 because the risk is low. Case (7): In particular, the condition output unit 25 may determine and output the lowest priority operation condition for a user whose causal score is less than the threshold C1 and whose risk index is less than the threshold R1 (see the area LL indicated by the dashed line in Figure 6). Case (8): However, for a user whose estimated risk index is equal to or greater than the threshold R4, the condition output unit 25 does not need to output an operation condition to which a low priority is assigned, since the user is likely to be defaulted in the first place. Rather, for a user whose risk index is equal to or greater than the threshold R4, it is preferable to determine and output an operation condition with a high priority, regardless of the causal score (the region LR indicated by the dashed line in FIG. 6 has a causal score less than C1, but the risk index is equal to or greater than R4, so the operation condition does not have a low priority).
[0045] As a result, in the example shown in Figure 6, the operation priority is increased for high-risk users and medium-risk and high-effectiveness users, and decreased for low-risk users and medium-risk and low-effectiveness users. Note that in the above example, an example was described in which areas are divided based on thresholds and common operation conditions are determined for users belonging to one area, but different operation conditions may be set for each user or for each user within one area. For example, even within the same area, the operation conditions may be varied to provide a gradation depending on the level of the causal score and / or risk index.
[0046] The amount of operations (total number of operations and number of users subject to operations) by the operation center management system 3 can be changed by adjusting the thresholds described above. For example, the amount of operations can be increased by lowering at least one of thresholds C1, C2, R1, and R2, and the amount of operations can be decreased by raising at least one of thresholds C1, C2, R1, and R2.
[0047] <Processing flow> Next, a flow of processing executed by the information processing system according to this embodiment will be described. Note that the specific content and processing order of the processing described below are an example for implementing the present disclosure. The specific content and processing order may be selected as appropriate depending on the embodiment of the present disclosure.
[0048] 7 is a flowchart showing the flow of the machine learning process according to this embodiment. The process shown in this flowchart is executed at a timing designated by the administrator of the operation center management system 3.
[0049] In steps S101 and S102, a machine learning model used for effect estimation is generated and / or updated. The machine learning unit 24 calculates a causal score for each of a plurality of user attributes based on user attribute data, operation history data, and credit card usage payment history data previously accumulated in the operation center management system 3 or the credit card management system 5, and creates training data including combinations of user attributes and causal scores (step S101). Here, the operation history data includes data that can be used to determine, for each user, whether an operation has been performed on that user, and the payment history data includes data that can be used to determine, for each user, whether that user has paid their credit card usage amount (whether or not there is default). Then, the machine learning unit 24 inputs the generated training data into the machine learning model and generates or updates a machine learning model (first model) used for effect estimation by the first effect estimation unit 21A (step S102). The process then proceeds to step S103.
[0050] In steps S103 and S104, a machine learning model used for risk estimation is generated and / or updated. The machine learning unit 24 calculates a risk index for each of multiple user attributes based on user attribute data, operation history data, and credit card usage payment history data previously accumulated in the operation center management system 3 or the credit card management system 5, and creates training data including combinations of user attributes and risk indexes (step S103). Then, the machine learning unit 24 inputs the created training data into the machine learning model, and generates or updates the machine learning model used for risk estimation by the risk estimation unit 23 (step S104). Thereafter, the processing shown in this flowchart ends.
[0051] 8 is a flowchart showing the flow of the correction function determination process according to this embodiment. The process shown in this flowchart is executed at a timing designated by the administrator of the operation center management system 3.
[0052] In step S901, a first causal score S is acquired. The first effect estimation unit 21A inputs user attribute data for each of a plurality of users to a first model generated and / or updated in the machine learning process described with reference to FIG. 7, and acquires a first causal score S corresponding to each of the plurality of users as an output from the first model. Here, the first model is a model trained using learning data from a certain period of time or more in the past (for example, up to last month), and the first causal score S, which is an output from the model, is the output of a model trained using learning data from a certain period of time or more in the past. Thereafter, the process proceeds to step S902.
[0053] In steps S902 and S903, a second causal score Eh is acquired. The information processing device 1 generates or updates a second model by allocating multiple users to multiple bins in a histogram according to the attributes of each user and calculating, for each bin, a causal score corresponding to the user group assigned to that bin (step S902). Here, fact data (e.g., a label indicating the presence or absence of a default) for the latest period (for example, this month) is used to generate or update the second model. When the second model is generated or updated, the second effect estimation unit 21B inputs the attributes of multiple users into the second model to acquire a second causal score Eh corresponding to each of the multiple users (step S903). Therefore, the second causal score Eh is output based on the fact data for the latest period. The process then proceeds to step S904.
[0054] In step S904, a correction function is determined. The correction function determination unit 22 determines a correction function and parameters that reduce (minimize, if possible) the L2 loss between the first causality score S and the second causality score Eh for multiple users. As described above, the first causality score S is the output of a model trained using learning data from a certain period of time or more in the past, and the second causality score Eh is the output based on fact data from the most recent period. Therefore, the correction function can correct the first causality score S estimated based on data from the past (e.g., up to last month) based on more accurate current fact data (e.g., this month). Then, the processing shown in this flowchart ends.
[0055] 9 is a flowchart showing the flow of the third causal score estimation process according to this embodiment. The process shown in this flowchart is executed at a preset timing every month. More specifically, the execution timing of the process is set to be after the specified date for payment of the credit card usage amount and before the scheduled date for execution of the operation for the unpaid user.
[0056] In step S1001, a first causal score S is acquired. For each of a plurality of users, the first effect estimation unit 21A inputs data of one or more user attributes related to the target user into a first model generated and / or updated in the machine learning process described with reference to Fig. 7, and acquires a first causal score S corresponding to the user as an output from the first model. Then, the process proceeds to step S1002.
[0057] In steps S1002 and S1003, a causal score to be used for outputting an operation condition is determined. The third effect estimation unit 21C determines a third causal score E (E=F(S)) for the target user by applying the first causal score S calculated for the target user in step S1001 to a correction function F(x) (step S1002). The effect estimation unit 21 determines the third causal score E calculated in step S1002 as the causal score to be used for outputting an operation condition for the target user in an operation condition output process described later with reference to FIG. 10 (step S1003). Thereafter, the process shown in this flowchart ends.
[0058] 10 is a flowchart showing the flow of the operation condition output process according to this embodiment. The process shown in this flowchart is executed at a preset time each month. More specifically, the execution timing of the process is set to be after the specified date for payment of the credit card usage amount and before the scheduled date for execution of the operation for the unpaid user.
[0059] In steps S201 and S202, a risk is estimated based on the effect of the operation and the probability that the user will not perform the action. For each of a plurality of users, the risk estimation unit 23 inputs data on one or more user attributes related to the target user into the machine learning model generated and / or updated in step S104, and acquires a risk index corresponding to the user as an output from the machine learning model (step S201). In addition, the effect estimation unit 21 executes the third causal score estimation process described with reference to FIG. 9 to acquire a causal score corresponding to the target user for each of a plurality of users (step S202). Thereafter, the process proceeds to step S203.
[0060] In step S203, operation conditions are determined and output. The condition output unit 25 determines operation conditions based on the causal scores and risk indexes estimated in steps S201 and S202, and outputs them to the operation center management system 3. In this embodiment, the condition output unit 25 identifies and outputs operation conditions that have been pre-mapped to the causal scores and risk indexes. However, the method for determining operation conditions is not limited to the example given in this embodiment. For example, the operation conditions may include values calculated by inputting the causal scores and risk indexes into a predetermined function. Thereafter, the processing shown in this flowchart ends.
[0061] When the operation conditions are output, the operation center management system 3 manages the operation for the target user in accordance with the operation conditions, and the operation terminal executes the operation in accordance with the instructions output by the operation center management system 3.
[0062] <Effects> According to this embodiment, more accurate scoring can be achieved by correcting the estimated result of the causal score. Furthermore, according to the present invention, by using a correction function that is simpler than the machine learning model itself, it becomes possible to respond to sudden behavioral changes, and a scoring technology that can quickly respond to changes in user behavior can be realized.
[0063] Furthermore, according to this embodiment, by setting priorities of operation conditions according to the effectiveness and risk of operations for each user and suppressing operations for users with low effectiveness or low risk, it is possible to suppress costs related to operations without reducing the debt collection rate. In other words, according to the present disclosure, it is possible to suppress costs for operations without reducing the effectiveness of operations that encourage users to take predetermined actions. Furthermore, according to this embodiment, by increasing operations for users with high effectiveness or high risk, it is expected that the debt collection rate will be increased while suppressing costs.
[0064] <Variations of operation condition output processing> In the embodiment described above, the flow of the operation condition output process has been roughly described with reference to FIG. 10, but more specifically, the operation condition output process may be processed as follows.
[0065] Fig. 11 is a flowchart showing the flow of the operation condition output process when the determination methods of cases (1) to (4) described with reference to Fig. 6 are adopted in this embodiment. According to the example shown in Fig. 11, when the risk index calculated by the risk estimation unit 23 (step S301) for a user having a certain attribute is less than the threshold R3 (third threshold) (NO in step S302), the calculation of the causal score by the effect estimation unit 21 is omitted, and an operation condition with low (or medium) priority is determined and output (step S303).
[0066] If the risk index calculated for the user is equal to or greater than threshold R3 (third threshold) (YES in step S302), a causal score is calculated (step S304). If the causal score is equal to or greater than threshold C2 and the risk index is equal to or greater than threshold R2 (YES in step S305), an operation condition with the highest priority is determined and output (step S306). If the causal score is equal to or greater than threshold C2 or the risk index is equal to or greater than threshold R2 (YES in step S307), an operation condition with a high priority is determined and output (step S308). Note that if the causal score is less than threshold C2 and less than threshold R2, an operation condition with a medium priority is determined and output (step S309).
[0067] Fig. 12 is a flowchart showing the flow of the operation condition output process when the determination methods of cases (5) to (8) described with reference to Fig. 6 are adopted in this embodiment. According to the example shown in Fig. 12, when the risk index calculated by the risk estimation unit 23 (step S401) for a user having a certain attribute is equal to or greater than the threshold R4 (NO in step S402), the calculation of the causal score by the effect estimation unit 21 is omitted, and an operation condition with a high (or medium) priority is determined and output (step S403).
[0068] If the risk index calculated for the user is less than threshold R4 (YES in step S402), a causal score is calculated (step S404), and if the causal score is less than threshold C1 and the risk index is less than threshold R1 (YES in step S405), an operation condition with the lowest priority is determined and output (step S406), and if the causal score is less than threshold C1 or the risk index is less than threshold R1 (YES in step S407), an operation condition with a low priority is determined and output (step S408).Note that if the causal score is equal to or greater than threshold C1 and threshold R1, an operation condition with a medium priority is determined and output (step S409).
[0069] <Other variations> In the above-described embodiment, an example was described in which the operation to the user was a call, but the type of operation to the user is not limited to a call. For example, sending a message may be adopted as the type of operation to the user. Note that the means for sending the message is not limited here, and an email system, a short message service (SMS), a message sending and receiving service of a social networking service (SNS), or the like may be used.
[0070] Furthermore, in the above-described embodiment, an example has been described in which an operation condition is determined based on an effect estimated for one type of operation (a call), but the operation condition may be determined based on an effect estimated for each of multiple types of operations (e.g., a call and sending a message). In this case, the effect estimation unit 21 estimates a first effect that a first operation (e.g., a call) to a user to prompt the user to perform a predetermined action has on whether the user will perform the action, and a second effect that a second operation (e.g., sending a message) to the user to prompt the user to perform the predetermined action has on whether the user will perform the action, and the condition output unit 25 outputs an operation condition for the user based on the estimated first effect and second effect.
[0071] In this case, a machine learning model for estimating the effect of an operation is also generated and updated for each type of operation. For example, if the first operation is making a call and the second operation is sending a message, a machine learning model for estimating the effect of making a call and a machine learning model for estimating the effect of sending a message may be generated and updated.
[0072] Furthermore, when operation conditions are determined based on the estimated effects of each of multiple types of operations, an operation with a high effect on the user may be selected from the multiple types of operations. In this case, the condition output unit 25 outputs operation conditions including whether to perform the first operation or the second operation on the user based on the estimated first and second effects. More specifically, the first effect (causal score related to the first operation) and the second effect (causal score related to the second operation) obtained for the target user may be compared, and the type of operation with the higher causal score may be selected as the type of operation with a high effect on the user.
[0073] Furthermore, in the above-described embodiment, an example was described in which two axes, a causal score indicating the effectiveness of the operation and a risk index indicating the risk, were used as evaluation axes for determining operation conditions, but in the technology disclosed herein, the evaluation axes for determining operation conditions need only include at least the effectiveness of the operation, and other evaluation axes may be adopted, or three or more evaluation axes may be adopted. For example, (1) a third index other than the risk index may be adopted as an index to be combined with the causal score, (2) a third index may be adopted in addition to the causal score and risk index, or (3) three axes, a causal score for calling, a causal score for sending a message, and a risk index, may be adopted.
[0074] Furthermore, in the above-described embodiment, an example was described in which the difference between the action execution rate of a sub-user group that received an operation and the action execution rate of a sub-user group that did not receive an operation was used as the causal score. However, the causal score may be calculated using other methods. For example, when generating and / or updating a machine learning model for effect estimation, the machine learning unit 24 may create a machine learning model based on training data in which, for each user attribute, a score based on statistics related to the action execution rate of users who performed a predetermined reaction to an operation among multiple users having the predetermined attribute and statistics related to the action execution rate of users who did not perform the predetermined reaction is defined as a causal score related to a user having the attribute. For example, in this variation, the causal score is calculated using the formula "(the debt collection rate when the user performed the predetermined reaction) - (the debt collection rate when the user did not perform the predetermined reaction)." In this case, the condition output unit 25 may output a condition related to the operation based on the presence or absence of a reaction by the user, the content of the reaction, etc., and the causal score corresponding to the reaction. At this time, the adjustment of the priority of the operation condition output by the condition output unit 25 may be performed using the relationship between the causal score and the priority described with reference to FIG.
[0075] Here, the predetermined reaction may be, for example, a user's response to a call by dialing, a conversation with an operator when the user calls back, a reply to a message, or marking a message as read. Furthermore, the content of the reaction may also be considered to be a positive response to payment, or whether or not a response was given regarding the due date of payment. If the user's reaction is a voice reaction, it is also possible to determine the user's emotions, etc., based on the user's voice and determine whether the reaction was positive. Furthermore, the priority of the next operation condition may be adjusted based on the user's emotions, etc., determined based on the voice.
[0076] The priority of the operation conditions may also be adjusted based on factors other than those described above. For example, the priority of the operation conditions may be higher for a user whose credit card payment settings are revolving payments or installment payments than for a user who pays in full, and the priority of the operation conditions may be higher for a user whose credit card usage includes cash advances than for a user whose credit card usage does not include cash advances. The priority of the operation conditions may also be adjusted based on the target user's transaction data other than credit card data (e.g., balance data of the credit card usage debit account, transaction history data at affiliated banks, etc.). Here, the condition output unit 25 may adjust the various thresholds corresponding to the various scores shown in FIG. 6 based on, for example, the credit card usage conditions. [Explanation of symbols]
[0077] 1. Information processing equipment
Claims
1. a first effect estimation means for inputting a user's attributes into a first model to obtain a first causal score indicating an effect of a predetermined operation on the user; a second effect estimation means for acquiring a second causal score indicating the effect by inputting the attributes of the user into a second model; a correction function determination means for determining a correction function for correcting the first causal score based on the first causal score and the second causal score calculated for each of a plurality of users; a third effect estimation means for determining a third causal score indicating the effect for the target user by applying the first causal score calculated for the target user to the correction function; the correction function determination means determines the correction function such that a difference between the third causal score and the second causal score obtained when the first causal score is applied to the correction function becomes smaller; the causality score is a score calculated based on a difference between a statistic related to a rate at which a predetermined action is performed by a user who has received the operation among a plurality of users and a statistic related to a rate at which the action is performed by a user who has not received the operation among the plurality of users; Information processing device.
2. the second effect estimation means acquires a second causal score indicating the effect by inputting the attributes of the user into the second model, the second model having a smaller bias and a larger variance than the first model; The information processing device according to claim 1 .
3. the correction function determining means determines a function that is a smooth function and a monotonic function as the correction function. The information processing device according to claim 1 .
4. the first effect estimation means estimates the first causality score using a machine learning model generated using a machine learning framework based on ensemble learning; The information processing device according to claim 1 .
5. the first effect estimation means estimates the first causal score using a machine learning model generated using a machine learning framework based on a gradient boosting decision tree; The information processing device according to claim 1 .
6. the effect is an effect that a predetermined operation for a user to prompt the user to perform a predetermined action has on whether or not the user performs the action; the first model is created based on training data in which a score based on a statistic relating to the rate of execution of the action by a user who has received the operation among a plurality of users having a predetermined attribute and a statistic relating to the rate of execution of the action by a user who has not received the operation among the plurality of users is defined as a score indicating the effect of the operation on the user who has the attribute; The information processing device according to claim 1 .
7. the second effect estimation means assigns a plurality of users to each of a plurality of bins in a histogram according to an attribute of each user, and in the second model generated by a method of calculating, for each bin, a causal score corresponding to a group of users assigned to the bin, identifies a bin to which a target user corresponds, and estimates the causal score calculated for the identified bin as a second causal score for the target user; The information processing device according to claim 1 .
8. the effect is an effect that a predetermined operation for a user to prompt the user to perform a predetermined action has on whether or not the user performs the action; the second model is generated by a method of calculating, for each bin, a causal score corresponding to a user group assigned to the bin, based on statistics relating to the rate at which the action is performed by users who have received the operation among a plurality of users having a predetermined attribute, and statistics relating to the rate at which the action is performed by users who have not received the operation among the plurality of users; The information processing device according to claim 7 .
9. A condition output means for outputting conditions regarding the operation for the user based on the estimated effect, further comprising condition output means for outputting conditions such that a higher priority is given to the operation for a user for whom the estimated effect is higher. The information processing device according to claim 1 .
10. The computer a first effect estimation step of inputting a user's attributes into a first model to obtain a first causal score indicating an effect of a predetermined operation on the user; a second effect estimation step of acquiring a second causal score indicating the effect by inputting the user's attributes into a second model; a correction function determination step of determining a correction function for correcting the first causal score based on the first causal score and the second causal score calculated for each of a plurality of users; a third effect estimation step of determining a third causal score indicating the effect for the target user by applying the first causal score calculated for the target user to the correction function; In the correction function determination step, a correction function is determined that reduces a difference between the third causal score and the second causal score obtained when the first causal score is applied to the correction function; the causality score is a score calculated based on a difference between a statistic related to a rate at which a predetermined action is performed by a user who has received the operation among a plurality of users and a statistic related to a rate at which the action is performed by a user who has not received the operation among the plurality of users; method.
11. Computer, a first effect estimation means for inputting a user's attributes into a first model to obtain a first causal score indicating an effect of a predetermined operation on the user; a second effect estimation means for acquiring a second causal score indicating the effect by inputting the attributes of the user into a second model; a correction function determination means for determining a correction function for correcting the first causal score based on the first causal score and the second causal score calculated for each of a plurality of users; and a third effect estimation means for determining a third causal score indicating the effect for the target user by applying the first causal score calculated for the target user to the correction function. the correction function determination means determines the correction function such that a difference between the third causal score and the second causal score obtained when the first causal score is applied to the correction function becomes smaller; the causality score is a score calculated based on a difference between a statistic related to a rate at which a predetermined action is performed by a user who has received the operation among a plurality of users and a statistic related to a rate at which the action is performed by a user who has not received the operation among the plurality of users; program.
Citation Information
Patent Citations
Method and system for managing overdue credit
JP2001282994A
System for distributing telephony work and its program
JP2007094998A
Large sum user management program, large sum user management system and large sum user management method
JP2007241714A
Reminder support system, method thereof, and program
JP2010224617A
Information processing device, prediction system, information processing method, and program
JP2020021151A