Examination apparatus, method, and program
The review apparatus uses dual machine learning models to assess users with limited credit data, enhancing examination accuracy by segmenting users and correlating scores, addressing the challenge of insufficient attribute data in conventional methods.
Patent Information
- Application Number
- JP2022100485
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-06-22
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-06-22
AI Technical Summary
Existing techniques for user examination in financing and credit transactions struggle to appropriately assess users with insufficient accumulation of attribute data.
A review apparatus utilizing two machine learning models to acquire first and second scores, identify user segments, and determine review results based on the correlation between these scores, even for users with limited credit information.
Enables accurate examination of users with insufficient credit data by leveraging multiple score correlations, improving assessment accuracy for younger generations and diverse transaction types.
Smart Images

Figure 0007713428000001 
Figure 0007713428000002 
Figure 0007713428000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technique for assisting in the examination of users.
Background Art
[0002] Conventionally, a financing information providing apparatus has been proposed that receives notifications such as transaction information of a store from a terminal device, inputs the transaction information and the like into a learned model to obtain a classification result, calculates a financing score based on this classification result, and then generates financing information based on the financing score and the financing condition information of each financial institution, and notifies the terminal device of the financing information of each financial institution and the recommendation information to be notified to the merchant (see Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventionally, in the examination for financing, credit transactions, etc. for users, the determination of the user's credit score using a machine learning model has been emphasized from the viewpoint of optimizing the examination. Also, conventionally, a technique has been proposed for selecting, as recommendation information, a combination of input data for a model that increases the financing score (see Patent Document 1). However, although the conventionally proposed techniques have a certain effect from the viewpoint of optimizing the examination, there is room for improvement in appropriately examining users with insufficient accumulation of attribute data such as credit information.
[0005] In view of the above problems, an object of the present disclosure is to provide an appropriate examination even for users with insufficient accumulation of attribute data.
Means for Solving the Problems
[0006] An example of the present disclosure includes: a first score acquisition means for acquiring a first score of a user based on an output obtained by inputting input data related to the user into a first machine learning model; a second score acquisition means for acquiring a second score of the user based on an output obtained by inputting the input data related to the user into a second machine learning model; a user segment identification means for identifying a user segment to which the target user belongs by performing segmentation of a group of users including the target user; and a review result determination means for determining a review result of the target user based on the first score and / or the second score of the target user according to the user segment to which the target user belongs. The disclosure relates to a review apparatus comprising these means.
[0007] The present disclosure can be understood as a review apparatus, a system, a method executed by a computer, or a program to be executed by a computer. Further, the present disclosure can also be understood as a recording medium in which such a program is recorded and can be read by a computer or other devices, machines, etc. Here, a computer-readable recording medium refers to a recording medium that accumulates information such as data and programs by an electrical, magnetic, optical, mechanical, or chemical action and can be read by a computer or the like.
Advantages of the Invention
[0008] According to the present disclosure, it is possible to provide an appropriate review even for a user with insufficient accumulation of attribute data.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Mode for Carrying Out the Invention
[0010] Hereinafter, embodiments of a review apparatus, method, and program according to the present disclosure will be described with reference to the drawings. However, the embodiments described below are merely illustrative of the embodiments, and do not limit the review apparatus, method, and program according to the present disclosure to the specific configurations described below. In practice, a specific configuration according to the implementation mode may be appropriately adopted, and various improvements and modifications may be made. In the technology according to the present disclosure, at least a part of the configurations in each of the embodiments and variations described later can be appropriately adopted with each other.
[0011] In the present embodiment, an embodiment in which the technology according to the present disclosure is implemented in a system that provides a review of financing for a user will be described. However, the technology according to the present disclosure can be widely used for a technology for supporting any review regarding a user, and the application target of the present disclosure is not limited to the example shown in the embodiment.
[0012] <Configuration of the System> FIG. 1 is a schematic diagram showing the configuration of an information processing system according to the present embodiment. In the information processing system according to the present embodiment, a review apparatus 1 and one or a plurality of service providing systems 5 are connected to each other so as to be communicable. The user is a user of the service provided by the service providing system 5, and receives the service by accessing the service providing system 5 from the user terminal.
[0013] The examination device 1 is a computer including a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, a storage device 14 such as an EEPROM (Electrically Erasable and Programmable Read Only Memory) or an HDD (Hard Disk Drive), a communication unit 15 such as a NIC (Network Interface Card), and the like. However, with regard to the specific hardware configuration of the examination device 1, appropriate omissions, replacements, and additions are possible according to the embodiments. Also, the examination device 1 is not limited to a device consisting of a single housing. The examination device 1 may be realized by a plurality of devices using technologies such as so-called cloud or distributed computing.
[0014] The examination device 1 performs an examination of transactions related to each user and provides the examination result to the service providing system 5. The service providing system 5 can determine whether to provide services such as financing (whether to conduct a transaction) to the user (hereinafter, the "target user") according to the result of the examination targeting an arbitrary user provided from the examination device 1.
[0015] The service providing system 5 is a computer including a CPU, a ROM, a RAM, a storage device, a communication unit, an input device, an output device, etc. (illustrations are omitted). Also, none of these systems and terminals are limited to devices consisting of a single housing. These systems and terminals may be realized by a plurality of devices using technologies such as so-called cloud or distributed computing.
[0016] The services provided by the service providing system 5 are, for example, financing review services such as loans, credit card / installment payment services, online shopping services, online reservation services, electronic money payment services, operation center services, or map information services, etc. Note that "installment payment" is not limited to the service called so-called Buy Now Pay Later (BNPL), and may include all purchases of goods / services by installment payment.
[0017] The services provided by the service providing system 5 are not limited to the examples in this embodiment. And the service providing system 5 notifies the review device 1 of the user's attribute data when providing the service. Here, the user's attribute data includes the service usage history data by the user. The content of the service usage history data varies according to the content of the service. For example, it may include the history data of the user's location information, the payment history data of the credit card usage amount / installment payment usage amount, the electronic money usage history data, the transaction history data (including the purchase history data of goods, etc.), the reservation history data, the operation history data for the user from the operation center, etc.
[0018] Here, the attribute data includes data represented in a data format such as a score (for example, a continuous value from 0 to 1) or a label (for example, a binary value according to presence / absence or yes / no). However, the format of the attribute data is not limited to the examples in this disclosure. Also, the attribute data may include, for example, the cancellation rate in online services such as the online service usage status, the usage status of electronic values including points, etc., and may also include data indicating default (default in debt repayment) in installment payment.
[0019] The attribute data may include factual attribute data and inferred attribute data. Here, the factual attribute data is attribute data that can be confirmed as facts about the user based on user-provided data obtained by being provided by the user himself / herself or historical data collected about the user, etc. Also, the inferred attribute data is attribute data obtained by being inferred by methods such as inputting user-provided data, historical data, factual attribute data, etc. into a VAE (Variational Autoencoder).
[0020] Note that the VAE used here is one that expresses the values input to the encoder (here, training data including factual attribute data) in a different form by inputting the latent vector output by the encoder in the first half of the VAE into the decoder in the second half of the VAE. Examples of user-provided data include registration data such as the name, email address, phone number, address, workplace, school attended, etc. registered by the user himself / herself, and data obtained as a result of the user's own answers to questionnaires, etc. Examples of historical data include, for example, the usage history data of the e-commerce service provided by the service providing system 5 described above. The factual attribute data is preferably data obtained by converting the aforementioned user-provided data and historical data into a data format suitable for marketing and / or analysis purposes. For example, as factual attribute data that can be obtained based on usage history data, in addition to the genre / category and brand of products / services frequently used by the user, commercial areas, pleasure spots, tourist spots, etc. frequently visited by the user can be mentioned. That is, in the present embodiment, the attribute data includes the personality, behavior tendency, user persona, etc. of the user estimated or predicted using machine learning techniques.
[0021] In this embodiment, weights are set for each piece of attribute data. The weight indicates the degree of correlation between the attribute data and the score when the attribute data is used in calculating the score, and each time the appropriateness of the score is evaluated by the machine learning unit 27 described later, the model parameters are adjusted so that the score becomes a more appropriate value. The weight corresponding to each piece of attribute data corresponds, for example, to the weight corresponding to each node (each regression tree) in a model for calculating a score such as a decision tree model described later, and is determined as appropriate in the process of calculating the score. Note that the score is determined based on, for example, the weights of the nodes.
[0022] Here, the attribute data group may include demographic attributes, behavioral attributes, or psychographic attributes. Demographic attributes are, for example, the gender of the user, family composition, age, etc. Behavioral attributes may be based on service usage history data, and for example, whether cash withdrawal is used, whether direct debit is used, the deposit and withdrawal history related to a predetermined account, the commercial transaction history related to any product / service including gambling or lottery (including the online transaction history in an online marketplace, etc.), the movement history of the user using location information and place information, etc. Psychographic attributes are, for example, the preferences related to gambling or lottery. However, the available user attributes are not limited to the examples in this embodiment. For example, "the time required for operation (such as making a call)" and "the amount of credit card usage / amount of post-payment settlement usage" from an operation center service, etc. may also be used as attributes. Note that attributes similar to demographic attributes may be attributes estimated based on attributes based on user-provided data or history data. Similarly, attributes similar to behavioral attributes may be attributes estimated based on attributes based on user-provided data or history data. Psychographic attributes may be attributes based on user-provided data including, as an example, the result of a user's voluntary input.
[0023] FIG. 2 is a diagram showing an outline of the functional configuration of the review apparatus 1 according to the present embodiment. In the review apparatus 1, a program recorded in the storage device 14 is read into the RAM 13 and executed by the CPU 11, and each hardware provided in the review apparatus 1 is controlled, so that an attribute determination unit 21, a first score acquisition unit 22, a second score acquisition unit 23, a user segment identification unit 24, a correlation determination unit 25, a review result determination unit 26, and a machine learning unit 27 are provided. In the present embodiment and the variations described later, each function provided in the review apparatus is executed by the CPU 11, which is a general-purpose processor, but some or all of these functions may be executed by one or more dedicated processors.
[0024] The attribute determination unit 21 determines fact attribute data that can be confirmed as a fact about the user based on user-provided data provided by the user himself / herself and / or the history data of the user. In the present embodiment, the attribute determination unit 21 uses methods such as aggregating user-provided data and / or history data, determining the corresponding attributes by referring to other data such as maps, and using the user-provided data and / or history data as they are, to determine the fact attribute data related to the user. In the present embodiment, a method of determining the fact attribute data related to the user based on the user-provided data and / or the history data of the user is adopted, but the fact attribute data related to the user may be acquired by other methods.
[0025] In addition, the attribute determination unit 21 determines the estimated attribute data estimated for the user based on the input data including at least one or more fact attribute data determined for the user. In the present embodiment, the attribute determination unit 21 may determine the estimated attribute data based on the output value obtained by inputting the input data including one or more fact attribute data related to the user into an attribute estimation model which is a machine learning model. Note that, in the present embodiment, the output value from the attribute estimation model is a value indicating the probability that the user has a predetermined attribute, and when the output value obtained from the attribute estimation model is within a predetermined range, the attribute determination unit 21 determines that the user has the attribute. When it is determined that the user has a predetermined attribute, the attribute determination unit 21 sets the label of the attribute data estimated for the user to a value indicating the presence or absence of the attribute or the type of the attribute. Further, the attribute data may be indicated by a score instead of a label. In this case, the attribute determination unit 21 sets a value indicating the degree (probability) to which the estimated attribute can be applied to the score of the estimated attribute data estimated for the user. The degree may be the output value of the attribute estimation model.
[0026] The first score acquisition unit 22 acquires the first score of the user based on the output (for example, a score normalized / standardized with 0 as the minimum value and 1 as the maximum value) obtained by inputting the input data related to the user into the first machine learning model. More specifically, the first score acquisition unit 22 acquires the first score of the user based on the output obtained by inputting the input data related to the user into the first machine learning model generated by learning the relationship between the attributes of the user and the creditworthiness of the user in the transactions of the first category (transactions of the first type). In the present disclosure, the transactions of the first category refer to transactions related to a period of a predetermined length and / or an amount of a predetermined amount or more. That is, examples of the transactions of the first category include relatively long-term and / or high-amount transactions such as long-term loans or high-amount loans.
[0027] In this embodiment, the first machine learning model is generated and / or updated using teacher data that takes, as input values, attribute data including the situations of other transactions of the same type as the transactions in the first category (for example, other long-term loans or high-value loans, etc.), and outputs, as output values, values (first scores) indicating the creditworthiness of the user. And the first machine learning model is a machine learning model for review that outputs, in transactions of the first category, whether financing can be provided or a review result corresponding to the default probability. Therefore, the types of input data in the previous high-value financing review model may be appropriately adopted for the first machine learning model.
[0028] The input data to the first machine learning model includes a group of attribute data including the above-described factual attribute data and / or estimated attribute data. As examples, the input data includes annual income, industry type of the workplace, date of birth, living situation (length of residence (e.g., in months), length of service (e.g., in months), borrowing / repayment situation (such as other loans with existing contracts), etc.).
[0029] The second score acquisition unit 23 acquires the second score of the user based on the output (for example, a score normalized / standardized with 0 as the minimum value and 1 as the maximum value) obtained by inputting the input data related to the user into the second machine learning model. More specifically, the second score acquisition unit 23 inputs the input data related to the user into the second machine learning model generated by learning the relationship between the attributes of the user and the risk of the user in transactions of a second category different from the transactions of the first category (second type of transactions), and acquires the second score of the user based on the output obtained thereby. In the present disclosure, the transactions of the second category refer to transactions related to a shorter period and / or a smaller amount compared to the transactions of the first category. That is, examples of the transactions of the second category include relatively short-term and / or small-amount transactions such as deferred payment settlement, credit card use, short-term loans, or small-amount loans.
[0030] In this embodiment, the second machine learning model uses, as input values, attribute data including the situations of other transactions of the same type as the transactions in the second category (for example, other deferred payments, credit card usage, short-term loans, small loans, etc.), and is generated and / or updated using teacher data with an output value being a value (second score) indicating the risk based on the probability of default for the user without making a payment (without the debt being recovered for the user). The method of expressing the risk is not limited, and various indicators may be adopted. For example, the risk can be expressed using the probability of default. Also, as the indicator indicating the risk, an indicator other than the probability of default may be adopted. For example, the risk may be classified (ranked), and which class (rank) it is can be used as an indicator. That is, the second machine learning model outputs, for example, a score that changes according to the risk that the settlement of a deferred payment is not properly fulfilled by the target user when the target user purchases goods and / or services in a deferred payment, which is a transaction in the second category, or a score related to the default probability when the target user purchases goods and / or services in a credit card payment, which is also a transaction in the second category. It is a machine learning model for review.
[0031] Therefore, for the second machine learning model, the types of input data in the review model of small loans such as credit card reviews may be appropriately adopted. Also, for the second machine learning model, a review model of BNPL (Buy-Now-Pay-Later, deferred payment) may be adopted. When adopting the BNPL review model as the second machine learning model, some kind of credit score or user attribute may be used for the determination of the final review result in the present disclosure.
[0032] Also in the second machine learning model, similar to the first machine learning model, a group of attribute data including factual attribute data and / or estimated attribute data can be appropriately adopted as input data. The input data includes, for example, annual income, living situation (industry type of the workplace, working form (whether remote work is possible, etc.)), credit card usage situation, and the like.
[0033] The user segment identification unit 24 identifies the user segment to which the target user belongs by performing segmentation on the user group including the target user. Here, the segmentation is performed by referring to the attribute data of the users included in the user group and putting the users having common attribute data into the same user segment. For example, users with the same annual income, industry type of workplace, gender, and age group are put into the same user segment. Here, the user attributes referred to during segmentation are not limited. However, for example, by using the attributes common between the input data of the first machine learning model and the input data of the second machine learning (in other words, the attributes used as input data in both the first machine learning model and the second machine learning model) for segmentation, it becomes easy to determine the correlation described later. Also, for segmentation, rule-based segmentation may be used, or inference-based segmentation using a machine learning model may be used. Also, the segmentation may be performed by executing clustering using a general clustering method for the user group and referring to the results of the clustering.
[0034] Based on the strength of the correlation between the first score and the second score in the user segment, the correlation determination unit 25 determines whether the user segment is a first user segment having a correlation of a predetermined standard or more between the first score and the second score, or a second user segment having a correlation less than the predetermined standard (in the present disclosure, those having no correlation are also included). That is, the correlation determination unit 25 determines, according to a predetermined standard, whether the user segment to which the user belongs is a user segment having a strong correlation between the first score and the second score, or a user segment having a weak correlation between the first score and the second score. More specifically, for example, the correlation determination unit 25 calculates the difference between the statistical value (for example, average value or median value) of the first score calculated in the past and the statistical value (for example, average value or median value) of the second score calculated in the past for any user segment, and determines that the user segment is a first user segment when the difference is equal to or more than a predetermined standard, and determines that the user segment is a second user segment when the difference is less than the predetermined standard. However, the specific determination means for the correlation between the first score and the second score in the user segment is not limited to the examples in the present disclosure.
[0035] For example, when the user group constituting the segment is men in their 50s to 60s, it is assumed that there is a strong correlation between the height of the first score, which is a score based on credit information determined by, for example, a credit information institution, and the height of the second score, which is a score based on the output from a review model for small payment, for example. On the other hand, when the user group constituting the segment is men in their 20s to 30s, it is assumed that the correlation between the first score and the second score is weak or there is no correlation. This is because in the case of a young age group such as those in their 20s to 30s, the accuracy of the first score is not sufficient due to insufficient accumulation of credit information, and the accuracy of the second score by a review model for applications with a high usage frequency by the young age group such as small payment is strongly likely to be sufficient.
[0036] The examination result determination unit 26 determines the examination result of the target user based on the first score and / or the second score of the target user according to the user segment to which the target user belongs. In the present embodiment, when it is determined that the user segment to which the target user belongs is the first user segment, for transactions in the first category, the examination result of the target user is determined based on the first score, and for transactions in the second category, the examination result of the target user is determined based on the second score. That is, in the present embodiment, in the case of a user belonging to the first user segment where the correlation between the first score and the second score is strong, when the first score is a value that works unfavorably for the examination, regardless of the second score, the final examination result is determined in a negative form (for example, the financing does not decline, etc.).
[0037] On the other hand, when it is determined that the user segment to which the target user belongs is the second user segment, the examination result determination unit 26 determines the examination result of the target user based on the score indicating a higher creditworthiness or a lower risk among the first score and the second score. In other words, in the present embodiment, when the specific user segment is the second user segment where the correlation between the first score and the second score is weak or there is no correlation, when a user belongs to the user segment, even when the first score is a value that works unfavorably for the examination, if the second score is a value that can be determined to have sufficient credit, the final examination result is determined in a positive form (the examination result is determined to work favorably for the examination). This means that even when the first score is high and it is determined on the surface that the credit is insufficient, if the second score, which is a more reliable score, indicates that the user has sufficient credit, the examination result is determined in a positive form.
[0038] FIG. 3 is a diagram showing an overview of the review process in the present embodiment. That is, in the present embodiment, when reviewing a target user, the loan review result of the user is determined based on the first score output by the machine learning model for review and the second score output by the machine learning model for other purposes. At this time, according to the strength of the correlation between the first score and the second score for each user segment, the determination logic of the review result by the review result determination unit 26 is determined, and the review result is output.
[0039] However, the specific method for determining the review result of the target user based on the first score and / or the second score is not limited to the disclosure in the present embodiment. For example, the review result determination unit 26 may determine the review result of the target user based on the third score calculated based on the first score and the second score. More specifically, for example, when it is determined that the user segment to which the target user belongs is the first user segment, for the transactions in the first category, the weight of the first score is set large and the weight of the second score is set small, and based on the third score calculated, for the transactions in the second category, the weight of the first score is set small and the weight of the second score is set large, and based on the third score calculated, the review result of the target user may be determined. Also, for example, when it is determined that the user segment to which the target user belongs is the second user segment, the weight of the score indicating a higher credit rating or a lower risk among the first score and the second score is set large, and the weight of the other score is set small, and based on the third score calculated, the review result of the target user may be determined.
[0040] The machine learning unit 27 generates and / or updates a machine learning model used for obtaining the first score by the first score acquisition unit 22 and a machine learning model used for obtaining the second score by the second score acquisition unit 23. The machine learning model for obtaining the first score is a machine learning model that outputs a first score indicating the creditworthiness of the user when data of one or more user attributes related to the target user is input. Also, the machine learning model for obtaining the second score is a machine learning model that outputs a second score indicating the degree of risk based on the probability that the user defaults without making a payment when data of one or more user attributes related to the target user is input.
[0041] In generating and / or updating the machine learning model for obtaining the first score, the machine learning unit 27 creates teacher data that defines, for each user, the attribute data group of the user as an input value and the score related to the user as an output value. Then, the machine learning unit 27 generates and / or updates the first machine learning model based on the teacher data. As described above, the attribute data group input to the first machine learning model includes factual attribute data and estimated attribute data, which are combined with the score of the corresponding user and input to the machine learning unit 27 as teacher data. The score set in the teacher data may be a score determined based on rules or a score set manually (annotated). Also, it may be a score that has been corrected by an administrator or the like after being output by the first machine learning model in the past.
[0042] In generating and / or updating the machine learning model for obtaining the second score, the machine learning unit 27 creates a machine learning model based on teacher data that defines, for each user attribute, a statistic related to the default occurrence rate of a plurality of users having a predetermined attribute (in this embodiment, the average value. However, statistical indicators such as the mode or median may be used, for example.) as the second score indicating the degree of risk of the users having the attribute. The calculated second score is combined with the attribute data group including the factual attribute data and estimated attribute data of the corresponding user and input to the machine learning unit 27 as teacher data.
[0043] When implementing the technology according to the present disclosure, the framework for generating / updating a machine learning model that can be adopted as a first machine learning model, a second machine learning model, etc. is based on, for example, an ensemble learning algorithm. For example, a machine learning framework based on a Gradient Boosting Decision Tree (GBDT) (e.g., LightGBM) may be adopted in the framework. In other words, the framework may adopt a machine learning framework based on a decision tree model that passes on the error between the correct answer and the predicted value between the previous and subsequent weak learners (weak classifiers). Here, the predicted value refers to, for example, the predicted value of the score. Note that, in addition to LightGBM, the framework may adopt boosting methods such as XGBoost and CatBoost. According to the framework using decision trees, a machine learning model with relatively high performance can be generated / updated with less effort in parameter adjustment compared to the framework using neural networks. However, the framework for generating / updating a machine learning model that can be adopted when implementing the technology according to the present disclosure is not limited to the examples in this embodiment. For example, another learner such as a random forest may be adopted instead of the gradient boosting decision tree as the learner, or a learner not called a so-called weak learner such as a neural network may be adopted. Also, particularly when a learner not called a so-called weak learner such as a neural network is adopted, ensemble learning may not be adopted.
[0044] FIG. 4 is a schematic diagram of the concept of a decision tree of the machine learning model adopted in the present embodiment. When adopting a machine learning framework of gradient boosting based on the decision tree algorithm, the branching conditions of each node of the decision tree are optimized. Specifically, in the machine learning framework of gradient boosting based on the decision tree algorithm, scores are calculated for user groups having attributes indicated by each of two child nodes branched from one parent node, and the difference between these scores is increased (for example, so that the difference is maximized, or so that it is equal to or greater than a predetermined threshold), that is, the branching condition of the parent node is optimized so that the two child nodes are cleanly branched. For example, when the attribute indicated as the branching condition of the node is age, the age set as the branching threshold may be changed, or the branching condition may be changed to an attribute other than age. In this way, by recursively optimizing the branching conditions of all nodes of the decision tree, the estimation accuracy of the score based on the attribute data group can be improved.
[0045] <Flow of processing> Next, the flow of processing executed by the examination apparatus according to the present embodiment will be described. Note that the specific content and processing order of the processing described below are examples for implementing the present disclosure. The specific processing content and processing order may be appropriately selected according to the embodiment of the present disclosure.
[0046] FIG. 5 is a flowchart showing the flow of machine learning processing according to the present embodiment. The processing shown in this flowchart is executed periodically or at a timing specified by an administrator.
[0047] In this embodiment, in the machine learning process, a first machine learning model and a second machine learning model are generated and / or updated. The machine learning unit 27 creates teacher data including a combination of the attribute data group for each user accumulated in the past and the first score determined in advance for the corresponding user (step S101). Then, the machine learning unit 27 generates and / or updates the first machine learning model using the created teacher data (step S102). Further, the machine learning unit 27 creates teacher data including a combination of the attribute data group for each user accumulated in the past and the second score determined in advance for the corresponding user (step S103). Then, the machine learning unit 27 generates and / or updates the second machine learning model using the created teacher data (step S104). Thereafter, the process shown in this flowchart ends.
[0048] FIG. 6 is a flowchart showing the flow of the review process according to this embodiment. The process shown in this flowchart is executed for each target user periodically or at a specified timing.
[0049] In steps S201 to S203, the first score and the second score are acquired. The attribute determination unit 21 acquires the attribute data group of the target user (step S201). Then, the first score acquisition unit 22 inputs the input data corresponding to the first machine learning model among the attribute data group acquired in step S201 to the first machine learning model, and based on the output value, acquires the first score set for the user (step S202). Also, the second score acquisition unit 23 inputs the input data corresponding to the second machine learning model among the attribute data group acquired in step S201 to the second machine learning model, and based on the output value, acquires the second score set for the user (step S203). Thereafter, the process proceeds to step S204.
[0050] In steps S204 and S205, the user segment to which the target user belongs is identified, and the correlation between the first score and the second score in the user segment is determined. The user segment identification unit 24 identifies the user segment to which the target user belongs by performing segmentation on the user group including the target user (step S204). When the user segment to which the target user belongs is identified, the correlation determination unit 25 determines whether the user segment is a first user segment having a correlation equal to or higher than a predetermined standard between the first score and the second score, or a second user segment having a correlation lower than the predetermined standard, based on the strength of the correlation between the first score and the second score in the identified user segment (step S205). When it is determined that the user segment is the first user segment (YES in step S205), the process proceeds to step S206. On the other hand, when it is determined that the user segment is the second user segment (NO in step S205), the process proceeds to step S207.
[0051] In step S206, based on the score related to the target category, the review result of the target user is determined. When it is determined that the user segment to which the target user belongs is the first user segment, the review result determination unit 26 determines the review result of the target user based on the first score if this review process is executed for transactions in the first category, and based on the second score if this review process is executed for transactions in the second category. Then, the process proceeds to step S208.
[0052] In step S207, based on the score indicating a higher credit rating or a lower risk, the review result of the target user is determined. When it is determined in step S205 that the user segment to which the target user belongs is the second user segment, the review result determination unit 26 determines the review result of the target user based on the score indicating a higher credit rating or a lower risk among the first score and the second score. Then, the process proceeds to step S208.
[0053] In step S208, the examination result is output. The examination result determination unit 26 outputs the examination result in step S206 or step S207. Thereafter, the processes shown in this flowchart end.
[0054] <Effect> According to this embodiment, for example, it is possible to support loan examinations and the like for any user belonging to a user segment with insufficient credit information accumulation, such as the younger generation, with higher accuracy. Further, according to this embodiment, by considering the correlation between different credit scores due to different combinations of input data and models, it is possible to more appropriately support the examination of users.
[0055] <Variation> Hereinafter, as a variation of the above-described embodiment, a variation in which the examination device 1b determines a second score based on the attribute data group of the target user will be described. Here, since the system configuration of the examination device 1b according to this variation and the examples of the attribute data of the user used in this variation are substantially the same as those of the above-described embodiment, the description thereof will be omitted (see FIG. 1).
[0056] FIG. 7 is a diagram showing an outline of the functional configuration of the examination device 1b according to the variation. In this variation, the examination device 1b includes a product-related data acquisition unit 28 in addition to the attribute determination unit 21, the first score acquisition unit 22, the second score acquisition unit 23, the user segment identification unit 24, the correlation determination unit 25, the examination result determination unit 26, and the machine learning unit 27 described in the above-described embodiment.
[0057] The product-related data acquisition unit 28 acquires product-related data of the product / service that the target user intends to purchase by receiving financing related to the transaction. Here, the product-related data acquisition unit 28 acquires at least either the attribute data of the product / service or the attribute data of the store (including online stores) that sells the product / service as the product-related data. Here, the product-related data only needs to be information related to the product / service, and the data acquired as the product-related data is not limited. Examples of the attribute data of the product / service include price, product / service category / genre, etc. Examples of the attribute data of the store include store price range, store category / genre, etc. Further, the product-related data may include business rules and transaction types associated with the transaction of the target product / service. Here, the business rules and transaction types may include business customs related to the product / service, the content of the contract associated with the transaction, and the delivery mode of the product / service (such as delivery method or handover at the store).
[0058] And in this variation, the second score acquisition unit 23 acquires a second score for determining approval or rejection of financing for the purchase of a product and / or service by the target user based on the output of the second machine learning model. Here, the second machine learning model may be appropriately learned for each product-related data.
[0059] More specifically, when the product-related data of the target product / service is acquired, the second score acquisition unit 23 inputs the attribute data of the target user into the latest second machine learning model generated or updated by the machine learning unit 27, and as an output value, obtains a second score that changes according to the risk that the target user will not correctly fulfill the financing settlement when purchasing the target product / service by financing. Since the processing after the second score is obtained is the same as the processing described in the above embodiment, the description is omitted.
[0060] <Other variations> As a variation of the above-described embodiment, the review apparatus 1 may determine a review result of a user based on outputs (first to nth scores) of each of a plurality of n machine learning models (first to nth machine learning models) including a first machine learning model and a second machine learning model (n > 1). Further, the input data to the first to nth machine learning models may be the input data to the above-described first machine learning model or the input data to the second machine learning model.
[0061] The first to nth machine learning models may be generated by learning the relationship between the attributes of a user and the creditworthiness of the user in first to nth categories of transactions classified according to the period and / or amount related to the transaction. Further, the first to nth machine learning models may be generated by learning the relationship between the attributes of a user and the creditworthiness of the user in first to nth categories of transactions classified according to other features such as interest rate, presence or absence of collateral, repayment method, etc., not limited to the period or amount.
[0062] The review apparatus 1 determines the review result of the target user based on the first to nth scores of the target user according to the user segment to which the target user belongs. The review apparatus 1 may determine the review result based on any one of the first to nth scores or may determine the review result based on the weighted first to nth scores according to the user segment to which the target user belongs.
Explanation of Signs
[0063] 1 Review apparatus
Claims
1. First score acquisition means for acquiring a first score of a user based on an output indicating the credit score of the user, obtained by inputting input data including the user's attribute data into a first machine learning model generated by learning the relationship between the user's attributes and the user's credit score in a first type of transaction; Second score acquisition means for acquiring a second score of the user based on an output indicating the risk of the user, obtained by inputting input data including the user's attribute data into a second machine learning model generated by learning the relationship between the user's attributes and the user's risk in a second type of transaction different from the first type of transaction; User segment identification means for identifying a user segment to which the target user belongs by performing segmentation of a user group including the target user; Correlation determination means for determining whether the user segment is a first user segment having a correlation equal to or higher than a predetermined standard or a second user segment having a correlation lower than the predetermined standard based on the strength of the correlation between the first score and the second score in the user segment; Review result determination means for determining a review result of the target user based on a score indicating a higher credit score or a lower risk among the first score and the second score when it is determined that the user segment to which the target user belongs is the second user segment; A review device comprising:
2. The first score acquisition means acquires the first score of the user based on an output obtained by inputting the input data related to the user into a first machine learning model which is a machine learning model for reviewing the first type of transaction related to a period of a predetermined length or more and / or an amount of a predetermined amount or more, The second score acquisition means acquires the second score of the user based on an output obtained by inputting the input data related to the user into a second machine learning model which is a machine learning model for reviewing the second type of transaction related to a period shorter and / or an amount less than that of the first type of transaction, The review device according to Claim 1.
3. The first score acquisition means acquires the first score of the user by using the first machine learning model generated and / or updated by using teacher data including attribute data of the user including the situation of other transactions of the same type as the first type of transaction and a value indicating the creditworthiness related to the user. The examination device according to claim 2.
4. The second score acquisition means acquires the second score of the user by using the second machine learning model generated and / or updated by using teacher data including attribute data of the user including the situation of other transactions of the same type as the second type of transaction and a value indicating the risk related to the user. The examination device according to claim 2.
5. The first score acquisition means estimates the first score by using the first machine learning model generated and / or updated by using a machine learning framework based on a gradient boosting decision tree. The second score acquisition means estimates the second score by using the second machine learning model generated and / or updated by using a machine learning framework based on a gradient boosting decision tree. The examination device according to claim 1.
6. For the user segment, the correlation determination means determines that the user segment is a first user segment when the difference between the statistical value of the first score calculated in the past and the statistical value of the second score calculated in the past is equal to or greater than a predetermined standard, and determines that the user segment is a second user segment when the difference is less than the predetermined standard. The examination device according to claim 1.
7. When it is determined that the user segment to which the target user belongs is the first user segment, the examination result determination means determines the examination result of the target user based on the first score for the first type of transaction. The examination device according to claim 1.
8. When it is determined that the user segment to which the target user belongs is the first user segment, the examination result determination means determines the examination result of the target user based on the second score for the second type of transaction. The examination device according to claim 1.
9. The first type of transaction is a type of transaction other than a transaction by postpaid settlement. The first score acquisition means acquires the first score of the user based on the output obtained by inputting the input data related to the user into the first machine learning model which is a machine learning model for reviewing the first type of transaction. The second score acquisition means acquires the second score of the user based on the output obtained by inputting the input data related to the user into the second machine learning model which is a machine learning model for reviewing transactions by deferred payment, and which is generated and / or updated using attribute data of the user including the status of transactions by deferred payment and teacher data including values indicating risks related to the user. The output value of the second machine learning model is a determination score that changes according to the risk that the deferred payment settlement will not be correctly fulfilled by the target user when the target user purchases goods and / or services by deferred payment. The review device according to claim 8.
10. A computer: A first score acquisition step of acquiring the first score of the user based on the output indicating the creditworthiness of the user, which is obtained by inputting input data including the user's attribute data into a first machine learning model generated by learning the relationship between the user's attributes and the user's creditworthiness in a first type of transaction; A second score acquisition step of acquiring the second score of the user based on the output indicating the risk of the user, which is obtained by inputting input data including the user's attribute data into a second machine learning model generated by learning the relationship between the user's attributes and the user's risk in a second type of transaction different from the first type of transaction; A user segment identification step of identifying the user segment to which the target user belongs by performing segmentation of a group of users including the target user; A correlation determination step of determining whether the user segment is a first user segment having a correlation equal to or greater than a predetermined standard or a second user segment having a correlation less than the predetermined standard based on the strength of the correlation between the first score and the second score in the user segment; When it is determined that the user segment to which the target user belongs is the second user segment, based on the score indicating a higher credit rating or a lower risk among the first score and the second score, a review result determination step of determining the review result of the target user; A method of execution.
11. A computer, A first score acquisition means for acquiring a first score of a user based on an output indicating the credit rating of the user obtained by inputting input data including the user's attribute data into a first machine learning model generated by learning the relationship between the user's attributes and the user's credit rating in a first type of transaction; A second score acquisition means for acquiring a second score of the user based on an output indicating the risk of the user obtained by inputting input data including the user's attribute data into a second machine learning model generated by learning the relationship between the user's attributes and the user's risk in a second type of transaction different from the first type of transaction; A user segment identification means for identifying the user segment to which the target user belongs by segmenting a group of users including the target user; Based on the strength of the correlation between the first score and the second score in the user segment, it is determined whether the user segment is a first user segment having a correlation equal to or higher than a predetermined standard between the first score and the second score, or a second user segment having a correlation lower than the predetermined standard. A correlation determination means; When it is determined that the user segment to which the target user belongs is the second user segment, based on the score indicating a higher credit rating or a lower risk among the first score and the second score, a review result determination means for determining the review result of the target user; A program that functions as.
Citation Information
Patent Citations
YINGPAN science and technology financial enterprise credit scoring algorithm model and application system
CN114037197A
Score calculating method, and score providing method
JP2002092305A
Credit risk management
JP2012500443A
Loan examination support device and system, and program
JP2017182284A
Information processing apparatus, information processing method, and information processing program
JP2019204208A