Method, apparatus and server for determining user risk
By using a preset multi-layer user risk prediction model to process user business data, the problem of large user risk prediction error in the prior art is solved, and more accurate risk level assessment and higher prediction accuracy are achieved.
Patent Information
- Application Number
- CN202110643176.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-06-09
AI Technical Summary
The prior art has large errors in predicting user risks, and cannot comprehensively and accurately determine the user's risk situation, and no effective solution has been proposed.
By obtaining the business data of the target user, calling the preset user risk prediction model to process this data, using multi-layer models (including the first layer model and the second layer model) to conduct risk assessment, and determining the user's risk level.
It realizes a more comprehensive and accurate determination of the user's risk level, effectively reducing the error in risk prediction and improving the accuracy of risk prediction.
Smart Images

Figure CN113379530B_ABST
Abstract
Description
Technical Field
[0001] This specification belongs to the field of artificial intelligence technology, and particularly relates to a method, apparatus, and server for determining user risks. Background Art
[0002] When a business handling institution handles a business involving user risks such as a credit business or a credit card business for a user, the handling institution often needs to first evaluate the risk situation of the user, and then determine whether to handle the corresponding business for the user based on the risk situation of the user.
[0003] However, based on the existing method for determining user risks, when specifically predicting user risks, there are often large errors, and it is impossible to comprehensively and accurately determine the risk situation of the user.
[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] This specification provides a method, apparatus, and server for determining user risks, which can comprehensively and accurately determine the comprehensive risk level of a target user, effectively reduce the error in determining user risks, and improve the risk prediction accuracy.
[0006] An embodiment of this specification provides a method for determining user risks, including:
[0007] Obtain the business data of the target user;
[0008] Call a preset user risk prediction model to process the business data of the target user to obtain a corresponding target processing result; wherein, the preset user risk prediction model includes a first-layer model and a second-layer model, and the first-layer model includes a plurality of sub-models, and the plurality of sub-models respectively correspond to a sub-business scenario;
[0009] Determine the risk level of the target user according to the target processing result.
[0010] In some embodiments, the method further includes:
[0011] Obtain a plurality of data tables of the business system; wherein, each of the plurality of data tables contains a plurality of business data of sample users;
[0012] Perform a preset clustering process on the plurality of data tables to obtain a plurality of sample data sets of the sample users; wherein, each of the plurality of sample data sets contains business data with business relevance corresponding to a sub-business scenario;
[0013] Use the plurality of sample data sets of the sample users to train an initial model to obtain the preset user risk prediction model.
[0014] In some embodiments, a preset clustering process is performed on the multiple data tables to obtain multiple sample data sets of sample users, including:
[0015] Based on the K-means clustering algorithm, the multiple data tables are clustered according to the identity identifiers of the sample users to obtain multiple aggregated tables; wherein, each of the aggregated tables contains business data corresponding to the identity identifier of a sample user and a sub-business scenario;
[0016] According to the multiple aggregated tables, multiple sample data sets of the sample users are constructed.
[0017] In some embodiments, the initial model is constructed in the following manner:
[0018] Based on the random forest algorithm, multiple initial sub-models for multiple sub-business scenarios are constructed; and the multiple initial sub-models are combined to obtain an initial first-layer model;
[0019] Based on the decision tree algorithm, an initial second-layer model is constructed;
[0020] The multiple initial sub-models in the initial first-layer model are connected to the initial second-layer model to obtain the initial model.
[0021] In some embodiments, after performing the preset clustering process on the multiple data tables to obtain multiple sample data sets of sample users, the method further includes:
[0022] According to the preset verification rules, data cleaning processes are respectively performed on the multiple sample data sets of the sample users to filter out invalid data.
[0023] In some embodiments, the invalid data includes at least one of the following: business data whose data generation time is greater than a preset time threshold, business data whose data format does not conform to the preset requirements, and business data whose data value is empty.
[0024] In some embodiments, the business data includes banking business data; correspondingly, the sub-business scenarios include: deposit and loan business scenarios, credit card business scenarios, and financial wealth management business scenarios.
[0025] In some embodiments, using the multiple sample data sets of the sample users to train the initial model to obtain the preset user risk prediction model, including:
[0026] Using the multiple sample data sets of the sample users to respectively train the corresponding initial sub-models in the initial first-layer model to obtain a first-layer model that meets the requirements;
[0027] Call the first-layer model to process multiple sample data sets of a sample user to obtain multiple intermediate processing results;
[0028] Use the multiple intermediate processing results to train the initial second-layer model to obtain a second-layer model that meets the requirements.
[0029] In some embodiments, while using the multiple intermediate processing results to train the initial second-layer model, the method further includes:
[0030] Associate with a credit investigation system to obtain the credit risk parameters of the sample user;
[0031] Use the credit risk parameters of the sample user to correct the second-layer model.
[0032] An embodiment of this specification also provides a training method for a preset user risk prediction model, including:
[0033] Obtain multiple data tables of a business system; wherein, each of the multiple data tables contains multiple business data of a sample user;
[0034] Perform a preset clustering process on the multiple data tables to obtain multiple sample data sets of the sample user; wherein, each of the multiple sample data sets contains business data with business relevance corresponding to a sub-business scenario;
[0035] Use the multiple sample data sets of the sample user to train an initial model to obtain the preset user risk prediction model; wherein, the initial model includes an initial first-layer model and an initial second-layer model, and the initial first-layer model includes multiple initial sub-models, and the initial sub-models respectively correspond to a sub-business scenario.
[0036] An embodiment of this specification also provides a device for determining user risk, including:
[0037] An acquisition module, configured to acquire the business data of a target user;
[0038] A call module, configured to call a preset user risk prediction model to process the business data of the target user to obtain a corresponding target processing result; wherein, the preset user risk prediction model includes a first-layer model and a second-layer model, and the first-layer model includes multiple sub-models, and the multiple sub-models respectively correspond to a sub-business scenario;
[0039] A determination module, configured to determine the risk level of the target user according to the target processing result.
[0040] An embodiment of this specification also provides a server, which includes a processor and a memory for storing instructions executable by the processor. When the processor executes the instructions, the following operations are implemented: obtaining service data of a target user; invoking a preset user risk prediction model to process the service data of the target user to obtain a corresponding target processing result; wherein the preset user risk prediction model includes a first-layer model and a second-layer model, the first-layer model includes a plurality of sub-models, and the plurality of sub-models respectively correspond to a sub-service scenario; determining the risk level of the target user according to the target processing result.
[0041] An embodiment of this specification also provides a computer-readable storage medium, on which computer instructions are stored. When the instructions are executed, the following operations are implemented: obtaining service data of a target user; invoking a preset user risk prediction model to process the service data of the target user to obtain a corresponding target processing result; wherein the preset user risk prediction model includes a first-layer model and a second-layer model, the first-layer model includes a plurality of sub-models, and the plurality of sub-models respectively correspond to a sub-service scenario; determining the risk level of the target user according to the target processing result.
[0042] A method, device and server for determining user risk provided in this specification. Before specific implementation, multiple data tables containing all service data can be subjected to a preset clustering process to obtain multiple sample datasets in which multiple sample users are clustered together based on service relevance; then, using the multiple sample datasets of multiple sample users, a preset user risk prediction model with a double-layer structure including both a first-layer model and a second-layer model and with high accuracy can be trained; wherein the above-mentioned first-layer model specifically includes a plurality of sub-models respectively corresponding to multiple sub-service scenarios; during specific implementation, after obtaining the service data of the target user, the above-mentioned preset user risk prediction model can be invoked to comprehensively process the service data of the target user, so as to more comprehensively and accurately determine the risk level of the target user, thereby effectively reducing the error when determining user risk and improving the risk prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] To more clearly illustrate the embodiments of this specification, the accompanying drawings required for the embodiments will be briefly introduced below. The accompanying drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0044] Figure 1 It is a schematic diagram of an embodiment of the system structure of the method for determining user risk provided by the embodiment of this specification;
[0045] Figure 2It is a schematic diagram of an embodiment of applying the method for determining user risk provided by the embodiments of this specification in a scenario example;
[0046] Figure 3 It is a schematic flowchart of the method for determining user risk provided by an embodiment of this specification;
[0047] Figure 4 It is a schematic flowchart of the method for training a preset user risk prediction model provided by an embodiment of this specification;
[0048] Figure 5 It is a schematic diagram of the structural composition of a server provided by an embodiment of this specification;
[0049] Figure 6 It is a schematic diagram of the structural composition of a device for determining user risk provided by an embodiment of this specification;
[0050] Figure 7 It is a schematic diagram of an embodiment of applying the method for determining user risk provided by the embodiments of this specification in a scenario example;
[0051] Figure 8 It is a schematic diagram of an embodiment of applying the method for determining user risk provided by the embodiments of this specification in a scenario example;
[0052] Figure 9 It is a schematic diagram of an embodiment of applying the method for determining user risk provided by the embodiments of this specification in a scenario example;
[0053] Figure 10 It is a schematic diagram of an embodiment of applying the method for determining user risk provided by the embodiments of this specification in a scenario example. Detailed implementation manners
[0054] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0055] The embodiments of this specification provide a method for determining user risk. The method for determining user risk can be specifically applied to a system including a server and a terminal device. Specifically, reference can be made to Figure 1 As shown, the terminal device and the server can be connected by wire or wirelessly to perform specific data interaction.
[0056] In this embodiment, the server may specifically include a background server applied to one side of the data processing system of Bank A's data center, which can implement functions such as data transmission and data processing. Specifically, the server may be, for example, an electronic device with data operation, storage, and network interaction functions. Or, the server may also be a software program running in this electronic device that provides support for data processing, storage, and network interaction. In this embodiment, the number of servers included in the server is not specifically limited. The server may specifically be one server, or several servers, or a server cluster formed by several servers.
[0057] In this embodiment, the terminal device may specifically include a front-end electronic device deployed at the counter of Bank A, which can implement functions such as data collection and data transmission. Specifically, the terminal device may be, for example, a desktop computer, a tablet computer, a laptop computer, a smart phone, etc. Or, the terminal device may also be a software application that can run in the above-mentioned electronic devices. For example, it may be a certain APP running on a desktop computer.
[0058] Current user B wants to apply for a credit loan service of Bank A in the business hall of Bank A. Staff member C responsible for handling the business can first interact with the server using the terminal device to determine the risk level of user B; then determine whether to handle the credit loan service for user B according to the risk level of user B, and specifically how to handle the credit loan service.
[0059] Specifically, staff member C can use the terminal device to generate a service data query request, where the service data query request may specifically carry the identity identifier of user B (for example, the name, account name, or ID number of user B, etc.); and send the service data query request to the database of Bank A's data center through the terminal device to query and obtain the service data of user B. Among them, the obtained service data of user B may specifically include a plurality of service data corresponding to the identity identifier of user B extracted from a plurality of data tables stored in the database of each business system of Bank A. For example, the financial management service currently handled by user B, the credit card historical repayment record of user B, the mortgage repayment record of user B, and so on.
[0060] Furthermore, staff member C can use the terminal device to generate a user risk prediction request, where the user risk prediction request may specifically carry the service data of user B; and send the above user risk prediction request to the server of Bank A's data center.
[0061] Correspondingly, the server receives the above user risk prediction request and parses and extracts the service data of user B from it.
[0062] Next, the server can use the above business data of User B as model input and input it into a preset user risk prediction model, and run the preset user risk prediction model to process the business data of User B.
[0063] Specifically, please refer to Figure 2 as shown. First, the business data of User B will be grouped according to business relevance, and multiple groups of business data will be respectively input into the corresponding multiple sub-models included in the first-layer model of the preset user risk prediction model. Among them, each of the above sub-models corresponds to a specific sub-business scenario. For example, in Figure 2 , the No. 1 sub-model corresponds to the deposit and loan business scenario, the No. 2 sub-model corresponds to the credit card business scenario, the No. 3 sub-model corresponds to the financial wealth management business scenario, the No. 4 sub-model corresponds to the bank VIP service business scenario, etc.
[0064] Next, control the multiple sub-models to respectively process the input groups of business data to output the risk probability values for the multiple sub-business scenarios as intermediate processing results; and output the multiple intermediate processing results from the first-layer model and input them into the second-layer model.
[0065] Furthermore, the second-layer model can be controlled to calculate the corresponding target processing result through synthesizing the multiple intermediate processing results as the final model output of the preset user risk prediction model.
[0066] The server can determine a relatively comprehensive and accurate risk level for the target user according to the above target processing result; and send the risk level to the terminal device.
[0067] The terminal device can display the risk level of User B to staff member C.
[0068] Staff member C can compare the risk level of User B with a preset first-level threshold. When it is determined that the risk level of User B is relatively high, for example, higher than the preset first-level threshold, it can be judged that there is a relatively high risk in handling a credit loan business for this user. At this time, staff member C can reject handling the credit loan business for this user to avoid Bank A undertaking non-performing assets.
[0069] When it is determined that the risk level of User B is relatively low, for example, lower than the preset first-level threshold, staff member C can judge that the risk of handling a credit loan business for this user is relatively low and can handle the applied loan business for this user.
[0070] Furthermore, staff member C can also compare the risk level of user B with a preset second-level threshold. When it is determined that the risk level of user B is higher than the preset second-level threshold, it can be determined to handle a small-amount credit loan business for this user.
[0071] When it is determined that the risk level of user B is lower than the preset second-level threshold, staff member C can determine to handle a large-amount credit loan business for this user.
[0072] Through the above embodiments, a preset user risk prediction model can be used to comprehensively process various business data of a user, so as to comprehensively and accurately determine the risk level of the user, and then corresponding services can be precisely handled for the user according to the risk level of the user.
[0073] Refer to Figure 3 As shown, the embodiments of this specification provide a method for determining user risk, where this method is specifically applied to the server side. Specifically implemented, this method may include the following:
[0074] S301: Obtain the business data of the target user;
[0075] S302: Invoke a preset user risk prediction model to process the business data of the target user to obtain a corresponding target processing result; where the preset user risk prediction model includes a first-layer model and a second-layer model, and the first-layer model includes multiple sub-models, and the multiple sub-models respectively correspond to a sub-business scenario;
[0076] S303: Determine the risk level of the target user according to the target processing result.
[0077] Through the above embodiments, a preset user risk prediction model with a pre-trained two-layer structure can be used to comprehensively process the business data of the target user, so that the risk level of the target user can be more comprehensively and accurately determined, effectively reducing the error when determining the user risk and improving the risk prediction accuracy.
[0078] In some embodiments, the above target user can be specifically understood as a user object for which business risk needs to be determined. According to specific situations and processing requirements, the above business risk can specifically be a transaction risk, a credit risk, or a health risk, etc., of various different types of risks.
[0079] In some embodiments, the business data of the above target user can specifically be one or more parameter data related to the target user that can reflect the risk situation of the target user directly or indirectly. Corresponding to different application scenarios, the above business data can include various different types of data.
[0080] Specifically, taking the scenario of banking business handling as an example, the above-mentioned business data includes banking business data. Further, the above-mentioned business data may specifically include: record data of the financial management business handled by the target user in the past, the historical credit card repayment records of the target user, the mortgage repayment records of the target user, and so on. Of course, it should be noted that the above-listed business data is only an illustrative description. In specific implementation, according to the specific application scenario and processing requirements, the above-mentioned business data may also include other types of data. This specification does not make any limitations in this regard.
[0081] Through the above embodiments, the method for determining user risk provided by the embodiments of this specification can be effectively applied to process the banking business data of the target user, so as to accurately predict the risk level of the target user in the scenario of banking business handling.
[0082] In some embodiments, in specific implementation, according to the identity identifier of the target user, by querying the database, the data that matches the identity identifier of the target user can be extracted from multiple data tables stored in the database as the business data of the target user.
[0083] In some embodiments, the above-mentioned preset user risk prediction model can be specifically understood as a pre-trained classification model that can comprehensively process various different types of business data of the user to more comprehensively and accurately determine the risk level of the user.
[0084] In some embodiments, the above-mentioned preset user risk prediction model can specifically be a classification model with a two-layer structure including a first-layer model and a second-layer model. Among them, the above-mentioned first-layer model can specifically include multiple sub-models. Each sub-model corresponds to a specific sub-business scenario. Each sub-model is used to predict the risk probability value in the corresponding sub-business scenario based on the corresponding business data of the user as an intermediate processing result. The second-layer model is used to comprehensively process multiple intermediate processing results to predict the overall risk level of the user as the final target processing result. Regarding how to specifically train and establish the above-mentioned preset user risk prediction model, it will be described separately later.
[0085] Specifically, in the scenario of banking business handling, the above-mentioned multiple sub-business scenarios may include: deposit and loan business scenarios, credit card business scenarios, financial management business scenarios, and so on. Of course, it should be noted that the above-listed sub-business scenarios are only an illustrative description. In specific implementation, according to the specific application scenario and processing requirements, the above-mentioned sub-business scenarios may also include other types of business scenarios. This specification does not make any limitations in this regard.
[0086] In some embodiments, during specific implementation, the business data of the target user can be used as a model input and input into a preset user risk prediction model; and by running the preset user risk prediction model, the business data of the target user is processed to output a corresponding target processing result.
[0087] During specific operation, the business data of the target user will first be divided into multiple groups of business data based on business relevance, where each group of business data corresponds to a sub-business scenario. Then, the multiple groups of business data will be separately input into the corresponding sub-models in the first-layer model for specific processing, and the risk probability values under the corresponding sub-business scenarios will be output as intermediate processing results. Further, the intermediate processing results output by the multiple sub-models will be uniformly input into the second-layer model for processing by the first-layer model. Finally, the second-layer model can output the risk level for the overall situation of the target user as the target processing result by synthesizing the multiple intermediate processing results. Accordingly, the risk level of the target user can be comprehensively and precisely determined based on the above target processing result.
[0088] In some embodiments, after determining the risk level of the target user, during the specific implementation of the method, the following content may further be included: performing corresponding risk marking on the target user according to the risk level of the target user; and / or, determining whether to respond to the target business handling application of the target user according to the risk level of the target user, handling the corresponding target business for the target user, and adopting what matching method to handle the target business for the target user.
[0089] Among them, the target business may specifically include at least one of the following: wealth management business, VIP service business, credit business, credit card business, and so on.
[0090] In some embodiments, during specific implementation, the risk level of the target user can be compared with a preset first-level threshold. When it is determined that the risk level of the target user is greater than or equal to the preset first-level threshold, it is determined that the overall risk of the target user is relatively large, and at this time, it can be determined to reject handling the target business for the target user. When it is determined that the risk level of the target user is less than the preset first-level threshold, it is determined that the overall risk of the target user is relatively small, and at this time, it can be determined to handle the target business for the target user.
[0091] In some embodiments, after determining to handle the target business for the target user, during the specific implementation of the method, the following content may further be included: comparing whether the risk level of the target user is less than a preset second-level threshold, where the preset second-level threshold is less than the preset first-level threshold.
[0092] In the case where it is determined that the risk level of the target user is lower than the preset second-level threshold, it can be judged that the target user has a relatively high credibility. At this time, a target service with a relatively higher permission level can be handled for the target user. For example, a credit service with a higher limit can be handled for the target user.
[0093] On the contrary, in the case where it is determined that the risk level of the target user is greater than or equal to the preset second-level threshold, it can be judged that the target user lacks credibility. At this time, a target service with a relatively lower permission level can be handled for the target user. For example, a credit service with a lower limit can be handled for the target user.
[0094] In some embodiments, after handling the target service for the target user, when the method is specifically implemented, the following content may further be included: at intervals of a preset time period (for example, every week), collect the service data of the target user within the preset time period; and call a preset user risk prediction model to process the service data of the target user within the preset time period, so that the change situation of the risk level of the target user can be tracked and analyzed in real time according to the target processing result corresponding to the preset time period; furthermore, the target service handled by the target user can be adjusted in a timely manner according to the change situation of the risk level of the target user.
[0095] Specifically, for example, when it is determined through tracking and analysis that the risk level of the target user is becoming higher and higher according to the change situation of the risk level of the target user, the permission level of the target user based on the target service can be gradually recovered and controlled until the target service of the target user is suspended, so that risks can be timely identified and discovered and avoided, and the risk loss of the service handling institution can be reduced.
[0096] In some embodiments, when the method is specifically implemented, the following content may further be included:
[0097] S1: Obtain a plurality of data tables of the service system; wherein, each of the plurality of data tables contains a plurality of service data of sample users;
[0098] S2: Perform a preset clustering process on the plurality of data tables to obtain a plurality of sample data sets of the sample users; wherein, each of the plurality of sample data sets contains service data with business relevance corresponding to a sub-business scenario;
[0099] S3: Use the plurality of sample data sets of the sample users to train an initial model to obtain the preset user risk prediction model.
[0100] Through the above embodiments, multiple data tables can be subjected to preset clustering processing first, and multiple sample data sets with business relevance and rich features of sample users can be clustered; furthermore, the initial model can be comprehensively trained by using the multiple sample data sets of the above sample users to obtain a preset user risk prediction model with high accuracy and good effect.
[0101] In some embodiments, a data table directly collected by a business system may contain business data for different sub-business scenarios at the same time, and different data tables may respectively contain business data corresponding to the same sub-business scenario. By performing clustering, multiple business data with high relevance for the same sub-business scenario of the same sample user can be clustered together to obtain the first sample data set of the sample user with rich data for this business scenario.
[0102] In some embodiments, to perform preset clustering processing on the multiple data tables to obtain multiple sample data sets of sample users, the specific implementation may include the following: Based on the K-means clustering algorithm, cluster the multiple data tables according to the identity identifier of the sample user to obtain multiple aggregated tables; wherein, the aggregated table contains business data corresponding to the identity identifier of a sample user and a sub-business scenario; according to the multiple aggregated tables, construct multiple sample data sets of the sample user.
[0103] Through the above embodiments, the preset clustering processing can be efficiently completed by using the K-means clustering algorithm, and multiple sample data sets of sample users with rich features and more suitable for model training can be obtained.
[0104] In some embodiments, the above aggregated table may specifically correspond to the identity identifier of a sample user and a specific sub-business scenario. Specifically, an aggregated table may contain one or more business data with business relevance for a certain sub-business scenario of the corresponding sample user.
[0105] In some embodiments, when specifically clustering the multiple data tables according to the identity identifier of the sample user based on the K-means clustering algorithm, first, according to the identity identifier of the sample user, retrieve the multiple data tables to find multiple business data corresponding to the identity identifier of the same sample user; then, calculate the field values of the multiple business data according to the semantic processing rules; furthermore, the K-means clustering algorithm can be used to cluster the field values of the above multiple business data; and according to the clustering result, combine the multiple business data clustered together into an aggregated table corresponding to the identity identifier of the sample user.
[0106] In some embodiments, the initial model can be specifically constructed in the following manner:
[0107] S1: Based on the random forest algorithm, construct multiple initial sub-models for multiple sub-business scenarios; and combine the multiple initial sub-models to obtain an initial first-layer model.
[0108] S2: Based on the decision tree algorithm, construct an initial second-layer model.
[0109] S3: Connect the multiple initial sub-models in the initial first-layer model to the initial second-layer model to obtain the initial model.
[0110] Through the above embodiments, an initial model with a better effect and a double-layer result including an initial first-layer model and an initial second-layer model can be constructed.
[0111] In some embodiments, after performing a preset clustering process on the multiple data tables to obtain multiple sample data sets of sample users, when the method is specifically implemented, the following content may further be included: According to a preset verification rule, perform data cleaning processing on the multiple sample data sets of the sample users respectively to filter out invalid data.
[0112] Among them, the above invalid data can be specifically understood as data that has little impact on the training of the preset user risk prediction model or is likely to introduce data errors.
[0113] Through the above embodiments, after obtaining the multiple sample data sets of sample users, before specifically training the model using the multiple sample data sets of the sample users, the multiple sample data sets can be first subjected to data cleaning processing to filter out invalid data, avoiding introducing model errors or causing overfitting in the subsequent model training process due to the use of invalid data, so that a preset user risk prediction model with relatively higher accuracy can be trained.
[0114] In some embodiments, when specifically implemented, the corresponding invalid data can be identified and determined according to a preset verification rule. Among them, the above preset verification rule can be specifically obtained by summarizing the historical business data in the corresponding application scenario in advance.
[0115] In some embodiments, specifically, the invalid data includes at least one of the following: business data whose data generation time is greater than a preset time threshold (for example, business data from ten years ago), business data whose data format does not conform to the preset requirements (for example, business data with incorrect formats), business data with null data values, and so on. Of course, the above-listed invalid data is only an illustrative description. In specific implementation, according to the specific application scenario and processing requirements, the above invalid data may further include other types of data, such as business data that conflicts with existing business data, or business data that duplicates existing business data.
[0116] Through the above implementation, various invalid data can be identified and filtered from the sample data set according to the preset verification rules, so that a sample data set with higher accuracy and better effect can be obtained.
[0117] In some embodiments, training the initial model using multiple sample data sets of the sample user to obtain the preset user risk prediction model may specifically include the following contents in specific implementation:
[0118] S1: Using multiple sample data sets of the sample user to separately train the corresponding initial sub-models in the initial first-layer model to obtain a first-layer model that meets the requirements;
[0119] S2: Invoking the first-layer model to process multiple sample data sets of the sample user to obtain multiple intermediate processing results;
[0120] S3: Using the multiple intermediate processing results to train the initial second-layer model to obtain a second-layer model that meets the requirements.
[0121] Through the above embodiments, a first-layer model and a second-layer model that meet the requirements can be sequentially trained, and thus a preset user risk prediction model that meets the requirements can be obtained.
[0122] In some embodiments, when specifically training the first-layer model, taking the current initial sub-model among the multiple initial sub-models included in the first-layer model as an example, the risk labels of the overall target user and the sub-risk labels of the sample data set of the target user for specific sub-business scenarios can be marked first to obtain multiple labeled sample data sets of the sample user; then, the current initial sub-model is continuously trained using the labeled sample data set corresponding to the current initial sub-model in the multiple labeled sample data sets of the sample user to determine multiple CART trees, where each CART tree can correspond to a type of feature extraction and processing structure in this sub-business scenario; then, the above multiple CART trees are combined to obtain the corresponding random forest model as the current sub-model.
[0123] Through the above embodiments, multiple initial sub-models can be respectively trained using multiple sample data sets of sample users to obtain corresponding multiple sub-models, thereby completing the training of the first-layer model.
[0124] In some embodiments, when specifically training the second-layer model, the already trained first-layer model can be first called to process multiple sample data sets of sample users to obtain risk parameters in multiple sub-business scenarios as multiple intermediate processing results. Then, the multiple intermediate processing results are combined with the risk labels of the sample users to obtain multiple combined sample data groups; wherein, the combined sample data groups each contain multiple intermediate results corresponding to one sample user; the multiple combined sample data groups are divided into a training set and a test set; and further, the training set and the test set can be used to train and test the initial second-layer model to obtain a second-layer model that meets the requirements.
[0125] In some embodiments, for the banking business handling scenario, when training the initial second-layer model using the multiple intermediate processing results, the method may further include the following when specifically implemented: associating with the credit investigation system to obtain the credit risk parameters of the sample users; and using the credit risk parameters of the sample users to correct the second-layer model.
[0126] Through the above embodiments, during the process of training the second-layer model, the credit investigation system can also be introduced and utilized to perform targeted optimization and correction on the second-layer model, so that a second-layer model with higher accuracy and better effect can be obtained.
[0127] In some embodiments, when specifically using the credit risk parameters of the sample users to correct the second-layer model, it may include: calling the second-layer model to determine the target risk level of the sample users; comparing the risk level of the sample users with the credit risk parameters to obtain a corresponding comparison result; and correcting the second-layer model according to the comparison result.
[0128] Specifically, according to the comparison result, when it is determined that the difference value between the target risk level and the credit risk parameters is greater than a preset difference value, the sample user can be marked as a positive training sample; and the second-layer model can be specifically trained using the positive training sample to make the training data of the second-layer model progress towards the correct training result, realizing the optimization and correction of the second-layer model.
[0129] In some embodiments, when training the initial second-layer model using the multiple intermediate processing results, the method may further include the following when specifically implemented: associating with the business system to obtain the business risk labels of the sample users; and using the business risk labels of the sample users to correct the second-layer model.
[0130] In some embodiments, when specifically using the business risk tags of sample users to correct the second-layer model, it may include: calling the second-layer model to determine the target risk level of the sample user; comparing the risk level of the sample user with the business risk tags to obtain a corresponding comparison result; and correcting the second-layer model according to the comparison result.
[0131] Specifically, according to the comparison result, when it is determined that the difference value between the target risk level and the business risk tag is greater than a preset difference value, this sample user can be marked as a positive training sample; and the second-layer model can be specifically trained using this positive training sample, so that the training data of the second-layer model progresses towards the correct training result, realizing the optimization and correction of the second-layer model.
[0132] In some embodiments, during specific implementation, the above two optimization methods can also be used simultaneously to correct and optimize the second-layer model, obtaining a second-layer model with relatively higher accuracy.
[0133] In some embodiments, when specifically implementing the above-mentioned process of using a preset user risk prediction model to process the business data of the target user, it may include the following contents: performing a preset clustering process on the business data of the target user to obtain multiple data sets of the target user; inputting the multiple data sets of the target user into the preset user risk prediction model according to a preset input rule; and running the preset user risk prediction model to obtain a corresponding risk level as the target processing result.
[0134] In some embodiments, specifically, after combining multiple data sets in a preset order according to a preset input rule, they can be input into the preset user risk prediction model. Among them, the above-mentioned preset order can specifically be the arrangement order of multiple sub-models in the first-layer model. The business data of the above-mentioned target user can specifically be the full-volume business data of the target user.
[0135] Through the above embodiments, a preset user risk prediction model can be used to comprehensively process the full-volume business data of the target user, so as to more comprehensively and accurately determine the risk level of the target user based on the overall dimension.
[0136] In some embodiments, when only the risk level of the target user for one or more target sub-business scenarios needs to be predicted, after performing a preset clustering process on the business data of the target user to obtain multiple data sets of the target user, the data sets corresponding to the target sub-business scenarios can be first screened out from the multiple data sets and recorded as target data sets;
[0137] According to the preset input rules, input the target data set into the preset user risk prediction model (for example, other data sets except the target data set can be set to be empty, and then combined according to the preset sorting and input into the preset user risk prediction model); and run the preset user risk prediction model to obtain the risk level for the target sub-business scenario as the target processing result.
[0138] In some embodiments, after handling the target business for the target user, when the method is specifically implemented, it may further include the following: at intervals of a preset time period, collect the business data of the target user in the preset time period; call the preset user risk prediction model to process the business data of the target user in the preset time period to track and analyze the change of the risk level of the target user in real time; according to the change of the risk level of the target user, adjust the execution of the target business for the target user.
[0139] As can be seen from the above, before the specific implementation of the method for determining user risk provided in the embodiments of this specification, through preset clustering processing on multiple data tables containing all business data, multiple sample data sets in which multiple sample users are clustered based on business relevance can be obtained; then, using the multiple sample data sets of multiple sample users, a preset user risk prediction model with a high accuracy and a double-layer structure including a first-layer model and a second-layer model can be trained; among them, the first-layer model further includes multiple sub-models corresponding to multiple sub-business scenarios respectively; when specifically implemented, after obtaining the business data of the target user, the above preset user risk prediction model can be called to comprehensively process the business data of the target user, so as to more comprehensively and accurately determine the risk level of the target user, thereby effectively reducing the error in determining user risk and improving the risk prediction accuracy.
[0140] Refer to Figure 4 As shown, the embodiments of this specification also provide a training method for a preset user risk prediction model. Among them, when the method is specifically implemented, it may include the following:
[0141] S401: Obtain multiple data tables of the business system; wherein, each of the multiple data tables contains multiple business data of sample users.
[0142] S402: Perform preset clustering processing on the multiple data tables to obtain multiple sample data sets of sample users; wherein, each of the multiple sample data sets contains business data with business relevance corresponding to a sub-business scenario.
[0143] S403: Train an initial model using multiple sample data sets of the sample user to obtain the preset user risk prediction model; wherein, the initial model includes an initial first-layer model and an initial second-layer model, the initial first-layer model includes multiple initial sub-models, and each of the initial sub-models corresponds to a sub-business scenario.
[0144] In some embodiments, the multiple business data of the sample user included in the above multiple data tables may constitute the full-volume business data of the sample user.
[0145] As can be seen from the above, based on the training method of the preset user risk prediction model provided in the embodiments of this specification, the full-volume business data of the sample user can be fully utilized to train a preset user risk prediction model with a wide application range, being relatively comprehensive and accurate.
[0146] The embodiments of this specification also provide a server, including a processor and a memory for storing processor-executable instructions. When specifically implemented, the processor may execute the following steps according to the instructions: obtain the business data of the target user; call the preset user risk prediction model to process the business data of the target user to obtain a corresponding target processing result; wherein, the preset user risk prediction model includes a first-layer model and a second-layer model, the first-layer model includes multiple sub-models, and each of the multiple sub-models corresponds to a sub-business scenario; determine the risk level of the target user according to the target processing result.
[0147] To be able to complete the above instructions more accurately, refer to Figure 5 As shown, the embodiments of this specification also provide another specific server. Among them, the server includes a network communication port 501, a processor 502, and a memory 503. The above structures are connected by internal cables so that each structure can perform specific data interactions.
[0148] Among them, the network communication port 501 can specifically be used to obtain the business data of the target user.
[0149] The processor 502 can specifically be used to call the preset user risk prediction model to process the business data of the target user to obtain a corresponding target processing result; wherein, the preset user risk prediction model includes a first-layer model and a second-layer model, the first-layer model includes multiple sub-models, and each of the multiple sub-models corresponds to a sub-business scenario; determine the risk level of the target user according to the target processing result.
[0150] The memory 503 can specifically be used to store the corresponding instruction programs.
[0151] In this embodiment, the network communication port 501 can be bound to different communication protocols, so as to send or receive different data through virtual ports. For example, the network communication port can be a port responsible for web data communication, or a port responsible for FTP data communication, or a port responsible for mail data communication. In addition, the network communication port can also be a physical communication interface or communication chip. For example, it can be a wireless mobile network communication chip, such as GSM, CDMA, etc.; it can also be a Wifi chip; it can also be a Bluetooth chip.
[0152] In this embodiment, the processor 502 can be implemented in any suitable manner. For example, the processor can be in the form of, for example, a microprocessor or a processor, and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, application specific integrated circuit (ASIC), programmable logic controller, and embedded microcontroller, etc. This specification does not make any limitations.
[0153] In this embodiment, the memory 503 can include multiple levels. In a digital system, anything that can store binary data can be a memory; in an integrated circuit, a circuit with a storage function without a physical form is also called a memory, such as RAM, FIFO, etc.; in a system, a storage device with a physical form is also called a memory, such as a memory module, TF card, etc.
[0154] The embodiment of this specification also provides a computer-readable storage medium based on the above user risk determination method. The computer-readable storage medium stores computer program instructions, which when executed, implement: obtaining service data of a target user; calling a preset user risk prediction model to process the service data of the target user to obtain a corresponding target processing result; where the preset user risk prediction model includes a first-layer model and a second-layer model, the first-layer model includes multiple sub-models, and the multiple sub-models respectively correspond to a sub-business scenario; determining the risk level of the target user according to the target processing result.
[0155] In this embodiment, the above storage medium includes, but is not limited to, random access memory (RAM), read-only memory (ROM), cache, hard disk drive (HDD), or memory card. The memory can be used to store computer program instructions. The network communication unit can be set according to the standards specified by the communication protocol and is used for the interface of network connection communication.
[0156] In this embodiment, the functions and effects specifically realized by the program instructions stored in the computer-readable storage medium can be explained by comparison with other embodiments and will not be elaborated here.
[0157] Refer to Figure 6 As shown, at the software level, an embodiment of this specification also provides a device for determining user risks. The device may specifically include the following structural modules:
[0158] An acquisition module 601, which can specifically be used to acquire the service data of the target user;
[0159] An invocation module 602, which can specifically be used to invoke a preset user risk prediction model to process the service data of the target user and obtain a corresponding target processing result; wherein, the preset user risk prediction model includes a first-layer model and a second-layer model, and the first-layer model includes multiple sub-models, and the multiple sub-models respectively correspond to a sub-service scenario;
[0160] A determination module 603, which can specifically be used to determine the risk level of the target user according to the target processing result.
[0161] It should be noted that the units, devices, or modules, etc. described in the above embodiments can specifically be implemented by computer chips or entities, or by products with certain functions. For the convenience of description, when describing the above devices, various modules are described separately according to their functions. Of course, when implementing this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0162] As can be seen from the above, the user risk determination device provided by the embodiments of this specification can comprehensively process the business data of the target user by invoking a preset user risk prediction model with high accuracy and wide application range, and can determine the risk level of the target user more comprehensively and accurately. Thus, the error in determining user risks can be effectively reduced, and the risk prediction accuracy can be improved.
[0163] In a specific scenario example, the user risk determination method provided by the embodiments of this specification can be applied. First, based on the full industry business data of the supervision system, the full industry business is clustered according to relevance to form several data modules with strong business relevance (for example, multiple sub-business scenarios). For example, deposit and loan data modules, customer information modules, financial wealth management modules, credit card modules, etc. A machine learning model (for example, a sub-model in the first layer model) can be trained for each data module. Then, an initial analysis of the customer (for example, the target user) can be performed from a multi-dimensional business perspective, and then each initial analysis result (for example, the intermediate processing result) is used as the input of the customer comprehensive analysis model (for example, the second layer model), so as to finally determine the risk level of the customer. At the same time, the credit investigation system and the business system can be linked, and the data can be corrected based on the data comparison results of the credit investigation system and the business system to form positive training samples and input them into the model to optimize the model algorithm, thereby realizing real-time, comprehensive, and accurate risk analysis of bank customers.
[0164] For the specific implementation process, the following content can be referred to for execution.
[0165] Based on the user risk determination method provided by the embodiments of this specification, a customer risk analysis system based on the supervision system can be constructed first. Specifically, reference can be made to Figure 7 as shown. This system can specifically include the following structures: a data preparation device 1, a risk level model device 2, and a model adjustment device 3.
[0166] When the above system is specifically implemented, first, the data preparation device 1 can perform clustering, cleaning, and feature extraction on the data collected by the supervision system (for example, multiple data tables) to complete data preparation. Then, the prepared data is input into the risk level model device 2 to perform risk analysis on the customer and complete the customer risk analysis. Finally, the model adjustment device 3 is used to optimize the model algorithm.
[0167] Refer to Figure 8 as shown. The data preparation device 1 can specifically include the following structures: a clustering module, a cleaning module, and a feature extraction module. The entire device realizes data clustering according to business relevance and data cleaning through a clustering algorithm and the verification rules preset in the supervision system. Then, feature extraction is realized according to the input requirements of the risk level model device 2. The core of this device is to perform grouped feature extraction on the data according to business relevance.
[0168] When the data preparation device 1 is specifically implemented, a clustering algorithm (for example, the K-means algorithm) can be used to group data with high business relevance. The purpose is to be able to extract feature data with rich enough data dimensions (for example, the approval information and card-issuing information of customer A's credit card were originally in different data tables. Through the clustering algorithm, the information in these two parts can be aggregated together, and then more rich feature data based on the business category dimension can be obtained during feature extraction). This not only ensures the comprehensiveness of the data but also improves the accuracy of subsequent model training, providing a solid data foundation for subsequent model algorithms. Among them, the data preparation process can specifically include the following contents:
[0169] 1) Obtain the full-business aggregation table data of the supervision system, and cluster the aggregated data according to business relevance through a clustering algorithm, dividing the data into several large groups (corresponding to multiple sample data sets). Among them, the clustering can be achieved through field information such as customer numbers, accounting subject numbers, and customer account numbers in the aggregation table. For example, for customer Zhang San in the following table, his customer number "001" indicates an individual customer, and the accounting subject number "010102" indicates a deposit subject. His relevant data is in aggregation table A, and there is also customer Zhang San in aggregation table B with the same account number "2347658389", and the subject "10102" indicates a wealth management business, indicating that this customer also has a wealth management business under the account "2347658389". Therefore, these two pieces of data in aggregation tables A and B are of business relevance and can be clustered into the deposit and loan data module. Similarly, the data of the supervision system can also be clustered into business categories such as financial wealth management modules and credit card modules, so as to obtain as comprehensive information as possible about each customer and classify the information according to established strategies. Specifically, refer to the content shown in Table 1.
[0170] Table 1
[0171] Name Customer Number Accounting Subject Number Customer Account Aggregate Table to Which Belongs Business Module Zhang San 001 010102 2347658389 A Deposit and Loan Data Module Zhang San 001 010103 2347658389 B Deposit and Loan Data Module Li Si 003 010104 2347658391 C Credit Card Module
[0172] 2) Input the clustered data into the cleaning module, and eliminate garbage data, error data, and null data in combination with the verification rules preset in the supervision system (for example, filter out invalid data) to ensure the integrity and accuracy of the training data required by the model as much as possible.
[0173] 3) Input the cleaned data into the feature selection module, and perform feature selection according to the large module categories after clustering. For example, the "deposit and loan data module" extracts information such as customer numbers, deposit amounts, monthly transaction volumes, loan amounts, and overdue amounts as input information for the deposit and loan data model; the "credit card module" extracts information such as customer numbers, our bank's approval amounts, other banks' approval amounts, overdraft amounts, and card statuses as input information for the credit card model.
[0174] Refer to Figure 9 As shown, the risk level model device 2 can specifically be composed of multiple sets of machine learning models. Among them, the risk level model device can be divided into a two-layer structure. The first layer (for example, the first layer model) consists of multiple sets of initial analysis models in parallel to form an initial analysis model group. Among them, each set of models (for example, sub-models) corresponds to a major business module (for example, a sub-business scenario), so that each model analyzes the risk level of customers from different business dimensions. This processing method improves the stability of the entire risk level model and the accuracy of analysis by dividing the responsibilities of the machine learning models according to business dimensions. The initial analysis model of the first layer is constructed using the random forest algorithm.
[0175] Specifically, for example, taking the major credit card business as an example, it corresponds to the credit card model in the first layer. The data of this major business is extracted according to three major feature categories (credit granting features, card-issuing features, and post-loan features). Each feature type corresponds to a CART tree, and the analysis examples of each CART tree can be referred to as shown in Tables 2, 3, and 4.
[0176] Table 2 Example Table of Credit Granting Features
[0177]
[0178]
[0179] Table 3 Example Table of Card-Issuing Features
[0180] Name Customer Number Overdraft Amount Card Status Invested Industry Number of Extension Periods Risk Level Zhang San 001 50 yuan Normal General Consumption 0 Risk-free Li Si 002 100 yuan Overdue General Consumption 1 Low Risk
[0181] Table 4 Example Table of Post-Loan Features
[0182] Name Customer Number Whether Asset Transfer Transfer Amount Whether Write-off Write-off Amount Risk Level Zhang San 001 No 0 0 0 Risk-free Li Si 002 No 0 Yes 100 Low Risk Wang Wu 003 Yes 100 Yes 100 High Risk
[0183] For the remaining initialization models, similar to the credit card model, select appropriate feature data for the major feature categories and complete the preliminary risk analysis for each major business.
[0184] The second layer (for example, the second layer model) is a comprehensive risk analysis model, which uses the decision tree algorithm to implement the model. The output results of multiple sets of initial analysis models in the first layer can be used as the input data for the second layer, and finally the comprehensive risk level of the customer is determined. The comprehensive risk analysis model can be linked to the credit investigation system and the business system. By comparing with the customer credit investigation data in the credit investigation system and the revised data in the business system, positive training samples are formed and input into the decision tree model, thereby realizing algorithm optimization.
[0185] The specific optimization process can be referred to Figure 10 As shown. The model adjustment device 3 can realize the algorithm self-optimization of the comprehensive risk analysis model through the following two ways respectively.
[0186] 1) Connect to the customer credit investigation system, compare the model analysis results with the risk levels of the corresponding customers in the credit investigation system. If the comparison results are at the two extremes of the risk levels (i.e., the model analysis results are very different from the credit investigation system results), then match the corresponding feature vectors to form positive training samples, input them into the comprehensive risk analysis model, and optimize the decision-making algorithm.
[0187] 2) Connect the customer risk analysis system to the business system. Business personnel can query the customer risk level in the business system. If they think the customer risk level is very different from the actual situation, they can mark the customer. At the same time, the system matches the corresponding feature vectors to form positive training samples, input them into the comprehensive risk analysis model, and optimize the decision-making algorithm.
[0188] Through the above scenario examples, based on the whole industry business data provided by the supervision system, the system can cluster the whole industry business according to the relevance, forming several data modules with strong business relevance, such as deposit and loan data modules, customer information modules, financial wealth management modules, credit card modules, etc. Each data module corresponds to a machine learning model, and then the initial analysis of customers can be carried out from multiple business perspectives. Then, each initial analysis result is used as the input scenario of the customer comprehensive analysis model, and finally the customer risk level is decided. Specifically, it can have the following multiple advantages:
[0189] 1. Through the clustering algorithm, the whole industry business data is accurately positioned according to the business relevance, enriching the dimension of the feature data, thus completing the data preparation work for improving the model analysis accuracy;
[0190] 2. Adopt multiple sets of machine learning models to jointly realize risk analysis, enabling each model to analyze the customer risk level from different business perspectives, increasing the stability of the model and the accuracy of the analysis;
[0191] 3. By linking the credit investigation system and the business system, dynamically adjust the model parameters, and then realize the algorithm optimization, so as to obtain a model with higher accuracy.
[0192] Although this specification provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or client product is executed, it can be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in a parallel processor or multi-threaded processing environment, or even in a distributed data processing environment). The terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, product or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device comprising the said elements. Words such as first, second, etc. are used to denote names and do not denote any particular order.
[0193] As is also known to those skilled in the art, in addition to implementing the controller in the form of pure computer-readable program code, the method steps can be logically programmed to enable the controller to implement the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. Therefore, such a controller can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0194] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer-readable storage media including storage devices.
[0195] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that this specification can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this specification can essentially be embodied in the form of a software product, which can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of this specification.
[0196] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. This specification can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.
[0197] Although this specification is depicted through embodiments, those of ordinary skill in the art know that this specification has many variations and changes without departing from the spirit of this specification. It is hoped that the appended claims will cover these variations and changes without departing from the spirit of this specification.
Claims
1. A method for determining user risk, characterized in that, it includes: Obtain the business data of the target user; Call a preset user risk prediction model to process the business data of the target user to obtain a corresponding target processing result; wherein, the preset user risk prediction model is a model with a two-layer structure including a first-layer model and a second-layer model, the first-layer model includes multiple sub-models, and the multiple sub-models respectively correspond to a sub-business scenario; the sub-business scenarios include: deposit and loan business scenarios, credit card business scenarios, and financial wealth management business scenarios; the first-layer model is a model structure constructed based on the random forest algorithm; the second-layer model is a model structure constructed based on the decision tree algorithm; Determine the risk level of the target user according to the target processing result; Among them, calling a preset user risk prediction model to process the business data of the target user to obtain a corresponding target processing result includes: performing a preset clustering process on the business data of the target user to obtain multiple groups of business data of the target user; each group of business data corresponds to a sub-business scenario; input the multiple groups of business data of the target user into the corresponding sub-model in the first-layer model of the preset user risk prediction model according to a preset input rule for processing to obtain corresponding multiple intermediate processing results; input the multiple intermediate processing results into the second-layer model through the first-layer model for processing to obtain a corresponding target processing result; Among them, the preset user risk prediction model is trained in the following manner: use multiple sample data sets of sample users to train the corresponding initial sub-models in the initial first-layer model respectively to obtain a first-layer model that meets the requirements; call the first-layer model to process multiple sample data sets of sample users to obtain multiple intermediate processing results; use the multiple intermediate processing results to train the initial second-layer model to obtain a second-layer model that meets the requirements; Among them, training the corresponding initial sub-models in the initial first-layer model respectively includes: training the current initial sub-model in the following manner: label the overall risk label of the sample user and the sub-risk label of the sample data set of the sample user for the sub-business scenario to obtain multiple labeled sample data sets of the sample user; continuously train the current initial sub-model using the labeled sample data sets corresponding to the current initial sub-model among the multiple labeled sample data sets of the sample user to determine multiple CART trees, where each CART tree is used to correspond to a feature extraction and processing structure in the sub-business scenario; combine the multiple CART trees to obtain a corresponding random forest model as the current sub-model in the first-layer model; Invoke the first-layer model to process multiple sample data sets of a sample user, and obtain multiple intermediate processing results; use the multiple intermediate processing results to train the initial second-layer model, including: Invoke the trained first-layer model to process multiple sample data sets of a sample user, and obtain risk parameters under multiple sub-business scenarios as multiple intermediate processing results; combine the multiple intermediate processing results with the risk labels of the sample user to obtain multiple combined sample data groups; wherein, the combined sample data group contains multiple intermediate results corresponding to one sample user; divide the multiple combined sample data groups into a training set and a test set; use the training set and the test set to train and test the initial second-layer model to obtain a second-layer model that meets the requirements.
2. The method according to claim 1, wherein, the method further includes: Obtain multiple data tables of the business system; wherein, each of the multiple data tables contains multiple business data of the sample user; Perform a preset clustering process on the multiple data tables to obtain multiple sample data sets of the sample user; wherein, each of the multiple sample data sets contains business data with business relevance corresponding to one sub-business scenario; Use the multiple sample data sets of the sample user to train the initial model to obtain the preset user risk prediction model.
3. The method according to claim 2, wherein, Performing a preset clustering process on the multiple data tables to obtain multiple sample data sets of the sample user includes: Based on the K-means clustering algorithm, cluster the multiple data tables according to the identity identifier of the sample user to obtain multiple aggregated tables; wherein, the aggregated table contains the identity identifier of one sample user and the business data of one sub-business scenario; Construct multiple sample data sets of the sample user according to the multiple aggregated tables.
4. The method according to claim 2, wherein, The initial model is constructed in the following manner: Based on the random forest algorithm, construct multiple initial sub-models for multiple sub-business scenarios; and combine the multiple initial sub-models to obtain the initial first-layer model; Based on the decision tree algorithm, construct the initial second-layer model; Connect the multiple initial sub-models in the initial first-layer model to the initial second-layer model to obtain the initial model.
5. The method according to claim 2, wherein, After performing a preset clustering process on the multiple data tables to obtain multiple sample data sets of the sample user, the method further includes: According to the preset verification rules, perform data cleaning on each of the multiple sample data sets of the sample user to filter out invalid data.
6. The method according to claim 5, wherein, The invalid data includes at least one of the following: business data whose data generation time is greater than the preset time threshold, business data whose data format does not meet the preset requirements, and business data whose data value is empty.
7. The method according to claim 1, wherein, While training the initial second-layer model by using the multiple intermediate processing results, the method further includes: Associating with a credit investigation system to obtain the credit risk parameters of sample users; Using the credit risk parameters of sample users to correct the second-layer model.
8. A method for training a preset user risk prediction model, Characterized in that, It includes: Obtaining multiple data tables of a business system; wherein, each of the multiple data tables contains multiple business data of sample users; Performing a preset clustering process on the multiple data tables to obtain multiple sample data sets of sample users; wherein, each of the multiple sample data sets contains business data with business relevance corresponding to a sub-business scenario; the sub-business scenarios include: deposit and loan business scenarios, credit card business scenarios, and financial wealth management business scenarios; Using the multiple sample data sets of sample users to train an initial model to obtain the preset user risk prediction model; wherein, the initial model includes an initial first-layer model and an initial second-layer model, the initial first-layer model includes multiple initial sub-models, and the initial sub-models respectively correspond to a sub-business scenario; the initial first-layer model is a model structure constructed based on the random forest algorithm; the initial second-layer model is a model structure constructed based on the decision tree algorithm; Wherein, the preset user risk prediction model is used to process the business data of a target user to obtain a corresponding target processing result, including: performing a preset clustering process on the business data of the target user to obtain multiple groups of business data of the target user; each group of business data corresponds to a sub-business scenario; inputting the multiple groups of business data of the target user into the corresponding sub-model of the first-layer model of the preset user risk prediction model according to a preset input rule for processing to obtain corresponding multiple intermediate processing results; inputting the multiple intermediate processing results through the first-layer model into the second-layer model for processing to obtain a corresponding target processing result; Wherein, using the multiple sample data sets of sample users to train the initial model to obtain the preset user risk prediction model includes: using the multiple sample data sets of sample users to respectively train the corresponding initial sub-models in the initial first-layer model to obtain a first-layer model that meets the requirements; calling the first-layer model to process the multiple sample data sets of sample users to obtain multiple intermediate processing results; using the multiple intermediate processing results to train the initial second-layer model to obtain a second-layer model that meets the requirements; Among them, the corresponding initial sub-models in the initial first-layer model are trained separately, including: training the current initial sub-model in the following manner: marking the risk labels of the overall sample user and the sub-risk labels of the sample data set of the sample user for the sub-business scenario, so as to obtain multiple labeled sample data sets of the sample user; continuously training the current initial sub-model with the labeled sample data set corresponding to the current initial sub-model in the multiple labeled sample data sets of the sample user to determine multiple CART trees, where each CART tree is used to correspond to a class of feature extraction and processing structures in the sub-business scenario; combining the multiple CART trees to obtain the corresponding random forest model as the current sub-model in the first-layer model; Call the first-layer model to process multiple sample data sets of the sample user to obtain multiple intermediate processing results; use the multiple intermediate processing results to train the initial second-layer model, including: calling the trained first-layer model to process multiple sample data sets of the sample user to obtain risk parameters in multiple sub-business scenarios as multiple intermediate processing results; combining the multiple intermediate processing results with the risk labels of the sample user to obtain multiple combined sample data groups; where the combined sample data group contains multiple intermediate results corresponding to one sample user; dividing the multiple combined sample data groups into a training set and a test set; using the training set and the test set to train and test the initial second-layer model to obtain a second-layer model that meets the requirements.
9. A device for determining user risk, characterized in that, it includes: an acquisition module for acquiring the business data of the target user; a call module for calling a preset user risk prediction model to process the business data of the target user to obtain a corresponding target processing result; where the preset user risk prediction model is a model with a two-layer structure including a first-layer model and a second-layer model, the first-layer model includes multiple sub-models, and the multiple sub-models respectively correspond to a sub-business scenario; the sub-business scenarios include: deposit and loan business scenarios, credit card business scenarios, and financial wealth management business scenarios; the first-layer model is a model structure constructed based on the random forest algorithm; the second-layer model is a model structure constructed based on the decision tree algorithm; a determination module for determining the risk level of the target user according to the target processing result; wherein, the call module is specifically used for: performing a preset clustering process on the business data of the target user to obtain multiple groups of business data of the target user; each group of business data corresponds to a sub-business scenario; inputting the multiple groups of business data of the target user into the corresponding sub-model in the first-layer model of the preset user risk prediction model according to a preset input rule for processing to obtain corresponding multiple intermediate processing results; inputting the multiple intermediate processing results into the second-layer model through the first-layer model for processing to obtain a corresponding target processing result; Among them, the preset user risk prediction model is trained as follows: using multiple sample data sets of sample users, respectively training the corresponding initial sub-models in the initial first-layer model to obtain a first-layer model that meets the requirements; calling the first-layer model to process multiple sample data sets of sample users to obtain multiple intermediate processing results; using the multiple intermediate processing results to train the initial second-layer model to obtain a second-layer model that meets the requirements. Among them, respectively training the corresponding initial sub-models in the initial first-layer model includes: training the current initial sub-model in the following manner: marking the risk labels of the overall sample user and the sub-risk labels of the sample data set of the sample user for the sub-business scenario to obtain multiple labeled sample data sets of the sample user; continuously training the current initial sub-model using the labeled sample data sets corresponding to the current initial sub-model in the multiple labeled sample data sets of the sample user to determine multiple CART trees, where each CART tree is used to correspond to a class of feature extraction and processing structures in the sub-business scenario; combining the multiple CART trees to obtain the corresponding random forest model as the current sub-model in the first-layer model. Calling the first-layer model to process multiple sample data sets of sample users to obtain multiple intermediate processing results; using the multiple intermediate processing results to train the initial second-layer model includes: calling the already trained first-layer model to process multiple sample data sets of sample users to obtain risk parameters in multiple sub-business scenarios as multiple intermediate processing results; combining the multiple intermediate processing results with the risk labels of the sample user to obtain multiple combined sample data groups; where the combined sample data group contains multiple intermediate results corresponding to one sample user; dividing the multiple combined sample data groups into a training set and a test set; using the training set and the test set to train and test the initial second-layer model to obtain a second-layer model that meets the requirements.
10. A server Characterized in that it includes a processor and a memory for storing processor-executable instructions, and when the processor executes the instructions, it implements the steps of the method according to any one of claims 1 to 7.
11. A computer-readable storage medium Characterized in that it stores computer instructions, and when the instructions are executed, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Big-data-based insurance policy underwriting model training method and underwriting risk assessment method
CN110516910A
Method and system for reducing pre-loan business risks
CN111325248A
Training method of an institution risk prediction model and institution risk prediction method and device
CN112561320A
Weight determination model training method, risk prediction method and device
CN112767128A