A Method and Device for Dynamically Identifying High-Risk Users in the Financial Field

Through the dynamic identification method of high-risk users in the financial field, and the debt grading and clustering algorithms are used to solve the problem of dynamic identification of high-risk large investors, achieving more efficient and accurate identification results.

CN113435987BActive Publication Date: 2025-06-17INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110704410.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-24
Publication Date
2025-06-17
Estimated Expiration
2041-06-24

AI Technical Summary

Technical Problem

In the financial field, the difficulty of dynamically identifying high-risk large investors lies in how to quantitatively evaluate the risks of large investors and how to dynamically identify high-risk large investors under different overall risk levels. The existing technology mainly relies on offline manual evaluation, and the efficiency and accuracy are not high.

Method used

By generating feature sets based on user debt ratings and using clustering algorithms to classify users, classification results are generated to dynamically identify high-risk users. Specific steps include data cleaning, risk value calculation, feature set generation, cluster analysis, etc.

Benefits of technology

It has achieved dynamic identification of high-risk large players under different overall risk levels, improved identification efficiency and accuracy, and changed the inefficient status of offline manual evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113435987B_ABST
    Figure CN113435987B_ABST
Patent Text Reader

Abstract

The present invention can be used in the financial field or other fields. The present invention provides a method and device for dynamically identifying high-risk users in the financial field. The method for dynamically identifying high-risk users in the financial field includes: generating a feature set of a user according to the user's debt grading; using a clustering algorithm to classify the user according to the feature set to generate a classification result; and dynamically identifying high-risk users according to the classification result. The dynamic identification of high-risk large customers is realized under different overall risk levels at different time points, and the current situation of low efficiency and low accuracy of offline manual evaluation of high-risk large customers can be changed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of big data, and particularly relates to service calls in a distributed system, and specifically relates to a method and device for dynamically identifying high-risk users in the financial field. Background Art

[0002] In the current financial field, regulatory agencies and in-house systems both require specific matter management for high-risk large customers. It can be understood that high-risk large customers are a relative concept, referring to large customers with relatively higher risks at a certain point in time among all large customers in the bank. That is, the overall risk levels are different at different points in time, and it is necessary to find out the corresponding large customers with relatively higher risks under different overall risk levels.

[0003] According to what criteria to evaluate the risks of large customers for the dynamically changing multi-dimensional large customer information (i.e., Difficulty 1: How to quantitatively evaluate the risks of large customers), and how to accurately identify high-risk large customers from nearly ten thousand large customers in the whole bank dynamically according to the overall risk level at the current point in time (i.e., Difficulty 2: How to dynamically identify high-risk large customers) is undoubtedly a very difficult thing. If too many high-risk large customers are selected, the management difficulty and cost will increase, and if too few are selected, the management risk will increase. Due to the lack of effective technical solutions currently, the current method still uses offline manual evaluation (i.e., based on factors such as customer ratings, guarantees, and business status, and judging by experience) for screening. This method is difficult to be satisfactory in terms of both efficiency and accuracy.

[0004] Therefore, it is particularly important to propose a method for dynamically and automatically identifying high-risk large customers. Summary of the Invention

[0005] It should be noted that a method and device for dynamically identifying high-risk users in the financial field disclosed by the present invention can be used in the financial field and can also be used in any field other than the financial field. The application fields of the method and device for dynamically identifying high-risk users in the financial field disclosed by the present invention are not limited.

[0006] The method and device for dynamically identifying high-risk users in the financial field provided by the present invention can realize the dynamic identification of high-risk large customers under different overall risk levels at the same point in time, and will change the current situation of low efficiency and low accuracy of offline manual evaluation of high-risk large customers.

[0007] To solve the above technical problems, the present invention provides the following technical solutions:

[0008] In the first aspect, the present invention provides a method for dynamically identifying high-risk users in the financial field, including:

[0009] Generating a feature set of a user according to the user's debt grading;

[0010] Using a clustering algorithm, classify the user according to the feature set to generate a classification result;

[0011] Dynamically identify high-risk users according to the classification result.

[0012] In one embodiment, the method for dynamically identifying high-risk users in the financial field further includes: determining the debt grading of the user according to the user's customer rating, guarantee method parameter, mortgage value, economic nature parameter, and business status parameter.

[0013] In one embodiment, the generating the feature set of the user according to the user's debt grading includes:

[0014] Clean the data of the user's debt grading;

[0015] Calculate the current risk value of the user according to the cleaned user's debt grading;

[0016] Generate the feature set according to the current risk value of the user.

[0017] In one embodiment, the clustering algorithm, which classifies the user according to the feature set to generate a classification result, includes:

[0018] Calculate the average value of the current risk values of all users according to the feature set;

[0019] Classify the users according to the average value to generate a first user set and a second user set;

[0020] Perform an iterative operation:

[0021] Randomly select a current user from the first user set;

[0022] Calculate a first difference between the current risk value of the current user and the average value of the current risk values of the first user set, and a second difference between the current risk value of the current user and the average value of the current risk values of the second user set respectively;

[0023] When the first difference is less than the second difference, move the current user to the second user set;

[0024] Until the first difference is not less than the second difference.

[0025] In a second aspect, the present invention provides a device for dynamically identifying high-risk users in the financial field, including:

[0026] A feature set generation module, configured to generate a feature set of a user according to the user's debt grading;

[0027] A classification result generation module, configured to classify the user according to the feature set by using a clustering algorithm to generate a classification result;

[0028] A user identification module, configured to dynamically identify high-risk users according to the classification result.

[0029] In one embodiment, the high-risk user dynamic identification device in the financial field further includes: a debt item grading determination module, configured to determine the debt item grading of the user according to the user's customer rating, guarantee method parameter, mortgage value, economic nature parameter, and business status parameter.

[0030] In one embodiment, the feature set generation module includes:

[0031] A data cleaning unit, configured to perform data cleaning on the user's debt item grading;

[0032] A risk value calculation unit, configured to calculate the current risk value of the user according to the cleaned user's debt item grading;

[0033] A feature set generation unit, configured to generate the feature set according to the current risk value of the user.

[0034] In one embodiment, the classification result generation module includes:

[0035] An average value calculation unit, configured to calculate the average value of the current risk values of all users according to the feature set;

[0036] A user classification unit, configured to classify the user according to the average value to generate a first user set and a second user set;

[0037] An iterative operation unit, configured to perform an iterative operation:

[0038] Randomly select a current user from the first user set;

[0039] Calculate a first difference between the current risk value of the current user and the average value of the current risk values of the first user set, and a second difference between the current risk value of the current user and the average value of the current risk values of the second user set, respectively;

[0040] When the first difference is less than the second difference, move the current user to the second user set;

[0041] Until the first difference is not less than the second difference.

[0042] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the high-risk user dynamic identification method in the financial field are implemented.

[0043] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for dynamically identifying high-risk users in the financial field are implemented.

[0044] As can be seen from the above description, for the method and device for dynamically identifying high-risk users in the financial field provided by the embodiments of the present invention, first, a feature set of a user is generated according to the user's debt rating; then, a clustering algorithm is used to classify the users according to the feature set to generate a classification result; finally, high-risk users are dynamically identified according to the classification result. The present invention realizes the dynamic identification of high-risk large customers under different overall risk levels at different time points by establishing a risk assessment model that can quantitatively evaluate the risks of large customers and a dynamic clustering algorithm, and can change the current situation of low efficiency and low accuracy of offline manual evaluation of high-risk large customers. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0046] Figure 1 Schematic diagram of the process of the method for dynamically identifying high-risk users in the financial field in the embodiments of the present invention Figure 1 ;

[0047] Figure 2 Schematic diagram of the process of the method for dynamically identifying high-risk users in the financial field in the embodiments of the present invention Figure 2 ;

[0048] Figure 3 Schematic diagram of the process of step 100 in the method for dynamically identifying high-risk users in the financial field in the embodiments of the present invention;

[0049] Figure 4 Schematic diagram of the process of step 200 in the method for dynamically identifying high-risk users in the financial field in the embodiments of the present invention;

[0050] Figure 5 Schematic diagram of the process of the method for dynamically identifying high-risk users in the financial field in a specific application example of the present invention;

[0051] Figure 6 Mind map of the method for dynamically identifying high-risk users in the financial field in a specific application example of the present invention;

[0052] Figure 7Schematic flowchart of the nested dynamic clustering analysis method in a specific application example of the present invention;

[0053] Figure 8 Schematic structure of the high-risk user dynamic identification device in the financial field in an embodiment of the present invention Figure 1 ;

[0054] Figure 9 Schematic structure of the high-risk user dynamic identification device in the financial field in an embodiment of the present invention Figure 2 ;

[0055] Figure 10 Schematic structure diagram of the feature set generation module 10 in an embodiment of the present invention;

[0056] Figure 11 Schematic structure diagram of the classification result generation module 20 in an embodiment of the present invention;

[0057] Figure 12 Schematic structure diagram of the electronic device in an embodiment of the present invention. Detailed implementation manners

[0058] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0060] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned accompanying drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0061] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The following will describe the present application in detail with reference to the accompanying drawings and in combination with the embodiments.

[0062] An embodiment of the present invention provides a specific implementation manner of a method for dynamically identifying high-risk users in the financial field. Refer to Figure 1 , and the method specifically includes the following content:

[0063] Step 100: Generate a feature set of the user according to the user debt grading.

[0064] Specifically, extract the total financing balance of all major risk customers and the 12-level classification (risk rating) at the current time point to form an initial feature set of the major risk customer object. Then, based on the initial feature set of the major risk customer object, use the risk evaluation model to quantitatively evaluate the risk value of each major risk customer object one by one, and create a risk value quantification feature for each major risk customer object one by one, thereby constructing a risk feature set of major risk customers.

[0065] Step 200: Use the clustering algorithm to classify the user according to the feature set to generate a classification result.

[0066] It can be understood that cluster analysis, also known as cluster analysis, is a statistical analysis method for studying (sample or index) classification problems, and is also an important algorithm in data mining. Cluster analysis is composed of several patterns. Usually, a pattern is a vector of a measurement, or a point in a multi-dimensional space. Cluster analysis is based on similarity. Patterns within a cluster have more similarity than patterns not in the same cluster.

[0067] Step 300: Dynamically identify high-risk users according to the classification result.

[0068] After completing the cluster analysis in Step 300, all users in the 'high-risk major customer set' are the relatively high-risk major customers corresponding to the current time point that need to be identified in this scenario.

[0069] As can be seen from the above description, the method for dynamically identifying high-risk users in the financial field provided by the embodiment of the present invention first generates a feature set of the user according to the user debt grading; then, uses the clustering algorithm to classify the user according to the feature set to generate a classification result; and finally dynamically identifies high-risk users according to the classification result. The present invention realizes the dynamic identification of high-risk major customers under different overall risk levels at different time points by establishing a risk evaluation model that can quantitatively evaluate the risk of major customers and a dynamic clustering algorithm, and will change the current situation of low efficiency and low accuracy of offline manual evaluation of high-risk major customers.

[0070] In one embodiment, referring to Figure 2 , the high-risk user dynamic identification method in the financial field further includes:

[0071] Step 400: Determine the debt grading of the user according to the user's customer rating, guarantee method parameter, mortgage value, economic nature parameter, and business status parameter.

[0072] In one embodiment, referring to Figure 3 , step 100 includes:

[0073] Step 101: Perform data cleaning on the user's debt grading;

[0074] Data cleaning refers to the process of reexamining and validating data, aiming to delete duplicate information, correct existing errors, and provide data consistency. Specifically, it includes:

[0075] Consistency check: Consistency check is to check whether the data meets the requirements according to the reasonable value range and mutual relationship of each variable, and find data that exceeds the normal range, is logically unreasonable, or is contradictory. For example, a variable measured on a 1-7 scale shows a value of 0, or the weight shows a negative number, which should be regarded as exceeding the normal value range. Computer software such as SPSS, SAS, and Excel can automatically identify each variable value that exceeds the defined range. Answers with logical inconsistencies may appear in various forms: for example, many respondents say they drive to work but also report having no car; or respondents report being heavy purchasers and users of a certain brand, but at the same time give a very low score on the familiarity scale. When inconsistencies are found, the questionnaire serial number, record serial number, variable name, error category, etc. should be listed for further verification and correction.

[0076] Handling of invalid values and missing values: Due to survey, coding, and entry errors, there may be some invalid values and missing values in the data, which need to be properly processed. Common processing methods include: estimation, case deletion, variable deletion, and pairwise deletion. Estimation. The simplest method is to replace invalid values and missing values with the sample mean, median, or mode of a certain variable. This method is simple but does not fully consider the existing information in the data, and the error may be relatively large. Another method is to estimate based on the answers of the respondents to other questions through correlation analysis or logical inference between variables. For example, the ownership of a certain product may be related to household income, and the possibility of owning this product can be estimated based on the household income of the respondents.

[0077] Casewise deletion is to remove the samples containing missing values. Since many questionnaires may have missing values, the result of this approach may lead to a significant reduction in the effective sample size and the failure to fully utilize the data that has been collected. Therefore, it is only suitable for cases where key variables are missing or the proportion of samples containing invalid or missing values is very small.

[0078] Variable deletion. If there are many invalid and missing values for a certain variable and this variable is not particularly important for the research problem, then this variable can be considered for deletion. This approach reduces the number of variables available for analysis but does not change the sample size.

[0079] Pairwise deletion is to use a special code (usually 9, 99, 999, etc.) to represent invalid and missing values while retaining all variables and samples in the dataset. However, in specific calculations, only the samples with complete answers are used, so the effective sample size will vary for different analyses due to different variables involved. This is a conservative treatment method that maximally retains the available information in the dataset.

[0080] Using different treatment methods may affect the analysis results, especially when the occurrence of missing values is not random and there is an obvious correlation between variables. Therefore, invalid and missing values should be avoided as much as possible in the survey to ensure the integrity of the data.

[0081] Step 102: Calculate the current risk value of the user according to the cleaned user debt classification;

[0082] Specifically, the current risk value of the user is calculated using formula (1):

[0083]

[0084] where X i is the financing balance corresponding to the customer under a certain 12-level classification, and if the financing balance of a large customer is 0, then the risk value of the large customer is 0. Therefore, the range of the risk value of the large customer is [0, 12].

[0085] For example: Large customer A has 3 financings, and the financing balances are all 1000, and the 12-level classifications are 1 to 3 levels respectively. Therefore, the risk value of this large customer is '(1×1000 + 2×1000 + 3×1000) / (1000 + 1000 + 1000) = 2'.

[0086] Step 103: Generate the feature set according to the current risk value of the user.

[0087] In one embodiment, referring to Figure 4 , step 200 includes:

[0088] Step 201: Calculate the average value of the current user risk values of all users according to the feature set;

[0089] Step 202: Classify the users according to the average value to generate a first user set and a second user set;

[0090] In Step 201 and Step 202, initially classify using the average risk value of all large customers in the whole bank at the current time point, into two categories: higher than the average value and lower than the average value.

[0091] Next, obtain the average risk value of each category (the centroid of this category, as the discrimination criterion for the next step).

[0092] Step 203: Perform an iterative operation:

[0093] Randomly select a current user from the first user set;

[0094] Calculate respectively the first difference between the current user and the average value of the current user risk values of the first user set, and the second difference between the current user and the average value of the current user risk values of the second user set;

[0095] When the first difference is less than the second difference, move the current user to the second user set;

[0096] Until the first difference is not less than the second difference.

[0097] Step 203 is essentially a clustering analysis process: In the clustering analysis process, use the average risk value (equivalent to the centroid of the class) as the discrimination criterion for the class, and sequentially judge all the risk values in the corresponding class. If the difference between this risk value in the class and the average risk value of the class where it is located is less than the difference between this risk value and the average risk value of another class (that is, this risk value in this class is closer to the centroid of the other class), then move this risk value out of the original class and put it into another class. On the contrary, then this risk value remains in the original class. After performing this judgment on all the risk values of the two classes respectively, one clustering analysis process is completed. Repeatedly perform this clustering analysis process until there are no more risk values that need to be moved from one class to another class (that is, the risk values in each class are tightly clustered around the centroid of the class where they are located, closer to the average risk value of the class where they are located). After completing the above clustering, the large customers corresponding to the class with a higher centroid are high-risk large customers.

[0098] To further illustrate this solution, the present invention also provides a specific application example of the high-risk user dynamic identification method in the financial field. For specific content, see Figure 5 and Figure 6 .

[0099] Step S101: Establish a risk assessment model.

[0100] The risk assessment model is used to accurately measure the risk level of large customers. (That is, first quantitatively evaluate how much risk each large customer has).

[0101] To determine the risk level of large customers, it is undoubtedly necessary to consider various factors such as 'customer rating, guarantee method, mortgage value, economic nature, and business status'. However, it is undoubtedly very difficult and inefficient to build a model for judgment based on many factors. Considering that the bank's unique 12-level classification for debt items has utilized the above-mentioned risk factors to determine the risk level of debt items, a weighted average algorithm based on the 12-level classification is proposed to measure the risk of each large customer:

[0102]

[0103] Where X i is the financing balance corresponding to the customer under a certain 12-level classification. And if the financing balance of a large customer is 0, then the risk value of the large customer is 0. Therefore, the range of the risk value of large customers is [0, 12].

[0104] For example: Large customer A has 3 financings, and the financing balances are all 1000, and the 12-level classifications are 1-3 levels respectively. Therefore, the risk value of this large customer is "(1×1000 + 2×1000 + 3×1000) / (1000 + 1000 + 1000) = 2".

[0105] Step S102: Use the clustering algorithm to identify risk users.

[0106] Affected by macro factors and others, the overall risk level of large customers in the whole bank must be different at different time points. It may be relatively high as a whole at some time points, relatively low as a whole at some time points, there are more large customers with high risks at some time points, and there are fewer large customers with high risks at some time points. Therefore, it is not appropriate to use common clustering algorithms such as machine learning or logistic regression models for judgment. For this reason, a dynamic clustering algorithm is proposed to dynamically screen high-risk large customers under different risk levels of large customers in the whole bank at different time points. (That is, according to the risk level of the large customer group at a certain time point, dynamically cluster and screen out high-risk large customers) to dynamically adapt to the screening needs of high-risk large customers in different scenarios. See Figure 7 , the steps are as follows:

[0107] S1: Initialize the characteristic data of risk large customers.

[0108] Extract the financing balances and 12-level classifications of all risk large customers at the current time point to form an initial characteristic set of risk large customer objects. Based on the initial characteristic set of risk large customer objects, use the risk assessment model to quantitatively evaluate the risk value of each risk large customer object one by one, and newly build a risk value quantization characteristic for each risk large customer object one by one. Construct an initialization set of risk large customer risk characteristic data.

[0109] S2: Initial clustering.

[0110] For all major risk accounts, calculate the average value of the quantified risk characteristics of major accounts. Based on the average value of the quantified risk characteristics of major accounts, conduct initial clustering on all major risk accounts: If the risk characteristic value of a major account is greater than or equal to this average value, include it in the 'high-risk major account set'; if the risk characteristic value of a major account is less than this average value, include it in the 'low-risk major account set'.

[0111] S3: Nested dynamic clustering analysis.

[0112] Reset the clustering centroid: For the 'high-risk major account set' and 'low-risk major account set', calculate the average value of the quantified risk characteristics of major accounts within the set respectively, as the new clustering centroids of the 'high-risk major account set' and 'low-risk major account set' respectively.

[0113] Re-cluster: Use the new clustering centroids as the new discrimination criteria for the 'high-risk major account set' and 'low-risk major account set', and re-cluster all major accounts in the original 'high-risk major account set' and 'low-risk major account set' in turn:

[0114] a. If the difference between the risk value of a major account in the original 'high-risk major account set' and the new clustering centroid of this set is greater than the difference between this risk value and the new clustering centroid of the 'low-risk major account set' (that is, this risk value in this category is closer to the centroid of the other set), then move this major account to the 'low-risk major account set'.

[0115] b. If the difference between the risk value of a major account in the original 'high-risk major account set' and the new clustering centroid of this set is less than or equal to the difference between this risk value and the new clustering centroid of the 'low-risk major account set' (that is, this risk value in this category is closer to the centroid of the original set), then keep this major account in the 'high-risk major account set'.

[0116] c. If the difference between the risk value of a major account in the original 'low-risk major account set' and the new clustering centroid of this set is greater than the difference between this risk value and the new clustering centroid of the 'high-risk major account set' (that is, this risk value in this category is closer to the centroid of the other set), then move this major account to the 'high-risk major account set'.

[0117] d. If the difference between the risk value of a major account in the original 'low-risk major account set' and the new clustering centroid of this set is less than or equal to the difference between this risk value and the new clustering centroid of the 'high-risk major account set' (that is, this risk value in this category is closer to the centroid of the original set), then keep this major account in the 'low-risk major account set'.

[0118] S4: Clustering end determination.

[0119] If the new centroid coincides with the original centroid after the clustering centroid is reset, there is no need to perform re-clustering. The nested dynamic clustering analysis should be exited. At this time, the risk values of each large customer in the 'high-risk large customer set, low-risk large customer set' are tightly around the centroid of the large customer risk value of the set where they belong, and are closer to the average large customer risk value of the set where they belong.

[0120] S5: Output of clustering results.

[0121] After the above clustering analysis, all large customers in the 'high-risk large customer set' are the relatively high-risk large customers corresponding to the current time point that need to be identified in this scenario.

[0122] In order to achieve automatic dynamic identification of high-risk large customers and change the current situation of low efficiency and low accuracy in the offline manual evaluation of high-risk large customers, in view of the two difficult problems to be solved urgently in the above background, the present invention provides a method for dynamically identifying high-risk large customers based on weighted average and dynamic clustering, which can dynamically and automatically identify the relatively high-risk large customers corresponding to different time points and different overall risk levels. Specifically, first, a feature set of users is generated according to the user debt grading; then, using a clustering algorithm, the users are classified according to the feature set to generate a classification result; finally, the high-risk users are dynamically identified according to the classification result. The present invention realizes the dynamic identification of high-risk large customers at different time points and different overall risk levels by establishing a risk evaluation model that can quantify the risk of large customers and a dynamic clustering algorithm, and will change the current situation of low efficiency and low accuracy in the offline manual evaluation of high-risk large customers.

[0123] Based on the same inventive concept, the embodiments of the present application also provide a device for dynamically identifying high-risk users in the financial field, which can be used to implement the method described in the above embodiments, as in the following embodiments. Since the principle of the device for dynamically identifying high-risk users in the financial field to solve problems is similar to that of the method for dynamically identifying high-risk users in the financial field, the implementation of the device for dynamically identifying high-risk users in the financial field can refer to the implementation of the method for dynamically identifying high-risk users in the financial field, and the repeated parts will not be described again. Hereinafter, the term "unit" or "module" may refer to a combination of software and / or hardware that can implement a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0124] An embodiment of the present invention provides a specific implementation manner of a device for dynamically identifying high-risk users in the financial field that can implement the method for dynamically identifying high-risk users in the financial field. Refer to Figure 8 , and the device for dynamically identifying high-risk users in the financial field specifically includes the following content:

[0125] A feature set generation module 10, configured to generate a feature set of users according to the user debt grading;

[0126] A classification result generation module 20, configured to classify the user according to the feature set by using a clustering algorithm to generate a classification result;

[0127] A user identification module 30, configured to dynamically identify high-risk users according to the classification result.

[0128] In one embodiment, referring to Figure 9 , the high-risk user dynamic identification device in the financial field further includes: a debt item grading determination module 40, configured to determine the debt item grading of the user according to the customer rating, guarantee method parameter, mortgage value, economic nature parameter, and business status parameter of the user.

[0129] In one embodiment, referring to Figure 10 , the feature set generation module 10 includes:

[0130] A data cleaning unit 101, configured to perform data cleaning on the debt item grading of the user;

[0131] A risk value calculation unit 102, configured to calculate the current risk value of the user according to the cleaned debt item grading of the user;

[0132] A feature set generation unit 103, configured to generate the feature set according to the current risk value of the user.

[0133] In one embodiment, referring to Figure 11 , the classification result generation module 20 includes:

[0134] An average value calculation unit 201, configured to calculate the average value of the current risk values of all users according to the feature set;

[0135] A user classification unit 202, configured to classify the user according to the average value to generate a first user set and a second user set;

[0136] An iterative operation unit 203, configured to perform an iterative operation:

[0137] Randomly select a current user from the first user set;

[0138] Calculate a first difference between the current risk value of the current user and the average value of the current risk values of the first user set, and a second difference between the current risk value of the current user and the average value of the current risk values of the second user set, respectively;

[0139] When the first difference is less than the second difference, move the current user to the second user set;

[0140] Until the first difference is not less than the second difference.

[0141] As can be seen from the above description, the high-risk user dynamic recognition device in the financial field provided by the embodiments of the present invention first generates a feature set of users according to user debt grading; then, uses a clustering algorithm to classify users according to the feature set to generate a classification result; and finally dynamically recognizes high-risk users according to the classification result. The present invention realizes the dynamic recognition of high-risk large customers at different time points and different overall risk levels by establishing a risk evaluation model that can quantitatively evaluate the risks of large customers and a dynamic clustering algorithm, which will change the current situation of inefficient and low-precision offline manual evaluation of high-risk large customers.

[0142] Next, refer to Figure 12 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing the embodiments of the present application.

[0143] As Figure 12 shown, the electronic device 600 includes a central processing unit (CPU) 601, which can perform various appropriate operations and processes according to the programs stored in the read-only memory (ROM) 602 or the programs loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the system 600 are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0144] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed in the storage section 608 as needed.

[0145] Specifically, according to the embodiments of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present invention include a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method for determining the distance between personnel in the data computer room scenario are implemented, and the steps include:

[0146] Step 100: Generate a feature set of users according to user debt grading;

[0147] Step 200: Use a clustering algorithm to classify the user according to the feature set to generate a classification result;

[0148] Step 300: Dynamically identify high-risk users according to the classification result.

[0149] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611.

[0150] For convenience of description, when describing the above device, various units are described separately according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0151] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0152] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0153] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the element.

[0154] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0155] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for dynamically identifying high-risk users in the financial field, characterized in that, Including: Generating a feature set of a user according to the user's debt grading; Using a clustering algorithm to classify the users according to the feature set to generate a classification result; Dynamically identifying high-risk users according to the classification result; The using a clustering algorithm to classify the users according to the feature set to generate a classification result includes: Calculating the average value of the current risk values of all users according to the feature set; Classifying the users according to the average value to generate a first user set and a second user set; Performing an iterative operation: Randomly selecting a current user from the first user set; Calculating a first difference between the current risk value of the current user and the average value of the current risk values of the users in the first user set, and a second difference between the current risk value of the current user and the average value of the current risk values of the users in the second user set respectively; When the first difference is less than the second difference, moving the current user to the second user set; Until the first difference is not less than the second difference, specifically: The average value of the current risk value is the centroid of the class, and the average value of the current risk value is used as the discrimination criterion of the class. All the risk values in the corresponding class are judged in turn. If the difference between this risk value in the class and the average risk value of the class where it is located is less than the difference between this risk value and the average risk value of another class, that is, this risk value in this class is closer to the centroid of another class, then move this risk value out of the original class and put it into another class; on the contrary, this risk value remains in the original class; after this judgment is completed for all the risk values of the two classes, a clustering analysis process is completed; this clustering analysis process is repeated until there are no more risk values that need to be moved from one class to another class, that is, the risk values in each class are tightly around the centroid of the class where they are located and closer to the average risk value of the class where they are located. After the above clustering is completed, the large users corresponding to the class with a higher centroid are high-risk large users.

2. The method for dynamically identifying high-risk users in the financial field according to claim 1, characterized in that, Also including: Determining the user's debt grading according to the user's customer rating, guarantee method parameter, mortgage value, economic nature parameter and business status parameter.

3. The method for dynamically identifying high-risk users in the financial field according to claim 2, characterized in that, The generating a feature set of a user according to the user's debt grading includes: Performing data cleaning on the user's debt grading; Calculating the current risk value of the user according to the cleaned user's debt grading; Generating the feature set according to the current risk value of the user.

4. A device for dynamically identifying high-risk users in the financial field, characterized in that, Including: A feature set generation module for generating a feature set of a user according to the user's debt grading; A classification result generation module for using a clustering algorithm to classify the users according to the feature set to generate a classification result; A user identification module for dynamically identifying high-risk users according to the classification result; The classification result generation module includes: An average value calculation unit for calculating the average value of the current risk values of all users according to the feature set; A user classification unit for classifying the users according to the average value to generate a first user set and a second user set; An iterative operation unit for performing an iterative operation: Randomly selecting a current user from the first user set; Calculate the first difference between the average of the current risk values of the users in the set of the current user and the first user, and the second difference between the average of the current risk values of the users in the set of the current user and the second user respectively; When the first difference is less than the second difference, move the current user into the set of the second user; Until the first difference is not less than the second difference, specifically: The average of the current risk values is the centroid of the class. Taking the average of the current risk values as the discrimination criterion of the class, judge all the risk values in the corresponding class in turn. If the difference between this risk value in the class and the average risk value of the class where it is located is less than the difference between this risk value and the average risk value of another class, that is, this risk value in this class is closer to the centroid of the other class, then move this risk value out of the original class and put it into another class; on the contrary, this risk value remains in the original class; after this judgment is completed for all the risk values of the two classes respectively, a clustering analysis process is completed; repeat this clustering analysis process until there are no more risk values that need to be moved from one class to another class, that is, the risk values in each class are tightly around the centroid of the class where they are located and closer to the average risk value of the class where they are located. After the above clustering is completed, the large customers corresponding to the class with a higher centroid, that is, the high-risk large customers.

5. The device for dynamically identifying high-risk users in the financial field according to claim 4, characterized in that, It further includes:A debt item grading determination module, configured to determine the debt item grading of the user according to the customer rating, guarantee method parameter, mortgage value, economic nature parameter, and operation status parameter of the user.

6. The high-risk user dynamic identification device in the financial field according to claim 5, wherein, The feature set generation module includes: A data cleaning unit, configured to perform data cleaning on the debt item grading of the user; A risk value calculation unit, configured to calculate the current risk value of the user according to the cleaned debt item grading of the user; A feature set generation unit, configured to generate the feature set according to the current risk value of the user.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the steps of the high-risk user dynamic identification method in the financial field according to any one of claims 1 to 3.

8. A computer-readable storage medium, having a computer program stored thereon, wherein, When the computer program is executed by the processor, it implements the steps of the high-risk user dynamic identification method in the financial field according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Risk identification method and system

    CN108629680A