Financial Risk Prediction Method, Device and Electronic Device Based on Financial Time Nodes
By extracting user data from multiple dimensions such as financial time nodes and establishing a sub-prediction model, the problem of not taking time factors into account in the existing technology is solved, and the accuracy of financial risk prediction and model accuracy are improved.
Patent Information
- Application Number
- CN202110166441.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-05
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-02-05
AI Technical Summary
The existing financial risk prediction methods do not consider time factors when used, resulting in inaccurate calculations of the model and low accuracy in risk assessment of users.
By extracting user data from multiple dimensions such as time dimensions, transaction dimensions, user resource usage performance data types and quantity, etc., a sub-prediction model corresponding to different user groups is established to achieve prediction of user groups within different time segments.
This improves the accuracy of the model, achieves more accurate user risk prediction, and reduces the financial risk losses of financial service institutions.
Smart Images

Figure CN112488865B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer information processing, and more particularly, to a financial risk prediction method, apparatus, and electronic device based on financial time nodes. Background Art
[0002] Risk prediction is the quantification of risks and is a key technology in risk management. Currently, risk prediction is generally carried out by means of modeling. In the process of establishing a model, there are mainly steps such as data extraction, feature generation, feature selection, algorithm model generation, and rationality evaluation.
[0003] In the prior art, the main purpose of financial risk prediction is how to distinguish good customers from bad customers, evaluate the risk situation of users, so as to reduce credit risks and achieve maximum profits. In addition, as the sources of data become more and more abundant, the data that can be used as risk feature variables is also increasing. However, many data such as user data and other related data do not consider the changes caused by time factors when used. Therefore, when using the above data for model calculation, the model calculation value is not accurate enough, and even the accuracy of risk assessment for some users is relatively low. Thus, there is still a large room for improvement in aspects such as model accuracy improvement, model optimization, and data extraction.
[0004] Therefore, it is necessary to provide a new financial risk prediction method to further improve the model accuracy and more accurately predict the risk situations of different users. Summary of the Invention
[0005] In order to more accurately predict the risk situations of each user in different user groups, improve the accuracy of the prediction model, and reduce the financial risk losses of financial service institutions. The present invention performs more effective data extraction from multiple dimensions such as the time dimension of financial time nodes, the transaction dimension, the data type and quantity of user resource usage performance, etc., to further segment user groups, and establish sub-prediction models corresponding to different user groups, so as to achieve the prediction of user groups within different time segments.
[0006] The present invention provides a financial risk prediction method based on financial time nodes, including: establishing a plurality of sub-training data sets according to the extraction rules of historical users at financial time nodes, where the historical users are historical users with resource quotas and resource usage behaviors, and the financial time nodes include resource quota granting nodes, resource usage nodes, and resource return nodes; establishing a plurality of sub-prediction models corresponding to each sub-training data set, and training the corresponding sub-prediction models using the corresponding sub-training data sets; judging the sub-prediction model matched with the current user according to the matching rules; using the matched sub-prediction model to perform financial risk prediction on the current user.
[0007] Preferably, it includes: setting a matching rule, where the matching rule includes the number of resource usage behaviors occurring within a specific time period from the resource quota granting node, the occurrence time of the first resource usage behavior, and the occurrence time of the second resource usage behavior.
[0008] Preferably, the matching rule further includes a first matching rule, a second matching rule, and a third matching rule. Among them, the first matching rule is to determine whether the current user has had a resource usage behavior within a specific time period from the resource quota granting node; the second matching rule determines whether the current user has had two resource usage behaviors within a specific time period from the resource quota granting node and whether the occurrence time of the second resource usage behavior is within a specific time period since the occurrence time of the first resource usage behavior; the third matching rule is to determine whether the number of times the current user has used resources within a specific time period from the resource quota granting node exceeds a specific number.
[0009] Preferably, it further includes: when the user data of the current user hits the first matching rule, then the current user is a first - type user; when the user data of the current user hits the second matching rule, then the current user is a second - type user; when the user data of the current user hits both the second matching rule and the third matching rule, then the current user is a third - type user.
[0010] Preferably, it further includes: the extraction rule includes time parameters, event parameters, and is extracted according to the time parameters and / or event parameters. The time parameters include within a specific time period from the resource quota granting node, within the time period from the resource quota granting node to the occurrence time of the first resource usage behavior, and within a specific time period since the occurrence time of the first resource usage behavior; the event parameters include determining whether it is a new user, whether there is overdue data, whether there is default data, whether there is collection data, and whether there is a multi - head user; through the extraction rule, the time feature data and event feature data of the user are extracted.
[0011] Preferably, it further includes: obtaining the user data of the current user, using the extraction rule to extract the time feature data and event feature data of the current user; using the matching rule to determine the sub - prediction model corresponding to the current user, and using the determined sub - prediction model to input the time feature data and event feature data of the current user to calculate the financial prediction value of the current user.
[0012] Preferably, it further includes: setting an evaluation index, and adjusting the model parameters of each sub - prediction model by calculating the evaluation index. The evaluation index includes ROC index and AUC index.
[0013] In addition, the present invention also provides a financial risk prediction device based on financial time nodes, including: a data processing module, which establishes a plurality of sub-training data sets according to extraction rules based on historical users at financial time nodes, where the historical users are historical users with resource quotas and resource usage behaviors, and the financial time nodes include resource quota grant nodes, resource usage nodes, and resource return nodes; a model establishment module, which is used to establish a plurality of sub-prediction models corresponding to each sub-training data set and train the corresponding sub-prediction models using the corresponding sub-training data sets; a judgment module, which judges the sub-prediction model that matches the current user according to the matching rules; and a prediction module, which uses the matched sub-prediction model to predict the financial risk of the current user.
[0014] Preferably, it includes a setting module, and the setting module is used to set the matching rules, and the matching rules include the number of resource usage behaviors occurring within a specific time period from the resource quota grant node, the occurrence time of the first resource usage behavior, and the occurrence time of the second resource usage behavior.
[0015] Preferably, the matching rules further include a first matching rule, a second matching rule, and a third matching rule. Among them, the first matching rule is to judge whether the current user has had a resource usage behavior within a specific time period from the resource quota grant node; the second matching rule judges whether the current user has had two resource usage behaviors within a specific time period from the resource quota grant node and the occurrence time of the second resource usage behavior is within a specific time period since the occurrence time of the first resource usage behavior; the third matching rule is to judge whether the number of times of resource usage by the current user within a specific time period from the resource quota grant node exceeds a specific number.
[0016] Preferably, it further includes a determination module, and the determination module is used to determine the user category to which the current user belongs; when the user data of the current user hits the first matching rule, the current user is a first type of user; when the user data of the current user hits the second matching rule, the current user is a second type of user; when the user data of the current user hits both the second matching rule and the third matching rule, the current user is a third type of user.
[0017] Preferably, the extraction rules include time parameters, event parameters, and are extracted according to the time parameters and / or event parameters. The time parameters include within a specific time period from the resource quota grant node, within the time period from the resource quota grant node to the occurrence time of the first resource usage behavior, and within a specific time period since the occurrence time of the first resource usage behavior; the event parameters include judging whether it is a new user, whether there is overdue data, whether there is default data, whether there is collection data, and whether there is a multi-headed user; through the extraction rules, the time feature data and event feature data of the user are extracted.
[0018] Preferably, it further includes: obtaining user data of the current user, using the extraction rule to extract the time feature data and event feature data of the current user; using the matching rule to determine the sub-prediction model corresponding to the current user, and using the determined sub-prediction model to input the time feature data and event feature data of the current user to calculate the financial prediction value of the current user.
[0019] Preferably, it further includes: setting evaluation indicators, and adjusting the model parameters of each sub-prediction model by calculating the evaluation indicators, where the evaluation indicators include ROC indicator and AUC indicator.
[0020] In addition, the present invention also provides an electronic device, where the electronic device includes: a processor; and a memory storing computer-executable instructions, and the executable instructions, when executed, cause the processor to execute the financial risk prediction method based on financial time nodes of the present invention.
[0021] In addition, the present invention also provides a computer-readable storage medium, where the computer-readable storage medium stores one or more programs, and the one or more programs, when executed by a processor, implement the financial risk prediction method based on financial time nodes of the present invention.
[0022] Beneficial effects
[0023] Compared with the prior art, the present invention can extract user data from multiple dimensions such as the time dimension, transaction dimension, and data dimension (especially the user resource usage performance data type and the quantity of the same type of user resource usage performance data) of financial time nodes, etc., and can more accurately achieve the fine classification of user groups; by using evaluation indicators and selecting a test data set to evaluate each sub-prediction model, the model parameters can be further optimized, and the model accuracy can be improved; by establishing sub-prediction models corresponding to different user groups, the risk situation of users can be predicted more accurately, the prediction accuracy of each sub-prediction model is improved, and the financial risk loss of financial service institutions is also reduced. Description of the drawings
[0024] In order to make the technical problems solved by the present invention, the technical means adopted, and the technical effects obtained clearer, the specific embodiments of the present invention will be described in detail below with reference to the drawings. It should be noted that the drawings described below are only the drawings of the exemplary embodiments of the present invention, and those skilled in the art can obtain the drawings of other embodiments without creative efforts based on these drawings.
[0025] Figure 1 It is a flowchart of an example of the financial risk prediction method based on financial time nodes in Embodiment 1 of the present invention.
[0026] Figure 2 It is a flowchart of another example of the financial risk prediction method based on financial time nodes in Embodiment 1 of the present invention.
[0027] Figure 3 It is a flowchart of yet another example of the financial risk prediction method based on financial time nodes in Embodiment 1 of the present invention.
[0028] Figure 4 It is a schematic diagram of an example of the financial risk prediction device based on financial time nodes in Embodiment 2 of the present invention.
[0029] Figure 5 It is a schematic diagram of another example of the financial risk prediction device based on financial time nodes in Embodiment 2 of the present invention.
[0030] Figure 6 It is a schematic diagram of yet another example of the financial risk prediction device based on financial time nodes in Embodiment 2 of the present invention.
[0031] Figure 7 It is a structural block diagram of an exemplary embodiment of an electronic device according to the present invention.
[0032] Figure 8 It is a structural block diagram of an exemplary embodiment of a computer-readable medium according to the present invention. Detailed Embodiments
[0033] Now, exemplary embodiments of the present invention will be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, providing these exemplary embodiments enables the present invention to be more complete and comprehensive, and more conveniently conveys the inventive concept to those skilled in the art. Identical reference numerals in the figures denote the same or similar elements, components, or parts, and thus their repeated description will be omitted.
[0034] On the premise of conforming to the technical concept of the present invention, the features, structures, characteristics, or other details described in a specific embodiment may not be excluded from being combined in a suitable manner in one or more other embodiments.
[0035] In the description of specific embodiments, the features, structures, characteristics, or other details described in the present invention are for enabling those skilled in the art to fully understand the embodiments. However, it does not exclude that those skilled in the art can practice the technical solutions of the present invention without one or more of the specific features, structures, characteristics, or other details.
[0036] The flowcharts shown in the accompanying drawings are merely illustrative and not necessarily include all content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.
[0037] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0038] It should be understood that although ordinal adjectives such as first, second, third, etc. may be used herein to describe various devices, elements, components or parts, this should not be limited by these ordinal adjectives. These ordinal adjectives are used to distinguish one from another. For example, the first device can also be called the second device without departing from the technical solution of the essence of the present invention.
[0039] The term "and / or" or "and / or" includes any one and all combinations of one or more of the associated listed items.
[0040] In view of the above problems, the present invention proposes a financial risk prediction method based on financial time nodes. This method extracts user data from multiple dimensions such as the time dimension, transaction dimension, and data dimension (especially the user resource usage performance data type and the quantity of the same type of user resource usage performance data) of financial time nodes, etc., to further segment the user group, and establishes sub-prediction models corresponding to different user groups to predict the risk situation of users, improving the prediction accuracy of each sub-prediction model and reducing the financial risk losses of financial service institutions. The specific process of the method of the present invention will be described in detail below.
[0041] It should be noted that in the present invention, resources refer to any available substances, information, and time. Information resources include computing resources and various types of data resources. Data resources include various special data in various fields. The innovation of the present invention lies in how to use the information interaction technology between the server and the client to make the process of resource allocation more automated, efficient and reduce labor costs. Thus, essentially, the present invention can be applied to the risk prediction when various resources are allocated and returned, not limited to financial resources, including physical goods, water, electricity, and meaningful materials, etc. However, for convenience, the present invention takes financial data resources as an example to illustrate the implementation of resource allocation, but those skilled in the art should understand that the present invention can also be used for the risk prediction of other resources.
[0042] Embodiment 1
[0043] Next, embodiments of the financial risk prediction method based on financial time nodes of the present invention will be described with reference to Figures 1 to 3 Figures 1 to 3
[0044] Figure 1 FIG. is a flowchart of an example of the financial risk prediction method based on financial time nodes of the present invention. As Figure 1 shown, the method includes the following steps.
[0045] Step S101, establish a plurality of sub-training data sets according to the extraction rules of historical users at financial time nodes, where the historical users are historical users with resource quotas and resource usage behaviors, and the financial time nodes include resource quota granting nodes, resource usage nodes, and resource return nodes.
[0046] Step S102, establish a plurality of sub-prediction models corresponding to each sub-training data set, and use the corresponding sub-training data set to train the corresponding sub-prediction model.
[0047] Step S103, determine the sub-prediction model that matches the current user according to the matching rule.
[0048] Step S104, use the matched sub-prediction model to perform financial risk prediction on the current user.
[0049] First, in step S101, a plurality of sub-training data sets are established according to the extraction rules of historical users at financial time nodes.
[0050] In this example, in the application scenario where a user uses resources for financial service products or financial wealth management products, for example, user data in the above application scenario is obtained from relevant databases such as financial institutions and third-party payment institutions.
[0051] Specifically, the financial time nodes include resource quota granting nodes, resource usage nodes, and resource return nodes.
[0052] For example, according to the extraction rules of historical users at the resource usage node, time feature data and event feature data of the user are extracted.
[0053] As Figure 2 shown, it further includes step S201 of determining the extraction rules corresponding to each financial time node.
[0054] In step S201, determine the extraction rules corresponding to each financial time node.
[0055] Specifically, for example, determine the extraction rules corresponding to the resource usage node, and the extraction rules include time parameters, event parameters, and are extracted according to time parameters and / or event parameters.
[0056] Further, the time parameters include a specific time period starting from the resource quota granting node, a time period from the resource quota granting node to the occurrence time of the first resource usage behavior, and a specific time period starting from the occurrence time of the first resource usage behavior. For example, within 30 to 120 days starting from the occurrence time of the first resource usage behavior.
[0057] Furthermore, the event parameters include determining whether it is a new user, whether there is overdue data, whether there is default data, whether there is collection data, and whether there is a multi - borrower user.
[0058] In this example, based on the time parameters and event parameters, time - feature data and event - feature data of historical users are extracted to establish multiple sub - training data sets corresponding to the resource usage nodes.
[0059] It should be noted that in this example, the historical users are those historical users who have resource quotas and resource usage behaviors.
[0060] Preferably, data extraction is performed in the same way as the extraction method in step S201, but the time parameters and event parameters will change accordingly with the changes of financial time nodes. For example, the time parameters and event parameters corresponding to the resource return node. The time parameter includes a specific time calculated backward for a period of time starting from each resource return node, and the event parameter includes whether the resource return has been completed at or before each resource return point and whether there is collection data.
[0061] Thus, user data can be extracted more effectively according to the changes of financial time nodes and performance data, and sub - training data sets are established based on the extracted user data for training the model, thereby improving the model accuracy.
[0062] Further, according to the corresponding time parameters and event parameters, time - feature data and event - feature data of historical users are extracted to establish multiple sub - training data sets corresponding to the resource return nodes or resource granting nodes.
[0063] For example, the multiple sub - training data sets include time - feature data, event - feature data, user resource usage behavior data within a specific time period, overdue probability, and / or default probability of historical users related to the resource usage nodes. Among them, the specific time period includes a specific time period starting from the resource quota granting node, a time period from the resource quota granting node to the occurrence time of the first resource usage behavior, a specific time period starting from the occurrence time of the first resource usage behavior, etc.
[0064] For another example, the multiple sub-training data sets include time feature data, event feature data, user resource usage behavior data within a specific time period, overdue probability, and / or default probability of historical users related to the resource return node. Among them, the specific time period includes a time period that is pushed forward for a certain period from the resource usage node and the time between the resource quota grant node and the current resource usage node.
[0065] It should be noted that the above is only for illustrative purposes as a preferred example and should not be construed as a limitation to the present invention. In other examples, the sub-training data set may further include user feature data, and the user feature data may further include user basic information data, social behavior data, etc. For example, user age, gender, occupation, monthly income / annual income, etc.
[0066] It should be noted that the above is only for illustrative purposes as a preferred example and should not be construed as a limitation to the present invention.
[0067] Next, in step S102, multiple sub-prediction models corresponding to the respective sub-training data sets are established, and the corresponding sub-training data sets are used to train the corresponding sub-prediction models.
[0068] In this example, first, based on the resource quota grant node, the resource usage node, and the resource return node, the sample data is divided into three segments, namely, the sample data between the resource quota grant nodes, the sample data between the resource quota grant node and the resource usage node, and the sample data with the resource usage times greater than a specific number and reaching the resource return specific period.
[0069] Furthermore, it further includes respectively defining positive samples and negative samples according to the sample data of the above three segments, with the labels being 0 and 1. Among them, 1 represents the sample with the user's overdue probability (or default probability) being above Y, and 0 represents the sample with the user's overdue probability (or default probability) being less than Y. Among them, the Y values in each segment are different. Generally, the lower the user's overdue probability (or default probability), the better the situation of recovering the principal of the loan, the better the utilization efficiency of the funds, and the lower the risk level of the assets, and vice versa.
[0070] Therefore, by giving the sample label value Y, classifying users into target users and non-target users, it is possible to achieve the classification of the user group, and extract user data from multiple dimensions such as the time dimension, transaction dimension, and data dimension (especially the user resource usage performance data type and the quantity of the same type of user resource usage performance data) of financial time nodes, etc., and it is possible to more accurately achieve the fine classification of the user group.
[0071] Specifically, using the sample data with the label value Y, multiple sub-training data sets corresponding to the financial time nodes are established.
[0072] Further, according to the sample data, its quantity, and the influencing factors of financial time nodes, one or a combination of algorithms such as logistic regression algorithm, Xgboost algorithm, TextCNN algorithm, and random forest algorithm are used to establish multiple sub-prediction models corresponding to each sub-training data set, and the corresponding sub-training data set is used to train the corresponding sub-prediction model.
[0073] It should be noted that the above is only described as a preferred example and should not be construed as a limitation to the present invention. In addition, the specific algorithms used can be determined according to the sampled data and / or business requirements.
[0074] Preferably, it further includes the step of evaluating the model using evaluation metrics.
[0075] Specifically, evaluation metrics are set, and by calculating the evaluation metrics, the model parameters of each sub-prediction model are adjusted. The evaluation metrics include ROC metric and AUC metric.
[0076] Further, it also includes establishing a test data set corresponding to each sub-training data set for model parameter adjustment.
[0077] Therefore, by using evaluation metrics and selecting a test data set to evaluate each sub-prediction model, the model parameters can be further optimized and the model accuracy can be improved.
[0078] It should be noted that the above is only described as a preferred example and should not be construed as a limitation to the present invention.
[0079] Next, in step S103, according to the matching rule, the sub-prediction model that matches the current user is determined.
[0080] Specifically, according to different application scenarios, for example, according to the application scenario of resource use, matching rules corresponding to resource use nodes are set to determine the sub-prediction model suitable for the current user.
[0081] Preferably, on the time dimension of financial time nodes such as resource quota granting nodes, resource use nodes, and resource return nodes, on the time dimension between adjacent time nodes, and on the time dimension of a specific time or specific time period that is pushed forward or backward from the financial time node by a certain period of time, whether there are performance feature data, the type and quantity of performance feature data, and the type and quantity of the same type of performance feature data are used to set the matching rules, further subdivide the user group, and obtain sub-prediction models that match each user group.
[0082] In this example, the matching rules include a first matching rule, a second matching rule, and a third matching rule. Among them, the first matching rule is to determine whether the current user has had a resource usage behavior within a specific time period from the resource quota granting node; the second matching rule is to determine whether the current user has had two resource usage behaviors within a specific time period from the resource quota granting node and whether the occurrence time of the second resource usage behavior is within a specific time period since the occurrence time of the first resource usage behavior; the third matching rule is to determine whether the number of times the current user has had resource usage within a specific time period from the resource quota granting node exceeds a specific number of times.
[0083] As Figure 3 shown, it further includes step S301 of determining the user category to which the current user belongs.
[0084] In step S301, determine the user category to which the current user belongs, and based on the determined user category, determine the sub-prediction model adapted to the current user.
[0085] In this example, obtain the user data of the current user, use the extraction rules determined in step S101 to extract the time feature data and event feature data of the current user, and use the matching rules to perform matching judgments on the extracted time feature data and event feature data.
[0086] Specifically, when the user data of the current user hits the first matching rule, then the current user is a first-type user; when the user data of the current user hits the second matching rule, then the current user is a second-type user; when the user data of the current user hits both the second matching rule and the third matching rule, then the current user is a third-type user.
[0087] It should be noted that in this example, the first-type users, second-type users, and third-type users refer to further subdivided user groups from the user group corresponding to the sample data between the resource quota granting node and the resource usage node and the user group corresponding to the sample data with the number of resource usages greater than a specific number and reaching a specific number of resource return periods.
[0088] Furthermore, based on the determined user category, determine the sub-prediction model corresponding to the current user.
[0089] Thus, it is possible to extract user data from multiple dimensions such as the time dimension, transaction dimension, and data dimension (especially the user resource usage performance data type and the quantity of user resource usage performance data of the same type) of financial time nodes, etc., and it is possible to more accurately divide users into different sub-user groups, achieving more precise user group subdivision.
[0090] It should be noted that the above is only for illustration purposes and should not be construed as a limitation to the present invention.
[0091] Next, in step S104, the matched sub - prediction model is used to perform financial risk prediction on the current user.
[0092] Specifically, the sub - prediction model corresponding to the current user determined in step S103 or S301 (i.e., the matched sub - prediction model) is used, and with the determined sub - prediction model, the time - feature data and event - feature data of the current user are input to calculate the financial prediction value of the current user.
[0093] Preferably, according to the calculated financial prediction value, the risk status of the current user is judged, and thus corresponding risk strategies are adopted for different users.
[0094] Specifically, the financial prediction value is a numerical value between 0 and 1.
[0095] For example, the risk strategies include prohibiting or restricting resource requests for the user, freezing the remaining resources, increasing resource requests, increasing resource quotas, etc.
[0096] Thus, by establishing sub - prediction models corresponding to different user groups, the risk situation of users can be predicted more accurately, the prediction accuracy of each sub - prediction model is improved, and the financial risk losses of financial service institutions are also reduced.
[0097] It should be noted that the above is only for illustration purposes and should not be construed as a limitation to the present invention.
[0098] Those skilled in the art can understand that all or part of the steps of implementing the above - mentioned embodiments are realized as a program (computer program) executed by a computer data - processing device. When this computer program is executed, the above - mentioned method provided by the present invention can be realized. Moreover, the computer program can be stored in a computer - readable storage medium, which can be a readable storage medium such as a disk, an optical disc, a ROM, a RAM, etc., or a storage array composed of multiple storage media, such as a disk or tape storage array. The storage medium is not limited to centralized storage, and it can also be distributed storage, such as cloud storage based on cloud computing.
[0099] Compared with the prior art, the present invention can extract user data from multiple dimensions such as the time dimension, transaction dimension, and data dimension (especially the data type of user resource usage performance and the quantity of user resource usage performance data of the same type) of financial time nodes, etc., and can more precisely achieve the fine classification of user groups; by using evaluation indicators and selecting a test data set to evaluate each sub-prediction model, the model parameters can be further optimized, and the model accuracy can be improved; by establishing sub-prediction models corresponding to different user groups, the risk situation of users can be predicted more precisely, the prediction accuracy of each sub-prediction model is improved, and the financial risk loss of financial service institutions is also reduced.
[0100] Embodiment 2
[0101] The device embodiments of the present invention are described below. The device can be used to execute the method embodiments of the present invention. For the details described in the device embodiments of the present invention, they should be regarded as a supplement to the above method embodiments; for the details not disclosed in the device embodiments of the present invention, they can be implemented with reference to the above method embodiments.
[0102] Referring to Figure 4 、 Figure 5 and Figure 6 , the present invention also provides a financial risk prediction device 400 based on financial time nodes. The financial risk prediction device 400 includes: a data processing module 401, which establishes a plurality of sub-training data sets according to the extraction rules of historical users at financial time nodes, where the historical users are historical users with resource quotas and resource usage behaviors, and the financial time nodes include resource quota granting nodes, resource usage nodes, and resource return nodes; a model establishment module 402, which is used to establish a plurality of sub-prediction models corresponding to each sub-training data set and train the corresponding sub-prediction models using the corresponding sub-training data sets; a judgment module 403, which judges the sub-prediction model that matches the current user according to the matching rules; and a prediction module 403, which uses the matched sub-prediction model to predict the financial risk of the current user.
[0103] As Figure 5 shown, it includes a setting module 501. The setting module 501 is used to set the matching rules, and the matching rules include the number of resource usage behaviors occurring within a specific time period from the resource quota granting node, the occurrence time of the first resource usage behavior, and the occurrence time of the second resource usage behavior.
[0104] Preferably, the matching rules further include a first matching rule, a second matching rule, and a third matching rule. Among them, the first matching rule is to determine whether the current user has had a resource usage behavior within a specific time period from the resource quota granting node; the second matching rule is to determine whether the current user has had two resource usage behaviors within a specific time period from the resource quota granting node, and the occurrence time of the second resource usage behavior is within a specific time period since the occurrence time of the first resource usage behavior; the third matching rule is to determine whether the number of times the current user has used resources within a specific time period from the resource quota granting node exceeds a specific number of times.
[0105] As Figure 6 shown, it further includes a determination module 601. The determination module 601 is used to determine the user category to which the current user belongs; when the user data of the current user hits the first matching rule, the current user is a first-type user; when the user data of the current user hits the second matching rule, the current user is a second-type user; when the user data of the current user hits both the second matching rule and the third matching rule, the current user is a third-type user.
[0106] Preferably, it further includes: the extraction rules include time parameters, event parameters, and are extracted according to the time parameters and / or event parameters. The time parameters include within a specific time period from the resource quota granting node, the time period from the resource quota granting node to the occurrence time of the first resource usage behavior, and within a specific time period since the occurrence time of the first resource usage behavior; the event parameters include determining whether it is a new user, whether there is overdue data, whether there is default data, whether there is collection data, and whether there is a multi-headed user; through the extraction rules, the time feature data and event feature data of the user are extracted.
[0107] In this example, it further includes: obtaining the user data of the current user, using the extraction rules to extract the time feature data and event feature data of the current user; using the matching rules to determine the sub-prediction model corresponding to the current user, and using the determined sub-prediction model to input the time feature data and event feature data of the current user to calculate the financial prediction value of the current user.
[0108] Furthermore, it further includes: setting evaluation indicators, and adjusting the model parameters of each sub-prediction model by calculating the evaluation indicators. The evaluation indicators include ROC indicators and AUC indicators.
[0109] It should be noted that in Embodiment 2, the description of the same parts as in Embodiment 1 is omitted.
[0110] Those skilled in the art can understand that the modules in the above device embodiments can be distributed in the device as described, or can be correspondingly changed and distributed in one or more devices different from the above embodiments. The modules of the above embodiments can be combined into one module, or can be further split into multiple sub-modules.
[0111] Compared with the prior art, the present invention can extract user data from multiple dimensions such as the time dimension, transaction dimension, and data dimension (especially the data types of user resource usage performance and the quantity of user resource usage performance data of the same type) of financial time nodes, etc., and can more accurately achieve the fine classification of user groups; by using evaluation indicators and selecting test data sets to evaluate each sub-prediction model, the model parameters can be further optimized, and the model accuracy can be improved; by establishing sub-prediction models corresponding to different user groups, the risk situation of users can be predicted more accurately, the prediction accuracy of each sub-prediction model is improved, and the financial risk losses of financial service institutions are also reduced.
[0112] Embodiment 3
[0113] The following describes an embodiment of the electronic device of the present invention, and this electronic device can be regarded as a specific physical implementation manner of the above method and device embodiments of the present invention. For the details described in the embodiment of the electronic device of the present invention, they should be regarded as a supplement to the above method or device embodiments; for the details not disclosed in the embodiment of the electronic device of the present invention, they can be implemented with reference to the above method or device embodiments.
[0114] Figure 7 is a structural block diagram of an exemplary embodiment of an electronic device according to the present invention. The following refers to Figure 7 to describe the electronic device 200 according to this embodiment of the present invention. Figure 7 The electronic device 200 shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.
[0115] As Figure 7 shown, the electronic device 200 is presented in the form of a general-purpose computing device. The components of the electronic device 200 may include but are not limited to: at least one processing unit 210, at least one storage unit 220, a bus 230 connecting different system components (including the storage unit 220 and the processing unit 210), a display unit 240, etc.
[0116] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 210, so that the processing unit 210 executes the steps according to various exemplary embodiments of the present invention described in the processing method part of the above electronic device in this specification. For example, the processing unit 210 can execute the steps as Figure 1 shown.
[0117] The storage unit 220 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 2201 and / or a cache storage unit 2202, and may further include a read-only storage unit (ROM) 2203.
[0118] The storage unit 220 may also include a program / utilities 2204 having a set (at least one) of program modules 2205. Such program modules 2205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0119] The bus 230 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.
[0120] The electronic device 200 may also communicate with one or more external devices 300 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 200, and / or may communicate with any device that enables the electronic device 200 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be through an input / output (I / O) interface 250. Moreover, the electronic device 200 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 260. The network adapter 260 may communicate with other modules of the electronic device 200 through the bus 230. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0121] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described in the present invention can be implemented by software, or can be implemented by the way of software combined with necessary hardware. Therefore, the technical solution according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a computer-readable storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which may be a personal computer, a server, or a network device, etc.) to execute the above method according to the present invention. When the computer program is executed by a data processing device, the computer-readable medium can implement the above method of the present invention.
[0122] As Figure 8 shown, the computer program can be stored on one or more computer-readable media. The computer-readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0123] The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.
[0124] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or, it can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0125] In summary, the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that general-purpose data processing devices such as microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (e.g., a computer program and a computer program product) for performing part or all of the methods described herein. Such a program for implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or in any other form.
[0126] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the present invention is not inherently related to any specific computer, virtual device, or electronic device, and various general-purpose devices can also implement the present invention. The above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for determining the user risk status based on the resource usage time nodes, characterized in that, it includes: According to the performance data of historical users at the resource usage time nodes, determine the extraction rules corresponding to each resource usage time node, and extract the time feature data and event feature data of each historical user from the time dimension, transaction dimension, and data dimension of the resource usage time nodes to establish multiple sub-training data sets; the historical users are historical users with resource quotas and resource usage behaviors, the resource usage time nodes include resource quota grant nodes, resource usage nodes, and resource return nodes, and the extraction rules include time parameters or event parameters; the time parameters include within a specific time period since the resource quota grant node, within the time period from the resource quota grant node to the occurrence time of the first resource usage behavior, within a specific time period since the occurrence time of the first resource usage behavior, and the event parameters include determining whether it is a new user, whether there is overdue data, whether there is default data, whether there is collection data, and whether there is a multi-headed user, and the time parameters and event parameters change with the change of the resource usage time nodes; Establish multiple sub-prediction models corresponding to each sub-training data set through a general computing device, and use the corresponding sub-training data sets to train the corresponding sub-prediction models. Based on the resource quota grant node, resource usage node, and resource return node, divide the sample data into three segments, namely: sample data between resource quota grant nodes, sample data between the resource quota grant node and the resource usage node, and sample data with the number of resource usage times greater than a specific number and reaching a specific number of resource return periods; according to the application scenario of resource usage, by judging whether there is performance feature data, the type and quantity of performance feature data, and the type and quantity of the same type of performance feature data, set the matching rules corresponding to each resource usage time node, determine the category to which the current user belongs, and further judge the sub-prediction model that matches the current user. The matching rules include judging that a resource usage behavior has occurred within a specific time period and the number of resource usage behaviors. Determining the category to which the current user belongs includes: when the user data of the current user meets the condition that a resource usage behavior has occurred once within a specific time period from the resource quota grant node, the current user is a first-type user; when the user data of the current user meets the condition that two resource usage behaviors have occurred within a specific time period from the resource quota grant node and the occurrence time of the second resource usage behavior is within a specific time period since the occurrence time of the first resource usage behavior, the current user is a second-type user; when the user data of the current user simultaneously meets the conditions that two resource usage behaviors have occurred within a specific time period from the resource quota grant node and the occurrence time of the second resource usage behavior is within a specific time period since the occurrence time of the first resource usage behavior and the number of resource usage times within a specific time period from the resource quota grant node exceeds a specific number, the current user is a third-type user; Using the matched sub-prediction model, input the time feature data and event feature data of the current user into a general computing device, calculate the predicted resource usage value of the current user, and determine the risk status of the current user.
2. The method according to claim 1, wherein, it further includes: The extraction rules include time parameters, event parameters, and are extracted according to time parameters and / or event parameters.
3. The method according to claim 2, wherein, it further includes: Obtain the user data of the current user, and use the extraction rules to extract the time feature data and event feature data of the current user; Use the matching rules to determine the sub-prediction model corresponding to the current user, and use the determined sub-prediction model to input the time feature data and event feature data of the current user, and calculate the predicted resource usage value of the current user.
4. The method according to claim 1, wherein, it further includes: Set evaluation indicators, and adjust the model parameters of each sub-prediction model by calculating the evaluation indicators. The evaluation indicators include ROC indicators and AUC indicators.
5. A device for determining the user risk status based on the resource usage time node, wherein, it includes: A data processing module, which determines the extraction rules corresponding to each resource usage time node according to the performance data of historical users at the resource usage time node, and extracts the time feature data and event feature data of each historical user from the time dimension, transaction dimension, and data dimension of the resource usage time node to establish multiple sub-training data sets; Historical users are historical users with resource quotas and resource usage behaviors. The resource usage time nodes include resource quota granting nodes, resource usage nodes, and resource return nodes. The extraction rules include time parameters or event parameters; The time parameters include within a specific time period since the resource quota granting node, within the time period from the resource quota granting node to the occurrence time of the first resource usage behavior, and within a specific time period since the occurrence time of the first resource usage behavior. The event parameters include determining whether it is a new user, whether there is overdue data, whether there is default data, whether there is collection data, and whether there are multiple borrowers. The time parameters and event parameters change with the change of the resource usage time node; A model building module, which is used to build multiple sub-prediction models corresponding to each sub-training data set, and use the corresponding sub-training data set to train the corresponding sub-prediction model. Based on the resource quota granting node, resource usage node, and resource return node, the sample data is divided into three segments, namely: sample data between resource quota granting nodes, sample data between resource quota granting nodes and resource usage nodes, and sample data with the number of resource usage times greater than a specific number and reaching the specific number of resource return periods; A judgment module, according to the application scenarios of resource usage, sets matching rules corresponding to each resource usage time node by judging the presence or absence of performance characteristic data, the type and quantity of performance characteristic data, and the type and quantity of performance characteristic data of the same type, determines the category to which the current user belongs, and further judges the sub-prediction model matching the current user. The matching rules include judging that a resource usage behavior has occurred within a specific time period and the number of resource usage behaviors. Determining the category to which the current user belongs includes: when the user data of the current user meets the condition that a resource usage behavior has occurred once within a specific time period from the resource quota granting node, the current user is a first-type user; when the user data of the current user meets the condition that two resource usage behaviors have occurred within a specific time period from the resource quota granting node and the occurrence time of the second resource usage behavior is within a specific time period since the occurrence time of the first resource usage behavior, the current user is a second-type user; when the user data of the current user simultaneously meets the conditions that two resource usage behaviors have occurred within a specific time period from the resource quota granting node and the occurrence time of the second resource usage behavior is within a specific time period since the occurrence time of the first resource usage behavior and the number of resource usage occurrences within a specific time period from the resource quota granting node exceeds a specific number, the current user is a third-type user. A prediction module, uses the matched sub-prediction model, inputs the time characteristic data and event characteristic data of the current user, calculates the resource usage prediction value of the current user, and judges the risk status of the current user.
6. An electronic device wherein the electronic device includes: a processor; and a memory storing computer-executable instructions, and the executable instructions, when executed, cause the processor to execute the method for determining the user risk status based on the resource usage time node according to any one of claims 1-4.
7. A computer-readable storage medium wherein the computer-readable storage medium stores one or more programs, and the one or more programs, when executed by a processor, implement the method for determining the user risk status based on the resource usage time node according to any one of claims 1-4.
Citation Information
Patent Citations
Resource limit adjusting method and device and electronic equipment
CN111598494A
Resource quota determination method and device based on tweedie distribution and electronic equipment
CN112017042A