A risk prediction method, device and electronic device based on two-party network diagram data
By building a network relationship diagram of both parties and extracting mismatched feature data, and establishing a fraud prediction model, the problem of low accuracy in identifying and preventing Internet loan fraud in the prior art is solved, and more efficient risk prediction and fraud identification are achieved.
Patent Information
- Application Number
- CN201911290909.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-12-16
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2039-12-16
AI Technical Summary
When identifying and preventing fraud in Internet loans, the existing technology faces problems such as sparse credit data, high-speed and frequent lending behaviors and iterative updates of fraud, resulting in low model accuracy and difficulty in effectively identifying fraud.
By obtaining the basic feature data and behavior feature data of historical users, a network relationship diagram of both parties is constructed, mismatched feature data is extracted, and fraud prediction models are established, and the model is trained for risk prediction.
It improves the accuracy of risk prediction, improves the accuracy of fraud model, optimizes data utilization, reduces online lending risks, and improves business level.
Smart Images

Figure CN111353871B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communications, and particularly to a risk prediction method, apparatus, electronic device, and storage medium based on two-party network graph data. Background Art
[0002] The credit model of applying for loans through the Internet has developed rapidly. However, compared with the traditional credit model, applying for loans online brings convenience to people while also increasing the risk of fraud and loan default to the credit business department. If fraud behaviors cannot be well identified and processed, it will bring inestimable losses to the Internet financial platform.
[0003] Recently, the scale of domestic online loans has gradually increased, and the risk control of online lending has become a top priority. It is reported that in 2017, more than 2,000 online loan companies completed more than 200 billion loan transactions in total. Such a large number of loan transactions have exposed these companies to frequent fraud behaviors. In order to identify these fraud behaviors, online loan companies have the following problems: sparse credit-related data; high frequency, high volume, and high speed of lending behaviors; iterative updates of fraud behaviors, etc. Therefore, how to monitor these fraud behaviors in real time and timely feedback them into the business process is very crucial.
[0004] The existing technology builds a special fraud model to score the applying users for fraud. However, due to the low probability of fraud behaviors occurring, the data is relatively single and the data volume is insufficient. Especially, there are a large number of missing or incomplete data in the relationship network graph data. Therefore, when using the existing fraud model for model optimization, it may not accurately and efficiently identify fraudsters or fraud behaviors, resulting in problems such as low model accuracy.
[0005] In summary, it is necessary to provide a more accurate risk prediction method. Summary of the Invention
[0006] To solve the above problems, the present invention provides a risk prediction method based on two-party network graph data, including: obtaining the basic feature data and behavioral feature data of historical users, and constructing a two-party network relationship graph, where the two-party network relationship graph includes two types of nodes, namely user nodes and information nodes, the user nodes are nodes representing users, and the information nodes are nodes that associate different users; extracting the mismatch feature data of historical users from the two-party network relationship graph, where the mismatch feature data includes first mismatch feature data and second mismatch feature data; establishing a fraud prediction model, and training the fraud prediction model using the mismatch feature data and fraud performance data of the historical users; obtaining the basic feature data and behavioral feature data of a target user, adding the target user to the two-party network relationship graph to extract the mismatch feature data of the target user, and inputting it into the fraud prediction model to calculate the fraud prediction value of the target user for risk prediction.
[0007] Preferably, the extracting the mismatch feature data of the target user includes: extracting the first mismatch feature data and the second mismatch feature of the target user.
[0008] Preferably, the extracting the mismatch feature data of the target user includes: extracting the first mismatch feature data or the second mismatch feature of the target user.
[0009] Preferably, the risk prediction method further includes: in the case of different data sources, calculating the similarity of the same user node or information node based on determining the Jaccard distance to determine the first mismatch feature data.
[0010] Preferably, the risk prediction method further includes: in the case of different data sources, calculating the matching degree of the same user node or information node based on determining the shortest distance between adjacent two nodes to determine the first mismatch feature data.
[0011] Preferably, the risk prediction method further includes: monitoring whether the information data of the user node and the user nodes associated with it match to determine the second mismatch feature data.
[0012] Preferably, the user nodes include user personal feature data and network feature data; and / or the information nodes include at least one of APP information, location information, address book information, call record information, device information, and operator information.
[0013] Preferably, the prediction method further includes: setting a risk threshold, comparing the calculated risk prediction value of the target user with the risk threshold to classify the risk of the target user.
[0014] In addition, the present invention also provides a risk prediction device based on graph data. The risk prediction device includes: a data acquisition module that acquires the basic feature data and behavioral feature data of historical users and constructs a bipartite network relationship graph. The bipartite network relationship graph includes two types of nodes, namely user nodes and information nodes. The user nodes are nodes representing users, and the information nodes are nodes that associate different users; a data processing module that extracts the mismatched feature data of historical users from the bipartite network relationship graph. The mismatched feature data includes first mismatched feature data and second mismatched feature data; a model establishment module that establishes a fraud prediction model and trains the fraud prediction model using the mismatched feature data and fraud performance data of historical users; a calculation module that acquires the basic feature data and behavioral feature data of a target user, adds the target user to the bipartite network relationship graph to extract the mismatched feature data of the target user, and inputs it into the fraud prediction model to calculate the fraud prediction value of the target user for risk prediction.
[0015] Preferably, the risk prediction device further includes an extraction module that is used to extract the first mismatched feature data and the second mismatched features of the target user.
[0016] Preferably, the risk prediction device further includes an extraction module that is used to extract the first mismatched feature data or the second mismatched features of the target user.
[0017] Preferably, the risk prediction device further includes a determination module that, when the data sources are different, calculates the similarity of the same user node or information node based on the determined Jaccard distance to determine the first mismatched feature data.
[0018] Preferably, the risk prediction device further includes a determination module that, when the data sources are different, calculates the matching degree of the same user node or information node based on the determined shortest distance between adjacent two nodes to determine the first mismatched feature data.
[0019] Preferably, the risk prediction device further includes a monitoring module that is used to monitor whether the information data of a user node and the user nodes associated with it match to determine the second mismatched feature data.
[0020] Preferably, the user nodes include user personal feature data and network feature data; and / or the information nodes include at least one of APP information, location information, address book information, call record information, device information, and operator information.
[0021] Preferably, the risk prediction device further includes a setting module configured to set a risk threshold, compare the calculated risk prediction value of the target user with the risk threshold, and classify the risk of the target user.
[0022] In addition, the present invention also provides an electronic device, which includes: a processor; and a memory storing computer-executable instructions, where the executable instructions, when executed, cause the processor to execute the risk prediction method of the present invention.
[0023] In addition, the present invention also provides a computer-readable storage medium, where the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by a processor, the risk prediction method of the present invention is implemented.
[0024] Beneficial effects
[0025] Compared with the prior art, the risk prediction method of the present invention has a wide range of applications and is suitable for large-scale data processing and data analysis. Especially when the characteristic data of the user node is incomplete or missing, it can also predict the risk of the user represented by the user node by using the mining of group graph features (quadrilateral feature data), improving the accuracy of risk prediction. In addition, the risk prediction method of the present invention predicts risks by using mismatched features, improving the accuracy of the fraud model; increasing the utilization rate of data, optimizing the target data; reducing the online lending risk; and improving the business level. Description of the drawings
[0026] In order to make the technical problems solved by the present invention, the technical means adopted, and the technical effects obtained more clear, the specific embodiments of the present invention will be described in detail below with reference to the drawings. However, it should be noted that the drawings described below are only the drawings of the exemplary embodiments of the present invention, and those skilled in the art can obtain the drawings of other embodiments without creative efforts.
[0027] Figure 1 is a flowchart of an example of the risk prediction method based on two-party network graph data of the present invention.
[0028] Figure 2 is a partial schematic diagram of the two-party network relationship diagram of Embodiment 1 of the present invention.
[0029] Figure 3 is a schematic block diagram of the construction process of the risk prediction model of Embodiment 1 of the present invention.
[0030] Figure 4 is a diagram of an example of the graph data extraction in the two-party network relationship diagram of the present invention.
[0031] Figure 5 It is a flowchart of another example of the risk prediction method based on two - party network diagram data of the present invention.
[0032] Figure 6 It is a flowchart of yet another example of the risk prediction method based on two - party network diagram data of the present invention
[0033] Figure 7 It is a structural block diagram of an example of the risk prediction device in Embodiment 2 of the present invention.
[0034] Figure 8 It is a structural block diagram of another example of the risk prediction device in Embodiment 2 of the present invention.
[0035] Figure 9 It is a structural block diagram of yet another example of the risk prediction device in Embodiment 2 of the present invention.
[0036] Figure 10 It is a structural block diagram of an exemplary embodiment of an electronic device according to the present invention.
[0037] Figure 11 It is a structural block diagram of an exemplary embodiment of a computer - readable medium according to the present invention. Detailed Embodiments
[0038] Now, exemplary embodiments of the present invention will be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, providing these exemplary embodiments enables the present invention to be more comprehensive and complete, and more conveniently conveys the inventive concept to those skilled in the art. Identical reference numerals in the figures denote the same or similar elements, components, or parts, and thus their repeated description will be omitted.
[0039] On the premise of conforming to the technical concept of the present invention, features, structures, characteristics, or other details described in a specific embodiment may not be excluded from being combined in a suitable manner in one or more other embodiments.
[0040] In the description of specific embodiments, the features, structures, characteristics, or other details described in the present invention are for enabling those skilled in the art to fully understand the embodiments. However, it does not exclude that those skilled in the art can practice the technical solutions of the present invention without one or more of the specific features, structures, characteristics, or other details.
[0041] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0042] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0043] It should be understood that although the terms "first", "second", "third", etc. may be used herein to describe various devices, elements, components or parts, this should not be limited by these terms. These terms are used to distinguish one from another. For example, a first device may also be called a second device without departing from the essential technical solution of the present invention.
[0044] The term "and / or" or "and / or" includes any one and all combinations of one or more of the associated listed items.
[0045] Example 1
[0046] Next, we will refer to Figures 1 to 6 The risk prediction method based on bilateral network graph data of the present invention is described.
[0047] First embodiment
[0048] Figure 1 FIG. 1 is a flow chart of an example of a risk prediction method based on bilateral network graph data of the present invention. Figure 1 As shown, a risk prediction method based on bilateral network graph data includes the following steps.
[0049] Step S101, obtain basic feature data and behavior feature data of historical users, and construct a bilateral network relationship diagram, wherein the bilateral network relationship diagram includes two types of nodes, namely user nodes and information nodes. The user nodes are nodes representing users, and the information nodes are nodes that associate different users.
[0050] Step S102: extracting non-matching feature data of historical users from the network relationship graph of both parties, wherein the non-matching feature data includes first non-matching feature data and second non-matching feature data.
[0051] Step S103: establishing a fraud prediction model, and using the mismatching feature data and fraud performance data of the historical users to train the fraud prediction model.
[0052] Step S104: Obtain the basic feature data and behavioral feature data of the target user, add the target user to the bilateral network relationship graph to extract the mismatched feature data of the target user, and input it into the fraud prediction model to calculate the fraud prediction value of the target user for risk prediction.
[0053] Specifically, in step S101, obtain the basic feature data and behavioral feature data of historical users, and construct a bilateral network relationship graph.
[0054] In this embodiment, the basic feature data of historical users includes feature data such as gender, age, and occupation. The behavioral feature data includes user association behavior data and financial behavior data. Among them, the user association behavior feature data includes the behavioral feature data between the user and the associated person, the behavioral feature data of the users associated with the same device, etc. The financial behavior data refers to the data related to the user's financial behavior, such as monthly income, annual income, borrowing information, repayment information, overdue information, etc.
[0055] For data acquisition, for example, SDK embedding based on user authorization is used to obtain all relevant legal data for data analysis. In some specific scenarios, the behavioral data in the APP can also be obtained through the method of data logging. After obtaining the data, a network relationship graph is constructed. See Figure 2 .
[0056] In the present invention, the inventor adopts a construction method different from the traditional method, called the bilateral network, that is, the nodes in the network relationship are divided into two types of nodes (user nodes and information nodes). One type of node is the node representing the user, and the other type of node is the node that associates different users. See Figure 3 . Specifically, for example, using the device id, wifi address, GPS range, etc. as nodes, the feature data that associates users is collected, and the above-mentioned nodes are called information nodes. In other words, all the data of the user is divided into the data of the user node and the data of the information node. Based on these two types of nodes, a bilateral network relationship graph is constructed.
[0057] Furthermore, the connection between the user node and the information node is through a connection (or edge). Among them, each node can include one or more labels, and each node can also contain attributes. The attributes can exist in any key-value pair form. For example, the key is a string, and the value is a Java string and primitive data, or an array of these data types. Similarly, the connection can also include characteristics and attributes.
[0058] Specifically, the user node includes user personal feature data and network feature data. The information node includes at least one of APP information, location information, address book information, call record information, device information, and operator information.
[0059] Next, in step S102, extract the mismatched feature data of historical users from the bilateral network relationship diagram, where the mismatched feature data includes first mismatched feature data and second mismatched feature data.
[0060] It should be noted that in many cases, information or data mismatches occur. Specifically, in the case where data comes from different channels, due to different data sources, data mismatches of the same user may occur. In the present invention, the above-mentioned data mismatch is referred to as first mismatched feature data.
[0061] Furthermore, step S102 further includes determining the first mismatched feature data.
[0062] Preferably, in the case of different data sources, based on determining the Jaccard distance, calculate the similarity (or dissimilarity) of the same user node or information node to determine the first mismatched feature data.
[0063] For example, the feature data of user node 1 obtained from channel Q1 is A, and the feature data of user node 1 obtained from channel Q2 is B. Determine the Jaccard distance through the Jaccard algorithm, as specifically shown in the following expression 1.
[0064] Expression 1
[0065] D j (A,B) = 1 - J(A,B)
[0066] Where D j (A,B) refers to the Jaccard distance, which is used to describe dissimilarity; A is sample set A; B is sample set B.
[0067] By determining D j (A,B), the dissimilarity (or mismatch degree) of user node 1 can be calculated to determine the first mismatched feature data and extract it from the bilateral network relationship diagram.
[0068] It should be noted that the method for calculating similarity is not limited to the Jaccard algorithm, and also includes the cosine angle method, etc.
[0069] In addition, preferably, based on determining the shortest distance between adjacent two nodes, calculate the matching degree (or mismatch degree) of the same user node (or information node) to determine the first mismatched feature data.
[0070] In this embodiment, all nodes in the constructed bipartite network relationship graph are traversed by means of, for example, breadth-first search. For example, using the Dijkstra algorithm, calculate and record the path lengths between each node and its neighbor nodes, and find out the shortest path lengths between each user node (or information node) and its neighbor nodes.
[0071] Further, in the case of different data sources, for example, using the shortest path lengths between user nodes (or information nodes) and their neighbor nodes, calculate the matching degrees of the same user node (or information node), thereby determining the first mismatch feature and extracting it from the bipartite network relationship graph.
[0072] In addition, in the case where personal information conflicts with other network information, for example, it is detected that a certain user has a group relationship with other users in terms of geographical location, but the actual location information of this user is completely different from the location information of other users. In the present invention, this kind of mismatch is called the second mismatch feature data, for example, the mismatch degree information feature data in the geographical information dimension.
[0073] More specifically, the second mismatch feature data includes at least one of the mismatch between personal information and network information of the same user node, the mismatch of geographical information between two mutually related user nodes, and the mismatch of information between a mutually related user node and an information node.
[0074] In this example, the risk prediction method further includes: monitoring whether the information data of a user node and its associated user node match to determine the second mismatch feature data.
[0075] Next, in step S103, a fraud prediction model is established, and the fraud prediction model is trained using the mismatch feature data and fraud performance data of historical users.
[0076] Specifically, for the creation of the fraud prediction model, algorithms such as the CART algorithm or the XGB algorithm can be used to create a model tree (ModelTree), etc. In this example, the XGB algorithm is used to create a model tree (Model Tree).
[0077] It should be noted that the above is only for illustration and should not be construed as a limitation to the present invention. In other examples, other algorithms can also be used, or two or more algorithms can be used in combination, etc.
[0078] In this example, the fraud prediction model is trained using the mismatch feature data and fraud performance data of historical users (as training data), where the mismatch feature data of historical users is used as the feature (X) of the input layer, and the fraud performance data of historical users is used as the feature (Y) of the output layer. In this example, the fraud performance data of historical users is, for example, the fraud probability.
[0079] In addition, training a fraud model using training data also includes defining good and bad samples. As a specific example, "whether a user has fraudulent behavior" can be used to define good and bad samples, that is, the label "whether a user has fraudulent behavior" has a label value specified as 0 or 1, where 1 indicates that the user has fraudulent behavior and 0 indicates that the user does not have fraudulent behavior.
[0080] For each target user, the fraud prediction values (fraud probabilities in this example) of each target product output by the fraud prediction model are usually a numerical value between 0 and 1. The closer it is to 1, the more likely the target user is to have fraudulent behavior.
[0081] Therefore, using a fraud prediction model can achieve fraud prediction of users.
[0082] Next, in step S104, obtain the basic feature data and behavioral feature data of the target user, add the target user to the bilateral network relationship graph to extract the mismatched feature data of the target user, and input it into the fraud prediction model to calculate the fraud prediction value of the target user for risk prediction.
[0083] Specifically, extracting the mismatched feature data of the target user includes: extracting the first mismatched feature data and / or the second mismatched feature of the target user.
[0084] It should be noted that the specific meanings of the mismatched feature data of the target user and the historical user are the same, so the specific description of the mismatched feature data of the target user is omitted.
[0085] Preferably, the prediction method further includes: setting a risk threshold, comparing the calculated risk prediction value of the target user with the risk threshold to classify the risk of the target user, for example, into fraudulent users and non-fraudulent users, and further dividing the fraudulent users into risk levels.
[0086] In other examples, corresponding processing is performed based on the calculated fraudulent users, such as "bad" marking, rejection, reduction of credit rating, etc.
[0087] It should be noted that the above are only preferred embodiments and should not be construed as limitations to the present invention.
[0088] Second Embodiment
[0089] When constructing a network relationship graph, the feature data (sample data) of some users is missing or incomplete. For example, after processing the feature data of a user, there are multiple missing values in the processed vector data, or the feature data of some users cannot be obtained. Therefore, this kind of feature data cannot be used in data analysis, or the effectiveness of using this kind of feature data is low.
[0090] In view of the above problems, the inventors of the present invention have proposed an improved solution, which is specifically as follows.
[0091] See Figure 4 and Figure 5 Describing the second embodiment, the difference between the second embodiment and the first embodiment is that in step S102, local graph feature data of historical users is extracted from the bilateral network relationship graph. However, this is not limited thereto, and local graph feature data and mismatched feature data of historical users can also be extracted.
[0092] In the second embodiment, the risk prediction method further includes determining local graph feature data. Specifically, the local graph feature data includes at least one of degree sequence feature data, polygon feature data, and local clustering coefficient. For specific graph data extraction, see Figure 4 .
[0093] Furthermore, the degree sequence feature data includes out-degree and in-degree. In addition, it also includes the number of associated application users, the proportion of associated fraud users, the weighted number of application users, and the weighted number of fraud users.
[0094] It should be noted that in the present invention, the polygon feature is a circular feature representing a path. In terms of the bilateral network relationship graph, a path in the bilateral network relationship graph will form a loop (or be closed).
[0095] In this embodiment, the polygon feature is a quadrilateral feature (also known as quadrilateral closure). Further, the polygon feature data is the quadrilateral feature data formed by two user nodes associating different information nodes, where the data of one user node corresponds to the data of the target user. Therefore, all the data of the nodes included in the quadrilateral feature is represented by the quadrilateral feature data. Even when the feature data of a certain user node in the quadrilateral feature is incomplete or missing, the user represented by the user node can still be predicted for risk using the quadrilateral feature data.
[0096] Specifically, the quadrilateral feature data includes basic statistical data, label data, and weighted coefficients. More specifically, the basic statistical data of the quadrilateral feature data includes the number of quadrilaterals, the average / maximum / median number of quadrilaterals of associated application users. The label data includes the proportion of quadrilaterals associated with fraud users, the average / maximum / median number of quadrilaterals associated with fraud users. The weighted coefficients include the weighted total number of quadrilaterals, the average / maximum / median number of quadrilaterals of weighted associated application users, and the average / maximum / median number of quadrilaterals of weighted associated fraud users. However, this is not limited thereto, and the above is only for illustrative purposes and should not be construed as a limitation to the present invention.
[0097] It should be noted that here, the weighting is actually an operation based on edges, and the strength also refers to the strength of edges. For example, in anti-fraud, the relationship based on the ID card is stronger than the relationship based on the device, and anti-fraud pays more attention to timeliness, and the timeliness factor of the above relationships is added. Therefore, this kind of weighting is comprehensively defined as a weighting coefficient.
[0098] In other examples, based on the quadrilateral feature data and the information data of the known user nodes and information nodes, the information data of the unknown user nodes (or information nodes) is calculated. For example, in a quadrilateral feature data, the information data of two information nodes and the information data of a user node are known. In this case, based on the known information data, the information data of the unknown user node can be calculated. Therefore, the effectiveness of using data is improved.
[0099] In addition, the polygon feature can also be a triangle feature (also known as triangle closure), or a triangle feature and a quadrilateral feature, but not limited to this. The above is only a preferred embodiment and should not be construed as a limitation to the present invention.
[0100] In the present invention, the clustering coefficient is a coefficient used to evaluate the aggregation degree of nodes in the relationship graph, and the clustering coefficient includes the local clustering coefficient and the global clustering coefficient. Specifically, the local clustering coefficient is a coefficient representing the aggregation degree between nodes in the quadrilateral feature data, and the global clustering coefficient is a coefficient representing the overall aggregation degree in the entire network relationship graph (i.e., the bidirectional network relationship graph).
[0101] More specifically, the local clustering coefficient includes clustering coefficients of different degrees, such as the 1-degree clustering coefficient, the 2-degree clustering coefficient, the 3-degree clustering coefficient, etc. Further, the 1-degree clustering coefficient refers to the coefficient that contains information nodes and user nodes and does not contain fraud user nodes in the user nodes, the 2-degree clustering coefficient refers to the coefficient that contains information nodes and user nodes and contains one fraud user node in the user nodes, and the 3-degree clustering coefficient refers to the coefficient that only contains information nodes and fraud user nodes. In addition, the local clustering coefficient also includes the connectivity reflecting the quadrilateral feature.
[0102] In addition, the difference between the second embodiment and the first embodiment is that in step S104, the basic feature data and the behavior feature data of the target user are obtained, the target user is added to the bilateral network relationship graph to extract the local graph feature data of the target user, and the local graph feature data is input into the fraud prediction model to calculate the fraud prediction value of the target user for risk prediction. Not limited to this, in other examples, the local graph feature data and the mismatched feature data can also be extracted.
[0103] It should be noted that the specific meanings of the local graph feature data of the target user and the historical user are the same, so the specific description of the local graph feature data of the target user is omitted.
[0104] In other examples, the prediction method further includes: a step of further updating the bipartite network relationship graph using the calculated data. Specifically, by calculating the information data of unknown user nodes or information nodes, and adding the calculated information data of the user nodes or information nodes to the bipartite network relationship graph, and using it as the known data in the quadrilateral feature data to calculate the information data of the next unknown user node (or information node).
[0105] Therefore, through the quadrilateral feature data, not only can the information data of unknown and incomplete nodes be calculated, but also the calculated information data of the nodes can be used as new known data to further update the graph data of the bipartite network relationship graph.
[0106] Preferably, for example, detecting the information data of the calculated user node (or information node), comparing the calculated information data of the user node (or information node) with the detected information data of the user node (or information node), and correcting the calculated information data based on the detected information data to make the graph data more accurate.
[0107] Furthermore, the corrected information data of the user node (or information node) is used to further update the graph data of the bipartite network relationship graph, and the above-mentioned graph data is stored for subsequent data analysis, etc.
[0108] It should be noted that since the second embodiment is the same as other parts of the first embodiment, the description of other parts is omitted.
[0109] Third Embodiment
[0110] In existing data extraction, usually nodes with no labels or few labels are not extracted. In order to extract and use the feature data of such nodes, the inventor of the present invention proposed global graph feature data, which is determined based on the data of all extracted nodes (including nodes with no labels or few labels). The specific solution is as follows.
[0111] See Figure 6 Describing the third embodiment, the difference between the third embodiment and the first embodiment is that in step S102, the global graph feature data of historical users is extracted from the bipartite network relationship graph, where the global feature data reflects the characteristics of the overall network, such as the characteristics representing overall connectivity or general characteristics. Without limitation, in other examples, local graph feature data and global graph feature data can also be extracted.
[0112] Specifically, the global feature data includes general features, global clustering coefficients, and / or the connectivity reflecting the overall graph.
[0113] It should be noted that in the present invention, the global clustering coefficient is a coefficient representing the overall aggregation degree in the entire network relationship graph (i.e., the bilateral network relationship graph).
[0114] In this example, for all user nodes and information nodes in the bilateral network relationship graph, according to the importance of in-degree and out-degree, the association degree of each node is calculated to determine the global graph feature data of historical users.
[0115] Preferably, the PageRank (web page ranking) algorithm is used to calculate the PR values of all user nodes and information nodes (the total number of nodes is N) in the bilateral network relationship graph. After iterative operations, the PR value of each node is obtained.
[0116] It should be noted that the PR value originally referred to the probability that a web page was accessed by other web pages. In the present invention, the PR value refers to the probability that a user node is associated with other nodes and represents the association degree of the node. The larger the PR value, the greater the possibility of fraud of the node.
[0117] In addition, for the calculation of the PR value, it also includes the preset of the damping coefficient d, which can be artificially set according to actual needs.
[0118] After calculation by the PageRank algorithm, the PR value (representing the association degree) of each node is compared with a preset threshold, and the feature data of the user node or information node corresponding to the association degree greater than the preset threshold is extracted as the global feature graph feature data.
[0119] In addition, the weights are adjusted according to the historical fraud performance data to obtain the fraud probability of each non-labeled user. For example, as time changes, the weights of historical nodes can be attenuated based on the calculated data.
[0120] In addition, the difference between the third embodiment and the first embodiment is that in step S104, the basic feature data and behavioral feature data of the target user are obtained, the target user is added to the bilateral network relationship graph to extract the global graph feature data of the target user, and the fraud prediction value of the target user is calculated for risk prediction. In other examples, the global graph feature data and mismatched feature data of the target user can be extracted. The above is only for illustration and should not be construed as a limitation of the present invention.
[0121] It should be noted that since the other parts of the third embodiment are the same as those of the first embodiment and the second embodiment, the description of the other parts is omitted.
[0122] Fourth Embodiment
[0123] The difference between the fourth embodiment and the first embodiment is that in step S102, local graph feature data, global graph feature data, and mismatch feature data of historical users are extracted from the bilateral network relationship graph, where the mismatch feature data includes first mismatch feature data and / or second mismatch feature data. Without limitation, in other examples, any combination of the above three types of feature data can also be extracted.
[0124] In addition, the difference between the fourth embodiment and the first embodiment is that in step S104, basic feature data and behavioral feature data of the target user are obtained, the target user is added to the bilateral network relationship graph to extract local graph feature data, global graph feature data, and mismatch feature data of the target user, and the fraud prediction model is input to calculate the fraud prediction value of the target user for risk prediction. In other examples, any combination of the above three types of feature data can be extracted.
[0125] It should be noted that since the other parts of the fourth embodiment are the same as those of the first embodiment, the second embodiment, and the third embodiment, the description of the other parts is omitted.
[0126] The above are only preferred embodiments and should not be construed as limitations on the present invention. In other examples, step S104 can also be split into two steps (S601 and S104), specifically refer to Figure 6 .
[0127] Those skilled in the art can understand that all or part of the steps of implementing the above embodiments are realized as a program (computer program) executed by a computer data processing device. When the computer program is executed, the above methods provided by the present invention can be realized. Moreover, the computer program can be stored in a computer-readable storage medium, which can be a disk, an optical disc, a ROM, a RAM, or other readable storage media, or a storage array composed of multiple storage media, such as a disk or tape storage array. The storage medium is not limited to centralized storage, and it can also be distributed storage, such as cloud storage based on cloud computing.
[0128] Next, the risk prediction method of the present invention is applied to a specific business for effect verification, specifically refer to Table 1 below.
[0129]
[0130] Note: Traditional model: An integrated learning model based on personal graph data and features
[0131] The risk prediction model of the present invention: An integrated learning model based on bilateral network graph data and features
[0132] As can be seen from Table 1, compared with the traditional method, the KS of the prediction model in the risk prediction method of the present invention has increased by 0.08, approaching 27%, and the AUC has increased by 4%. Therefore, the accuracy of the fraud model has been greatly improved, and the first default rate has decreased, improving the business level.
[0133] Compared with the prior art, the risk prediction method of the present invention is widely applicable, suitable for large-scale data processing and data analysis. Especially when the feature data of user nodes is incomplete or missing, it can also predict the risk of the user represented by the user node by using the mining of group graph features (quadrilateral feature data), improving the accuracy of risk prediction. In addition, the risk prediction method of the present invention predicts risks by using mismatched features, improving the accuracy of the fraud model; improving the utilization rate of data, optimizing the target data; reducing the risk of online loans; and improving the business level.
[0134] Embodiment 2
[0135] The device embodiments of the present invention are described below. The device can be used to execute the method embodiments of the present invention. For the details described in the device embodiments of the present invention, they should be regarded as a supplement to the above method embodiments; for the details not disclosed in the device embodiments of the present invention, they can be implemented with reference to the above method embodiments.
[0136] Referring to Figure 7 、 Figure 8 and Figure 9 The present invention also provides a risk prediction device 700 based on graph data. The risk prediction device 700 includes: a data acquisition module 701, which is used to acquire the basic feature data and behavioral feature data of historical users and construct a bilateral network relationship graph. The bilateral network relationship graph includes two types of nodes, namely user nodes and information nodes. The user node is a node representing a user, and the information node is a node that associates different users; a data processing module 702, which extracts the mismatched feature data of historical users from the bilateral network relationship graph. The mismatched feature data includes first mismatched feature data and second mismatched feature data; a model establishment module 703, which establishes a fraud prediction model and trains the fraud prediction model using the mismatched feature data and fraud performance data of historical users; a calculation module 704, which acquires the basic feature data and behavioral feature data of a target user, adds the target user to the bilateral network relationship graph to extract the mismatched feature data of the target user, and inputs the mismatched feature data into the fraud prediction model to calculate the fraud prediction value of the target user for risk prediction.
[0137] Preferably, the risk prediction device further includes an extraction module, and the extraction module is used to extract the first mismatched feature data and the second mismatched feature of the target user.
[0138] Preferably, the risk prediction device further includes an extraction module, which is configured to extract the first mismatched feature data or the second mismatched feature of the target user.
[0139] Preferably, as Figure 8 shown, the risk prediction device further includes a determination module 801. When the data sources are different, the determination module calculates the similarity (dissimilarity) of the same user node or information node based on the determined Jaccard distance to determine the first mismatched feature data.
[0140] Preferably, the risk prediction device further includes a determination module. When the data sources are different, the determination module calculates the matching degree (mismatching degree) of the same user node or information node based on the determined shortest distance between two adjacent nodes to determine the first mismatched feature data.
[0141] Preferably, as Figure 9 shown, the risk prediction device further includes a monitoring module 901, which is configured to monitor whether the information data of the user node and its associated user node match to determine the second mismatched feature data.
[0142] Preferably, the user node includes user personal feature data and network feature data; and / or the information node includes at least one of APP information, location information, address book information, call record information, device information, and operator information.
[0143] Preferably, the risk prediction device further includes a setting module, which is configured to set a risk threshold, compare the calculated risk prediction value of the target user with the risk threshold, and classify the risk of the target user.
[0144] It should be noted that in Embodiment 2, the description of the parts identical to those in Embodiment 1 is omitted.
[0145] Those skilled in the art can understand that the modules in the above device embodiments can be distributed in the device according to the description, or can be correspondingly changed and distributed in one or more devices different from the above embodiments. The modules in the above embodiments can be combined into one module, or further split into multiple sub-modules.
[0146] Embodiment 3
[0147] Embodiments of the electronic device of the present invention are described below. The electronic device can be regarded as a specific physical implementation manner of the above-described method and device embodiments of the present invention. For the details described in the embodiments of the electronic device of the present invention, they should be regarded as a supplement to the above-described method or device embodiments; for the details not disclosed in the embodiments of the electronic device of the present invention, reference can be made to the above-described method or device embodiments for implementation.
[0148] Figure 10 is a structural block diagram of an exemplary embodiment of an electronic device according to the present invention. The following will be described with reference to Figure 10 the electronic device 200 according to this embodiment of the present invention. Figure 10 The illustrated electronic device 200 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0149] As Figure 10 shown, the electronic device 200 is presented in the form of a general-purpose computing device. The components of the electronic device 200 may include, but are not limited to: at least one processing unit 210, at least one storage unit 220, a bus 230 connecting different system components (including the storage unit 220 and the processing unit 210), a display unit 240, etc.
[0150] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 210, so that the processing unit 210 executes the steps according to various exemplary embodiments of the present invention described in the above-mentioned electronic prescription transfer processing method part of this specification. For example, the processing unit 210 can execute the steps as Figure 1 shown.
[0151] The storage unit 220 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 2201 and / or a cache storage unit 2202, and may further include a read-only storage unit (ROM) 2203.
[0152] The storage unit 220 may further include a program / utility 2204 having a set (at least one) of program modules 2205. Such program modules 2205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0153] The bus 230 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.
[0154] The electronic device 200 can also communicate with one or more external devices 300 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 200, and / or communicate with any device that enables the electronic device 200 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 250. Moreover, the electronic device 200 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 260. The network adapter 260 can communicate with other modules of the electronic device 200 through the bus 230. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0155] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described in the present invention can be implemented by software, or can be implemented by the way of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present invention can be embodied in the form of a software product, and the software product can be stored in a computer-readable storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the present invention. When the computer program is executed by a data processing device, the computer-readable medium can implement the above method of the present invention, that is: using the APP download sequence vector data and overdue information of historical users as training data to train the created user risk control model, and using the created user risk control model to calculate the financial risk prediction value of the target user.
[0156] As Figure 11 shown, the computer program can be stored on one or more computer-readable media. The computer-readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0157] The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.
[0158] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0159] In summary, the present invention may be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that general-purpose data processing devices such as microprocessors or digital signal processors (DSPs) may be used in practice to implement some or all of the functions of some or all of the components in accordance with the embodiments of the present invention. The present invention may also be implemented as a device or apparatus program (e.g., a computer program and a computer program product) for performing part or all of the methods described herein. Such a program implementing the present invention may be stored on a computer-readable medium, or may be in the form of one or more signals. Such signals may be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0160] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the present invention is not inherently related to any specific computer, virtual device, or electronic device, and various general-purpose devices can also implement the present invention. The above description is only for the specific embodiments of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc., made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A risk prediction method based on two - party network diagram data, characterized in that, Including: Obtain the basic feature data and behavioral feature data of historical users, and construct a bilateral network relationship graph, which includes two types of nodes: user nodes and information nodes; the behavioral feature data includes user association behavior data and financial behavior data; the user nodes are nodes representing users, including user personal feature data and network feature data; the information nodes are nodes that connect different users; among them, for the local graph feature data, calculate the information data of unknown user nodes or information nodes through the information data of known user nodes and information nodes, and use the calculated information data to further update the bilateral network relationship graph; Extract the global graph feature data, mismatch feature data, and local graph feature data of historical users from the bilateral network relationship graph; among them, the global graph feature data reflects the characteristics of the overall network, including characteristics indicating overall connectivity or general characteristics, and is determined based on extracting the data of all nodes; the mismatch feature data includes first mismatch feature data and second mismatch feature data; the local graph feature data includes at least one of degree sequence feature data, quadrilateral feature data, and / or triangle feature data, local clustering coefficient, and global clustering coefficient, where the quadrilateral feature data is formed by two user nodes associating with different information nodes, and represents the data of all nodes included in the quadrilateral feature; Establish a fraud prediction model, use the mismatch feature data of the historical users as the features of the input layer and the fraud performance data of the historical users as the features of the output layer, define the quality of the samples, and train the fraud prediction model; Obtain the basic feature data and behavioral feature data of the target user, add the target user to the bilateral network relationship graph to extract the mismatch feature data, local graph feature data, and global graph feature data of the target user, and input them into the fraud prediction model to calculate the fraud prediction value of the target user for risk prediction.
2. The risk prediction method according to claim 1, wherein Also including: Based on determining the Jaccard distance or the shortest distance between adjacent two nodes when the data sources are different, calculate the similarity of the same user node or information node to determine the first mismatch feature data; Monitor whether the information data of a user node and the user nodes associated with it match to determine the second mismatch feature data, where the second mismatch feature data includes at least one of the mismatch between the personal information and network information of the same user node, the mismatch of geographical information between two mutually associated user nodes, and the mismatch of information between a mutually associated user node and information node; The degree sequence feature data includes out-degree, in-degree, and also includes the number of associated applying users, the proportion of associated fraudulent users, the weighted number of applying users, and the weighted number of fraudulent users; The quadrilateral feature data is the quadrilateral feature data formed by two user nodes associating with different information nodes, representing the data of all nodes included in the quadrilateral feature, making the data of one user node correspond to the data of the target user, and it includes: basic statistical data, label data, and weighted coefficient; And, calculating information data of an unknown user node or information node based on the quadrilateral feature data and the information data of known user nodes and information nodes.
3. The risk prediction method according to claim 1, wherein Using the calculated information data to further update the bilateral network relationship graph, including: Detecting the information data of the calculated user node or information node, comparing the information data of the calculated user node or information node with the detected information data of the user node or information node, correcting the calculated information data based on the detected information data, and using the corrected information data of the user node or information node to update the graph data of the bilateral network relationship graph.
4. The risk prediction method according to claim 1, wherein The risk prediction method further includes: Monitoring whether the information data of a user node and its associated user nodes match to determine second mismatch feature data.
5. The risk prediction method according to claim 1, wherein The prediction method further includes: Setting a risk threshold, comparing the calculated risk prediction value of the target user with the risk threshold to classify the risk of the target user.
6. A risk prediction device based on graph data, characterized in that, The risk prediction device includes: A data acquisition module, which acquires the basic feature data and behavioral feature data of historical users and constructs a bilateral network relationship graph. The bilateral network relationship graph includes two types of nodes: user nodes and information nodes; the behavioral feature data includes user association behavior data and financial behavior data; the user node is a node representing a user, including user personal feature data and network feature data; the information node is a node that associates different users; among them, for the local graph feature data, the information data of unknown user nodes or information nodes is calculated through the information data of known user nodes and information nodes, and the calculated information data is used to further update the bilateral network relationship graph; A data processing module, which extracts the global graph feature data, mismatch feature data, and local graph feature data of historical users from the bilateral network relationship graph; among them, the global graph feature data reflects the characteristics of the overall network, including features representing overall connectivity or general features, and is determined based on the data of all extracted nodes; the mismatch feature data includes first mismatch feature data and second mismatch feature data; the local graph feature data includes at least one of degree sequence feature data, quadrilateral feature data, and / or triangle feature data, local clustering coefficient, and global clustering coefficient. The quadrilateral feature data is formed by two user nodes associating different information nodes, and the data of all nodes included in the quadrilateral feature is represented by the quadrilateral feature data; A model establishment module, which establishes a fraud prediction model, uses the mismatch feature data of historical users as the features of the input layer and the fraud performance data of historical users as the features of the output layer, defines the quality of the samples, and trains the fraud prediction model; A calculation module, which acquires the basic feature data and behavioral feature data of a target user, adds the target user to the bilateral network relationship graph to extract the mismatch feature data, local graph feature data, and global graph feature data of the target user, and inputs them into the fraud prediction model to calculate the fraud prediction value of the target user for risk prediction.
7. The risk prediction device according to claim 6, wherein It further includes: Calculating the similarity of the same user node or information node based on determining the Jaccard distance or the shortest distance between two adjacent nodes in different data sources to determine the first mismatched feature data; Monitoring whether the information data of a user node and its associated user nodes match to determine the second mismatched feature data, where the second mismatched feature data includes at least one of the following: the personal information and network information of the same user node do not match, the geographical information of two mutually associated user nodes does not match, and the information of mutually associated user nodes and information nodes does not match; The degree sequence feature data includes the out-degree, in-degree, and also includes the number of associated applying users, the proportion of associated fraudulent users, the weighted number of applying users, and the weighted number of fraudulent users; The quadrilateral feature data is the quadrilateral feature data formed by two user nodes associating with different information nodes, representing the data of all nodes included in the quadrilateral feature, such that the data of one user node corresponds to the data of the target user, and it includes: basic statistical data, label data, and weighted coefficients; And, calculating the information data of unknown user nodes or information nodes based on the quadrilateral feature data and the known information data of user nodes and information nodes.
8. The risk prediction device according to claim 6, wherein: Further updating the bilateral network relationship graph using the calculated information data, including: Detecting the information data of the calculated user node or information node, comparing the calculated information data of the user node or information node with the detected information data of the user node or information node, correcting the calculated information data based on the detected information data, and using the corrected information data of the user node or information node to update the graph data of the bilateral network relationship graph.
9. The risk prediction device according to claim 6, wherein It further includes a monitoring module for monitoring whether the information data of a user node and its associated user nodes match to determine the second mismatched feature data.
10. The risk prediction device according to claim 6, wherein It further includes a setting module for setting a risk threshold, comparing the calculated risk prediction value of the target user with the risk threshold to classify the risk of the target user.
11. An electronic device, wherein, The electronic device includes: A processor; and, A memory storing computer-executable instructions, and the executable instructions, when executed, cause the processor to execute the risk prediction method according to any one of claims 1-5.
12. A computer-readable storage medium, wherein, The computer-readable storage medium stores one or more programs, and when the one or more programs are executed by the processor, the risk prediction method according to any one of claims 1-5 is implemented.