A financial fraud behavior prediction method, device, medium and product based on a complex network

CN121190068BActive Publication Date: 2026-09-11SHANGHAI FANGFUTONG TECH SERVICES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511345727.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-09-11
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

但是,之前的复杂网络金融欺诈行为预测模型,并未对线下群体欺诈行为进行优化,因此对此群体欺诈行为预测效果有欠缺

Benefits of technology

[0036] This application provides a method, device, medium, and product for predicting financial fraud based on complex networks. The method first constructs a complex network, treating participating entities in loan orders as nodes, and establishing connections between nodes based on order relationships. Targeting the characteristics of group fraud, the method performs specific optimization extraction of node features, including using business address tile encoding technology to process geographic location information, and classifying and marking suspected fraudulent small group networks based on order four-category labels. Subsequently, the association features between nodes and small groups are calculated, and prediction is performed using an XGBoost model. This application can effectively identify potential fraud risk groups and key nodes, providing target clues for risk verification, while improving the prediction accuracy of offline organized fraudulent activities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121190068B_ABST
    Figure CN121190068B_ABST
Patent Text Reader

Abstract

The application discloses a financial fraud behavior prediction method and device based on a complex network, a medium and a product, relates to the field of financial risk control, and comprises the following steps: obtaining a to-be-predicted loan order; the to-be-predicted loan order is an online loan order or an offline loan order; updating a complex network of historical loan orders based on the to-be-predicted loan order, to obtain an updated complex network; obtaining a node feature set of the to-be-predicted loan order in the updated complex network and a topological distance feature from a suspect small group of a customer node and a customer spouse node of the to-be-predicted loan order; and determining a fraud behavior category of the to-be-predicted loan order by using a financial fraud behavior prediction model according to the node feature set and the topological distance feature. The application can effectively identify potential fraud risk groups and key nodes, and meanwhile, the prediction accuracy of offline organized fraud behaviors is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of financial risk control, and in particular to a method, device, medium and product for predicting financial fraud based on complex networks. Background Technology

[0002] For business loans, banks currently combine offline data collection by account managers with online risk control models. However, organized fraud is prevalent offline, making it difficult to detect through online order financial characteristics. These fraudulent activities often exhibit key node misuse features (misuse of customer and spouse relationships, business address, business location, business license, etc.). Using a four-category labeling system based on historical orders (0: no fraud and no overdue payments, 1: overdue, 2: fraud, 3: fraud and overdue), algorithms can be trained to identify related fraudulent activities and groups within complex network relationships through community segmentation, key edge betweenness, and node-specific features. However, previous models predicting complex network financial fraud have not been optimized for offline group fraud, thus their predictive effectiveness against such groups is lacking. Summary of the Invention

[0003] The purpose of this application is to provide a method, device, medium, and product for predicting financial fraud based on complex networks, which can effectively identify potential fraud risk groups and key nodes, while improving the prediction accuracy of organized offline fraud.

[0004] To achieve the above objectives, this application provides the following solution:

[0005] Firstly, this application provides a method for predicting financial fraud based on complex networks, including:

[0006] Obtain loan orders to be predicted; the loan orders to be predicted can be online loan orders or offline loan orders.

[0007] Based on the loan order to be predicted, the complex network of historical loan orders is updated to obtain the updated complex network; the complex network is constructed based on the historical loan orders; the entities of the historical loan orders are the nodes of the complex network, and the relationships between the historical single orders are the edges of the complex network; the entities include account managers, marketing managers, reviewers, customers, customer spouses, business license numbers, and business locations;

[0008] The node feature set of the loan order to be predicted in the updated complex network is obtained, along with the topological distance features from the customer node and customer spouse node of the loan order to the suspected small group. The node feature set includes customer manager node features, marketing manager node features, reviewer node features, customer node features, customer spouse node features, business license node features, and address tile node features. The node features include node salience index, weighted maximum salience index of surrounding nodes, node k-core index, surrounding node k-core index, node betweenness, and node knn index. The suspected small group includes hidden small groups related to fraud and delinquency, fraud small groups, and delinquency small groups. The suspected small groups are determined based on the complex network using the k-core hierarchical small group algorithm, the ECG-LA method, and the label propagation-based LPA algorithm.

[0009] Based on the node feature set and the topological distance feature, a financial fraud behavior prediction model is used to determine the fraud behavior category of the loan order to be predicted; the fraud behavior category is no fraud and no overdue payment, overdue payment, fraud, or fraud and overdue payment; the financial fraud behavior prediction model is obtained by training an XGBoost model.

[0010] In one implementation, based on the loan order to be predicted, the complex network of historical loan orders is updated to obtain the updated complex network, specifically including:

[0011] The loan orders to be predicted are anonymized and tile-encoded to obtain the processed loan orders;

[0012] Based on the processed loan orders, the complex network of historical loan orders is updated to obtain the updated complex network.

[0013] In one embodiment, the process of constructing the financial fraud prediction model specifically includes:

[0014] Obtain a number of historical loan orders; the historical loan orders include online historical single-item orders and offline historical loan orders;

[0015] The historical loan orders are desensitized and tile-encoded to obtain the processed historical loan orders;

[0016] A complex network is constructed by using the entities of the processed historical loan orders as nodes and the relationships between the processed historical single orders as edges.

[0017] Based on the complex network, the k-core hierarchical small group algorithm, ECG-LA method, and LPA algorithm based on label propagation are used to identify suspected small groups and label them with four categories of labels: no fraud and no overdue, overdue, fraud, or fraud and overdue.

[0018] Extract the node feature set of training loan orders in the complex network, as well as the topological distance features from the customer nodes and customer spouse nodes of the training loan orders to the suspected small group, and label the training loan orders with four-class labels; the training loan orders are obtained from the historical loan orders;

[0019] The XGBoost model is trained using the node feature set and topological distance features of the training loan orders as input and the four-class classification labels of the training loan orders as output, to obtain the financial fraud behavior prediction model.

[0020] In one embodiment, the historical loan orders are de-identified and tile-encoded to obtain processed historical loan orders, specifically including:

[0021] The account manager, marketing manager, reviewer, customer, customer's spouse and business license number in the historical loan order are de-identified using the MD5 encryption algorithm to obtain the de-identified historical loan order.

[0022] The business locations in the de-identified historical loan orders are then subjected to tile encoding to obtain the processed historical loan orders.

[0023] In one embodiment, extracting the node feature set of training loan orders in the complex network specifically includes:

[0024] Calculate the betweenness number of each edge in the complex network, as well as the node degree, node kNN index, and node k-core index of each node;

[0025] The node salience index of each node in the training loan order is calculated based on the betweenness of the edge, the degree of the node, and the k-core index of the node.

[0026] Based on the betweenness of edges and the salience index of nodes, the weighted maximum salience index of the surrounding nodes of each node in the training loan order is calculated.

[0027] Extract the node kNN index, node k-core index, node betweenness number, and the maximum k-core index of surrounding nodes for each node in the training loan order to obtain the node feature set of the training loan order.

[0028] In one implementation, the node salience index of each node in the training loan order is calculated based on the betweenness of edges, the degree of nodes, and the k-core index of nodes, specifically including:

[0029] Using formula V s =log(v j *v k *v d *v ldmax ) Calculate the node salience index for each node in the training loan order; where V s For node prominence indicators; v j v is the betweenness number of the edge; k v is the node's k-core index; d The degree of the node; v ldmax This represents the maximum degree value of the surrounding nodes.

[0030] In one implementation, based on the betweenness of edges and node salience indices, the weighted maximum salience index of the surrounding nodes of each node in the training loan order is calculated, specifically including:

[0031] Using the formula N s =MAX(E b ×V s Calculate the weighted maximum salience index of the surrounding nodes for each node in the training loan order; where N s E is the weighted maximum salience index of the surrounding nodes associated with the current node. b Let V be the betweenness number of the edge. s It serves as a prominent indicator for surrounding nodes.

[0032] Secondly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the financial fraud prediction method based on complex networks as described above.

[0033] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the financial fraud prediction method based on complex networks described above.

[0034] Fourthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the financial fraud prediction method based on complex networks described above.

[0035] According to the specific embodiments provided in this application, this application has the following technical effects:

[0036] This application provides a method, device, medium, and product for predicting financial fraud based on complex networks. The method first constructs a complex network, treating participating entities in loan orders as nodes, and establishing connections between nodes based on order relationships. Targeting the characteristics of group fraud, the method performs specific optimization extraction of node features, including using business address tile encoding technology to process geographic location information, and classifying and marking suspected fraudulent small group networks based on order four-category labels. Subsequently, the association features between nodes and small groups are calculated, and prediction is performed using an XGBoost model. This application can effectively identify potential fraud risk groups and key nodes, providing target clues for risk verification, while improving the prediction accuracy of offline organized fraudulent activities. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 A flowchart illustrating a method for predicting financial fraud based on complex networks, provided as an embodiment of this application;

[0039] Figure 2 A schematic diagram of tile coding;

[0040] Figure 3 This is a diagram illustrating the node relationships within an order.

[0041] Figure 4 A schematic diagram of a complex network;

[0042] Figure 5 This is a flowchart of the data processing process;

[0043] Figure 6 This is a schematic diagram illustrating the distance characteristics from a node to a suspected small group.

[0044] Figure 7 Flowchart for training a risk prediction model (financial fraud prediction model);

[0045] Figure 8 A flowchart for the risk prediction model;

[0046] Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0047] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0048] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0049] This application relates to the field of financial fraud prediction. For business loans, banks currently use a combination of offline data collection by account managers and online risk control models. In view of the possibility of organized fraud in offline data collection, this application proposes a method for predicting organized financial fraud based on a complex network model and using node features and community partitioning algorithms.

[0050] This application addresses group fraud by optimizing feature extraction and algorithms. During algorithm development, knowledge from fraud detection experts was incorporated and integrated into the algorithm, resolving the issue of poor prediction performance.

[0051] In one exemplary embodiment, such as Figure 1 As shown, a method for predicting financial fraud based on complex networks is provided, including the following steps:

[0052] S1: Obtain the loan orders to be predicted; the loan orders to be predicted are online loan orders or offline loan orders.

[0053] S2: Based on the loan order to be predicted, update the complex network of historical loan orders to obtain the updated complex network; the complex network is constructed based on the historical loan orders; the entities of the historical loan orders serve as nodes of the complex network, and the relationships between the historical single orders serve as edges of the complex network; the entities include account managers (collection managers), marketing managers, reviewers (review positions), customers, customer spouses, business license numbers, and business premises (business addresses).

[0054] As an optional implementation, S2 specifically includes:

[0055] S21: Desensitize and tile-encode the loan order to be predicted to obtain the processed loan order.

[0056] S22: Based on the processed loan orders, update the complex network of historical loan orders to obtain the updated complex network.

[0057] S3: Obtain the node feature set of the loan order to be predicted in the updated complex network, as well as the topological distance features from the customer node and customer spouse node of the loan order to be predicted to the suspected small group; the node feature set includes customer manager node features, marketing manager node features, reviewer node features, customer node features, customer spouse node features, business license node features, and address tile node features; the node features include node salience index, weighted maximum salience index of surrounding nodes, node k-core index, surrounding node k-core index, node betweenness, and node knn index; the suspected small group includes hidden small groups related to fraud and delinquency, fraud small groups, and delinquency small groups; the suspected small groups are determined based on the complex network using the k-core hierarchical small group algorithm, the ECG-LA method, and the label propagation-based LPA algorithm.

[0058] S4: Based on the node feature set and the topological distance feature, the fraud behavior category of the loan order to be predicted is determined using a financial fraud behavior prediction model; the fraud behavior category is no fraud and no overdue payment, overdue payment, fraud, or fraud and overdue payment; the financial fraud behavior prediction model is obtained by training an XGBoost model.

[0059] In one embodiment, the process of constructing the financial fraud prediction model specifically includes:

[0060] Step 1: Obtain several historical loan orders; the historical loan orders include online historical single-item orders and offline historical loan orders.

[0061] Step 2: Desensitize and tile-encode the historical loan orders to obtain the processed historical loan orders.

[0062] As an optional implementation, step 2 specifically includes:

[0063] Step 21: Use the MD5 encryption algorithm to de-identify the account manager, marketing manager, reviewer, customer, customer's spouse, and business license number in the historical loan order to obtain the de-identified historical loan order.

[0064] In this embodiment, each order includes node numbers for five personnel types: customer, customer's spouse, marketing manager, account manager, and reviewer, as well as a business license number. Based on the order's node type, a single de-identification task is performed on the personnel and business license nodes: a unique de-identification code is generated using a hash algorithm to ensure security when nodes are merged across networks. This step independently addresses data privacy issues.

[0065] The node encoding is de-identified using the MD5 encryption algorithm (MD5 is a type of hash algorithm) to avoid leaking personal information.

[0066] Step 22: Perform tile encoding on the business locations in the de-identified historical loan orders to obtain the processed historical loan orders.

[0067] After encoding, the node only contains the order-related employee ID, identity, business license, and business address information, excluding financial information, thus eliminating the risk of financial information leakage. The identity, business license, and business location GPS information have also been anonymized, further preventing the risk of information leakage.

[0068] For the GPS nodes of the business premises in step 21, special processing is performed: the GPS coordinates of the business premises are divided into tiles of 20 meters each, and adjacent tiles are numbered uniformly using a proprietary algorithm (e.g., Figure 2 The tile partitioning scheme completely resolves the address spoofing vulnerability by using AD1 as the association logic between two adjacent nodes and AD2 as the association logic between non-adjacent nodes. The tile partitioning scheme is as follows:

[0069] 1. Determine the longitude and latitude of the geometric center point of the region.

[0070] 2. Calculate the longitude difference delta_long and latitude difference delta_lat for a 20m x 20m area in this region.

[0071] 3. Divide the grid according to the longitude difference delta_long and the latitude difference delta_lat to form tiles.

[0072] 4. Determine which tile the business address is located on, and determine the tile's horizontal and vertical numbers, tile_x and tile_y.

[0073] 5. The geographical locations form a sparse matrix, with 1 for locations with business addresses and 0 for locations without business addresses.

[0074] 6. Scan the sparse matrix formed by the tile codes of the business address, and use a 3×3 adjacency matrix to determine whether there are other business address tiles around a tile with a business address. If so, adjust them to have the same code. Finally, assign a unique code to each tile.

[0075] 7. For new orders, the tile where the business address is located will be assigned a tile number if there are surrounding tiles with similar business addresses. Otherwise, a new tile number will be assigned. A tile coding diagram is shown below. Figure 2 As shown.

[0076] Step 3: Construct a complex network by using the entities of the processed historical loan orders as nodes and the relationships between the processed historical single orders as edges.

[0077] In this embodiment, the composition rules of nodes and edges are first clarified: nodes encompass all participating entities in a loan order, including customers, customer spouses, marketing managers, investigation managers, reviewers, tile codes (converted through GPS location conversion of the business premises), and business licenses, totaling seven node types; edges establish connections between nodes through node relationships, forming the initial network skeleton. This step provides the framework for subsequent feature calculations. The edge number is the order number. Based on the relationships between order nodes, such as... Figure 3 As shown, a complex network is constructed sequentially based on a series of order information. The resulting complex network diagram and node relationship diagram are as follows. Figure 4 As shown.

[0078] Step 4: Based on the complex network, using the k-core hierarchical small group algorithm, ECG-LA method and label propagation-based LPA algorithm, identify the suspected small groups and label them with four categories of labels; the four categories of labels are no fraud and no overdue, overdue, fraud, or fraud and overdue.

[0079] Step 5: Extract the node feature set of the training loan order in the complex network, as well as the topological distance features from the customer node and customer spouse node of the training loan order to the suspected small group, and label the training loan order with a four-classification label; the training loan order is obtained from the historical loan orders.

[0080] As an optional implementation, extracting the node feature set of training loan orders in the complex network specifically includes:

[0081] Calculate the betweenness number of each edge in the complex network, as well as the node degree, node kNN index, and node k-core index of each node.

[0082] Based on the betweenness of edges, the degree of nodes, and the k-core index of nodes, the node salience index of each node in the training loan order is calculated, specifically including:

[0083] Using formula V s =log(v j *v k *v d *v ldmax ) Calculate the node salience index for each node in the training loan order; where V s For node prominence indicators; v j v is the betweenness number of the edge; k v is the node's k-core index; d The degree of the node; v ldmax This represents the maximum degree value of the surrounding nodes.

[0084] Based on the edge betweenness and the node prominence indicators, the weighted maximum prominence indicator of neighboring nodes for each node in the training loan orders is calculated, which specifically comprises:

[0085] The weighted maximum prominence indicator of neighboring nodes for each node in the training loan orders is calculated by using the formula MAX(betweenness of edge × prominence indicator of neighboring node).

[0086] Extracting the node knn index, node k-core index, node betweenness and maximum k-core index of neighboring nodes of each node in the training loan orders to obtain a node feature set of the training loan orders.

[0087] In this embodiment, based on the complex network structure constructed in step 3, a feature extraction task is performed: calculating the betweenness of the order number edge as a core parameter for subsequent weighting; extracting basic attributes such as node degree, knn index and k-core index of nodes; finally, aggregating neighbor node features to generate a new feature family (such as peripheral weighted k-core value) by taking the edge betweenness as the weight, so as to complete the derivation of deep network features. In current calculation, the edge betweenness function of python igraph library is used for calculation.

[0088] 1. Calculate the betweenness of each edge.

[0089] Edge betweenness reflects the probability that this edge is passed by the shortest path between two nodes. A larger betweenness indicates that the probability that this edge lies on a shortest path is higher.

[0090] Edge betweenness measures the importance of an edge as a "bridge" in the network, defined as the sum of the proportions of the shortest paths between all pairs of vertices that pass through the edge.

[0091] For edge e, the calculation formula of its betweenness B(e) is:

[0092]

[0093] Wherein: V is the set of all nodes in the complex network; σ(s,t) represents the total number of shortest paths from node s to node t; σ(s,t|e) represents the number of shortest paths from node s to node t that pass through edge e; the summation is performed for all vertex pairs with s<t (to avoid repeated calculation).

[0094] Large edge betweenness also means that the node associated with the edge has a higher probability of becoming a key node exploited by group fraud cliques.

[0095] 2. First calculate basic attributes such as degree, knn index and k-core index of each node.

[0096] Note: a. The degree of a node is the number of edges directly associated with that node.

[0097] b. The k-nearest neighbors' average degree (also known as "average neighbor degree") of a node is used to describe the average degree of all its neighboring nodes.

[0098] For a node v, its knn exponent (denoted as knn(v)) is defined as the average degree of all its direct neighbors (nodes directly connected to v).

[0099]

[0100] Let N(v) be the set of neighbors of node v (i.e. all nodes directly connected to v), deg(u) represent the degree of neighbor node u, and |N(v)| represent the number of neighbors (i.e. the degree of node v itself, |N(v)| = deg(v)).

[0101] c. The k-core index of a node:

[0102] For a non-negative integer k, the k-core of a graph G is the largest subgraph H in G that satisfies:

[0103] H is a subgraph of G (containing the edges of all nodes in subgraph H in the original graph); the degree of each node in H is at least k.

[0104] 3. Aggregate the edges surrounding a node and the features of its neighboring nodes (edge ​​betweenness, node degree, k-nn index, k-core index), and weight them using the edges between the node and its neighboring nodes to generate a new feature family. This reflects the probability that the node is a key node in a group fraud.

[0105] 4. Ultimately, each node uses the following six features:

[0106] Node prominence metric: A custom prominence metric for nodes, which aggregates the node's own degree, the degree of surrounding nodes, and the betweenness of surrounding edges, and takes the logarithm of the weighted average.

[0107] Calculation formula: V s =log(v j *v k *v d *v ldmax ).

[0108] Weighted maximum salience index of surrounding nodes: A custom weighted maximum salience index of surrounding nodes.

[0109] Calculation formula: N s=MAX(E b ×V s ); where N s E is the weighted maximum salience index of the surrounding nodes associated with the current node. b Let V be the betweenness number of the edge. s It serves as a prominent indicator for surrounding nodes.

[0110] k_core index: The k_core value of a node.

[0111] k_core index of surrounding nodes: the maximum k_core value of surrounding nodes.

[0112] Node betweenness: The betweenness of a node reflects the probability that the node lies on the shortest path. Currently, it is calculated using the `betweenness` function from the `igraph` library.

[0113] KNN index for a node: The KNN index of a node.

[0114] The above features are normalized, and the node features are stored in the node attributes of the complex network. The complex network is then output for later use.

[0115] Based on the complex network structure and calculated node characteristics from step 3, a community detection task is performed: the k-core hierarchical small group algorithm and the ECG-LA method are used to divide the small group structure. Then, based on the fraud and delinquency status of historical loan orders, small groups are labeled with four categories (0 to 3, 0: no fraud and no delinquency, 1: delinquent, 2: fraud, 3: fraud and delinquency). The label propagation-based LPA algorithm is used to find hidden small groups associated with fraudulent and delinquent orders.

[0116] Output the relevant diagrams for the small group, to be used later.

[0117] A k-core clique is a maximal complete subgraph in a graph (where any two distinct nodes in the subgraph are directly connected), where each node has a degree of at least k.

[0118] The ECG-LA (Ensemble Clustering for Graphs) method is an improvement on the Louvain community partitioning method, developed by V. Pouliny and F. Théberge in 2019. The Louvain community partitioning method is a complex network community detection algorithm whose core objective is to identify tightly connected communities in the network by maximizing modularity.

[0119] The Label Propagation Algorithm (LPA) is a graph-based semi-supervised learning algorithm initially used for community detection, and later widely applied in semi-supervised classification, network analysis, and other fields. Its core idea is to "propagate" labels through the connections between nodes in the graph, ultimately resulting in closely connected nodes having the same or similar labels, thus achieving community segmentation or label prediction.

[0120] Overall data processing, complex network construction, node characteristic formation, and small group partitioning process are as follows: Figure 5 As shown.

[0121] Extract the node features of the loan orders used for training, as well as the distance indicators (topological distance features) from the customer nodes and customer spouse nodes of the orders to the suspected small group.

[0122] Based on the different node types of the loan orders used for training (a total of 7 node types: user: account manager, qrcode: marketing manager, review: reviewer, channel: customer, slave: customer's spouse, lc: business license node, ad: address tile node), the features of the above 6 types are extracted and then prefixed with the node type (user, qrcode, review, channel, slave, lc, ad), such as:

[0123] user_outstand_index: A prominent indicator for the account manager node (user).

[0124] channel_knn: The knn metric for the client (channel).

[0125] Finally, based on the node type of the order, 6×7=42 features were extracted.

[0126] Simultaneously, the topological distance features (the number of edges traversed from a node to a suspect group) from the customer node and its spouse node are extracted, and their relationships are as follows: Figure 6 As shown, the area within the dashed line represents small groups. The values ​​in the nodes indicate the distance from the node to the small group.

[0127] Based on this calculation method, six new characteristics are formed (two customer types: customer or spouse; three suspected group types: hidden group related to fraud and delinquency, fraud group, and delinquency group; therefore, there are 2 × 3 = 6 new characteristics), such as:

[0128] Channel_dist_hidden: Distance characteristic of the customer to a hidden group of people associated with fraud and delinquency.

[0129] slave_dist_fraud is the distance feature of a spouse to a fraudulent group.

[0130] Channel_dist_overdue: Distance characteristics of the customer to the overdue group.

[0131] Finally, the features are integrated into 42 + 6 = 48 features for the order.

[0132] Step 6: Using the node feature set and topological distance features of the training loan orders as input and the four-class label of the training loan orders as output, train the XGBoost model to obtain the financial fraud behavior prediction model.

[0133] In this embodiment, based on the loan order node information used for training, the aforementioned 48 features are extracted from the complex network and the distance to small groups using step 5, and then combined with the known four-class classification labels of the loan orders used for training. These features are then input into the XGBoost model for training.

[0134] Training Process: The training and validation sets are divided in a 6:4 ratio. The hyperparameters of the XgBoost model are set, and the input data for the training set is randomly arranged. Training is then performed, and model weights are obtained. The trained XgBoost model is used to predict on the validation set, and the prediction results are evaluated. The prediction results are comprehensively evaluated using the product of the precision and repeat metrics. The model weights and corresponding comprehensive evaluation metrics are saved. Based on the comprehensive evaluation metrics, the best-performing model weights are selected, and the model weights are output. Figure 7 As shown.

[0135] After training, using detailed information about order nodes, the aforementioned 48 features are extracted from the complex network and distance information of suspected small groups. These features are then input into the trained XGBoost model for prediction, and the output prediction category is 0-3. Figure 8 As shown.

[0136] Finally, output a small group of results and present them to the business personnel for further evaluation:

[0137] Output the small group to which the order belongs to to the graph database neo4j, and present it to the financial loan review personnel using the Cypher query language.

[0138] Financial loan review personnel make further judgments based on the connections between nodes in the small group.

[0139] Establish a dynamic update mechanism:

[0140] Every so often (e.g., once a week), new loan order data is incorporated, and the network structure, network node characteristic attributes, and information on suspected small groups are incrementally updated using the methods described in steps 1 to 4 above.

[0141] Every so often (e.g., January), based on the tags of new loan orders (combining new manually verified fraud tags and new overdue tags to form a four-category tag), the model parameters are retrained using the aforementioned step 6 to form a risk warning closed loop.

[0142] The financial fraud prediction method based on complex networks proposed in this application has the following effects:

[0143] It combines a complex network model (graph model) with the XGBoost model, resulting in low computational complexity.

[0144] By extracting features from the complex network and using distance features to specific small groups, the complex network is decoupled from the XGBoost model, simplifying the model training process and accelerating model prediction speed.

[0145] To address the impact of group fraud on individual node characteristics, specific node encoding and feature extraction methods were designed to achieve better prediction results.

[0146] Complex network modeling methods that are independent of specific financial characteristics facilitate the fusion of complex networks from different regions and data sources.

[0147] The prediction method based on the XGBoost model is easy to integrate with the XGBoost prediction model based on the financial model.

[0148] The four-category labeling method for orders more closely reflects the relationship between actual fraudulent behavior and overdue payments, and can better distinguish different behavioral characteristics. (The motivations behind fraudulent behavior and overdue payments are related but different; they have distinct behavioral characteristics.)

[0149] The problematic groups and key account manager nodes identified by the model can be used for further in-depth on-site investigations.

[0150] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method for predicting financial fraud based on complex networks.

[0151] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the above-described method for predicting financial fraud based on complex networks.

[0152] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described method for predicting financial fraud based on complex networks.

[0153] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 9 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and databases. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media to run. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for predicting financial fraud based on complex networks.

[0154] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0155] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0156] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0157] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0158] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0159] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for predicting financial fraud based on complex networks, characterized in that, include: Obtain loan orders to be predicted; The loan orders to be predicted can be online loan orders or offline loan orders; Based on the loan order to be predicted, the complex network of historical loan orders is updated to obtain the updated complex network; the complex network is constructed based on the historical loan orders; the entities of the historical loan orders are the nodes of the complex network, and the relationships between the historical single orders are the edges of the complex network; the entities include account managers, marketing managers, reviewers, customers, customer spouses, business license numbers, and business locations; Obtain the node feature set of the loan order to be predicted in the updated complex network, as well as the topological distance features from the customer node and customer spouse node of the loan order to be predicted to the suspected small group; the node feature set includes customer manager node features, marketing manager node features, reviewer node features, customer node features, customer spouse node features, business license node features, and address tile node features. Node features include node saliency index, weighted maximum saliency index of surrounding nodes, node k-core index, surrounding node k-core index, node betweenness, and node knn index; the suspected groups include hidden groups related to fraud and delinquency, fraud groups, and delinquency groups; the suspected groups are determined based on the complex network using the k-core hierarchical group algorithm, the ECG-LA method, and the label propagation-based LPA algorithm; Obtaining the node feature set of the loan order to be predicted in the updated complex network specifically includes: Calculate the betweenness number of each edge in the complex network, as well as the node degree, node kNN index, and node k-core index of each node; Based on the betweenness of edges, the degree of nodes, and the k-core index of nodes, the node salience index of each node in the loan order to be predicted is calculated. Based on the betweenness of edges and the salience index of nodes, the weighted maximum salience index of the surrounding nodes of each node in the loan order to be predicted is calculated; the surrounding nodes are nodes that are connected to the current node. Extract the node knn index, node k-core index, node betweenness, and the maximum k-core index of surrounding nodes for each node in the loan order to be predicted, and obtain the node feature set of the loan order to be predicted in the updated complex network. Based on the betweenness of edges, the degree of nodes, and the k-core index of nodes, the node salience index of each node in the loan order to be predicted is calculated, specifically including: Using formula Calculate the node salience index for each node in the loan order to be predicted; wherein, For the prominence of nodes; Let be the betweenness of the edge; The node's k-core index; The degree of a node; The maximum degree value among the surrounding nodes; Based on the betweenness of edges and node salience indices, the weighted maximum salience index of the surrounding nodes of each node in the loan order to be predicted is calculated, specifically including: Using formula N s =MAX(E b ×V s Calculate the weighted maximum salience index of the surrounding nodes for each node in the loan order to be predicted; where N s E is the weighted maximum salience index of the surrounding nodes associated with the current node. b Let be the betweenness of the edge; Based on the node feature set and the topological distance feature, a financial fraud behavior prediction model is used to determine the fraud behavior category of the loan order to be predicted; the fraud behavior category is no fraud and no overdue payment, overdue payment, fraud, or fraud and overdue payment; the financial fraud behavior prediction model is obtained by training an XGBoost model.

2. The financial fraud prediction method based on complex networks according to claim 1, characterized in that, Based on the loan orders to be predicted, the complex network of historical loan orders is updated to obtain the updated complex network, specifically including: The loan orders to be predicted are anonymized and tile-encoded to obtain the processed loan orders; Based on the processed loan orders, the complex network of historical loan orders is updated to obtain the updated complex network.

3. The financial fraud behavior prediction method based on complex networks according to claim 1, characterized in that, The construction process of the financial fraud prediction model specifically includes: Obtain a number of historical loan orders; the historical loan orders include online historical single-item orders and offline historical loan orders; The historical loan orders are desensitized and tile-encoded to obtain the processed historical loan orders; A complex network is constructed by using the entities of the processed historical loan orders as nodes and the relationships between the processed historical single orders as edges. Based on the complex network, the k-core hierarchical small group algorithm, ECG-LA method, and LPA algorithm based on label propagation are used to identify suspected small groups and label them with four categories of labels: no fraud and no overdue, overdue, fraud, or fraud and overdue. Extract the node feature set of training loan orders in the complex network, as well as the topological distance features from the customer nodes and customer spouse nodes of the training loan orders to the suspected small group, and label the training loan orders with four-class labels; the training loan orders are obtained from the historical loan orders; The XGBoost model is trained using the node feature set and topological distance features of the training loan orders as input and the four-class classification labels of the training loan orders as output, to obtain the financial fraud behavior prediction model.

4. The financial fraud prediction method based on complex networks according to claim 3, characterized in that, The historical loan orders are anonymized and tile-encoded to obtain the processed historical loan orders, specifically including: The account manager, marketing manager, reviewer, customer, customer's spouse and business license number in the historical loan order are de-identified using the MD5 encryption algorithm to obtain the de-identified historical loan order. The business locations in the de-identified historical loan orders are then subjected to tile encoding to obtain the processed historical loan orders.

5. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the financial fraud prediction method based on complex networks as described in any one of claims 1-4.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the financial fraud prediction method based on complex networks as described in any one of claims 1-4.

7. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the financial fraud prediction method based on complex networks as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Online loan anti-fraud method based on network embedding technology

    CN111429249A

  • Abnormal transaction prevention and control method and device, computer equipment and readable storage medium

    CN120494971A

  • Anti-fraud method and system based on complex relation network

    CN120579974A