Credit fraud data identification method and system based on artificial intelligence and graph algorithm
Through the credit fraud data identification method based on artificial intelligence and graph algorithms, the problem that the existing technology is difficult to identify a few types of fraud samples is solved, and the effect of accurately identifying complex fraud patterns and strong dynamic adaptability is achieved, reducing the missed detection rate and fraud losses.
Patent Information
- Application Number
- CN202510498103.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult for the existing technology to adjust category weights and focus parameters, enhance the model's attention to a few types of fraud samples, and easily ignore fraud samples, affecting the recognition effect.
The credit fraud data identification method based on artificial intelligence and graph algorithms is adopted, including collecting user information and behavioral data from multiple sources, building relationship network graphs, identifying dense communities, quantifying behavior patterns, combining deep learning to cluster similar groups, using focus loss function to optimize the fraud identification model, and classifying user risk levels based on scoring thresholds, triggering risk control measures.
Accurately identify complex fraud patterns, explore hidden relationships through graph algorithms, combine community portraits and behavior clustering to improve identification accuracy, strong dynamic adaptability, real-time feedback and incremental learning ensure rapid response to new fraud, reduce missed detection rates, reduce fraud losses, and achieve high return on investment.
Smart Images

Figure CN120013664A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of credit risk control technology, and in particular to a credit fraud data identification method and system based on artificial intelligence and graph algorithms. Background Art
[0002] Existing anti-fraud technologies mainly rely on the feature scoring and rule judgment of a single user, and lack the ability to analyze the implicit associations between multiple users and cross-accounts. This makes it difficult to detect intermediary operations or group fraud.
[0003] At present, the Chinese invention patent with application number CN202410512842.9 discloses a credit fraud data identification method and system using machine learning. By determining the target fraud behavior identification model involved in fraud behavior identification in the target risk control service system, and using the behavior-dependent knowledge network to accurately extract the business operation path data and the corresponding business operation impact weights involved in the credit page behavior data, by comprehensively considering the prior importance weight of each fraud behavior identification model, the preset influence coefficient of the credit page behavior data, and the business operation impact weight of the business operation path data, the fraud risk prediction parameters of each business operation path data can be accurately calculated. Finally, based on the fraud risk objects in the risk control results and the fraud risk prediction parameters of each business operation path data, high-risk credit operation behaviors can be effectively identified, thereby reducing the risk of credit fraud and improving the robustness and security of credit business.
[0004] The above technology makes it difficult to adjust category weights and focus parameters, which enhances the model's attention to minority fraud samples and makes it easy to ignore fraud samples, affecting the recognition effect. Summary of the invention
[0005] The technical problem solved by the present invention is that it is difficult to adjust the category weights and focus parameters in the prior art to enhance the model's attention to minority fraud samples, and it is easy to ignore fraud samples, affecting the recognition effect.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: The credit fraud data identification method based on artificial intelligence and graph algorithm includes the following steps: Step S1, collecting user information and behavior data from multiple sources, and constructing a relationship network diagram after cleaning; Step S2, using modularity optimization algorithm to identify dense communities, locating cross-community intermediary nodes by node importance ranking, and generating community portraits; Step S3, using a node vector embedding algorithm to quantify the behavior pattern, combining deep learning to cluster similar groups, and using a focal loss function to optimize the fraud identification model; Step S4, classifying the user risk level based on the scoring threshold, triggering risk control measures, and generating an explainable report; Step S5, adjusting the model and rules according to real-time data, and monitoring the false alarm rate and missed alarm rate.
[0007] Preferably, the step S1 includes the following sub-steps: Step S101, collecting user information and corresponding behavior data from databases, transaction platforms and social interfaces, wherein the user information includes user attributes, user devices, contacts and contact information, and the behavior data includes transaction relationships, transaction amounts and transaction timing; Step S102, cleaning the user information and behavior data, wherein the cleaning process includes missing value interpolation, outlier truncation, and structured unstructured data; Step S103, defining a node association rule, wherein the node association rule is: Filtering the user information, selecting contacts and user devices shared by the user information and establishing bidirectional edges, wherein the weighted frequencies of the bidirectional edges are shared times; Select the same transaction relationship and establish a weighted edge, wherein the weight frequency of the weighted edge is set based on the corresponding transaction amount and a preset correspondence between the transaction amount and the weight frequency; Construct a relationship network diagram and store it in a graph database.
[0008] Preferably, step S2 includes the following sub-steps: Step S201, using a modularity optimization algorithm to iteratively divide the relationship network graph and calculate the modularity; Step S202, calculating the node centrality score through the node importance ranking algorithm, marking the intermediary accounts across communities and with centrality scores higher than the threshold; Step S203, extracting community features, including the number of accounts, device distribution density, and transaction time concentration, and outputting them as a community portrait.
[0009] Preferably, step S3 includes the following sub-steps: Step S301, using a node vector embedding algorithm, generating a node sequence by random walk, inputting it into a skip-word model training, and obtaining a d-dimensional embedding vector of the node; Step S302, using a deep clustering model to group the d-dimensional embedding vectors, calculating the silhouette coefficient to evaluate the clustering quality, and outputting similar behavior groups; Step S303: input the historical fraud labels and behavior features into the extreme gradient boosting algorithm model, and use the focal loss function for optimization training.
[0010] Preferably, step S4 includes the following sub-steps: Step S401, calculate the risk score for the user node and community based on the following indicators: Individual behavior deviation: Euclidean distance from the cluster center; Community risk index: the percentage of fraudulent accounts in the community; Intermediary association strength: edge weight with intermediary nodes; Step S402: If the risk score is greater than 0.8, it is marked as high risk, triggering manual review and outputting a signal to freeze the account; If 0.5≤risk score≤0.8, it is marked as medium risk and a supplementary information signal is output; If the risk score is <0.5, it is marked as low risk and automatically passed; Step S403, using Shapley additive interpretation technology, calculate the feature contribution and output an interpretable report, wherein the interpretable report includes a topological path in the association graph database and annotates a multi-hop association chain of high-risk users.
[0011] Preferably, step S5 includes the following sub-steps: Step S501, input the newly added fraud samples into the graph database, update the node embedding vector and community division, and use a sliding time window to eliminate expired data, the sliding time window is 30 days; Step S502, monitor the false alarm rate and the missed alarm rate. If the false alarm rate and the missed alarm rate exceed the preset false alarm rate threshold and the missed alarm rate threshold for three consecutive days, adjust the decision threshold through reinforcement learning.
[0012] Preferably, the following data transfer association is also included: The graph data outputted from step S1 is passed to step S2 via the API; The community portrait outputted in step S2 is transmitted to step S3 via a message queue; The model prediction result of step S3 and the decision result of step S4 are written into the risk control database together for calling by step S5.
[0013] Preferably, the deployment includes: The initial model was trained based on historical data from the past year, and the recall rate was verified to be improved by ≥15% through A / B testing; The false alarm rate and missed alarm rate are counted in real time. If the missed alarm rate is >5% within 24 hours, the model retraining process is triggered.
[0014] Preferably, the mathematical expression of the modularity in step S201 is: ; in, is the modularity, is the total edge weight, is the edge weight, and is any node in the relationship network graph, For Node The node degree, For Node The node degree, is the community indicator function; The mathematical expression of the node centrality score in step S202 is: ; in, and is any node in the relationship network graph, For Node The node centrality score of For Node The node centrality score of is the damping coefficient, is the total number of all nodes in the relationship network graph, is a node set pointing to a node, For Node The number of outgoing links; The mathematical expression of the focus loss function in step S303 is: ; in, is the focal loss function, is the predicted probability of the focal loss function for the true category, is the category weight parameter, is the focus parameter; The mathematical expression of the feature contribution in step S403 is: ; in, is the set of all user characteristics, is a feature subset of the set, is the feature contribution, is the focal loss function on the feature subset The predicted value under ; The mathematical expression for adjusting the decision threshold by reinforcement learning in step S502 is: ; in, is the current risk score threshold, is the adjusted risk score threshold, is the learning rate, is the proportion of correctly identified fraud samples to all actual fraud samples, The proportion of correctly identified fraudulent samples to all samples marked as fraudulent.
[0015] A credit fraud data identification system based on artificial intelligence and graph algorithms, which is applied to the credit fraud data identification method based on artificial intelligence and graph algorithms, and includes a data acquisition module, a portrait generation module, a fraud identification module, a risk quantification module, and an adjustment monitoring module; The data collection module is used to collect user information and behavior data from multiple sources, and to construct a relationship network diagram after cleaning; The portrait generation module is used to identify dense communities using a modularity optimization algorithm, locate cross-community intermediary nodes by ranking node importance, and generate community portraits; The fraud identification module is used to quantify the behavior pattern using a node vector embedding algorithm, cluster similar groups in combination with deep learning, and optimize the fraud identification model using a focal loss function; The risk quantification module is used to classify user risk levels based on scoring thresholds, trigger risk control measures, and generate explainable reports; The adjustment monitoring module is used to adjust the model and rules according to real-time data and monitor the false alarm rate and the missed alarm rate.
[0016] The beneficial effects of the present invention are as follows: the present invention can accurately identify complex fraud patterns, mine hidden associations through graph algorithms, improve recognition accuracy by combining community portraits and behavioral clustering, has strong dynamic adaptability, and ensures rapid response to new types of fraud through real-time feedback and incremental learning. Highly explainable reports meet compliance requirements and enhance user trust. It can efficiently process large-scale data, has a modular design that adapts to different credit scenarios, reduces costs through automated decision-making, and reduces the missed detection rate, directly reducing fraud losses and achieving a high return on investment. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A flowchart of the steps of a credit fraud data identification method based on artificial intelligence and graph algorithms provided by one embodiment of the present invention; Figure 2 A schematic diagram of the basic flow of a credit fraud data identification system based on artificial intelligence and graph algorithms provided for one embodiment of the present invention. DETAILED DESCRIPTION
[0018] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.
[0019] Example 1, reference Figure 1 , provides a credit fraud data identification method based on artificial intelligence and graph algorithm, including the following steps: Step S1, collect user information and behavior data from multiple sources, and build a relationship network diagram after cleaning.
[0020] Step S2: Use modularity optimization algorithm to identify dense communities, locate cross-community intermediary nodes by ranking the importance of nodes, and generate community portraits.
[0021] In step S3, the node vector embedding algorithm is used to quantify the behavior pattern, and deep learning is combined to cluster similar groups, and the focal loss function is used to optimize the fraud identification model.
[0022] Step S4: classify the user risk level based on the scoring threshold, trigger risk control measures, and generate an explainable report.
[0023] Step S5, adjusting the model and rules according to real-time data, and monitoring the false alarm rate and missed alarm rate.
[0024] Step S1 includes the following sub-steps: Step S101, collect user information and corresponding behavior data from the database, transaction platform and social interface, the user information includes user attributes, user equipment, contacts and contact information, and the behavior data includes transaction relationship, transaction amount and transaction sequence.
[0025] Step S101 comprehensively collects user information and behavior data to ensure the diversity and integrity of the data and provide a rich data source for subsequent analysis.
[0026] Step S102, cleaning the user information and behavior data, the cleaning process includes missing value interpolation, outlier truncation and structured unstructured data.
[0027] Step S102 effectively cleans the user information and behavior data to improve data quality and reduce analysis errors caused by data problems.
[0028] Step S103, defining node association rules, the node association rules are: The user information is screened, contacts and user devices that are shared by the user information are selected, and bidirectional edges are established, wherein the weight frequency of the bidirectional edges is the number of times shared.
[0029] The same transaction relationship is selected and a weighted edge is established. The weight frequency of the weighted edge is set based on the corresponding transaction amount and the preset corresponding relationship between the transaction amount and the weight frequency.
[0030] Construct a relationship network diagram and store it in a graph database.
[0031] Step S103 defines node association rules based on user information and behavior data, constructs a relationship network diagram that reflects the relationship between users, and provides an intuitive graphical tool for subsequent identification of fraudulent behavior. At the same time, the relationship network diagram is stored in a graph database to facilitate subsequent data query and analysis.
[0032] Step S1 collects and cleans user information and behavior data from multiple sources, builds a relationship network diagram by defining node association rules, and provides basic data support for subsequent credit fraud identification.
[0033] Step S2 includes the following sub-steps: Step S201, using a modularity optimization algorithm to iteratively partition the relationship network graph and calculate the modularity.
[0034] The mathematical expression of the modularity in step S201 is: ; in, is the modularity, is the total edge weight, is the edge weight, and is any node in the relationship network graph, For Node The node degree, For Node The node degree, is the community indicator function.
[0035] Step S201 uses a modularity optimization algorithm to iteratively divide the relationship network graph, which can more accurately identify the community structure in the network and provide clear community boundaries for subsequent analysis.
[0036] Step S202, calculate the node centrality score through the node importance ranking algorithm, and mark the intermediary accounts that cross the community and whose centrality score is higher than the threshold.
[0037] The mathematical expression of the node centrality score in step S202 is: ; in, and is any node in the relationship network graph, For Node The node centrality score of For Node The node centrality score of is the damping coefficient, is the total number of all nodes in the relationship network graph, is a node set pointing to a node, For Node The number of outgoing links.
[0038] Step S202 calculates the node centrality score through the node importance ranking algorithm and marks the intermediary accounts across communities with centrality scores above the threshold. These intermediary accounts play an important role between communities and are key clues for identifying potential fraudulent behavior.
[0039] Step S203, extracting community features, including the number of accounts, device distribution density, and transaction time concentration, and outputting them as a community portrait.
[0040] Step S203 extracts community features and outputs them as community portraits, which can intuitively display the basic attributes and behavioral characteristics of the community, providing a deeper analysis perspective for credit fraud identification. At the same time, community portraits also provide strong support for subsequent risk assessment and risk control measures.
[0041] Step S2 divides the relationship network graph through the modularity optimization algorithm, identifies the community structure, and locates the key intermediary accounts by sorting the node importance. Finally, the community features are extracted and output as a community portrait, providing in-depth community analysis for credit fraud identification.
[0042] Step S3 includes the following sub-steps: Step S301, using a node vector embedding algorithm, generating a node sequence through random walk, inputting it into a skip-gram model training, and obtaining a d-dimensional embedding vector of the node.
[0043] Step S301 uses a node vector embedding algorithm to convert the nodes in the relationship network diagram into d-dimensional embedding vectors. These vectors can capture the local and global structural information of the nodes and provide effective feature representation for subsequent clustering analysis.
[0044] Step S302: Use the deep clustering model to group the d-dimensional embedding vectors, calculate the silhouette coefficient to evaluate the clustering quality, and output similar behavior groups.
[0045] Step S302 uses a deep clustering model to group the d-dimensional embedding vectors, and evaluates the clustering quality by calculating the silhouette coefficient to ensure that the obtained similar behavior groups have high internal consistency and low external similarity, providing accurate group division for fraud identification.
[0046] Step S303: input the historical fraud labels and behavior features into the extreme gradient boosting algorithm model, and use the focal loss function for optimization training.
[0047] The mathematical expression of the focus loss function in step S303 is: ; in, is the focal loss function, is the predicted probability of the focal loss function for the true category, is the category weight parameter, is the focus parameter.
[0048] Step S303 inputs the historical fraud labels and behavior features into the extreme gradient boosting algorithm model, and optimizes the model using the focus loss function, so that the model can more accurately identify potential fraudulent behaviors and improve the accuracy and efficiency of credit fraud identification. At the same time, this step also provides strong support for subsequent risk assessment and risk control measures.
[0049] Step S3 uses a node vector embedding algorithm and a deep clustering model to conduct an in-depth analysis of user behavior, identify similar behavior groups, and use the extreme gradient boosting algorithm model combined with historical fraud labels to perform fraud recognition training to improve the accuracy and efficiency of credit fraud recognition.
[0050] Step S4 includes the following sub-steps: Step S401, calculate the risk score for the user node and community based on the following indicators: Individual behavior deviation: Euclidean distance from the cluster center.
[0051] Community risk index: the percentage of fraudulent accounts in the community.
[0052] Intermediary association strength: the edge weight with the intermediary node.
[0053] Step S401 conducts a comprehensive risk assessment of user nodes and communities by integrating multiple indicators such as individual behavior deviation, community risk index, and intermediary association strength. These indicators can reflect the degree of abnormality of user behavior, the prevalence of fraudulent behavior in the community, and the strength of association with potential fraud intermediaries, providing a strong basis for subsequent risk classification and risk control measures.
[0054] Step S402: If the risk score is greater than 0.8, it is marked as high risk, triggering manual review and outputting a signal to freeze the account.
[0055] If 0.5≤risk score≤0.8, it is marked as medium risk and a supplementary information signal is output.
[0056] If the risk score is <0.5, it is marked as low risk and automatically passed.
[0057] Step S402 divides user nodes and communities into three categories: high risk, medium risk, and low risk, based on the calculated risk score, and takes corresponding risk control measures. High-risk users trigger manual review and output a freeze account signal to ensure that potential fraud is handled in a timely manner; medium-risk users output a supplementary information signal, requiring the user to provide more information to further verify their credit status; low-risk users are automatically approved to improve processing efficiency. This step achieves differentiated processing of users of different risk levels, improving the pertinence and effectiveness of risk control measures.
[0058] Step S403, using Shapley additive interpretation technology, calculate the feature contribution and output an interpretable report, wherein the interpretable report includes a topological path in the association graph database and annotates a multi-hop association chain of high-risk users.
[0059] The mathematical expression of the feature contribution in step S403 is: ; in, is the set of all user characteristics, is a feature subset of the set, is the feature contribution, is the focal loss function on the feature subset The predicted value below.
[0060] Step S403 calculates the feature contribution through Shapley additive interpretation technology and outputs an interpretable report. The report not only shows the user's risk score and risk control measures, but also associates the topological path in the graph database and marks the multi-hop association chain of high-risk users. This enables risk control personnel to more intuitively understand the user's risk source and potential fraud behavior patterns, and improves the transparency and accuracy of credit fraud identification. At the same time, the report also provides strong support for subsequent risk assessment and risk control strategy adjustment.
[0061] Step S4 conducts risk assessment on user nodes and communities based on multiple risk indicators, takes different risk control measures according to the risk points, and outputs an explainable report through Shapley additive interpretation technology to improve the transparency and accuracy of credit fraud identification.
[0062] Step S5 includes the following sub-steps: Step S501: input the newly added fraud samples into the graph database, update the node embedding vector and community division, and use a sliding time window to eliminate expired data, where the sliding time window is 30 days.
[0063] Step S501 inputs the newly added fraud samples into the graph database, which can reflect the latest fraud behavior patterns in real time and maintain the accuracy and timeliness of node embedding vectors and community divisions. The sliding time window is used to eliminate expired data, which can effectively manage the data scale, avoid the impact of redundant data on performance, and ensure that the analyzed data is the latest valid data within 30 days.
[0064] Step S502, monitor the false alarm rate and the missed alarm rate. If the false alarm rate and the missed alarm rate exceed the preset false alarm rate threshold and the missed alarm rate threshold for three consecutive days, adjust the decision threshold through reinforcement learning.
[0065] The mathematical expression for adjusting the decision threshold by reinforcement learning in step S502 is: ; in, is the current risk score threshold, is the adjusted risk score threshold, is the learning rate, is the proportion of correctly identified fraud samples to all actual fraud samples, The proportion of correctly identified fraudulent samples to all samples marked as fraudulent.
[0066] Effect of step S502 By monitoring the false alarm rate and missed alarm rate, problems in the fraud identification system can be discovered in a timely manner. If the false alarm rate and missed alarm rate exceed the preset threshold for three consecutive days, it indicates that the current decision threshold may no longer be applicable. At this time, by adjusting the decision threshold through reinforcement learning, the performance can be automatically optimized, false alarms and missed alarms can be reduced, and the accuracy and stability of credit fraud identification can be improved. This step achieves self-optimization and continuous improvement.
[0067] Step S5 continuously updates the fraud samples in the graph database to maintain the timeliness of the node embedding vector and community division, while monitoring and adjusting the decision threshold to optimize the fraud identification performance and ensure the accuracy and stability of the credit fraud identification system.
[0068] This method includes the following data transfer associations: The graph data output from step S1 is passed to step S2 via the API.
[0069] The community portrait outputted in step S2 is transmitted to step S3 via a message queue.
[0070] The model prediction result of step S3 and the decision result of step S4 are written into the risk control database together for calling by step S5.
[0071] This method includes: The initial model was trained based on historical data from the past year, and the recall rate was verified to be improved by ≥15% through A / B testing.
[0072] The false alarm rate and missed alarm rate are counted in real time. If the missed alarm rate is >5% within 24 hours, the model retraining process is triggered.
[0073] The present invention uses graph algorithms to deeply mine hidden associations, improving the accuracy and depth of fraud identification. Combining community portraits with behavioral clustering, it can more comprehensively understand user behavior patterns, thereby accurately identifying abnormal behaviors. Real-time data feedback and incremental learning mechanisms ensure that the system can respond quickly to new fraud methods, usually completing model optimization within 24 hours to maintain timeliness and accuracy. The rule base is dynamically adjusted to avoid the lag problem of manual intervention and can be flexibly adjusted according to actual conditions. The interpretable report clearly marks the abnormal characteristics of high-risk users, such as sudden increases in transaction frequency and intermediary association strength, meeting the needs of financial regulators. Transparency requirements enhance compliance. The topological path in the report marks the multi-hop association chain of high-risk users and provides detailed explanations, which helps to improve user trust and efficiently process large-scale data. The combination of graph database and distributed computing framework supports efficient storage and query of hundreds of millions of nodes, improves processing capabilities and response speed, and modular design allows horizontal expansion to adapt to credit scenarios of different scales, such as consumer finance, small and micro enterprise loans, etc., enhancing the flexibility and applicability of the system. Automated decision-making reduces the workload of manual review and reduces operating costs. The reduction in missed detection rate directly reduces fraud losses and improves the return on investment.
[0074] Example 2, reference Figure 2 , provides a credit fraud data identification system based on artificial intelligence and graph algorithms, including data collection module, portrait generation module, fraud identification module, risk quantification module and adjustment monitoring module.
[0075] The data collection module is used to collect user information and behavior data from multiple sources, and build a relationship network diagram after cleaning.
[0076] The portrait generation module is used to identify dense communities using a modularity optimization algorithm, locate cross-community intermediary nodes by ranking the importance of nodes, and generate community portraits.
[0077] The fraud identification module is used to quantify behavior patterns using a node vector embedding algorithm, cluster similar groups using deep learning, and optimize the fraud identification model using a focal loss function.
[0078] The risk quantification module is used to classify user risk levels based on scoring thresholds, trigger risk control measures, and generate explainable reports.
[0079] The adjustment monitoring module is used to adjust the model and rules according to real-time data and monitor the false alarm rate and missed alarm rate.
[0080] This system collects and cleans user information and behavior data from multiple sources to build an accurate relationship network diagram. This system can use modularity optimization algorithms to identify dense communities, and accurately locate intermediary nodes across communities by sorting node importance, thereby generating a detailed community portrait. The node vector embedding algorithm is used to quantify behavior patterns, combined with deep learning to cluster similar groups, and the focal loss function is used to optimize the fraud identification model, which significantly improves the accuracy and efficiency of fraud identification. At the same time, this system can classify users by risk level according to the scoring threshold, trigger corresponding risk control measures, and generate interpretable reports with topological path associations to enhance the transparency of decision-making. In addition, this system can flexibly adjust models and rules based on real-time data, effectively monitor false alarm rates and missed alarm rates, ensure the stability and continuous optimization capabilities of this system, and provide strong security guarantees for credit business.
[0081] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. Among them, the storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Read Only Memory, referred to as PROM), read-only memory (Read Only Memory, referred to as ROM), magnetic memory, flash memory, magnetic disk or optical disk. These computer program instructions may also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which is implemented in the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0082] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A credit fraud data identification method based on artificial intelligence and graph algorithms, characterized in that: The steps include: Step S1, collecting user information and behavior data from multiple sources, and constructing a relationship network diagram after cleaning; Step S2, using modularity optimization algorithm to identify dense communities, locating cross-community intermediary nodes by node importance ranking, and generating community portraits; Step S3, using a node vector embedding algorithm to quantify the behavior pattern, combining deep learning to cluster similar groups, and using a focal loss function to optimize the fraud identification model; Step S4, classifying the user risk level based on the scoring threshold, triggering risk control measures, and generating an explainable report; Step S5, adjusting the model and rules according to real-time data, and monitoring the false alarm rate and missed alarm rate.
2. The credit fraud data identification method based on artificial intelligence and graph algorithm as claimed in claim 1, characterized in that: The step S1 includes the following sub-steps: Step S101, collecting user information and corresponding behavior data from databases, transaction platforms and social interfaces, wherein the user information includes user attributes, user devices, contacts and contact information, and the behavior data includes transaction relationships, transaction amounts and transaction timing; Step S102, cleaning the user information and behavior data, wherein the cleaning process includes missing value interpolation, outlier truncation, and structured unstructured data; Step S103, defining a node association rule, wherein the node association rule is: Filtering the user information, selecting contacts and user devices shared by the user information and establishing bidirectional edges, wherein the weighted frequencies of the bidirectional edges are shared times; Select the same transaction relationship and establish a weighted edge, wherein the weight frequency of the weighted edge is set based on the corresponding transaction amount and a preset correspondence between the transaction amount and the weight frequency; Construct a relationship network diagram and store it in a graph database.
3. The credit fraud data identification method based on artificial intelligence and graph algorithm as claimed in claim 2, characterized in that: The step S2 includes the following sub-steps: Step S201, using a modularity optimization algorithm to iteratively divide the relationship network graph and calculate the modularity; Step S202, calculating the node centrality score through the node importance ranking algorithm, marking the intermediary accounts across communities and with centrality scores higher than the threshold; Step S203, extracting community features, including the number of accounts, device distribution density, and transaction time concentration, and outputting them as a community portrait.
4. The credit fraud data identification method based on artificial intelligence and graph algorithm as claimed in claim 3, characterized in that: The step S3 includes the following sub-steps: Step S301, using a node vector embedding algorithm, generating a node sequence by random walk, inputting it into a skip-word model training, and obtaining a d-dimensional embedding vector of the node; Step S302, using a deep clustering model to group the d-dimensional embedding vectors, calculating the silhouette coefficient to evaluate the clustering quality, and outputting similar behavior groups; Step S303: input the historical fraud labels and behavior features into the extreme gradient boosting algorithm model, and use the focal loss function for optimization training.
5. The credit fraud data identification method based on artificial intelligence and graph algorithm as claimed in claim 4, characterized in that: The step S4 includes the following sub-steps: Step S401, calculate the risk score for the user node and community based on the following indicators: Individual behavior deviation: Euclidean distance from the cluster center; Community risk index: the percentage of fraudulent accounts in the community; Intermediary association strength: edge weight with intermediary nodes; Step S402: If the risk score is greater than 0.8, it is marked as high risk, triggering manual review and outputting a signal to freeze the account; If 0.5≤risk score≤0.8, it is marked as medium risk and a supplementary information signal is output; If the risk score is <0.5, it is marked as low risk and automatically passed; Step S403, using Shapley additive interpretation technology, calculate the feature contribution and output an interpretable report, wherein the interpretable report includes a topological path in the association graph database and annotates a multi-hop association chain of high-risk users.
6. The credit fraud data identification method based on artificial intelligence and graph algorithm as claimed in claim 5, characterized in that: The step S5 includes the following sub-steps: Step S501, input the newly added fraud samples into the graph database, update the node embedding vector and community division, and use a sliding time window to eliminate expired data, the sliding time window is 30 days; Step S502, monitor the false alarm rate and the missed alarm rate. If the false alarm rate and the missed alarm rate exceed the preset false alarm rate threshold and the missed alarm rate threshold for three consecutive days, adjust the decision threshold through reinforcement learning.
7. The credit fraud data identification method based on artificial intelligence and graph algorithm as claimed in claim 6, characterized in that: The following data transfer associations are also included: The graph data outputted from step S1 is passed to step S2 via the API; The community portrait outputted in step S2 is transmitted to step S3 via a message queue; The model prediction result of step S3 and the decision result of step S4 are written into the risk control database together for calling by step S5.
8. The credit fraud data identification method based on artificial intelligence and graph algorithm as claimed in claim 7, characterized in that: Included during deployment: The initial model was trained based on historical data from the past year, and the recall rate was verified to be improved by ≥15% through A / B testing; The false alarm rate and missed alarm rate are counted in real time. If the missed alarm rate is >5% within 24 hours, the model retraining process is triggered.
9. The credit fraud data identification method based on artificial intelligence and graph algorithm as claimed in claim 8, characterized in that: The mathematical expression of the modularity in step S201 is: ; in, is the modularity, is the total edge weight, is the edge weight, and is any node in the relationship network graph, For Node The node degree, For Node The node degree, is the community indicator function; The mathematical expression of the node centrality score in step S202 is: ; in, and is any node in the relationship network graph, For Node The node centrality score of For Node The node centrality score of is the damping coefficient, is the total number of all nodes in the relationship network graph, is a node set pointing to a node, For Node The number of outgoing links; The mathematical expression of the focus loss function in step S303 is: ; in, is the focal loss function, is the predicted probability of the focal loss function for the true category, is the category weight parameter, is the focus parameter; The mathematical expression of the feature contribution in step S403 is: ; in, is the set of all user characteristics, is a feature subset of the set, is the feature contribution, is the focal loss function on the feature subset The predicted value under ; The mathematical expression for adjusting the decision threshold by reinforcement learning in step S502 is: ; in, is the current risk score threshold, is the adjusted risk score threshold, is the learning rate, is the proportion of correctly identified fraud samples to all actual fraud samples, The proportion of correctly identified fraudulent samples to all samples marked as fraudulent.
10. A credit fraud data identification system based on artificial intelligence and graph algorithms, which is applied to a credit fraud data identification method based on artificial intelligence and graph algorithms as claimed in any one of claims 1 to 9, characterized in that: It includes data collection module, portrait generation module, fraud identification module, risk quantification module and adjustment monitoring module; The data collection module is used to collect user information and behavior data from multiple sources, and to construct a relationship network diagram after cleaning; The portrait generation module is used to identify dense communities using a modularity optimization algorithm, locate cross-community intermediary nodes by ranking node importance, and generate community portraits; The fraud identification module is used to quantify the behavior pattern using a node vector embedding algorithm, cluster similar groups in combination with deep learning, and optimize the fraud identification model using a focal loss function; The risk quantification module is used to classify user risk levels based on scoring thresholds, trigger risk control measures, and generate explainable reports; The adjustment monitoring module is used to adjust the model and rules according to real-time data and monitor the false alarm rate and the missed alarm rate.
Citation Information
Patent Citations
Credit fraud data identification method and system applying machine learning
CN118333750A
Method and system for mining and checking fraud gang relationship in Internet
CN110413707A
Entity relationship graph display method and system
CN111309824A
Information interception method and device, computer equipment and storage medium
CN111835622A
Unbalanced network flow classification method and device based on cost-sensitive and gradient boosting algorithm
CN112272147A
Cited By
Renewable resource sorting method based on adaptive visual learning and user portrait linkage
CN120612560A
Multi-modal social network public opinion hidden danger checking method and system
CN120765004A
A multi-modal social network public opinion hidden danger checking method and system
CN120765004B
Real-time intention recognition method and system based on streaming incremental reasoning
CN121011177A
Network intelligence clue mining and key target identification method
CN121614738A