Unsupervised intelligent fraud detection system integrating multi-level graph embedding and approximate calculation
By introducing multi-level graph embedding and approximate computing technology into the fraud detection system, the problem of traditional methods relying on labeled data and high computational complexity is solved, and unsupervised efficient fraud detection and real-time monitoring are achieved.
Patent Information
- Application Number
- CN202411804086.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional fraud detection methods based on supervised learning rely on a large amount of labeled data, resulting in high time cost and resource consumption, making it difficult to adapt to changing fraud patterns, and at the same time, the computational complexity makes it difficult to process large-scale data in real time.
An unsupervised fraud detection system based on multi-level graph embedding and approximate computing is adopted to mine potential fraud behaviors from complex graph structures through graph embedding technology, and ambient computing technology is used to reduce the computational complexity and realize real-time detection.
Without the need for a large amount of labeled data, the system can efficiently process large-scale data, monitor user behavior in real time, identify potential fraudulent behaviors, and provide risk scores to significantly reduce the risk of fraud.
Smart Images

Figure CN119939041A_ABST
Abstract
Description
1. Technical Field
[0001] This invention relates to the fields of big data analysis, graph computing, and network security, particularly anti-fraud detection in finance, e-commerce, and social networks. By constructing and analyzing large-scale networks of accounts and behaviors, the invention achieves efficient identification and prevention of potential fraudulent behavior. Specifically, the invention proposes an unsupervised fraud detection system based on multi-level graph embedding and approximate computing. This system can efficiently process large amounts of data without requiring extensive labeled data, and uses graph embedding techniques to identify potential fraudulent behavior within complex graph structures.
[0002] In recent years, with the rapid increase in the number of internet users, the complexity of online transactions and social activities has also continued to rise. At the same time, online fraud has become more complex and covert, posing a significant threat to the economic interests of businesses and individual users. Therefore, establishing an effective and rapid fraud detection method has become particularly important. However, traditional fraud detection methods based on supervised learning often rely on large amounts of labeled data, which not only increases time costs and resource consumption but also makes it difficult to adapt to evolving fraud patterns. To address these issues, this paper utilizes unsupervised learning and graph embedding techniques to construct a flexible and highly scalable detection system.
[0003] This invention combines the latest technologies from multiple fields, including graph embedding, deep learning, and approximate computing, to effectively address the diversity and hidden nature of fraudulent behavior by exploring the characteristics and correlations of user behavior. This method is not only suitable for large-scale data processing but also maintains high-precision fraud detection capabilities with low computational resource consumption. Furthermore, the proposed system exhibits excellent scalability, allowing the model structure and parameters to be flexibly adjusted according to different application scenarios and requirements, thereby improving detection effectiveness and application value. 2. Background Technology
[0004] With the rapid development of the internet, online transactions and services have exploded in popularity. Online fraud methods are becoming increasingly diverse, complex, and covert, making traditional rule-based detection systems increasingly difficult to detect. To protect users and businesses from fraudulent activity, many companies and organizations have developed machine learning-based fraud detection systems. However, most of these systems rely on supervised learning and require large amounts of labeled data as training sets. However, obtaining labeled data is expensive and time-consuming. Especially given the ever-changing nature of fraudulent activity, supervised learning methods often lack sufficient flexibility to adapt to emerging fraud patterns.
[0005] To address this issue, unsupervised fraud detection methods have gained increasing attention in recent years. Unsupervised learning, which does not require labeled data, can identify potentially anomalous behavior by analyzing patterns and structures in unlabeled data. However, unsupervised learning methods typically face two challenges: first, how to extract meaningful features from massive amounts of data that accurately reflect user behavior; and second, how to efficiently compute across vast user networks to detect anomalous behavior in real time. These two issues directly impact the performance of unsupervised fraud detection systems.
[0006] While traditional graph computing methods can effectively represent relationships between users, their high computational complexity, particularly in large-scale data environments, often prevents them from meeting the requirements of real-time detection. Furthermore, the data in graph computing is often sparse, with certain user behavior patterns occurring only in specific circumstances. Effectively capturing these sparse patterns presents a significant challenge. Therefore, combining multiple technologies to build efficient unsupervised fraud detection systems has become a hot topic in current research.
[0007] This invention addresses the computational bottleneck in large-scale graph data processing by introducing multi-level graph embedding technology and approximate computing. By embedding the graph, complex graph structures can be converted into low-dimensional vector representations, significantly reducing computational complexity. Furthermore, approximate computing techniques are used to accelerate similarity calculations and community detection, ensuring the system's real-time performance and efficiency. Compared to traditional methods, the proposed system is more capable of capturing sparse user behavior patterns, particularly those concealed fraudulent behaviors that are difficult to detect using traditional methods. 3. Summary of the Invention
[0008] The purpose of this invention is to provide an unsupervised fraud detection system based on multi-level graph embedding and approximate computing. This system aims to address the existing problems of high computational complexity, lack of labeled data, low detection accuracy, and difficulty processing large amounts of data. This system can operate efficiently in large-scale data environments, monitor user behavior in real time, detect potential fraudulent activity, and issue early warnings, helping businesses and users reduce risk. Specifically, the system is implemented through the following key steps:
[0009] Data Processing and Feature Engineering: The system first acquires a large amount of user behavior data from a distributed database, such as account registration information, login records, geographic location, and device information. This data is then preprocessed, including missing value handling, outlier detection, feature construction, and dynamic feature extraction. To capture user behavior patterns, the system employs a sliding window strategy and a Pattern of Missing Values (POMV) strategy to generate more accurate behavioral features. Furthermore, through feature selection and engineering, the system extracts features that can effectively identify abnormal behavior, ensuring that the model can accurately determine user behavior patterns in subsequent analysis.
[0010] Constructing a heterogeneous graph of account-feature associations: The system constructs user behavior data into a heterogeneous graph of account-feature associations, where account nodes represent users and feature nodes represent user behavior characteristics. Account and feature nodes are connected by edges, representing the association between accounts and behaviors. This approach allows the system to clearly represent complex relationships between user behaviors and provides a foundation for subsequent graph computations.
[0011] Converting to a homogeneous graph and calculating similarity: Based on the heterogeneous graph, the system calculates similarity among user behavioral characteristics, generating a user similarity matrix. By converting the heterogeneous graph to a homogeneous graph, the system can connect accounts with similar behavioral patterns to form a user network. In the homogeneous graph, edge weights represent behavioral similarities between users. The system uses a similarity threshold to filter out highly correlated user pairs for further analysis.
[0012] Approximate Computing and Graph Embedding: To improve the efficiency of large-scale data processing, the system incorporates approximate computing and graph embedding technologies. Approximate computing reduces the dimensionality of raw high-dimensional data through methods such as random sampling and matrix decomposition, significantly reducing the computational effort. Graph embedding, meanwhile, converts complex graph structures into low-dimensional vector representations, significantly improving computational efficiency. The embedded vectors not only preserve the structural characteristics of the graph but also facilitate subsequent clustering and classification.
[0013] Unsupervised community discovery and fraud detection: The system uses the Louvain algorithm to detect communities, clustering similar accounts into the same community. The system then uses Graph2Vec technology to embed the community structure into a low-dimensional vector representation and classify the accounts using clustering algorithms such as KMeans. By analyzing the clustering results, the system can identify potentially fraudulent accounts and generate a risk score based on their behavior patterns. The system also provides detailed analysis reports to help users better understand the characteristics and patterns of fraudulent behavior.
[0014] System Demonstration and Application: The system features a user-friendly interface, providing functions such as data selection, parameter setting, and result display. Users can freely select analysis time periods, adjust algorithm parameters, and view system-generated fraud detection results. The system supports both real-time monitoring and offline analysis, making it suitable for a variety of application scenarios, including financial institutions, e-commerce platforms, and social networks.
[0015] The beneficial effects of this invention include: through unsupervised learning, the system can automatically detect potential fraud without the need for labeled data; through graph embedding and approximate computing, the system can maintain high efficiency when processing large amounts of data; and through community discovery and cluster analysis, the system can effectively identify complex fraud patterns, especially those that are more concealed. Furthermore, the system is highly scalable and can be flexibly adjusted and expanded according to specific application scenarios. 4. Description of the Figures
[0017] Figure 1 This diagram illustrates the overall framework of the proposed system, from data preprocessing and feature extraction to graph construction and embedding to fraud detection. The diagram clearly illustrates the interrelationships and data flows between each module, helping to understand how they work together.
[0019] Figure 2 A diagram of the sliding window strategy, showing how to dynamically capture user behavior characteristics through time window division. The system uses the sliding window strategy to extract user behavior characteristics over different time periods and analyze changes in their behavior patterns.
[0021] Figure 3 : A diagram of the Pattern of Missing Values (POMV) strategy, showing how to capture potential abnormal behaviors by analyzing missing value patterns. The POMV strategy embeds users' abnormal behaviors into the feature construction process by marking missing values and generating missing value pattern vectors.
[0023] Figure 4 A schematic diagram of the approximate computing process, comparing full and approximate calculations. Approximate computing techniques can significantly reduce the amount of computation while maintaining accuracy, adapting to large-scale data processing needs.
[0025] Figure 5 This chart compares model performance under different evaluation metrics, demonstrating the performance of the proposed system compared to other traditional fraud detection models. The chart evaluates the fraud detection performance of different models using metrics such as precision, recall, and F1-score.
[0027] Figure 6 A screenshot of the system's main panel interface, showing functions such as data selection, feature distribution, and parameter setting. Users can use the main panel to select a dataset for analysis, set corresponding parameters, and view the system's detection results.
[0029] Figure 7 A diagram showing community discovery and fraud detection results, showing potentially fraudulent accounts identified through cluster analysis. The diagram shows the distribution of accounts in different communities and the risk scores generated by the system. 5. Specific implementation methods
[0030] The following will further describe the implementation of the present invention in detail with reference to the accompanying drawings. It should be understood that the following specific embodiments are only used to illustrate the present invention and do not limit the scope of the present invention.
[0031] The present invention relates to an unsupervised intelligent fraud detection system integrating multi-level graph embedding and approximate computing, comprising a data processing and feature engineering module, a multi-level graph embedding module, an approximate computing module, an unsupervised community discovery module, a fraud detection module, and a system implementation and display module;
[0032] The data acquisition and preprocessing module batch-collects account registration and login information from a distributed database. The data set is divided by date, and each record is identified by a unique request ID. The record contains multiple features such as account ID, login time, geographic location, device information, operating system, browser type, and IP address.
[0033] In the data preprocessing stage, the data is first cleaned and organized; for records with missing key fields (such as account ID, login time, login IP, etc.), these records are directly deleted;
[0034] For missing values in non-key fields, fill in the values (such as mean, median, mode) or mark them with special values (such as -1 or NA) to avoid affecting subsequent analysis.
[0035] Outlier detection uses statistical methods and business rules to detect and handle outliers in the data, such as whether the login time is within a reasonable range, whether the geographical location is within the normal area, and whether there are abnormal changes in device information.
[0036] Next, standardize the numerical features, including normalization or standard deviation standardization, to eliminate the dimensional differences between different features and prevent certain features from having too great an impact on the model;
[0037] Based on data preprocessing, the system further constructs dynamic features to capture changing trends and abnormal patterns of user behavior;
[0038] The construction of dynamic features is accomplished through the sliding window strategy and the pattern of missing values (POMV) strategy;
[0039] The sliding window strategy divides time series data into fixed time windows (such as hourly, daily, or weekly) and counts the behavioral characteristics of accounts within each time window. For example, it counts the number of logins per account within an hour, the number of unique IP addresses, and the number of unique devices.
[0040] This method can capture short-term behavioral changes and help discover abnormal high-frequency operations or behavioral patterns. The missing value pattern (POMV) strategy captures possible fraudulent behavior by analyzing the distribution of missing values in the data. The specific implementation steps are:
[0041] Step 1: Mark missing values (NaN or None) in the data as 0 and non-missing values as 1 to form a binary vector.
[0042] Step 2: compress the above binary vector to generate the missing value pattern vector F′ MissValue .
[0043] Step 3: Convert the missing value pattern vector into frequency feature F″ MissValue , count the frequencies of different missing value patterns.
[0044] The POMV strategy can be used to identify accounts with specific missing value patterns, which may indicate abnormal behavior or fraud risks.
[0045] In the feature engineering stage, a large number of original features are screened and processed to improve the performance and stability of the model;
[0046] Statistical methods (such as information gain and chi-square tests) combined with business knowledge are used to select features that are highly correlated with fraudulent behavior. For example, the diversity of login IP addresses, the frequency of device fingerprint changes, and the abnormality of login times.
[0047] At the same time, new combined features are constructed to capture the interactive relationships between features; for example, calculating the matching degree between login IP and geographic location, and the correlation between device fingerprint and operating system.
[0048] For continuous features, discretization is performed and they are divided into multiple intervals to meet the requirements of some models for discrete features;
[0049] For categorical features, use appropriate encoding methods, such as One-Hot Encoding or Target Encoding;
[0050] In order to improve the robustness and generalization ability of the model, the present invention also introduces data augmentation technology. When there are fewer fraud samples, the oversampling method is used to generate more fraud samples and balance the data set.
[0051] In addition, the original data is modified to a certain extent through data perturbation methods, such as adding noise or randomly deleting and replacing certain feature values, so as to generate new data samples, increase data diversity and improve the generalization ability of the model.
[0052] The multi-level graph embedding module uses graph embedding methods to reveal the complex relationships between accounts and their behavioral characteristics, thereby improving the accuracy of fraud detection. The implementation of this module is divided into the following steps:
[0053] First, the present invention represents the complex association of account behaviors by constructing a heterogeneous graph between accounts and features; each account ID corresponds to an account node V A Each value of the selected important feature corresponds to a feature node V F , such as the device type is "iPhone" or the segment to which the IP address belongs;
[0054] Then in account node a i With feature node f j Create an edge E between AF ,setting the edge weight according to the strength of the association between the account and the feature, such as the number of occurrences or association frequency of the feature;
[0055] Through these nodes and edges, the account-feature association heterogeneous graph G is constructed. het =(V A , V F , E AF ), intuitively represents the complex relationship between accounts and various features;
[0056] In order to calculate the similarity between accounts, the present invention further converts the heterogeneous graph into a homogeneous graph; first, for each account a i Calculate its weight W on each feature ik ,The weight is determined based on the importance of the feature or the strength of the association between the account and the feature;
[0057] Subsequently, these weights are used to calculate the similarity sim(a i , a j), the similarity measurement methods used include cosine similarity or Jaccard similarity. In order to control the number of edges, the present invention sets a similarity threshold θ, which is only used when the similarity sim(a i , a j )≥θ in account a i and a i Create edges between
[0058] Connect all eligible account pairs through these edges to generate an isomorphic graph G het =(V A , E AA ), where E AA is the edge set between account nodes;
[0059] Based on the constructed isomorphic graph, this paper further performs multi-level graph embedding to capture the complex relationships between accounts. First, using graph embedding algorithms such as Node2Vec or DeepWalk, account nodes are mapped into low-dimensional vector representations. These algorithms capture the local and global structural features of nodes through random walks or word embedding training.
[0060] Then, on this basis, the present invention also performs community-level embedding, using methods such as Graph2Vec to further embed the community structure information of the account group into a vector representation;
[0061] To form a richer account representation, the present invention also fuses the node embedding vector with the account's feature vector, using fusion methods such as vector concatenation or weighted averaging to ultimately generate a comprehensive feature vector that includes account behavior and its associated structure.
[0062] This multi-level graph embedding method not only fully utilizes the complex associations between accounts and features, but also enhances the model's ability to detect abnormal account behavior through the comprehensive representation of the embedding vectors, providing strong support for subsequent fraud risk identification.
[0063] When performing calculations on large-scale graph data, computing resources and time consumption often become limiting factors. This paper uses approximate calculation methods to significantly improve computing efficiency while ensuring the accuracy of the results. The main aspects include the following:
[0064] First, in similarity calculation, this invention uses an approximation method to reduce the amount of calculation. To reduce computational complexity, the screening of candidate account pairs uses technologies such as MinHash and LSH (Locality Sensitive Hashing). These technologies can quickly screen out account pairs that may have high similarity, thereby reducing the number of account pairs that require detailed similarity calculations.
[0065] In addition, for the account-feature matrix, the present invention performs low-rank approximation through matrix decomposition technology, such as using SVD (singular value decomposition) or NMF (non-negative matrix factorization) methods to reduce the matrix dimension and accelerate the similarity calculation process;
[0066] Secondly, the present invention also uses sampling technology to reduce the amount of calculation when processing large-scale graph data. When the number of nodes is huge, node sampling technology is used to perform calculations only on the sampled subgraph. Common sampling methods include random sampling, degree sampling, PageRank sampling, etc.
[0067] For graphs with a huge number of edges, the present invention adopts an edge sampling method to retain important edges in the graph and discard those with smaller weights or unimportant edges to reduce the complexity of the graph.
[0068] In addition, to further improve computing efficiency, the present invention also adopts parallel computing technology; under the distributed computing framework, Spark, Flink and other big data computing frameworks are used to distribute computing tasks to multiple nodes for parallel execution, thereby accelerating the overall computing process;
[0069] At the same time, the powerful parallel computing capabilities of the GPU are used to accelerate matrix operations and embedding algorithms, greatly shortening processing time. This makes it suitable for large-scale data scenarios that require high-performance computing.
[0070] Through the above-mentioned approximate calculation method, the present invention can effectively improve computing efficiency and reduce the consumption of computing resources when processing large-scale graph data, providing an efficient solution for graph analysis in large-scale fraud detection.
[0071] The present invention also further improves the accuracy and efficiency of fraud identification through unsupervised community discovery and fraud detection, which mainly includes community detection, graph embedding and cluster analysis, and finally fraud detection;
[0072] First, in terms of community detection, this paper uses the Louvain algorithm, which is an efficient community detection algorithm based on modularity optimization. The core idea of the Louvain algorithm is to divide the nodes in the network into multiple closely connected communities by maximizing modularity. The specific implementation steps are as follows:
[0073] Node initialization and community division: First, the system initializes each account node as an independent community. That is, each account node initially belongs to its own community.
[0074] Calculation of Modularity Gain and Community Movement: Next, the algorithm traverses all nodes, attempting to move a node to the community of its neighboring nodes and calculating the modularity gain of this move. If moving a node increases modularity (i.e., the node's joining a new community increases the density within the community), the move is performed, adding the node to the neighboring node's community.
[0075] Community aggregation and graph simplification: After the initial community division is completed, the algorithm enters the community aggregation phase. In this phase, the algorithm aggregates the nodes in each community into a supernode and reconstructs the graph structure for these supernodes. This simplifies the graph structure and provides a foundation for subsequent iterations.
[0076] Iteration: The above steps are repeated until the modularity no longer increases significantly. Through multiple iterations, the algorithm can eventually divide the nodes into multiple communities, each containing a group of structurally similar account nodes.
[0077] After completing community detection, the present invention further analyzes the community structure through graph embedding and cluster analysis technology. Specifically, Graph2Vec technology is used, which is an important method in the field of graph embedding and can convert complex graph structures into fixed-length vector representations. Graph2Vec can effectively capture the global structural information of the community, including the association pattern and connection density between community nodes, by traversing and extracting features from subgraphs within the community.
[0078] These embedding vectors are analyzed for similarity using clustering algorithms such as KMeans, which clusters communities based on structural similarity. The clustering results can be used to classify communities into different categories, some of which correspond to normal account behavior, while others may indicate abnormal or fraudulent behavior.
[0079] In the fraud detection phase, the present invention detects fraudulent accounts by identifying abnormal communities. Based on the results of cluster analysis, the system can identify communities that are significantly different from normal communities.
[0080] These abnormal communities often have unusual structural characteristics, such as high-density connections between accounts, shared device fingerprints, identical IP addresses, or frequent operations within a specific time period. We then conduct a deeper feature analysis of these communities to further uncover the account behavior patterns within them.
[0081] Ultimately, by conducting risk assessments on each account within the community and combining business rules with expert knowledge, we can effectively determine whether there is fraudulent activity in the community and provide accurate risk predictions and preventive measures.
[0082] The architecture of the system of the present invention is divided into data layer, computing layer, application layer and presentation layer; the data layer includes a distributed database, data storage and cache system, which is used to store and manage account data, feature data, graph data, etc.
[0083] The computing layer includes a data preprocessing module, a feature engineering module, a graph construction module, an approximate computing module, a community detection module, and a graph embedding module. It uses a distributed computing framework and GPU acceleration technology to achieve efficient computing.
[0084] The application layer includes a fraud detection engine, risk assessment module, and alarm system, responsible for real-time fraud detection and risk assessment;
[0085] The presentation layer provides a user-friendly interface with functions such as data visualization, result display, and parameter setting, making it easy for users to operate and analyze.
[0086] In terms of system functions, users can select the data range for analysis through the system interface and set parameters such as similarity threshold, fault tolerance threshold, number of clusters, etc.
[0087] The system provides a variety of data visualization functions, including the usage frequency and distribution of account features, the geographical distribution of account logins, and the connection diagram between account nodes and feature nodes, to help users understand the relationship between accounts and features;
[0088] The system also displays the results of community detection, the distribution of accounts within the community, identifies high-risk communities and accounts, and provides detailed risk scores and explanations. Users can click on a specific account or community to view detailed information and characteristics, supporting in-depth analysis of suspicious accounts.
[0089] In terms of system performance, the system of the present invention can be deployed on the cloud or on a local server, supports horizontal expansion, and adapts to data processing needs of different scales; through approximate computing, parallel computing and caching mechanisms, the system can significantly improve processing speed and responsiveness;
[0090] In actual business scenarios, the system can effectively identify fraudulent behavior, reduce economic losses, and improve risk control capabilities. Experiments and tests have shown that the method of the present invention has achieved high scores on multiple evaluation indicators (such as accuracy, precision, recall rate, and F1-score), with high accuracy and high efficiency.
[0091] At the same time, this method is highly robust to data noise and missing data, and can adapt to different data quality and scales. The system enhances the interpretability of the detection results through graph structure and feature interpretation, and improves ease of use by providing a friendly user interface and visualization tools. Even users without a technical background can easily operate and analyze.
[0092] 6. Application prospects
[0093] The method and system of the present invention can be widely used in financial anti-fraud, network security monitoring, e-commerce platform anti-cheating, social network anomaly detection and other fields, and have important commercial value and social significance.
Claims
1. An unsupervised intelligent fraud detection system combining multi-level graph embedding and approximate computing, characterized in that: The system includes: data processing and feature engineering module, multi-level graph embedding module, approximate computing module, unsupervised community discovery and fraud detection module and system algorithm display module. The data processing and feature engineering module is used to obtain batch account registration information data from the distributed database, perform missing value processing, data cleaning, feature construction and aggregation to meet the needs of downstream algorithms; The multi-level graph embedding module constructs an account-feature association heterogeneous graph and further converts it into a homogeneous graph to achieve effective calculation of similarity between accounts; The approximate calculation module is used to reduce the complexity of graph calculation by using approximate technology. By setting the fault tolerance threshold ε and using methods such as sampling or matrix decomposition, the similarity between nodes is approximated and the consumption of computing resources is reduced. The unsupervised community discovery and fraud detection module combines the Louvain algorithm and graph embedding technology to perform community segmentation and fraud account detection; The system algorithm display module is based on the Flask framework, and displays the algorithm's interaction and operation results to users through a visual interface, including data feature distribution, time and geographic information distribution, community discovery and fraud detection performance, etc.
2. The unsupervised intelligent fraud detection system according to claim 1, characterized in that: The data processing and feature engineering module includes: Obtain batch account registration information data D from a distributed database, where rows represent account nodes and columns represent feature nodes; The data preprocessing process includes deleting records containing key missing values, filling regular missing values, and building dynamic features using sliding window strategy and missing value pattern (POMV) strategy; The dynamic features are extracted by using feature statistics and feature aggregation methods within a time window to ensure the time series consistency and rationality of the features.
3. The unsupervised intelligent fraud detection system according to claim 2, characterized in that: Building a multi-level graph embedding module includes: By creating nodes and edges based on non-empty data items, we construct an account-feature association heterogeneous graph G het =(V A , V F ,E AF ), where V A Represents the account node, V F represents a feature node, E AF Represents the edge between an account and a feature; By calculating the feature weights of the accounts, the heterogeneous graph G het Convert to isomorphic graph G hom =(V A , E AA ), where E AA Represents the edge between similar accounts, and the weight of the edge is the similarity between the accounts; Specifically, it includes initializing the nodes and edges in the dataset, calculating the feature weights of the accounts, and adding similarity-based edges to convert heterogeneous graphs into homogeneous graphs.
4. Based on the method of claim 3, the process of constructing the isomorphic graph G_hom further comprises: Initialize account nodes and feature nodes, calculate the feature weights of each account, and calculate the similarity between account nodes through the similarity measurement function; Based on the similarity of account nodes, a sparse isomorphic graph is constructed, in which the existence of edges depends on the set similarity threshold, ensuring the scalability and efficiency of graph computing in large-scale scenarios.
5. An unsupervised intelligent fraud detection system according to claims 1, 3 and 4, characterized in that: The approximate computing module includes: By setting the fault tolerance threshold ε, using random sampling or sparse matrix approximation methods, the similarity between nodes is approximated and the computational complexity is reduced to O(kn), where n is the number of nodes and k is a constant much smaller than n. Approximate calculation node a i and a j The similarity between: where sim′(a i , a j ) is the similarity calculated by the approximate method, βk is the weight coefficient; The complexity of graph computation is reduced through approximation techniques, enabling the system to process large-scale graphs containing millions of nodes.
6. The unsupervised intelligent fraud detection system according to claim 5, characterized in that: The unsupervised community discovery and fraud detection module includes: Use the Louvain algorithm to perform community detection on the homogeneous graph Ghom and cluster similar accounts into the same community; Using graph embedding technology Graph2vec, the high-dimensional structure of the community graph is converted into a low-dimensional vector representation; Apply KMeans clustering algorithm to low-dimensional embedding to identify fraudulent accounts, reduce misjudgment of normal accounts, and improve detection accuracy; Through unsupervised learning methods, user behavior patterns are systematically captured and analyzed to identify overt and covert fraudulent activities.
7. An unsupervised intelligent fraud detection system according to claims 1-6, characterized in that: The system has a user-friendly interface and is suitable for distributed deployment to process large-scale data. The system algorithm display module includes: An interactive visualization interface based on the Flask framework, which displays data feature distribution, time and geographic information distribution, community division results, fraud detection performance, etc. The visual interface supports dynamic updates and real-time data feedback, and managers can monitor and analyze the system operation status through the graphical interface; The system can be expanded to distributed environments, supports parallel graph computing and large-scale data processing, and provides a load balancing mechanism for multi-node environments; The interface supports real-time detection and result display. Users can dynamically monitor community divisions and fraud detection results through the interface to further optimize operational strategies and ensure efficient fraud detection.