A method and system for detecting abnormal users in a social network based on external graph data

By introducing external graph data of diversity and diversity of scale and adopting a selection mechanism for representative and diversity indicators, the problem of difficult to identify low-frequency legal user behavior in social networks in the prior art is solved, which significantly improves the accuracy and robustness of abnormal user detection.

CN119762262BActive Publication Date: 2025-05-30ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510253932.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-05-30
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

Existing graph anomaly detection methods are difficult to effectively identify low-frequency but legal user behavior in social networks, and are easily misjudged as abnormal users. When normal data is insufficient, the normal distribution learned by the model may be biased, weakening the effect of abnormal detection.

Method used

An external graph data set of field diversity and scale diversity is introduced, candidate graphs are generated through data enhancement, and an external graph data selection mechanism based on representative indicators and diversity indicators is adopted to filter out candidate graphs that are sufficiently similar to the target graph data and can fully cover various normal modes, thereby enhancing the generalization ability and detection accuracy of the model.

Benefits of technology

It significantly improves the accuracy and robustness of abnormal user detection tasks in social networks, avoids overfitting problems in traditional graph anomaly detection models, and enhances the generalization ability of the model through diversified external graph data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119762262B_ABST
    Figure CN119762262B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for detecting abnormal users in a social network based on external graph data, belonging to the technical field of abnormal detection of graph-structured data. An external graph data set and at least one target social network graph data are obtained; after data augmentation operations, they are used as candidate graphs, and node features are aligned; a spherical space is defined, and at least one original target social network graph data is used to pre-train a graph model to generate target graph node representations and their position coordinates in the spherical space; the candidate graphs are input into the pre-trained graph model to generate respective candidate graph node representations and their position coordinates in the spherical space; candidate graphs are screened according to the node representations and their position coordinates in the spherical space; the graph model is retrained using the screened candidate graphs, and the retrained graph model is used to detect abnormal users in the target social network to be detected. The present invention enhances the generalization ability of the model through diverse external graph data and improves the detection accuracy of abnormal users in the social network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graph-structured data anomaly detection, and particularly to a method and system for detecting abnormal users in a social network based on external graph data. Background Art

[0002] Graph Anomaly Detection (GAD) aims to identify abnormal instances or outliers in a graph that deviate from the main patterns of normal nodes. In social networks, abnormal users (such as fake accounts, spam dissemination, malicious advertisements, etc.) often exhibit interaction patterns different from those of normal users. Graph anomaly detection can help identify these abnormal accounts, for example, by analyzing the interaction relationships between users, the comment or post dissemination paths, and timely detecting and preventing potential malicious behaviors. For example, a certain social platform commonly uses graph anomaly detection technology to identify robot accounts and malicious groups manipulating the information flow. The present invention focuses on the detection of abnormal user behaviors in social networks.

[0003] With the wide application of Graph Neural Networks (GNNs) in various fields, Graph Anomaly Detection (GAD) has also made remarkable progress. Supervised and semi-supervised GAD methods use labeled abnormal data as supervision signals, enabling the model to learn the distributions of normal and abnormal nodes, and thus effectively distinguish abnormal nodes from normal nodes. However, anomalies in graph-structured data have some unique characteristics, which make supervised and semi-supervised graph anomaly detection methods based on limited labels often constrained by these challenges, and may face overfitting problems or be unable to effectively generalize to unseen abnormal behaviors. To overcome these limitations, many studies have proposed using unsupervised strategies, which focus on learning the distribution of comprehensive and robust normal behaviors rather than the distribution of abnormal behaviors. By effectively learning normal behavior patterns, the model can naturally distinguish abnormal instances that deviate significantly from these patterns. However, these methods rely to a large extent on sufficient and representative normal patterns in the target data, but in practical applications, this condition may not always be met. When the normal data is insufficient, the learned normal distribution may be biased, thereby weakening the effect of anomaly detection.

[0004] Taking the detection of abnormal users in social networks as an example: In social graph analysis, 90% of normal users exhibit stable social patterns: the number of friends gradually increases, the interaction time conforms to the circadian rhythm, and the content dissemination shows community aggregation. However, some scholars in certain professional fields (users with low activity frequency but legal behavior) may show different patterns, such as: frequent interactions during sudden academic conferences (usually 2 to 3 times a year), a high proportion of professional term dissemination, cross-time zone collaboration characteristics, etc. Traditional graph anomaly models usually focus on identifying high-frequency behavior patterns, which easily leads to misjudgment of these legal low-frequency users and mistakenly regards them as abnormal users. However, real anomalies are not simply characterized by low-frequency behavior, but are reflected in obvious irregular behaviors. For example, abnormal users with forged accounts often form social relationships with an obvious star topology structure, and such structures are relatively rare among real users. Therefore, providing a more comprehensive and accurate perspective of normal behavior helps the model better identify abnormal users in social networks. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a method and system for detecting abnormal users in social networks based on external graph data. An external graph data set with domain diversity and scale diversity is introduced and data augmentation is performed to generate candidate graphs. A selection mechanism for external graph data based on representative indicators and diversity indicators is adopted to screen candidate graphs that have sufficient similarity with the target graph data and can comprehensively cover various normal patterns. The generalization ability of the model is enhanced through diverse external graph data, and the accuracy of detecting abnormal users in social networks is improved.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] In the first aspect, the present invention proposes a method for detecting abnormal users in social networks based on external graph data, including the following steps:

[0008] (1) Obtain an external graph data set based on domain diversity and scale diversity, and obtain at least one target social network graph data; each graph data is composed of a node feature matrix and an adjacency matrix;

[0009] (2) Perform data augmentation operations on the original graph data as candidate graphs, generate the node feature matrix and adjacency matrix of the candidate graphs using the node feature matrix and adjacency matrix of the original graph data, and perform feature space alignment on the node feature matrix of the candidate graphs;

[0010] (3) Define a spherical space, pre-train a graph model using at least one original target social network graph data, generate target graph node representations and determine the position coordinates of each node representation in the spherical space; and input the candidate graphs into the pre-trained graph model to generate the node representations of each candidate graph and their position coordinates in the spherical space;

[0011] (4) Adopt an external graph data selection mechanism based on representative indicators and diversity indicators, and screen candidate graphs according to node representations and their position coordinates in the spherical space;

[0012] (5) Use the selected candidate Figure 2 to train the graph model for the second time, and use the graph model trained for the second time to detect abnormal users in the target social network to be detected.

[0013] As a preference of the present invention, the external graph data set includes two or more of e-commerce network graph data, social network graph data, academic citation network graph data, and hyperlink network graph data.

[0014] As a preference of the present invention, the data augmentation operations include feature masking, node deletion, edge perturbation, and subgraph extraction. Multiple intensity parameters are set for each data augmentation operation, and candidate graphs are generated by traversing each data augmentation operation under different intensity parameters.

[0015] As a preference of the present invention, the method for feature space alignment of the node feature matrix of the candidate graph is:

[0016] Judge the type of the node features of the candidate graph, and the type is one of text-form features and table-form features;

[0017] If it is table-form features, convert them to text-form features based on a rule-based conversion method; otherwise, no processing is required;

[0018] Input the text-form features into a multi-language pre-trained language model, and use the text embedding generated after the language model encodes the text as the node features after feature space alignment.

[0019] As a preference of the present invention, the pre-training of the graph model using at least one original target social network graph data includes:

[0020] Input the node feature matrix and adjacency matrix corresponding to the original target social network graph into the graph model to generate target graph node representations;

[0021] The training objective is that the target graph node representations output by the graph model are distributed in the spherical space as much as possible.

[0022] As a preference of the present invention, the representative indicators in the external graph data selection mechanism include the center similarity deviation based on the center stability constraint and the distribution similarity deviation based on the distribution similarity constraint;

[0023] The center similarity deviation is the L2 norm between the mean of the target graph node representations and the candidate graph node representations and the mean of the target graph node representations;

[0024] The described distribution similarity deviation is the distance between the probability distribution of the spherical coordinate sets of the target graph node representations and the candidate graph node representations, and the probability distribution of the spherical coordinate sets of the target graph node representations.

[0025] Preferably, in the present invention, the diversity index in the external graph data selection mechanism adopts the integral value of the minimum hypersphere energy with different scale radii.

[0026] Preferably, in the present invention, the final scores of each candidate graph are calculated by fusing the representativeness index and the diversity index, and a preset number of candidate graphs are selected in ascending order of the scores.

[0027] Preferably, in the present invention, the graph model using quadratic training is used to detect abnormal users in the target social network to be detected, including:

[0028] Input the node feature matrix and the adjacency matrix of the target social network to be detected into the graph model with quadratic training, generate node representations and determine the position coordinates of each node representation in the spherical space;

[0029] If the node representations are distributed in the spherical space, it is determined that the node corresponding to the node representation is a normal user, otherwise it is determined that the node corresponding to the node representation is an abnormal user.

[0030] In a second aspect, the present invention proposes a social network abnormal user detection system based on external graph data for implementing the above-mentioned social network abnormal user detection method based on external graph data.

[0031] The beneficial effects of the present invention are as follows:

[0032] By introducing external graph data and an accurate data selection strategy, the present invention significantly improves the accuracy and robustness in the task of detecting abnormal users in social networks. This method can effectively avoid the overfitting problem in traditional graph anomaly detection models and enhance the generalization ability of the model through diverse external graph data. Description of the Drawings

[0033] Figure 1 is the overall block diagram of a social network abnormal user detection method based on external graph data shown in an embodiment of the present invention.

[0034] Figure 2 is the schematic diagram of the feature mask shown in an embodiment of the present invention.

[0035] Figure 3 is the schematic diagram of node deletion shown in an embodiment of the present invention.

[0036] Figure 4 is the schematic diagram of edge perturbation shown in an embodiment of the present invention.

[0037] Figure 5 It is a schematic diagram of sub - graph extraction shown in the embodiments of the present invention. Detailed implementation manners

[0038] The present invention will be further described and explained below in conjunction with the detailed implementation manners. The embodiments are only examples of the disclosed content and do not delimit the scope of limitation. Without conflict, the technical features of each embodiment of the present invention can be combined accordingly.

[0039] The accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0040] The flowcharts shown in the accompanying drawings are only illustrative and do not necessarily include all steps. For example, some steps can be decomposed, while some steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.

[0041] Figure 1 The overall framework of the graph anomaly detection method based on external graph data of the present invention is shown. It mainly includes the construction of an external graph database and the screening of external graph data, and a graph anomaly detection model is trained using the screened external graphs. The functions and implementation details of each part are introduced separately below.

[0042] I. Construction of the external graph database:

[0043] It mainly includes three parts: collection of original external graph data, graph data augmentation, and feature space alignment, which are specifically as follows:

[0044] S11. Consider the diversity of the collection of original external graph data from two dimensions: domain diversity and scale diversity.

[0045] (1) Regarding domain diversity, to ensure a comprehensive representation of the normal patterns of cross - domain graph data, the present invention collects external graph data from four different domains: e - commerce networks, social networks, academic citation networks, and hyperlink networks. Among them:

[0046] In an e-commerce network, nodes represent users or products; the features of user nodes can include age, gender, geographical location, historical purchase records, user activity (such as the number of comments or purchases); the features of product nodes can include price, category, brand, mean or variance of user ratings, text embeddings of product descriptions, etc.; edges are used to represent associations between users (such as having commented on the same product), or associations between products (such as being purchased by the same user or being recommended to be purchased together).

[0047] In a social network, nodes represent individual users or organizations; node features can include basic attributes of users (such as age, gender, interest tags), embeddings of user-generated content (such as text features of profiles or posts); edges are used to represent social relationships between users such as friend / follow relationships, etc.

[0048] In an academic citation network, nodes represent academic papers, and node features are usually text embeddings of papers (such as titles, abstracts, keywords), and edges represent citation relationships between papers.

[0049] In a hyperlink network, nodes represent web pages, and node features are embeddings of web page content (such as keywords, topic model representations), and edges represent hyperlink relationships between web pages.

[0050] In a specific implementation of the present invention, a total of twelve graph datasets were sorted out for the above-mentioned fields, including P-Learning, P-Electronic, P-Entertainment, P-Household, P-Fashion, P-YelpRes, YelpNYC, Instagram, Reddit, Cora, PubMed, arXiv, and WikiCS. The graph data in the above graph datasets only contains normal nodes and no abnormal nodes.

[0051] In addition to the above external graph data, the graph database also needs to include target social network graph data containing a small number of abnormal nodes. Abnormal nodes refer to abnormal users, and the remaining nodes are normal users. The node features are the basic attributes of users (such as age, gender, interest tags) and / or embeddings of user-generated content (such as text features of profiles or posts), which are used to provide a more accurate and realistic reference for the model.

[0052] (2) Regarding scale diversity, the graph data collected by the present invention also covers a wide range in terms of graph scale size to capture diverse graph structures. For example, the smallest graph dataset Cora contains 2,708 nodes and 10,858 edges, while the largest graph dataset P-Learning contains 651,762 nodes and 24,599,888 edges. This diversity in scale enhances the model's ability to recognize and represent normal samples of graphs of different scales, thus more effectively capturing normal patterns.

[0053] S12. Perform data augmentation on the original external graph data and the original target graph data.

[0054] The present invention introduces graph data augmentation technology to diversify the candidate dataset, thereby achieving a wider coverage of normal patterns in graph data. The data augmentation technology helps the model better adapt to the changing data distribution, thus enhancing the generalization ability and robustness of the model. This technology is particularly beneficial in cross-dataset application scenarios. Therefore, the present invention applies graph data augmentation technology to expand the original graph data to obtain effective new samples.

[0055] In a specific implementation of the present invention, a set of data augmentation operations is designed to enhance the diversity and expressive ability of each graph data. The data augmentation operations include:

[0056] Feature masking: Set some node feature values to 0 or the average of the feature values of its neighbors; as Figure 2 shown, set the feature value of node 1 to the average of the feature values of nodes 2, 3, and 5.

[0057] Node deletion: Randomly remove a certain proportion of nodes from the original graph data; as Figure 3 shown, remove node 1 and its related edges.

[0058] Edge perturbation: Modify the structure of the original graph by adding or removing edges to generate graphs with different connection relationships; as Figure 4 shown, delete the edge between node 1 and node 2, and the edge between node 1 and node 5, and add the edge between node 2 and node 5.

[0059] Subgraph extraction: Extract some nodes from the original graph and some of their related edges and features to obtain a new subgraph; as Figure 5 shown, extract nodes 1, 3, 4, 5 from the original graph, and the edges between node 1 and node 3, between node 3 and node 4, between node 4 and node 5, and between node 5 and node 1.

[0060] To control the degree of modification of the original graph by the data augmentation method, the present invention presets an intensity parameter for each augmentation method, and its value is selected from {0.2, 0.4, 0.6, 0.8}. Specifically, for feature masking, the intensity is defined as the proportion of nodes that mask the features; for node deletion, the intensity is defined as the proportion of nodes randomly removed from the graph; for edge perturbation, the intensity is defined as the proportion of edges added or deleted from the original graph; for subgraph extraction, the intensity is defined as the difference between the proportion of subgraph nodes extracted from the original graph and 1. The data augmentation operations under different intensities determine the degree of modification of the original graph by the data augmentation method, thus affecting the structure and feature information of the graph data. The original graph data is augmented by traversing all data augmentation operations under different intensities. Here, subgraph extraction can be implemented by a conventional random walk algorithm, which will not be elaborated here.

[0061] In this embodiment, by integrating all the original graph data and their augmented samples, a dataset containing 221 candidate graphs is finally constructed, which provides rich samples and diverse feature distributions for model training. Here, the candidate graphs include the original external graph data and their augmented samples, and the original target graph data and their augmented samples.

[0062] S13, perform feature space alignment on the augmented graph data.

[0063] Since the collected external graph data comes from different fields and channels, its node features are often distributed in different feature spaces, with different dimensions and complex properties. The present invention converts the node features of these heterogeneous attributes into text form, and uses a language model to understand and process them in a unified semantic space. The node features of graph data can be divided into two categories: text features and table features. Text features are directly expressed as text information. For example, the node features in the academic citation network are usually text embedded descriptions of papers, and these features can be directly used as inputs of language models. Tabular features usually exist in the form of classification or numerical vectors. For example, in social networks, node features include basic attributes of users (such as age, gender, and interest tags), which are generally recorded in the form of tables. For such tabular features, they must first be converted into text descriptions. In this embodiment, a rule-based conversion method is used to textualize tabular features. For example, a node feature in the Tolokers data set represents a worker, and its tabular features are approved_rate: 0.8, skipped_rate:0.2, expired_rate: 0.2, rejected_rate:0.1. Through the rule-based conversion method, these table features are converted into the following text format: "The probability of this worker being supported is 0.8, the probability of being ignored is 0.2, the probability of being expired is 0.2, and the probability of being rejected is 0.1". After completing the text conversion of the table features, the present invention uses the powerful language understanding ability of the language model to input the node features in the text format of each graph data into the language model, and the text embedding generated by the language model after encoding the text is used as the aligned node feature, thereby realizing the alignment of the heterogeneous feature semantic space. In addition, in response to the multilingual problem existing in the collected extensive candidate graph dataset (for example, the dataset contains English and Chinese), the present invention adopts a multilingual pre-trained language model, such as mBERT, to ensure effective processing and semantic understanding between different languages. By aligning the feature space of the graph data after data enhancement, a candidate graph dataset containing rich and diverse graph data is obtained, and the node features of each candidate graph in the dataset are the node features after spatial alignment using the above method.

[0064] 2. External graph data screening:

[0065] The lack of normal patterns and poor quality in the external graph data may cause the representation space generated by the model to be irregular and deviate from the standard hypersphere, which may cause errors in the stage of using the hypersphere to evaluate normal and abnormal samples. To overcome this problem, the present invention specifically integrates external graph data that can assist model training and help optimize the representation space, in order to provide a more comprehensive depiction of normal patterns. Specifically, by defining a spherical space, suitable external graph data is screened out to enhance the performance of anomaly detection tasks.

[0066] It mainly includes four parts: defining the spherical space, training the GNN model, calculating the node representations of the candidate graph and determining the position coordinates of each node representation in the spherical space, and designing an external graph data selection mechanism, which are specifically as follows:

[0067] S21. Define the spherical space based on the following three core elements:

[0068] (1) Spherical center (abbreviated as the center of the sphere): The spherical center is defined as the mean of all node representations, and the calculation formula is , where represents the center of the sphere, represents the number of nodes in the graph data, represents the i-th node representation in the graph data.

[0069] (2) Radial distance: The Euclidean distance between each node representation and the center of the sphere, and the calculation formula is , where represents the radial distance of the i-th node representation in the graph data, represents the calculation of the Euclidean distance.

[0070] (3) Relative direction: The unit vector pointing from each node to the center of the sphere, and its calculation formula is , where represents the relative direction of the i-th node representation in the graph data.

[0071] In order to more comprehensively and discriminatively represent the normal graph data samples, the present invention proposes the above spherical space for analyzing the representation of nodes.

[0072] S22. Use any original target graph (containing a small number of abnormal nodes) to pre-train the GNN model, so that the node representations of the original target graph output by the GNN model are mostly distributed in the above spherical space.

[0073] In this embodiment, by combining the information in terms of both radial and angular aspects, each multi-dimensional node representation located on the hypersphere can be represented by the spherical coordinates .

[0074] By using any original target graph to train a basic model to learn its corresponding spherical space and obtain the representation of each node in the target graph and its spherical coordinates , the training objective is that the node representations of the original target graph output by the basic model are mostly distributed in the above spherical space, and the trained basic model reflects the current model's understanding of the normal mode in the target graph.

[0075] S23. Input the candidate graphs into the above GNN model to generate node representations for each candidate graph and determine the position coordinates of each node representation in the spherical space. This provides a key basis for subsequent screening of external graph data.

[0076] S24. Use the external graph data selection mechanism to screen graph data based on the position coordinates of each node representation in the spherical space.

[0077] In order to select external graph data that can enhance model training and enrich the representation space, the external graph data selection mechanism designed in the present invention follows two key criteria: the representativeness index and the diversity index. Among them, representativeness means that the selected graph data should be highly consistent with the potential pattern of the target graph, indicating a high similarity with the target graph. Diversity means that the selected graph data should exhibit sufficient diversity to cover a more comprehensive range in the representation space. These two criteria complement each other. Representativeness ensures the relevance of the external graph data, while diversity aims to maximize the coverage of normal pattern graph data. Ideal external graph data should satisfy both of these criteria. By comprehensively considering the above criteria, the present invention can calculate a final score for each candidate graph to evaluate its applicability and potential in improving the performance of the anomaly detection task.

[0078] The following separately introduces the quantitative evaluation of the representativeness index and the diversity index:

[0079] (1) Representativeness index

[0080] The similarity between the selected external graph data and the target graph data is crucial for the effective fusion of graph data. If the similarity between the external graph data and the target graph data is insufficient, it may introduce noise, resulting in a shift in the data distribution and having a negative impact on the model performance. The present invention conducts a quantitative evaluation of the representativeness index from the following two key perspectives.

[0081] First, the spherical center stability constraint:

[0082] After integrating the external graph data and the target graph data, the key is to maintain the stability of the spherical center position in the representation space. That is to say, the spherical center should not have a significant movement. To evaluate this stability, the Euclidean distance between the spherical center of the original representation space and the spherical center of the new representation space after combining the external graph data can be calculated.

[0083] Calculate the spherical center similarity of each candidate graph based on the node representation of the candidate graph :

[0084]

[0085]

[0086] Among them, represents the center of the sphere of the original representation space of the target graph, calculated as the mean of the node representations of the target graph; is the center of the sphere of the new representation space combining a candidate graph and the target graph, calculated as the mean of the node representations of the target graph and the node representations of a candidate graph; represents the set of node representations of the target graph, represents the set of node representations of a candidate graph, represents a node representation in the graph, respectively represent the number of nodes in the target graph and the candidate graph, represents the L2 norm.

[0087] Center similarity A smaller value of indicates a smaller displacement of the center of the sphere, which means that the external graph data can be effectively fused and does not distort the representation space. On the contrary, if this value is large, it indicates that the center of the sphere has a large displacement, which may have an adverse impact on the performance of the model in the anomaly detection task.

[0088] Second, distribution similarity constraint:

[0089] Evaluate the similarity between the distribution of the target graph node representations and the distribution of the representations after integrating the external graph data. In this embodiment, the Wasserstein distance is used as the metric. This distance measures the minimum cost required to transform one distribution into another, taking into account both the values and directions of the representation vectors.

[0090] Calculate the distribution similarity of each candidate graph according to the position coordinates of the node representations of the candidate graph in the spherical space :

[0091]

[0092] Among them, represents the set of spherical coordinates of the target graph node representations, represents the set of spherical coordinates after combining a candidate graph and the target graph, represents calculating the probability distribution.

[0093] Distribution similarity The smaller it is, the more similar the distribution of the mixed external graph data is to the distribution of the target graph, thus ensuring good alignment between the external graph data and the original graph data.

[0094] (2) Diversity index

[0095] Although representative metrics play an important role in ensuring a high degree of consistency between the selected external graph data and the target dataset, their single application may not be sufficient to comprehensively improve the performance of the model. Relying solely on representativeness may lead to overfitting of the model to specific patterns. In such cases, the model may overly focus on these specific patterns and ignore other potential patterns present in the graph data, thereby affecting the generalization ability of the model. Therefore, the present invention introduces a diversity metric to enable the model to learn and identify a wider range of normal sample features and patterns by promoting the uniform distribution of graph data samples in the representation space, enhancing the generalization ability of the model.

[0096] Specifically, the realization of diversity in the spherical space depends on the uniform distribution of node representations over the entire spherical space, that is, they should not be overly concentrated in a specific region of the hypersphere but should be widely distributed over the entire hypersphere. The present invention uses the Minimum Hyperspherical Energy (MHE) as a metric to measure the degree of uniformity of node representation distribution.

[0097] In the spherical space Define the hyperspherical energy as follows:

[0098]

[0099] where, and represent the minimum and maximum radii of the distribution of the target graph node representations and the candidate graph node representations in the spherical space after mixing, represents the set of node representations at radius , represents the set of node pairs in , represents a pair of node representations in ; represents the energy of a pair of node representations, which is defined using a Gaussian potential kernel

[0100]

[0101] where, represents the energy function of a pair of node representations in the spherical space, is the hyperparameter of the Gaussian potential kernel. This function is designed as a distance-based decreasing real-valued function, that is, in the representation space, a pair of node representations with a closer distance is assigned a higher energy, while a pair of node representations with a farther distance is assigned a lower energy.

[0102] The minimum and maximum radii and are defined as follows:

[0103]

[0104] The present invention introduces the integration of energies with different scale radii, which ensures that diversity is considered at multiple scales of the hypersphere, not limited to the case of a fixed radius (e.g., r = 1). This method of multi-scale analysis provides a more comprehensive measure for the evaluation of diversity.

[0105] III. Training the graph anomaly detection model:

[0106] The representativeness and diversity metrics proposed by the present invention together constitute an effective metric for evaluating the quality of external graph data. By calculating the final score of each candidate graph, the problem of selecting graph data can be transformed into the following optimization problem:

[0107]

[0108] where ∈[0,1] is a trade-off coefficient used to balance the weights between representativeness and diversity, respectively represent the center similarity, distribution similarity, and hypersphere energy after z-score normalization.

[0109] Given a budget limit of external graph data, by calculation, the top external graph data with the lowest scores are determined, and then the GNN model is retrained using these selected external graph data.

[0110] The GNN model obtained from the retraining is used as the graph anomaly detection model to implement the anomaly detection of the target graph data. The detection process is as follows:

[0111] The GNN model obtained from the retraining generates the node representations of the target graph data;

[0112] If the node representations are distributed in the spherical space, the node represents a normal user; otherwise, the node is an abnormal user.

[0113] The graph anomaly detection model of the present invention is first trained on the target graph data for initialization to establish the initial parameters of the model . On this basis, is applied to each candidate external graph. Given a budget limit of graphs, by calculation, the top external graphs with the lowest scores are determined to retrain the model, which ensures that the model can obtain the most effective and diverse knowledge from these external graphs.

[0114] This embodiment verifies the implementation effect of the present invention through a specific experiment.

[0115] I. Data Description

[0116] The present invention uses four social network target graph datasets with real anomalies from different fields, including YelpHotel, Amazon_cn, C15, and Twi20.

[0117] The specific introductions of these datasets are as follows:

[0118] The nodes in YelpHotel represent users of the Yelp website. The node features are all the review texts of hotels on Yelp.com by these users. The edges between users represent that two users have reviewed the same hotel. The abnormal users are malicious reviewers, and the abnormal labels are sorted and marked by Yelp's proprietary filtering mechanism.

[0119] Amazon is data collected from Amazon China. Its nodes represent users of this website, and the node features are the Chinese reviews of product information by users. The edges between users represent that two users have reviewed the same product. The abnormal users are malicious reviewers, and the abnormal labels are sorted and marked by the website's proprietary filtering mechanism.

[0120] C15 is based on a dataset consisting of real and fake Twitter users. The nodes represent Twitter users, the node features are the basic attributes of users, and the edges represent the mutual following relationships between users. The abnormal users are robot accounts maliciously placed on Twitter.

[0121] Twi20 is based on the Twitter platform. The nodes represent Twitter users, the node features are the basic attributes of users, and the edges represent the mutual following relationships between users. The abnormal users are robot accounts maliciously placed on Twitter.

[0122] II. Baseline Models

[0123] To comprehensively verify the effectiveness of the model of the present invention, this experiment compared it with several different types of baseline models. The baseline models compared with the model of the present invention include DOMINANT, AnomalyDAE, AdONE, GAAN, DONE, GAE, OCGNN, TAM, and RAND.

[0124] The specific introductions of these benchmark methods are as follows:

[0125] DOMINANT is a graph anomaly detection framework based on deep learning. It uses a graph autoencoder to learn node representations, reconstructs the input graph, and then measures the reconstruction error (including topological structure reconstruction error and attribute reconstruction error) to detect abnormal nodes.

[0126] AnomalyDAE is an anomaly detection model designed based on the reconstruction error of a graph autoencoder. The difference from DOMINANT is that it trains the structure and attribute autoencoders separately.

[0127] AdONE is an anomaly detection method based on One-Class Nearest Neighbor. It determines whether a node is an anomaly by calculating the distance between nodes.

[0128] GAAN is a graph anomaly detection method based on the attention mechanism. It uses a graph attention network (GAT) to learn the local and global features of nodes, and enhances the node representation ability by weighted aggregation of neighbor node information; GAAN can effectively capture hidden anomaly patterns in graph data and detect anomaly nodes by calculating anomaly scores.

[0129] GAE is a graph embedding method based on autoencoders. It can perform anomaly detection based on the reconstruction error of node embeddings by learning the low-dimensional representations of nodes.

[0130] OCGN is a model that transforms anomaly detection into a one-class classification problem by learning the feature distribution of normal patterns. OCGNN pays special attention to the representation of normal nodes in the graph. By excluding the interference of anomaly nodes, it can more accurately capture the distribution range of normal patterns and detect anomaly nodes that deviate from these distributions.

[0131] TAM is a method based on the characteristic that normal nodes in graph data usually have high similarity with neighbor nodes, while anomaly nodes deviate from this pattern. This model accurately captures the relationship between normal nodes by truncating affinity optimization and reduces the interference of anomaly nodes on the overall modeling.

[0132] RAND is a graph anomaly detection method that dynamically selects the neighbors of each node through reinforcement learning, so that the neighborhood features of nodes can more accurately reflect normal behavior while reducing the interference of anomaly nodes.

[0133] III. Evaluation Metrics

[0134] The present invention uses two evaluation metrics to evaluate the proposed method:

[0135] AUROC (Area Under the Receiver Operating Characteristic Curve) is a commonly used metric to evaluate the performance of binary classification models. It reflects the ability of the model to distinguish positive and negative samples at different thresholds by calculating the area under the ROC curve. The ROC curve takes the true positive rate (TPR) as the vertical axis and the false positive rate (FPR) as the horizontal axis. The value range of AUROC is from 0 to 100%, and the larger the value, the better the model performance.

[0136] AUCPR (Area Under the Precision-Recall Curve) is a commonly used evaluation metric, especially suitable for imbalanced datasets. By calculating the area under the Precision-Recall Curve (PR curve), AUCPR reflects the model's ability to detect positive samples. The PR curve has the recall rate on the horizontal axis and the precision rate on the vertical axis. The value range of AUCPR is from 0 to 100%, and a larger value indicates better model performance.

[0137] IV. Comparative Experiments

[0138] The experimental results are shown in Table 1 (AUCROC) and Table 2 (AUCPR). The values in parentheses are the metric values, and outside the parentheses are the standard deviations of multiple experiments.

[0139] Table 1

[0140]

[0141] Table 2

[0142]

[0143] It can be clearly observed from Table 1 and Table 2 that the proposed Wild-GAD framework in the present invention significantly outperforms the existing baseline methods in the field of detecting abnormal users in social networks. In the two key metrics of AUCROC and AUCPR for measuring model performance, Wild-GAD achieved significant improvements of 18% and 32% respectively. This performance improvement is mainly due to the ingenious introduction of external graph data in the framework. These data greatly enrich the information source of the model, enabling it to model the normal patterns more comprehensively and deeply, and thus significantly improving the ability to identify and distinguish abnormal patterns.

[0144] In addition, it can also be found that in most cases, adding more external graph data can bring more significant performance improvement to the model compared to adding only one type of external graph data. The logic behind this phenomenon is that as the amount of external graph data increases, the diverse normal patterns that the model encounters during training also increase. This enrichment of diversity enables the model to learn and understand the distribution characteristics of the data more comprehensively, thereby achieving more accurate judgments in the abnormal detection task.

[0145] Based on the same inventive concept, in this embodiment, a social network abnormal user detection system based on external graph data is also provided. The system includes:

[0146] An external graph database, which is used to obtain an external graph dataset based on domain diversity and scale diversity, and obtain at least one target social network graph data; each graph data consists of a node feature matrix and an adjacency matrix.

[0147] A graph data node feature alignment module, which is used to perform data augmentation operations on the original graph data as candidate graphs, generate the node feature matrix and adjacency matrix of the candidate graphs by using the node feature matrix and adjacency matrix of the original graph data, and perform feature space alignment on the node feature matrix of the candidate graphs;

[0148] A graph model pre-training module, which is used to define a spherical space, pre-train a graph model by using at least one original target social network graph data, generate target graph node representations and determine the position coordinates of each node representation in the spherical space; and input the candidate graphs into the pre-trained graph model to generate the node representations of each candidate graph and their position coordinates in the spherical space;

[0149] A candidate graph screening module, which is used to adopt an external graph data selection mechanism based on representative indicators and diversity indicators to screen candidate graphs according to the node representations and their position coordinates in the spherical space;

[0150] A graph model secondary training module, which is used to utilize the selected candidate Figure 2 to secondarily train the graph model;

[0151] An abnormal user detection module, which is used to detect abnormal users in the target social network to be detected by using the graph model trained secondly.

[0152] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment, and the implementation methods of the remaining modules will not be elaborated here. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0153] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The system embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory and running them.

[0154] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the present invention. For those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention.

Claims

1. A method for detecting abnormal users in social networks based on external graph data, characterized in that: The following steps are involved: (1) obtaining an external graph dataset based on domain diversity and scale diversity, and obtaining at least one target social network graph data; Each graph data consists of a node feature matrix and an adjacency matrix; (2) After data augmentation, the original graph data is used as a candidate graph. The node feature matrix and adjacency matrix of the original graph data are used to generate the node feature matrix and adjacency matrix of the candidate graph, and the node feature matrix of the candidate graph is aligned in feature space. (3) defining a spherical space, using at least one original target social network graph data to pre-train a graph model, generating a target graph node representation and determining the position coordinates of each node representation in the spherical space; And, inputting the candidate graph into the pre-trained graph model to generate the representation of each candidate graph node and its position coordinates in the spherical space; (4) Adopting an external graph data selection mechanism based on representativeness and diversity indicators to screen candidate graphs according to node representations and their position coordinates in spherical space; Representative indicators in the external graph data selection mechanism include sphere center similarity deviation based on sphere center stability constraint and distribution similarity deviation based on distribution similarity constraint; The spherical center similarity deviation is the L2 norm between the mean of the target graph node representation and the candidate graph node representation and the mean of the target graph node representation; the distribution similarity deviation is the distance between the probability distribution of the spherical coordinate set of the target graph node representation and the candidate graph node representation and the probability distribution of the spherical coordinate set of the target graph node representation; The diversity index in the external graph data selection mechanism adopts the integral value of the minimum hypersphere energy of different scale radii; The external graph data selection mechanism based on representativeness and diversity indicators refers to calculating the final score of each candidate graph by integrating the representativeness and diversity indicators, and screening a preset number of candidate graphs in order of scores from low to high; (5) The screened candidate graph is used to retrain the graph model, and the retrained graph model is used to detect abnormal users in the target social network to be detected.

2. The method for detecting abnormal users in social networks based on external graph data according to claim 1, characterized in that: The external graph data set includes two or more of e-commerce network graph data, social network graph data, academic citation network graph data, and hyperlink network graph data.

3. The method for detecting abnormal users in social networks based on external graph data according to claim 1, characterized in that: The data enhancement operations include feature masking, node deletion, edge perturbation and subgraph extraction. Each data enhancement operation sets multiple strength parameters, and each data enhancement operation under different strength parameters is traversed to generate a candidate graph.

4. The method for detecting abnormal users in social networks based on external graph data according to claim 1, characterized in that: The method for aligning the feature space of the node feature matrix of the candidate graph is: Determine the type of node features of the candidate graph, where the type is one of a text form feature and a table form feature; If it is a tabular feature, the rule-based conversion method converts it into a textual feature; Otherwise, no processing is required; The text form features are input into the multi-language pre-trained language model, and the text embedding generated by the language model after encoding the text is used as the node feature after feature space alignment.

5. The method for detecting abnormal users in social networks based on external graph data according to claim 1, characterized in that: The method of pre-training a graph model using at least one original target social network graph data includes: Input the node feature matrix and adjacency matrix corresponding to the original target social network graph into the graph model to generate the node representation of the target graph; The training goal is to distribute as many target graph node representations output by the graph model as possible in the spherical space.

6. The method for detecting abnormal users in social networks based on external graph data according to claim 1, characterized in that: The method of using a secondary trained graph model to detect abnormal users in a target social network to be detected includes: Input the node feature matrix and adjacency matrix of the target social network to be detected into the secondary trained graph model, generate node representations and determine the position coordinates of each node representation in the spherical space; If the node representations are distributed in the spherical space, the node corresponding to the node representation is judged to be a normal user, otherwise the node corresponding to the node representation is judged to be an abnormal user.

7. A social network abnormal user detection system based on external graph data, used to implement the method of claim 1; characterized in that: The system comprises: An external graph database, which is used to obtain an external graph data set based on domain diversity and scale diversity, and to obtain at least one target social network graph data; each graph data is composed of a node feature matrix and an adjacency matrix; A graph data node feature alignment module is used to perform data enhancement operations on the original graph data as a candidate graph, generate the node feature matrix and adjacency matrix of the candidate graph using the node feature matrix and adjacency matrix of the original graph data, and perform feature space alignment on the node feature matrix of the candidate graph; A graph model pre-training module is used to define a spherical space, pre-train a graph model using at least one original target social network graph data, generate a target graph node representation and determine the position coordinates of each node representation in the spherical space; and input a candidate graph into the pre-trained graph model to generate each candidate graph node representation and its position coordinates in the spherical space; A candidate graph screening module, which is used to select candidate graphs according to node representations and their position coordinates in spherical space by using an external graph data selection mechanism based on representative indicators and diversity indicators; A graph model secondary training module, which is used to secondary train the graph model using the screened candidate graphs; The abnormal user detection module is used to detect abnormal users in the target social network to be detected by using a secondary trained graph model.