An Online Social Network Harmful Node Identification and Community Partitioning Method and System

By building a directed graph structure of social networks, calculating edge weights and key nodes, and using Louvain algorithm and SVM model, the problem of difficult to identify social network manipulation behavior in the existing technology is solved, and efficient and precise detection of manipulation behavior and community division are achieved.

CN119850357BActive Publication Date: 2025-07-22WUHAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510330195.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-22
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

When detecting social network manipulation behavior, it is difficult to fully capture coordinated unreal behavior, and ignore real-time interaction between accounts in dynamic social networks, resulting in poor detection results.

Method used

By constructing a directed graph structure of social networks, the edge evaluation index is introduced to calculate the edge weight, the key nodes are determined based on the degree centering, the median centering and the Kshell value, and the community division is used using the Louvain algorithm, and the malicious collaboration behavior is determined in combination with the SVM support vector machine model.

Benefits of technology

It improves the ability to identify social network manipulation behavior, can accurately identify key nodes and efficiently divide communities, and realizes intelligent and refined detection and visual analysis of harmful nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850357B_ABST
    Figure CN119850357B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for identifying harmful nodes and community division in an online social network. The method includes: obtaining a dataset of user speeches and interaction information on a social platform; based on the dataset, constructing a directed graph structure of the social network with users as nodes and interaction relationships as edges; based on the directed graph structure of the social network, introducing an evaluation index for edges to calculate and determine the weights of the edges; based on the directed graph structure of the social network, introducing degree centrality, betweenness centrality, and Kshell value to calculate all nodes in the graph and determine the key nodes in the social network propagation; based on the weights of the edges and the key nodes, using the Louvain algorithm to achieve community division with the optimal modularity. The present invention can improve the ability to identify manipulation behaviors in social networks by analyzing the graph structure of dynamic social networks and identifying potential manipulation behaviors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of social network analysis and data mining, and particularly to a method and system for identifying harmful nodes and community division in an online social network based on the Louvain algorithm. Background Art

[0002] Currently, the detection theory of social network manipulation behavior mainly relies on methods such as user behavior analysis, content feature analysis, and social relationship detection. The main detection means include user portrait description, content review, and user association degree calculation (Reference 1). Specifically, (1) User portrait description identifies abnormal activities by analyzing users' interaction behaviors, such as likes, comments, and forwards. This method usually uses machine learning algorithms to model behavior patterns and detect frequent or unnatural interaction behaviors. (2) Content review performs text analysis on the published content and uses natural language processing technology to identify malicious information or fake news. This method evaluates the authenticity and tendency of the content through keyword monitoring and sentiment analysis. (3) User association degree calculation uses the graph structure analysis of the social network to analyze the relationships between accounts and identify key nodes and tightly connected groups.

[0003] However, the above methods have obvious limitations (Reference 2). First, user behavior monitoring often only focuses on the abnormal behaviors of a single account and cannot comprehensively capture the complexity of coordinated inauthentic behavior (CIB) because these behaviors usually involve collaboration between multiple accounts. Second, the content review method may be affected by content diversity and information complexity, resulting in inaccurate detection of fake news and malicious information (Reference 3). In addition, when facing a dynamic social network, user association degree calculation may ignore the real-time interaction relationships between accounts, greatly reducing the monitoring effect.

[0004] [1] Luo Xin. Computational Propaganda: A New Form of Public Opinion in the Era of Artificial Intelligence [J]. Academic Frontiers, 2020(15):13.DOI:10.16619 / j.cnki.rmltxsqy.2020.15.003.

[0005] [2] Magelinski, T., Ng, L.H.,&Carley, K. (2021). A SynchronizedAction Framework for Responsible Detection of Coordination on Social Media.ArXiv, abs / 2105.07454.

[0006] [3] Luis Vargas, Patrick Emami, and Patrick Traynor. 2020. On the Detection of Disinformation Campaign Activity with Network Analysis. In Proceedings of the 2020 ACM SIGSAC Conference on Cloud Computing Security Workshop (CCSW'20). Association for Computing Machinery, New York, NY, USA, 133–146. https: / / doi.org / 10.1145 / 3411495.3421363。 Summary of the Invention

[0007] To overcome the deficiencies of the above prior art, the present invention provides a method and system for identifying harmful nodes and community division in an online social network, which can identify potential manipulation behaviors by analyzing the graph structure of a dynamic social network, and can improve the ability to identify social network manipulation behaviors.

[0008] According to one aspect of the specification of the present invention, there is provided a method for identifying harmful nodes and community division in an online social network, including:

[0009] Obtain a dataset of user speeches and interaction information on a social platform;

[0010] Based on the dataset, construct a directed graph structure of the social network with users as nodes and interaction relationships as edges;

[0011] Based on the directed graph structure of the social network, introduce an evaluation index of the edge for calculation to determine the weight of the edge;

[0012] Based on the directed graph structure of the social network, introduce degree centrality, betweenness centrality, and K-shell value for full-graph node calculation to determine the key nodes in social network propagation;

[0013] Based on the weight of the edge and the key nodes, use the Louvain algorithm to achieve community division with the optimal modularity.

[0014] As a further technical solution, the method further includes:

[0015] Based on the trained SVM support vector machine model, determine a group of users with strong internal interactions to determine whether there is malicious promotion in their collaboration.

[0016] As a further technical solution, based on the social network directed graph structure, evaluation indicators of edges are introduced for calculation to determine the weights of the edges, including:

[0017] Calculate the text similarity based on the TF-IDF method;

[0018] Calculate the active time similarity based on the time normalization mechanism;

[0019] Calculate the forwarding similarity based on the Jaccard similarity coefficient;

[0020] Based on the text similarity, active time similarity, and forwarding similarity, obtain the weight of the edge by weighting.

[0021] As a further technical solution, based on the social network directed graph structure, degree centrality, betweenness centrality, and K-shell value are introduced for calculating all nodes in the graph to determine the key nodes in the social network propagation, including:

[0022] Construct the adjacency matrix of the social network directed graph structure;

[0023] Based on the adjacency matrix, calculate the degree centrality, betweenness centrality, and K-shell value of the nodes in the social network directed graph structure respectively;

[0024] Based on the calculated degree centrality, betweenness centrality, and K-shell value, combine the PageRank algorithm to propagate the importance of the nodes to determine the key nodes that promote the development of topics in the social network propagation.

[0025] As a further technical solution, when using the Louvain algorithm to achieve community division with the optimal modularity based on the edge weights and key nodes, perform the following maximization solution of the modularity Q:

[0026] ,

[0027] where m represents the total number of all edges in the graph; represents the actual edge weight between node i and node j; represents the indicator function, which takes the value of 1 when node and node belong to the same community, otherwise 0; represents node 's degree centrality; represents node 's degree centrality.

[0028] As a further technical solution, the method further includes:

[0029] The degree centrality, betweenness centrality, or K-shell value is in the topθ% Screen the key points to the key node set, and screen the remaining nodes to the non-key node set;

[0030] Construct the key point feature vector , the non-key node feature vector , the edge similarity feature vector , and the time development feature vector , combine the said feature vectors and to obtain the final discrimination vector for each community :

[0031] ,

[0032] Use the final discrimination vector to train the SVM support vector machine model to determine whether there is malicious collaboration behavior in the community.

[0033] According to one aspect of the specification of the present invention, there is provided an online social network harmful node identification and community division system, including:

[0034] A data acquisition module for acquiring a dataset of social platform user speeches and interaction information;

[0035] A graph construction module for constructing a directed graph structure of a social network with users as nodes and interaction relationships as edges based on the dataset;

[0036] A weight determination module for calculating based on the directed graph structure of the social network and introducing an evaluation index of the edge to determine the weight of the edge;

[0037] A key node determination module for calculating all nodes in the graph based on the directed graph structure of the social network and introducing degree centrality, betweenness centrality and Kshell value to determine the key nodes in the social network propagation;

[0038] A community division module for realizing community division with the optimal modularity based on the edge weight and the key nodes using the Louvain algorithm.

[0039] As a further technical solution, it further includes:

[0040] A collaboration behavior determination module for determining a group of users with strong internal interaction based on the trained SVM support vector machine model to determine whether there is malicious promotion in their collaboration.

[0041] According to one aspect of the specification of the present invention, there is provided an online social network harmful node identification and community division device, including: one or more processors; a memory for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the steps of the online social network harmful node identification and community division method described above.

[0042] According to one aspect of the specification of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions that cause a computer to execute the steps of the online social network harmful node identification and community division method described above.

[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0044] The present invention identifies potential manipulation behaviors by analyzing the graph structure of a dynamic social network. The method provided focuses on the relationships and interaction patterns between accounts in the social network, and reveals the potential network of manipulation behaviors by identifying key nodes and community structures. The graph structure-based method provided by the present invention can integrate user behaviors, content features, and network relationships, providing stronger support for identifying and dealing with complex manipulation behaviors, thereby enhancing the ability to identify manipulation behaviors in social networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings used in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0046] Figure 1 It is a schematic flowchart of an online social network harmful node identification and community division method provided by an embodiment of the present invention.

[0047] Figure 2 It is a schematic structural diagram of an online social network harmful node identification and community division system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The present invention provides a method for identifying harmful user nodes and community division in an online social network: First, collect the user speech information and user interaction relationships of a specific social platform, and construct a directed graph structure of the social network with users as nodes and the interaction relationships between users as edges; Second, quantify and score information such as text similarity, active time similarity, and forwarding similarity as the weight values of the adjacent edges between nodes; Third, statistically calculate indicators such as node degree, Kshell, and betweenness in the social network graph structure, and calculate the key nodes affecting community development through the PageRank algorithm; Finally, apply the Louvain algorithm to divide the graph structure into communities, calculate the modularity to evaluate the effect of community division, and visualize and make the community division results interactive for easy viewing of the composition of the social network community.

[0049] By constructing a social network graph with multi-dimensional similarity indicators and applying the Louvain algorithm, the present invention realizes the accurate identification of key nodes and the efficient division of communities from the perspective of the social network propagation structure, achieves the intelligent, refined detection and visual analysis of harmful nodes, and has the advantages of high efficiency, accuracy, automation, etc.

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. In addition, the technical features in each embodiment or a single embodiment provided by the present invention can be combined with each other arbitrarily to form a new technical solution. This combination is not restricted by the order of steps and / or the structure composition mode, but must be based on the fact that those of ordinary skill in the art can implement it. When the combination of technical solutions conflicts or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0051] Please refer to Figure 1 , the embodiments of the present invention provide a method for identifying harmful nodes and community division in an online social network, including:

[0052] Step 1, obtain a dataset of social platform user speech and interaction information. In this step, first collect the user speech and interaction information of a specific topic discussion of a certain social user platform to obtain a preliminary dataset; then preprocess the preliminary dataset using means such as missing value filling and non-text character filtering to obtain the final dataset.

[0053] In step 1, the constructed data structure includes the following fields:

[0054]

[0055] Among them, i represents the index, u represents the user identifier, t represents the name of the user who replies or interacts, represents the unique number of the user who replies or interacts, s represents the name of the user being replied to or interacted with, represents the unique number of the name of the user being replied to or interacted with, r represents the name of the user who creates the topic tag, represents the unique number of the user who creates the topic tag, p represents the time when the reply or interaction event occurs, o represents the original text being replied to or interacted with, represents the text of the reply or interaction, represents the interval between the occurrence of the reply or interaction event and the creation event of the topic tag, d represents the cascade depth of the reply or interaction event relative to the creation event of the topic tag, represents the unique number of the event, represents the forwarding link of the social user platform, represents the cascade width of the reply or interaction event.

[0056] Step 2: Based on the dataset, construct a directed graph structure of the social network with users as nodes and interaction relationships as edges. This step maps user interactions to a weighted directed graph structure based on the preliminary dataset. This graph structure uses users as nodes and interaction relationship information as edges to construct directed propagation relationship edges.

[0057] In Step 2, the constructed directed graph structure of the social network is as follows:

[0058]

[0059] Among them, V represents the set of nodes. The nodes represent users in the social network and include user identifiers and the corresponding user names . E represents the set of edges. The edges represent the interaction relationships between users. The edge direction of the directed graph is from the user who replies or interacts to the user being replied to or interacted with.

[0060] Step 3: Based on the directed graph structure of the social network, introduce evaluation indicators for edges to calculate and determine the weights of the edges. This step first introduces the TF-IDF method for text similarity analysis, introduces the time normalization mechanism for active time similarity analysis, and introduces the Jaccard similarity coefficient for forwarding similarity analysis; then uses these indicators as evaluation indicators for edges, evaluates the interaction behaviors between nodes based on specific weights, and takes the evaluation results as the total weights of the edges.

[0061] Step 3 further includes the following steps:

[0062] Step 3.1: Calculate the TF-IDF vector: For the original text content and the forwarded text (i.e., the text of the reply or interaction) Perform word segmentation and convert it into a vector representation using TF-IDF:

[0063]

[0064] Calculate the text similarity using cosine similarity :

[0065]

[0066] Step 3.2, Active time similarity analysis: Assume that the edge E(i, j) represents the user 's original text publication time and the user 's forwarding time , then the active time similarity is calculated as follows:

[0067]

[0068] where max(p)) and min(p) are the maximum and minimum publication times in the dataset respectively, ensuring the normalization of the time difference.

[0069] Step 3.3, Forwarding similarity analysis: Calculate the similarity between all the forwarding user sets and using the Jaccard similarity coefficient : :

[0070]

[0071] where represents the number of intersections of the forwarding user sets, represents the number of unions of the forwarding user sets.

[0072] Step 3.4, Comprehensive weight calculation:

[0073] The weight of the edge is obtained by weighting the three similarity values according to the weights , , :

[0074]

[0075] where , and The preset weight coefficients for text similarity, active time similarity, and repost similarity respectively.

[0076] Step 4: Based on the directed graph structure of the social network, introduce degree centrality, betweenness centrality, and K-shell value to calculate all graph nodes and determine the key nodes in the social network propagation. This step first extracts degree, betweenness, and K-shell metrics for all graph nodes based on the above graph structure to determine the key nodes in the social network propagation; then analyzes the graph structure through the PageRank algorithm to calculate the key nodes that promote the development of the topic.

[0077] The specific implementation of Step 4 includes the following sub-steps:

[0078] Step 4.1: Construct the adjacency matrix in the directed graph of the social network.

[0079]

[0080] Among them, A is an N×N matrix representing the connection relationship between nodes in the network, and N is the total number of nodes.

[0081] Step 4.2: Calculate the degree centrality of nodes based on the weighted directed graph structure: For node , the calculation formula for degree centrality is:

[0082]

[0083] Among them, represents the degree of node , that is, the number of nodes connected to it.

[0084] Step 4.3: Calculate the betweenness centrality of nodes based on the weighted directed graph structure: For node , the calculation formula for betweenness centrality is:

[0085]

[0086] Among them, represents the total number of shortest paths from node to node , represents the total number of shortest paths from node to node and passing through node .

[0087] Step 4.4: Construct a KShell decomposition (k-shell decomposition) and solve the KShell value of each node. Given a graph G=(V,E), the node The Kshell value of is

[0088]

[0089] It can be solved by iterating over the directed graph G as follows:

[0090] Initialize k = 1;

[0091] Iteratively remove nodes For the removed nodes Set its Kshell value to k;

[0092] Increase the value of k and repeat the above steps until all nodes are removed under the convergence condition.

[0093] Step 4.5, Propagate the importance of nodes based on the PageRank propagation formula.

[0094]

[0095] Where:

[0096] represents the degree centrality PageRank value of node at the (t + 1)-th iteration, represents the degree centrality PageRank value of node at the t-th iteration, represents the number of edges starting from as the starting point.

[0097] represents the betweenness centrality PageRank value of node at the (t + 1)-th iteration. represents the betweenness centrality PageRank value of node at the t-th iteration.

[0098] represents the Kshell PageRank value of node at the (t + 1)-th iteration, represents the Kshell PageRank value of node at the t-th iteration.

[0099] represents the degree centrality PageRank value of node at the initial stage. represents the betweenness centrality PageRank value of node at the initial stage. represents the Kshell PageRank value of node The Kshell PageRank value at the beginning.

[0100] For the PageRank value, continuously iterate until the convergence condition:

[0101] 。

[0102] Step 5: Based on the weights of the edges and the key nodes, use the Louvain algorithm to achieve community division with the optimal modularity. This step is based on the above graph structure and the analysis results of the key nodes. Use the Louvain algorithm to perform community division on the graph structure, and continuously iterate until the modularity index meets the convergence to achieve the discovery of user groups with strong internal interactions.

[0103] In step 5, in the case of achieving community division with the optimal modularity through the Louvain algorithm, to achieve the maximization solution of the modularity Q:

[0104]

[0105] Among them, m represents the total number of all edges in the graph, represents the actual edge weight between node i and node j, is an indicator function. When nodes and node belong to the same community, it takes the value of 1; otherwise, it is 0.

[0106] After the embodiment of the present invention performs community division and discovers user groups with strong internal interactions, it further includes:

[0107] Step 6: Based on the trained SVM support vector machine model, determine the user groups with strong internal interactions to determine whether there is malicious promotion in their collaboration. In this step, for the user groups with strong internal interactions, based on the key node indicators and some labeled tags, train the SVM support vector machine model to determine the user groups with strong internal interactions to determine whether there is malicious promotion in their collaboration. The labels here can be manually marked.

[0108] In step 6, when training the SVM support vector machine model, first screen to obtain the set of key nodes K where the Keshell value, degree centrality, or betweenness centrality is among the top θ% , and the set of non-key nodes N, and then execute the following sub-steps:

[0109] Step 6.1: Construct the key point feature vector , the non-key node feature vector , the edge similarity feature vector , and the time development feature vector 。

[0110]

[0111] Among them, represents the average value of feature m in the set of key nodes, represents the variance of feature m in the set of key nodes, represents the maximum value of feature m in the set of key nodes, represents the minimum value of feature m in the set of key nodes.

[0112] Step 6.2, the feature vector of non-key nodes also includes statistical features of Kshell value, degree centrality, and betweenness centrality, which are defined as follows:

[0113]

[0114] Among them, respectively represent the average value, variance, maximum value, and minimum value of the feature in the set of non-key nodes.

[0115] Step 6.3, the similarity feature vector of edges includes content similarity , time similarity and forwarding similarity of the statistical features, which are defined as follows:

[0116]

[0117] Among them, respectively represent the average value, variance, maximum value, and minimum value of the similarity feature m in the set of edges.

[0118] Step 6.4, count the number of replies and interaction events in the community hourly to obtain a time series , where represents the number of interactions in the i-th hour. The time development feature vector obtained by fitting with a 10th-order polynomial is expressed as: is expressed as:

[0119]

[0120] Step 6.5, combine the above feature vectors and to obtain the final discrimination vector X of each community:

[0121]

[0122] Use the above final discrimination vector X to train an SVM support vector machine model to determine whether there is malicious collaboration behavior in the community.

[0123] The implementation basis of each embodiment of the present invention is achieved through programmed processing by a device with processor functions. Therefore, in engineering practice, the technical solutions and functions of each embodiment of the present invention are encapsulated into various modules. Based on this actual situation, on the basis of the above embodiments, an embodiment of the present invention provides an online social network harmful node identification and community division system, which is used to execute an online social network harmful node identification and community division method in the above method embodiments.

[0124] See Figure 2 , the system includes: a data acquisition module, which is used to acquire a dataset of social platform user speech and interaction information; a graph construction module, which is used to construct a directed graph structure of the social network with users as nodes and interaction relationships as edges based on the dataset; a weight determination module, which is used to calculate by introducing an evaluation index of the edge based on the directed graph structure of the social network and determine the weight of the edge; a key node determination module, which is used to calculate all graph nodes by introducing degree centrality, betweenness centrality and Kshell value based on the directed graph structure of the social network and determine the key nodes in the social network propagation; a community division module, which is used to implement community division with the optimal modularity based on the edge weight and key nodes using the Louvain algorithm.

[0125] An online social network harmful node identification and community division system provided by an embodiment of the present invention aims at the limitations of existing social network manipulation behavior detection methods and adopts Figure 2 several modules among them to identify potential manipulation behaviors by analyzing the graph structure of the dynamic social network, which can improve the ability to identify social network manipulation behaviors.

[0126] It should be noted that the system embodiment provided by the present invention, in addition to being used to implement the method in the above method embodiment, is also used to implement the methods in other method embodiments provided by the present invention. The difference is only in setting corresponding function modules, and its principle is basically the same as that of the above system embodiment provided by the present invention. As long as those skilled in the art, on the basis of the above system embodiment, refer to the specific technical solutions in other method embodiments, obtain corresponding technical means by combining technical features, and the technical solutions composed of these technical means, and on the premise of ensuring the practicability of the technical solution, improve the modules in the above system embodiment to obtain corresponding system class embodiments for implementing the methods in other method class embodiments. For example:

[0127] Based on the content of the above system embodiment, as a preferred embodiment, an online social network harmful node identification and community division system provided by an embodiment of the present invention further includes:

[0128] A collaborative behavior determination module, which is used to determine the group of users with strong internal interaction based on the trained SVM support vector machine model to determine whether there is malicious promotion in their collaboration.

[0129] Further, the collaborative behavior determination module is used to execute the following instructions:

[0130] Filter the key points with degree centrality, betweenness centrality or Kshell value in the top θ% to the key node set, and filter the remaining nodes to the non-key node set;

[0131] Construct the key point feature vector , the non-key node feature vector , the edge similarity feature vector , and the time development feature vector , and combine the said feature vectors and to obtain the final discrimination vector of each community :

[0132] ,

[0133] Use the final discrimination vector to train the SVM support vector machine model to judge whether there is malicious collaboration behavior in the community.

[0134] Based on the content of the above system embodiment, as a preferred embodiment, in the online social network harmful node identification and community division system provided in the embodiment of the present invention, the weight determination module is further used to execute the following instructions:

[0135] Calculate the text similarity based on the TF-IDF method;

[0136] Calculate the active time similarity based on the time normalization mechanism;

[0137] Calculate the forwarding similarity based on the Jaccard similarity coefficient;

[0138] Based on the text similarity, active time similarity and forwarding similarity, weight to obtain the weight of the edge.

[0139] Based on the content of the above system embodiment, as a preferred embodiment, in the online social network harmful node identification and community division system provided in the embodiment of the present invention, the key node determination module is further used to execute the following instructions:

[0140] Construct the adjacency matrix of the directed graph structure of the social network;

[0141] Based on the adjacency matrix, calculate the degree centrality, betweenness centrality, and Kshell value of the nodes in the directed graph structure of the social network respectively;

[0142] Based on the calculated degree centrality, betweenness centrality, and Kshell value, combine with the PageRank algorithm to propagate the importance of the nodes, and determine the key nodes that promote the development of topics in the social network propagation.

[0143] Based on the content of the above system embodiments, as a preferred embodiment, in the embodiment of the present invention, an online social network harmful node identification and community division system is provided, and the community division module is further used to execute the following instructions:

[0144] When using the Louvain algorithm to achieve community division with the optimal modularity based on the weight of the edge and the key nodes, perform the maximization solution of the following modularity Q:

[0145] ,

[0146] where m represents the total number of all edges in the graph; represents the actual edge weight between node i and node j; represents the indicator function, which takes the value of 1 when node and node belong to the same community, otherwise 0; represents node 's degree centrality; represents node 's degree centrality.

[0147] Based on the same inventive concept as the above embodiments, the embodiment of the present invention also provides an online social network harmful node identification and community division device, including: one or more processors; a memory for storing one or more programs, when one or more of the said programs are executed by one or more of the said processors, enabling one or more of the said processors to implement the steps of the said method for identifying harmful nodes and dividing communities in an online social network.

[0148] In an embodiment of the present invention, the memory may be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), etc., or may also be a volatile memory, such as a random-access memory (RAM). The memory is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in the embodiment of the present invention may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0149] In an embodiment of the present invention, the processor may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0150] Based on the same inventive concept as the above embodiments, the embodiments of the present invention further provide a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions cause the computer to execute the steps of the method for identifying harmful nodes and community partitioning in an online social network.

[0151] When the above computer instructions are implemented in the form of software functional units and sold or used as independent products, they are stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, is embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (a personal computer, a server, or a network device) to execute all or part of the steps of the methods described in the various method embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random-access memories (RAMs), magnetic disks, or optical discs, and various media for storing program codes.

[0152] In summary, through the construction of a multi-dimensional similarity index social network graph and the application of the Louvain algorithm, the present invention can accurately identify key nodes and efficiently partition communities from the perspective of the social network propagation structure, realizing the intelligent, refined detection and visualization analysis of harmful nodes, and having the advantages of high efficiency, accuracy, automation, etc.

[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. An online social network malicious collaboration behavior recognition and community division method, characterized in that Including: Obtain a dataset of user speeches and interaction information on the social platform; Based on the dataset, construct a directed graph structure of the social network with users as nodes and interaction relationships as edges; Based on the directed graph structure of the social network, introduce an evaluation index of the edge for calculation to determine the weight of the edge; Based on the directed graph structure of the social network, introduce degree centrality, betweenness centrality, and K-shell value for full-graph node calculation to determine the key nodes in social network propagation; Based on the weight of the edge and the key nodes, use the Louvain algorithm to achieve community division with the optimal modularity; The method further includes: Based on the trained SVM support vector machine model, determine the group of users with strong internal interactions to determine whether there is malicious promotion in their collaboration. Further including: Select the key points with degree centrality, betweenness centrality or Kshell value in the top θ% to the key node set, and the remaining nodes to the non-key node set; Construct key point feature vectors , non-key node feature vectors , edge similarity feature vectors , and time development feature vectors , the feature vectors , all contain statistical features of Kshell value, degree centrality, and betweenness centrality. The feature vectors contain statistical features of content similarity, time similarity, and forwarding similarity. The feature vectors contain the coefficients of the time series obtained by statistically counting the number of reply and interaction events in the community. Combine the feature vectors , , and to obtain the final discrimination vector of each community , , Use the final discrimination vector to train the SVM support vector machine model to determine whether there is malicious collaboration behavior in the community.

2. The method for identifying malicious collaboration behavior and community division in an online social network according to claim 1, wherein Based on the directed graph structure of the social network, introduce an evaluation index of the edge for calculation to determine the weight of the edge, including: Calculate the text similarity based on the TF-IDF method; Calculate the active time similarity based on the time normalization mechanism; Calculate the forwarding similarity based on the Jaccard similarity coefficient; Based on the text similarity, active time similarity, and forwarding similarity, obtain the weight of the edge by weighting.

3. The method for identifying malicious collaborative behaviors and community division in an online social network according to claim 1, characterized in that Based on the directed graph structure of the social network, introduce degree centrality, betweenness centrality, and K-shell value for full-graph node calculation to determine the key nodes in social network propagation, including: Construct the adjacency matrix of the directed graph structure of the social network; Based on the adjacency matrix, calculate the degree centrality, betweenness centrality, and K-shell value of the nodes in the directed graph structure of the social network respectively; Based on the calculated degree centrality, betweenness centrality, and K-shell value, combine the PageRank algorithm to propagate the importance of the nodes to determine the key nodes that promote the development of topics in social network propagation.

4. An online social network malicious collaboration behavior recognition and community division method according to claim 1, characterized in that, When using the Louvain algorithm to achieve community division with the optimal modularity based on the weight of the edge and the key nodes, perform the following maximization solution of the modularity Q: , Among them, m represents the total number of all edges in the directed graph structure of the social network; represents the node and the node the actual edge weight between them; represents the indicator function, which takes the value of 1 when the node and the node belong to the same community, and 0 otherwise; represents the degree centrality of the node ; represents the degree centrality of the node ; 5. An online social network malicious collaboration behavior recognition and community division system, characterized in that, Including: A data acquisition module for obtaining a dataset of user speeches and interaction information on the social platform; A graph construction module for constructing a directed graph structure of the social network with users as nodes and interaction relationships as edges based on the dataset; A weight determination module for introducing an evaluation index of the edge for calculation based on the directed graph structure of the social network to determine the weight of the edge; A key node determination module for introducing degree centrality, betweenness centrality, and K-shell value for full-graph node calculation based on the directed graph structure of the social network to determine the key nodes in social network propagation; A community division module for using the Louvain algorithm to achieve community division with the optimal modularity based on the weight of the edge and the key nodes; The collaborative behavior determination module is used to determine the group of users with strong internal interactions based on the trained SVM support vector machine model to determine whether there is any malicious promotion in their collaboration. It further includes: screening the key points with degree centrality, betweenness centrality or Kshell value in the top θ% to the key node set, and screening the remaining nodes to the non-key node set; constructing the key point feature vector , the non-key node feature vector , the edge similarity feature vector , and the time development feature vector . The said feature vectors 、 all contain the statistical features of Kshell value, degree centrality and betweenness centrality. The said feature vector contains the statistical features of content similarity, time similarity and forwarding similarity. The said feature vector contains the coefficients of the time series obtained by counting the number of reply and interaction events in the community. Combining the said feature vectors 、 、 and to obtain the final discrimination vector for each community. , Use the final discrimination vector to train the SVM support vector machine model to determine whether there is malicious collaboration behavior in the community.

6. An online social network malicious collaboration behavior recognition and community division device, characterized in that, Including: One or more processors; A memory for storing one or more programs, which when executed by one or more of the processors, cause the one or more processors to implement the steps of a method for identifying malicious collaboration behaviors and community partitioning in an online social network according to any one of claims 1 to 4.

7. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions that cause the computer to execute the steps of a method for identifying malicious collaboration behaviors and community partitioning in an online social network according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Abnormal social account identification method and device, computer equipment and storage medium

    CN111708823A