Key user identification method and system in social network
By combining social network graph models and graph theory algorithms, global and local centrality indicators, and dynamically adjusting weights, we can identify key users in social networks, solving the problem of inaccurate identification in existing technologies and achieving stable and efficient identification in dynamic environments.
Patent Information
- Application Number
- CN202510780825.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Existing technologies have difficulty in dynamically adapting to topological evolution in social networks, resulting in inaccurate identification of key users. Especially when community structure changes or malicious interference occurs, key nodes that transmit information across communities are difficult to detect.
It adopts social network graph model construction and graph theory algorithms, combines global and local centrality indicators, and dynamically adjusts weights through hierarchical clustering and bridge node analysis to generate global and local scores, screen key users, and detect malicious behavior through time series characteristics.
It improves the comprehensiveness and accuracy of key user identification, reduces the misjudgment rate, can maintain stable recognition performance under changes in social network structure and interference from malicious behavior, and accurately captures the structural characteristics of nodes within the community.
Smart Images

Figure CN120632226A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of user identification, and in particular to a method and system for identifying key users in a social network. Background Art
[0002] Identifying key users in social networks is a crucial branch of graph theory and complex network research, with its roots in the exploration of the complex structural properties and information dissemination patterns of social networks. With the rapid development of internet technology, social networks have evolved into extremely large-scale, heterogeneous graphs comprising billions of nodes. Users, through complex networks of relationships, form a dynamic system of information dissemination, opinion exchange, and behavioral influence. Within this context, identifying key users has become a core issue for understanding network topology, optimizing information dissemination efficiency, and predicting group behavior.
[0003] Existing technologies often use fixed thresholds or simple clustering algorithms to divide communities, failing to dynamically adapt to evolving social network topologies. When community structures split or merge due to changes in user behavior or malicious interference, traditional methods struggle to adjust community boundaries in a timely manner, leading to underestimation or misjudgment of the local importance of key players. Existing technologies often separate the strength of inter-community connections from the role of bridging nodes when calculating node importance, analyzing only internal metrics within each community. This makes it difficult to identify key nodes that transmit information or coordinate actions across communities.
[0004] In order to solve the above-mentioned defects in the prior art, the present technical solution proposes a method and system for identifying key users in a social network. Summary of the Invention
[0005] The present invention provides a method and system for identifying key users in a social network, so as to solve the defects in the prior art.
[0006] In one aspect, the present invention provides a method for identifying key users in a social network, comprising: Construct a social network graph model based on social network data; social network data includes user nodes; Combined with the social network graph model, the importance of user nodes is quantified through graph theory algorithms to calculate basic centrality indicators; Divide social network data into multiple communities and identify the local roles of user nodes in multiple communities; Calculate the local centrality index of local roles within each community; Perform weighted summation on the basic centrality index and local centrality index to generate global scores and local scores; rank the global scores and local scores respectively, and filter key user nodes based on the score rankings to output the key user node set; Perform score mutation detection on key user nodes in the key user node set to identify malicious behavior, and output the final key user node set based on the identification results.
[0007] According to a method for identifying key users in a social network provided by the present invention, the steps of constructing a social network graph model include: Extract user nodes and relationship edges from social network data and create user node sets and relationship edge sets; Construct an adjacency matrix based on the user node set and the relationship edge set; Generate a graph structure based on the adjacency matrix, user node set and edge set.
[0008] According to a method for identifying key users in a social network provided by the present invention, the step of calculating a basic centrality index includes: According to the graph structure, configure the basic indicator algorithm library and determine the indicator weight distribution rules; Combined with the basic indicator algorithm library, the indicator value of each user node is calculated by traversing the graph structure; Normalize the index values to a unified dimension and output basic centrality index and standardized basic score.
[0009] According to a method for identifying key users in a social network provided by the present invention, the step of identifying local roles includes: Use hierarchical clustering algorithms to divide social network data into multiple communities; Combine multiple communities and user nodes to output the local user nodes within each community and the connection strength between communities; Analyze the node location characteristics of local user nodes within their community; Determine the bridge nodes between communities based on the strength of the connections; Based on node location characteristics and bridge nodes, local roles within multiple communities are determined.
[0010] According to a method for identifying key users in a social network provided by the present invention, the steps of generating a global score and a local score include: Calculate the local centrality index of local roles within each community independently; Adjust the weight of local centrality index according to node location characteristics; Determine the role type of the local role based on the bridge node and assign a weight coefficient based on the role type; Combined with the weight coefficients, a weighted sum is performed to generate a local score; According to the standardized basic score and local score, a weighted fusion function is constructed and the weight parameters are output; The global score is calculated based on the weight parameters, the standardized basic score and the data score.
[0011] According to a method for identifying key users in a social network provided by the present invention, the step of screening a set of key user nodes includes: According to the global score and local score, perform dual-dimensional descending sorting and output the sorting results; According to the sorting results, calculate the comprehensive ranking of local roles; The top 10% of the comprehensive ranking is used as the pre-selected set of key user nodes, low-scoring user nodes are filtered out, and the key user node set is output.
[0012] According to a method for identifying key users in a social network provided by the present invention, the step of performing score mutation detection on a set of key user nodes includes: Based on social network data, extract the behavioral time series features of all key user subsets in the key user node set; Combined with behavioral temporal features, a change point detection algorithm is used to identify abnormal mutation scores; According to the abnormal mutation score, mark the relevant mutation key user nodes and output the marking information; Based on the tag information and log data in the social network data, the rule engine is used to verify and output abnormal key user nodes. Eliminate abnormal key user nodes in the key user node set and output the final key user node set.
[0013] The present invention also provides a key user identification system in a social network, comprising: Graph data modeling module, used to build a social network graph model based on social network data; Centrality calculation engine, which is used to combine social network graph models, quantify the importance of user nodes through graph theory algorithms, and calculate basic centrality indicators; Community division module, used to divide social network data into multiple communities; Role identification module, used to identify the local roles of user nodes in multiple communities; The local index calculation module is used to calculate the local centrality index of local roles in each community; The score calculation and fusion module is used to calculate the global score and local score based on the basic centrality index and the local centrality index; and to perform score ranking, screen key user nodes, and output the key user node set; The anomaly detection module performs score mutation detection on the key user nodes in the key user node set, and outputs the final key user node set based on the detection results.
[0014] According to a key user identification system in a social network provided by the present invention, the centrality calculation engine includes an algorithm selection unit and an indicator calculation unit; the algorithm selection unit is used to combine the social network graph model and select different quantitative algorithms according to different centrality indicators; the indicator calculation unit is used to calculate multiple centrality indicators of user nodes according to different quantitative algorithms.
[0015] According to a key user identification system in a social network provided by the present invention, the anomaly detection module includes a feature extraction unit, an anomaly identification unit, an anomaly marking unit, and an anomaly feedback unit; the feature extraction unit is used to extract behavioral time series features of all key user subsets in a key user node set; the anomaly identification unit is used to combine the behavioral time series features and use a change point detection algorithm to identify anomaly mutation scores; the anomaly marking unit is used to mark relevant mutation key user nodes based on the anomaly mutation scores and output marking information; the anomaly feedback unit is used to exclude mutation key user nodes in the key user node set based on the marking information and output a final key user node set.
[0016] The present invention provides a method and system for identifying key users in a social network. By integrating a dual-dimensional analysis of global centrality indicators and local role characteristics, it not only retains the advantages of traditional global influence assessment, but also accurately captures the structural characteristics of nodes within the community, effectively solving the problem of insufficient identification of cross-community bridging nodes and local core users by traditional methods, and significantly improving the comprehensiveness and accuracy of key user identification. The weights of local indicators are dynamically adjusted through node position characteristics, combined with the role type allocation coefficient, to achieve scenario-adaptive optimization of the indicator system. This enables the system to maintain stable recognition performance in scenarios such as changes in social network structure and interference from malicious behavior. Through dual verification of time series feature analysis and rule engine verification, abnormal states disguised as high-influence users can be effectively identified. Compared with traditional single-dimensional detection methods, the misjudgment rate is greatly reduced, significantly improving the reliability of the results. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 This is a flow chart of a method for identifying key users in a social network provided by an embodiment of the present invention; Figure 2 This is a flow chart of score mutation detection for a key user node set provided by an embodiment of the present invention; Figure 3This is a structural diagram of a key user identification system in a social network provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0020] Example 1: The following combination Figure 1-Figure 3 The present invention describes a method and system for identifying key users in a social network.
[0021] like Figure 1-Figure 2 As shown, an embodiment of the present invention provides a method for identifying key users in a social network, including: Based on the social network data, a social network graph model is constructed. The social network data includes user nodes. The steps of constructing the social network graph model include: Extract user nodes and relationship edges from social network data to create sets of user nodes and relationship edges. User nodes contain unique identifiers and attribute fields (such as registration time and location), while relationship edges contain interaction types (follow / followed, like / comment), timestamps, and weights. Separate individual users and their associated behaviors from the data to form sets of users and relationships.
[0022] An adjacency matrix is constructed based on a set of user nodes and a set of relationship edges. A sparse matrix storage structure is used to process large-scale network data, and weighted relationship edges are normalized. Connections between users are recorded in a tabular format, with matrix elements representing the intensity of interaction between users (e.g., number of mutual friends or number of messages sent).
[0023] Generate a graph structure based on the adjacency matrix, user node set, and edge set. Visualize and dynamically update the graph structure using a graph database (such as Neo4j) or an in-memory computing framework (such as NetworkX).
[0024] Combined with the social network graph model, the importance of user nodes is quantified through graph theory algorithms to calculate basic centrality indicators. The steps for calculating basic centrality indicators include: Based on the graph structure, a basic indicator algorithm library is configured and the indicator weight assignment rules are determined. The basic indicator algorithm library includes core indicators such as degree centrality (number of direct connections), closeness centrality (average distance to other nodes), and betweenness centrality (path control capability). Weight assignment is dynamically calibrated using the Analytic Hierarchy Process (AHP) combined with the entropy weight method.
[0025] Degree centrality is used to measure local influence based on the number of neighbors and is calculated as follows:
[0026] Where C D (v) represents the degree centrality score of node v, deg(v) represents the number of direct neighbors of node v, that is, the number of relationship edges. n is the total number of user nodes in the social network, and (n-1) is the normalization factor, indicating that a node may have at most n-1 neighbors.
[0027] Closeness centrality is used to reflect the average closeness of node v to other user nodes. The formula is expressed as:
[0028] Where C C (v) represents the closeness centrality score of node v, V represents the set of all user nodes in the social network, and d(v,u) represents the shortest path length from node v to node u.
[0029] Betweenness centrality is used to measure the intermediary role of node v in the shortest path of all user-node pairs. The higher the value, the closer the node v is to the key hub of the social network. The formula is expressed as:
[0030] Where C B (v) is the betweenness centrality score of node v, which represents the sum of the mediating effects of v in all node pairs on the shortest path (the larger the value, the stronger the mediating effect of v). ∑ s,v,t Represents the cumulative mediating effect of v for all possible node pairs (s, t). Traverse all possible node triplets (s, v, t), where s and t are any two different nodes in the network other than v (i.e., s, t∈V\{v}), and s=t (avoiding self-loop paths). σ st is the total number of shortest paths from node s to t (regardless of whether the path passes through v). st (v) represents the number of paths that pass through v in the shortest path from node s to t.
[0031] Eigenvector centrality is usually also included to measure the global influence of user nodes, which takes into account the topological structure of the network. The formula is expressed as:
[0032] The above formula is solved by the power iteration method. A represents the adjacency matrix. If A ij = 1, indicating that user nodes i and j are connected. x represents the eigenvector centrality score vector, where each element corresponds to a user node score. λ is the maximum value of the corresponding eigenvector centrality score. The score xv of node v is proportional to the sum of the scores of its neighboring nodes, meaning that nodes connected to highly influential neighbors have higher scores.
[0033] Combined with the basic indicator algorithm library, the indicator value of each user node is calculated by traversing the graph structure. The breadth-first search (BFS) algorithm is used to optimize traversal efficiency and a distributed computing framework is enabled for ultra-large-scale networks (with more than 1M nodes).
[0034] Normalize the index values to a unified dimension and output the basic centrality index and standardized basic score. Use the Min-Max normalization method and set a dynamic correction coefficient to deal with outlier interference.
[0035] Divide social network data into multiple communities and identify the local roles of user nodes in multiple communities. The steps include: A hierarchical clustering algorithm is used to partition social network data into multiple communities. The hierarchical clustering algorithm is an unsupervised learning method that gradually merges or splits clusters based on similarity between clusters. It groups data by constructing a tree-like hierarchical structure. The modularity optimization function is used to evaluate the quality of the partitioning, and the resolution parameter is used to control the granularity of the communities. Modularity is a measure of the quality of community partitioning, reflecting the difference between the density of connections within a community and that expected in a random network. The resolution parameter adjusts the coarseness of the community partitioning; higher values tend to generate more small communities.
[0036] Combine multiple communities and user nodes to output the connection strength between local user nodes within each community and between communities. Calculate the Jaccard similarity of bridging edges between communities as a measure of connection strength. Jaccard similarity measures similarity by calculating the ratio of the intersection to the union of two sets. Here, it is used to quantify the degree of association between bridging edges between communities.
[0037] Analyze the node location characteristics of local user nodes within their communities. Node location characteristics include topological properties such as degree centrality and betweenness centrality, reflecting their influence and intermediary role within the community.
[0038] Based on the connection strength, bridge nodes between communities are identified. Bridge nodes are nodes that connect different communities, and their removal will lead to community isolation.
[0039] Determine local roles within multiple communities based on node location characteristics and bridge nodes. Build a role classification model (e.g., decision tree + cluster fusion) to categorize roles such as core hub nodes, boundary propagation nodes, and isolated response nodes.
[0040] Calculate local centrality metrics for local roles within each community. Differentiated metrics are designed for different role types: Flow Betweenness Centrality is calculated for hub nodes, and Semi-Local Centrality is calculated for boundary nodes. Flow Betweenness Centrality measures a node's ability to control network traffic; Semi-Local Centrality is calculated for boundary nodes, focusing on the strength of a node's connections with neighboring communities.
[0041] The basic centrality index and local centrality index are weighted and summed to generate global scores and local scores. The global scores and local scores are ranked respectively, and based on the score ranking, key user nodes are screened and the key user node set is output. The steps to generate global scores and local scores include: Calculate the local centrality index of local roles within each community. Set up a local centrality index system, including quantifying the dispersion of node closeness centrality within the community, using the percentile of betweenness centrality within the community to measure node control ability, and evaluating the efficiency of information diffusion based on the average path length of the community network.
[0042] Adjust the weight of local centrality indicators based on node location characteristics. Dynamically adjust indicator contributions by analyzing the depth and breadth of nodes in the hierarchy. Set hierarchical weight generation rules, including constructing a depth-breadth adjustment matrix where depth weight decays exponentially with increasing hierarchy and breadth weight increases linearly with increasing coverage. Example adjustment rule: For each additional level of depth, the weight coefficient decreases by 15%; for each doubling of breadth, the weight coefficient increases by 10%.
[0043] The role type of local roles is determined based on the bridging nodes, and a weight coefficient is assigned based on the role type. Role types include core hub nodes, edge coordinators, and structural intermediary nodes. Core hub nodes are nodes with a centrality index in the top 10% within a community and the highest depth weight. Edge coordinators are nodes with a bridging index exceeding the threshold and located at the edge of the community. Structural intermediary nodes are nodes with connectivity across more than three communities and an unusual HDA value.
[0044] Combined with the weight coefficients, a weighted summation is performed to generate a local score. Min-Max normalization is used to eliminate dimensional differences in local scores, and Z-score normalization is used for outlier-sensitive indicators. A weighted fusion function is constructed based on the standardized base score and local scores, outputting weight parameters. The weighted fusion function generates a comprehensive evaluation value by linearly combining and normalizing the multi-dimensional scores. The global score = α × standardized base score + β × local score. The NSGA-II multi-objective optimization algorithm is used to determine the optimal α / β ratio. A regularization term is then used to constrain the weight fluctuation range (0.4 ≤ α ≤ 0.6, 0.4 ≤ β ≤ 0.6).
[0045] Finally, the global score is calculated based on the weight parameters, the standardized basic score and the data score.
[0046] The steps to screen the key user node set include: Based on the global and local scores, a two-dimensional descending sort is performed, and the sorting results are output. The global score reflects the strategic position of the node in the overall network topology, and the local score reflects the degree of connectivity within a specific community. The sorting results are output as a structured data table containing the node ID, global score, local score, and community identifier.
[0047] Based on the ranking results, a comprehensive ranking of local roles was calculated. An analytic hierarchy process (AHP) decision model was constructed, with a global weight coefficient of 0.6 and a local weight coefficient of 0.4. Multi-attribute decision-making was performed on the ranking results. The entropy method was used to eliminate dimensionality effects and generate a comprehensive ranking table with confidence intervals.
[0048] The top 10% of the overall rankings are used as a pre-selected set of key user nodes. Low-scoring user nodes are filtered out to output the key user node set. Based on the Pareto optimality principle, the top 10% of nodes are selected as the pre-selected set. A dynamic threshold filtering mechanism is implemented: when the network size is less than 100,000, the top 10% is retained. When the size is 100,000 ≤ ≤ 500,000, top 10% + standard deviation filtering is used. When the size is greater than 500,000, additional centrality cross-validation is performed. Elimination criteria include: local scores below the community average, activity decay exceeding 35% in the past 30 days, and zombie nodes with isolated edge connections.
[0049] Perform score mutation detection on key user nodes in the key user node set to identify malicious behavior, and output the final key user node set based on the identification results. The steps of performing score mutation detection on the key user node set include: Based on social network data, we extract behavioral time series features for all key user subsets within the key user node set. We extract four-dimensional time series features from social network data: behavioral pattern features (time series autocorrelation of posts / comments / likes), social relationship features (Hurst exponent of friend growth rate), content features (volatility of text sentiment polarity), and device features (entropy of IP address change frequency). We use the TSFEL toolkit for feature extraction, generating a rolling dataset with a 72-hour time window.
[0050] Combined with behavioral temporal features, a change point detection algorithm was used to identify abnormal mutation scores. A hybrid model of Facebook Prophet and LSTM was deployed for change point detection, and Bayesian change point analysis was performed with a set significance level. Detected abnormal mutation points were then secondary validated using the isolation forest algorithm, calculating the anomaly score S = 1 / (1 + e^(-x)), where x is the model output value.
[0051] Based on the abnormal mutation score, the relevant mutation key user nodes are labeled and the label information is output. Based on the label information, combined with the log data in the social network data, the rule engine is run to verify and output the abnormal key user nodes. The abnormal key user nodes in the key user node set are eliminated, and the final key user node set is output.
[0052] like Figure 3 As shown, the present invention also provides a key user identification system in a social network, comprising: The graph data modeling module is used to construct social network graph models based on social network data. It uses a heterogeneous graph construction framework to integrate user attribute data, social relationship data, and time series data. It implements a multidimensional relationship encoding mechanism and supports the unified modeling of directed weighted graphs and multi-relationship networks through a hybrid storage structure of adjacency and incidence matrices.
[0053] The centrality calculation engine combines social network graph models, quantifies the importance of user nodes through graph theory algorithms, and calculates basic centrality metrics. Basic centrality metrics typically include traditional metrics, advanced metrics, and dynamic metrics. Traditional metrics include degree centrality, closeness centrality, and betweenness centrality. Advanced metrics include eigenvector centrality, PageRank, and Katz centrality. Dynamic metrics include time decay centrality and communication influence centrality.
[0054] The community partitioning module is used to partition social network data into multiple communities. Specifically, it first uses the Louvain algorithm for fast coarse-grained partitioning, then uses the label propagation algorithm for fine-grained optimization, and finally introduces the modularity density optimization criterion to correct the partitioning results.
[0055] The role identification module is used to identify the local roles of user nodes in multiple communities. Role types include core hub nodes, edge coordinators, and structural intermediary nodes.
[0056] The local index calculation module is used to calculate the local centrality index of local roles in each community.
[0057] The score calculation and fusion module is used to calculate the global score and local score based on the basic centrality index and the local centrality index. It also performs score ranking, filters key user nodes, and outputs the key user node set.
[0058] The anomaly detection module performs score mutation detection on the key user nodes in the key user node set, and outputs the final key user node set based on the detection results.
[0059] According to a key user identification system in a social network provided by the present invention, the centrality calculation engine includes an algorithm selection unit and an indicator calculation unit; the algorithm selection unit is used to combine the social network graph model and select different quantitative algorithms according to different centrality indicators; the indicator calculation unit is used to calculate multiple centrality indicators of user nodes according to different quantitative algorithms.
[0060] In summary, the present invention provides a method and system for identifying key users in a social network. By integrating a dual-dimensional analysis of global centrality indicators and local role characteristics, it not only retains the advantages of traditional global influence assessment, but also accurately captures the structural characteristics of nodes within the community, effectively solving the problem of traditional methods' insufficient identification of cross-community bridge nodes and local core users, and significantly improving the comprehensiveness and accuracy of key user identification. By dynamically adjusting the weights of local indicators through node position characteristics and combining the role type allocation coefficient, the scenario-adaptive optimization of the indicator system is achieved. This enables the system to maintain stable recognition performance in scenarios such as changes in social network structure and malicious behavior interference. Through dual verification of time series feature analysis and rule engine verification, it effectively identifies abnormal states disguised as high-influence users. Compared with traditional single-dimensional detection methods, the error rate is greatly reduced, significantly improving the reliability of the results. Adopting a three-level evaluation framework of role-community-global, through role-aware community division, dynamic weight fusion mechanism, and adversarial anomaly detection, it solves the technical defects of traditional methods in cross-community influence assessment, dynamic role evolution capture, and false node identification.
[0061] Example 1: The graph structure is a directed weighted graph A→B(0.5),B→C(0.8),A→C(0.3),C→D(0.7).
[0062] The shortest paths and traffic distribution between nodes are shown in Table 1.
[0063] Table 1:
[0064] The calculation steps are as follows: 1. Enumerate the shortest paths between all pairs of nodes.
[0065] A to D: two paths (A→B→C→D and A→C→D); B to D: one path (B→C→D); A to C: one path (A→C); C to D: one path (C→D).
[0066] 2. Traffic is distributed in proportion to the total weight of the path.
[0067] The total weight from A to D is 2.0 + 1.0 = 3.0, of which the proportion of traffic passing through C is: (2.0 / 3.0) + (1.0 / 3.0) = 1.0 (all traffic passes through C).
[0068] 3. Accumulate the traffic contribution of node C.
[0069] The A→D path contributes 1.0; the B→D path contributes 1.0 (all traffic flows through C); the total flow betweenness score is: 1.0+1.0=2.0 (normalized to 0.5).
[0070] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for identifying key users in a social network, characterized in that: include: Build a social network graph model based on social network data; The social network data includes user nodes; In combination with the social network graph model, the importance of the user nodes is quantified through graph theory algorithms to calculate basic centrality indicators; dividing the social network data into a plurality of communities, and identifying the local roles of the user node in the plurality of the communities; calculating a local centrality index of the local role in each of the communities; Performing weighted summation on the basic centrality index and the local centrality index to generate a global score and a local score; ranking the global score and the local score, and screening key user nodes based on the score ranking to output a key user node set; Score mutation detection is performed on the key user nodes in the key user node set to identify malicious behavior, and a final key user node set is output based on the identification result.
2. The method for identifying key users in a social network according to claim 1, characterized in that: The steps of constructing the social network graph model include: Extracting user nodes and relationship edges from the social network data, and creating a user node set and a relationship edge set; Constructing an adjacency matrix according to the user node set and the relationship edge set; A graph structure is generated according to the adjacency matrix, the user node set, and the edge set.
3. The method for identifying key users in a social network according to claim 2, wherein: The steps of calculating the basic centrality index include: According to the graph structure, configure the basic indicator algorithm library and determine the indicator weight distribution rules; In combination with the basic indicator algorithm library, the indicator value of each user node is calculated by traversing the graph structure; The index values are normalized to a unified dimension, and the basic centrality index and the standardized basic score are output.
4. The method for identifying key users in a social network according to claim 3, wherein: The steps to identify local roles include: dividing the social network data into a plurality of communities using a hierarchical clustering algorithm; combining the plurality of communities and the user nodes, and outputting the connection strength between the local user nodes in each community and each community; Analyzing node location characteristics of the local user node within the community to which it belongs; determining bridge nodes between the communities according to the connection strength; A plurality of local roles within the communities are determined according to the node location characteristics and the bridge nodes.
5. The method for identifying key users in a social network according to claim 4, characterized in that: The steps to generate global and local scores include: Calculating the local centrality index of the local role within the independent scope of each of the communities; Adjusting the weight of the local centrality index according to the node position characteristics; Determining the role type of the local role according to the bridge node, and assigning a weight coefficient according to the role type; In combination with the weight coefficients, performing weighted summation to generate the local score; Constructing a weighted fusion function according to the standardized basic score and the local score, and outputting a weight parameter; The global score is calculated according to the weight parameter, the normalized basic score and the data score.
6. The method for identifying key users in a social network according to claim 5, characterized in that: The steps to screen the key user node set include: Performing bi-dimensional descending sorting based on the global score and the local score, and outputting the sorting result; Calculating the comprehensive ranking of the local roles according to the sorting results; The top 10% of the comprehensive ranking is used as the pre-selected set of key user nodes, low-scoring user nodes are filtered out, and the key user node set is output.
7. The method for identifying key users in a social network according to claim 1, characterized in that: The step of performing score mutation detection on the key user node set includes: Extracting behavioral time series features of all key user subsets in the key user node set based on the social network data; Combining the behavioral temporal features, a change point detection algorithm is used to identify abnormal mutation scores; Marking the relevant mutation key user nodes according to the abnormal mutation score and outputting the marking information; Based on the tag information and in combination with the log data in the social network data, rule engine verification is performed to output abnormal key user nodes; The abnormal key user nodes in the key user node set are excluded, and the final key user node set is output.
8. A key user identification system in a social network, which adopts a key user identification method in a social network according to any one of claims 1 to 7, characterized in that: include: Graph data modeling module, used to build a social network graph model based on social network data; A centrality calculation engine, which is used to combine the social network graph model, quantify the importance of user nodes through graph theory algorithms, and calculate basic centrality indicators; A community division module, configured to divide the social network data into a plurality of communities; A role identification module, configured to identify the local roles of the user node in the plurality of communities; A local index calculation module, used to calculate the local centrality index of the local role in each of the communities; A score calculation and fusion module, configured to calculate a global score and a local score based on the basic centrality index and the local centrality index; And perform score ranking, filter key user nodes, and output the key user node set; The anomaly detection module performs score mutation detection on the key user nodes in the key user node set, and outputs a final key user node set based on the detection result.
9. The key user identification system in a social network according to claim 8, characterized in that: The centrality calculation engine includes an algorithm selection unit and an indicator calculation unit; the algorithm selection unit is used to combine the social network graph model and select different quantitative algorithms according to different centrality indicators; the indicator calculation unit is used to calculate multiple centrality indicators of the user node according to different quantitative algorithms.
10. The key user identification system in a social network according to claim 8, characterized in that: The anomaly detection module includes a feature extraction unit, an anomaly identification unit, an anomaly marking unit and an anomaly feedback unit; the feature extraction unit is used to extract the behavioral time series features of all key user subsets in the key user node set; the anomaly identification unit is used to combine the behavioral time series features and use a change point detection algorithm to identify anomaly mutation scores; the anomaly marking unit is used to mark relevant mutation key user nodes according to the abnormal mutation scores and output marking information; the anomaly feedback unit is used to exclude mutation key user nodes in the key user node set according to the marking information and output the final key user node set.
Citation Information
Patent Citations
Intelligent enterprise network key participant identification method and system based on hierarchical process analysis
CN118278615A
Systems and methods for determining influencers in a social data network and ranking data objects based on influencers
US20150120717A1
Cited By
Method and system for carbon emission accounting and emission reduction optimization in plateau mountain area highway construction period
CN121810318A
Pyramid structure-based node centrality determination method and device, and electronic device
CN122470958A
Method, apparatus, and electronic equipment for determining node centrality based on pyramid structure
CN122470958B