Social network key node identification method and system fusing propagation characteristics

By integrating the social network key node identification method with propagation characteristics, using the sand cat swarm optimization algorithm and reinforced recurrent neural network, a node evaluation system is constructed, which solves the problems of local optimality and inaccurate identification in traditional methods and achieves more accurate key node identification.

CN120632225AActive Publication Date: 2025-09-12School of Political Science, National Defense University of the Chinese People's Liberation Army

Patent Information

Application Number
CN202510780707.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-12
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

In the existing technology, traditional key node identification methods are prone to fall into local optimal solutions and cannot fully consider the multi-dimensional characteristics of the information propagation process, resulting in inaccurate identification results.

Method used

A social network key node identification method that integrates propagation characteristics is adopted. By collecting external factors and social network data in real time, a reinforced recurrent neural network based on the sand cat swarm optimization algorithm is constructed. The key nodes are screened out by combining the multi-factor influence coefficient and the node evaluation system.

Benefits of technology

It improves the understanding and prediction capabilities of the propagation process, can more accurately identify key nodes, and comprehensively reflect the role and influence of nodes in the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632225A_ABST
    Figure CN120632225A_ABST
Patent Text Reader

Abstract

The invention provides a social network key node identification method and system fusing propagation characteristics, and relates to the technical field of node identification, and the method comprises the steps: collecting various data in real time; the method comprises the following steps: identifying network propagation data to obtain key information data, constructing an enhanced recurrent neural network optimized based on a sodat swarm optimization algorithm, inputting the key information data, outputting to obtain propagation network characteristics, and analyzing the key information data to obtain node propagation performance data. And extracting a multi-factor influence coefficient from the external factor data, constructing a node evaluation system, and screening the preliminary nodes according to the multi-factor influence coefficient to obtain key nodes. According to the method, the recurrent neural network is enhanced, the propagation network features are extracted, the node evaluation system is constructed, the topological structure data are classified, and the key nodes are determined in combination with the multi-factor influence coefficient, so that the understanding and prediction capabilities of the propagation process are improved, and the key nodes are identified more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of node identification technology, and in particular to a method and system for identifying key nodes in a social network by integrating propagation characteristics. Background Art

[0002] Social networks play a vital role in modern society. They encompass a variety of online platforms where people exchange information, interact socially, and share content. Social networks have not only changed the way people communicate but have also had a profound impact on many aspects of information dissemination, business marketing, and more.

[0003] Information dissemination is a core phenomenon in social networks. Information can spread rapidly among users, but its speed, reach, and impact are constrained by a variety of factors. Understanding the patterns of information dissemination is crucial for predicting public opinion trends and promoting products. For example, in marketing, companies hope to identify key nodes in social networks and leverage their influence to spread product information more widely. Key nodes play a crucial role in information dissemination on social networks. These nodes may be information initiators, key disseminators, or possess significant influence, driving widespread information dissemination within the network. Accurately identifying key nodes can help us better understand the structure and function of social networks and support various applications. For example, in viral marketing, focusing resources on key nodes can improve marketing effectiveness. In cybersecurity, monitoring and protecting key nodes can enhance the security of the entire network.

[0004] Traditional key node identification methods often use conventional optimization algorithms such as gradient descent to adjust model parameters when building models. When searching in complex, high-dimensional parameter spaces, these algorithms tend to converge to local optimal solutions rather than global optimal solutions. This limits the model's ability to identify key nodes and fails to accurately reflect the nodes' actual role in network communication. Most existing methods identify key nodes based solely on partial node characteristics, such as the number of connections or activity, ignoring the combined influence of multi-dimensional information such as text features, user behavior characteristics, and external factors during information dissemination. This results in an incomplete assessment of key nodes, and the identification results deviate from the actual dissemination situation.

[0005] In view of the above-mentioned deficiencies in the prior art, the present invention aims to propose a method and system for identifying key nodes in a social network by integrating propagation features. Summary of the Invention

[0006] The present invention provides a method and system for identifying key nodes in a social network by integrating propagation features, so as to solve the defects of the parameter optimization method in the prior art that the method is prone to fall into local optimum and the key nodes are not accurately identified.

[0007] In one aspect, the present invention provides a method for identifying key nodes in a social network by integrating propagation features, comprising: Collect external factor data and social network topology data in real time, and collect network communication data.

[0008] According to the Wenxingchuanhe method, the network propagation data is identified to obtain the key information data, and a reinforced recurrent neural network optimized based on the sand cat swarm optimization algorithm is constructed. The key information data is input and the propagation network characteristics are output. The key information data is then analyzed to obtain the node propagation performance data.

[0009] The multi-factor influence coefficient is extracted from the external factor data, and a node evaluation system is constructed according to the propagation network characteristics and node propagation performance data. The topological structure data is classified to obtain preliminary nodes, and the preliminary nodes are screened according to the multi-factor influence coefficient to obtain key nodes.

[0010] According to a method for identifying key nodes in a social network by integrating propagation features provided by the present invention, the steps of identifying and obtaining key information data include: Keywords, words and subject features are extracted from the text content of network communication data to form text features. User features are obtained by analyzing user behavior characteristics, and communication features are obtained by studying the speed, scope and path of the information dissemination process.

[0011] According to the text features, the subject text information related to the topic is identified, and by analyzing the user characteristics, the key users and the information to be disseminated are determined to obtain user information data, and the dissemination characteristics are analyzed to identify the key dissemination information with satisfactory preset dissemination effects.

[0012] Integrate the subject text information, user information data and key communication information to obtain key information data.

[0013] According to a method for identifying key nodes in a social network by integrating propagation features provided by the present invention, the steps of constructing a reinforced recurrent neural network include: The size of the sand cat population is defined. The position vector of each sand cat in the population represents the parameters to be optimized in the recurrent neural network. The sand cat population is initialized, and the position of each sand cat is randomly generated in the search space.

[0014] The fitness is determined according to the loss function value, and the fitness value of each sand cat is calculated, and the position of the sand cat with the highest fitness value is selected as the optimal individual position.

[0015] The position of the sand cat is updated according to the current position and random angle to obtain the corresponding updated position of the sand cat.

[0016] The fitness value of each sand cat's updated position is calculated, and the sand cat position with the highest fitness value is selected to replace the optimal individual position until the preset number of iterations is reached. The parameters to be optimized corresponding to the optimal individual position are selected to construct a recurrent neural network to obtain a reinforced recurrent neural network.

[0017] According to a method for identifying key nodes in a social network by integrating propagation features provided by the present invention, the step of outputting propagation network features includes: The key information data is converted into a time series feature matrix, where each time step corresponds to the feature vector of a node.

[0018] The input vector of each time step is used as the step vector. At each time step, the step vector and the hidden state of the previous time step are input into the hidden layer of the reinforced recurrent neural network, and the updated hidden state is calculated.

[0019] The hidden state is passed to the output layer for calculation to obtain the propagation network features.

[0020] According to a method for identifying key nodes in a social network by integrating propagation features provided by the present invention, the steps of analyzing and obtaining node propagation performance data include: A unique identifier is assigned to each node, and the topic text information is associated with the propagation events in the propagation key information to obtain integrated event data.

[0021] Count the total number of initial information sources and subsequent nodes involved in the dissemination, calculate the average number of disseminations for each node in different levels, draw a dissemination time series graph, and observe how the speed of information dissemination changes over time to obtain the dissemination time characteristics.

[0022] A communication network graph is constructed based on the integrated event data, where data points represent users and subject-related entities, edges represent communication relationships, and node centrality indicators are calculated from degree centrality, closeness centrality, and betweenness centrality.

[0023] The relationship between keywords, sentiment tendencies and node propagation performance in the main text information is analyzed to obtain the text propagation relationship.

[0024] The user's activity index and node centrality index are correlated and analyzed, and the node propagation performance data is obtained by combining the text propagation relationship and the influence of the user's social relationship on the propagation path.

[0025] According to a method for identifying key nodes in a social network by integrating propagation features provided by the present invention, the step of extracting multi-factor influence coefficients includes: Calculate the Pearson correlation coefficient between external factor data and communication indicators.

[0026] The influence propagation characteristics are extracted from the external factor data based on the Pearson correlation coefficient.

[0027] A multi-factor model was constructed, and the influencing propagation characteristics and propagation indicators were used as target variables to train the multi-factor model and obtain the multi-factor influence coefficient.

[0028] According to a method for identifying key nodes in a social network by integrating propagation features provided by the present invention, the steps of constructing a node evaluation system include: The evaluation objectives are determined based on the importance, influence and activity of the nodes in information dissemination, and target indicators are selected from the propagation network characteristics and node propagation performance data based on the evaluation objectives.

[0029] By constructing a judgment matrix to compare the relative importance of each indicator, and calculating the weight of each indicator, a node evaluation model is constructed based on the evaluation objectives and weights.

[0030] The node evaluation data is obtained by calculating the comprehensive score of each node according to the node evaluation model.

[0031] According to a method for identifying key nodes in a social network by integrating propagation features provided by the present invention, the step of classifying and obtaining preliminary nodes includes: The classification criteria are set according to the target features and weights, and the nodes in the topological structure data are classified to obtain the initial nodes.

[0032] The evaluation data corresponding to the initial node is obtained according to the node evaluation data, and the evaluation data is classified according to the hierarchical clustering method to obtain the category label of each node.

[0033] Based on multiple category labels, the distribution of classified nodes is analyzed to understand the position, quantity and role of nodes of different categories in the network, thereby determining the preliminary nodes.

[0034] According to a method for identifying key nodes in a social network by integrating propagation features provided by the present invention, the step of obtaining key nodes includes: Set screening thresholds and screening targets based on target requirements and multi-factor influencing coefficients.

[0035] Arrange all preliminary nodes in ascending order according to the multi-factor influence coefficient to obtain node arrangement data, and filter nodes that meet the threshold from the node arrangement data according to the screening threshold.

[0036] The activity and credibility of the nodes that meet the threshold are further evaluated, and the nodes that meet the threshold that meet the screening target are extracted to obtain key nodes.

[0037] In another aspect, the present invention further provides a social network key node identification system integrating propagation features, comprising: The multi-data collection module is used to collect external factor data and social network topology data in real time, and collect network communication data.

[0038] The network communication analysis module is used to identify network communication data according to the text-line-communication method to obtain key information data, build a reinforced recurrent neural network optimized based on the sand cat swarm optimization algorithm, input key information data, output the communication network characteristics, and analyze the key information data to obtain node communication performance data.

[0039] The information dissemination evaluation and screening module is used to extract the multi-factor influence coefficient from the external factor data, build a node evaluation system based on the propagation network characteristics and node propagation performance data, classify the topological structure data to obtain preliminary nodes, and screen the preliminary nodes according to the multi-factor influence coefficient to obtain key nodes.

[0040] The present invention provides a method and system for identifying key nodes in a social network that integrates propagation features. By utilizing the sand cat swarm optimization algorithm to optimize the parameters of an enhanced recurrent neural network, an efficient neural network model is constructed to extract propagation network features. This solves the problem that traditional parameter optimization methods may be inefficient or prone to falling into local optimality, and achieves the beneficial effect of more effectively extracting propagation network features and improving the understanding and prediction capabilities of the propagation process.

[0041] The present invention provides a method and system for identifying key nodes in social networks that integrates communication characteristics. By combining communication network characteristics with node communication performance data to construct a node evaluation system, the system determines evaluation objectives, selects corresponding indicators and calculates weights, and then calculates comprehensive node scores. It then uses methods such as hierarchical clustering to classify topological structure data to obtain preliminary nodes. These preliminary nodes are then screened based on multi-factor influence coefficients to ultimately identify key nodes. This method addresses the problem of inaccurate identification results caused by a failure to fully consider both node communication characteristics and network structure features, achieving a beneficial effect that facilitates a deeper understanding of node behavior and influence in social networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 This is one of the flow charts of a method for identifying key nodes in a social network by integrating propagation features provided by an embodiment of the present invention; Figure 2 This is a second flow chart of a method for identifying key nodes in a social network by integrating propagation features provided by an embodiment of the present invention; Figure 3This is one of the flow charts of a method for identifying key nodes in a social network by integrating propagation features provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0044] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0045] The following combination Figure 1-Figure 3 The present invention describes a method and system for identifying key nodes in a social network by integrating propagation features.

[0046] like Figure 1 As shown, an embodiment of the present invention provides a method, device, and storage medium for identifying key nodes in a social network by integrating communication features. The execution subject may be a method for identifying key nodes in a social network by integrating communication features, including: Collect external factor data and social network topology data in real time, and collect network communication data. External factor data can include social events, etc. Network communication data can be obtained from social media platforms, news websites, forums, etc., collecting relevant data including text content, user information, communication time, and communication path.

[0047] According to the Wenxingchuanhe method, the network propagation data is identified to obtain the key information data, and a reinforced recurrent neural network optimized based on the sand cat swarm optimization algorithm is constructed. The key information data is input and the propagation network characteristics are output. The key information data is then analyzed to obtain the node propagation performance data.

[0048] The steps to identify key information data include: Keywords, words and subject features are extracted from the text content of network communication data to form text features. User features are obtained by analyzing user behavior characteristics, and communication features are obtained by studying the speed, scope and path of the information dissemination process.

[0049] Based on text features, topical text information related to the topic is identified. By analyzing user characteristics, key users and disseminated information are identified to obtain user information data. Dissemination characteristics are analyzed to identify key dissemination information with satisfactory predetermined dissemination effects. Key information related to the topic is identified based on the extracted text features. For example, through keyword matching and semantic analysis, content closely related to a specific event or topic can be found. Key users and the information they disseminate can be identified by analyzing user behavioral characteristics. For example, node users who play a key role in the dissemination process are identified; the information they disseminate is likely to be of high importance. Dissemination characteristics are analyzed to identify information with significant dissemination effects. For example, information that disseminates quickly and widely may be key information.

[0050] Integrate the subject text information, user information data and key communication information to obtain key information data.

[0051] like Figure 2 As shown in Figure 2, the steps to construct a reinforced recurrent neural network include: The size of the sand cat population is defined. The position vector of each sand cat in the population represents the parameters to be optimized in the recurrent neural network. The sand cat population is initialized, and the position of each sand cat is randomly generated in the search space.

[0052] The fitness is determined according to the loss function value, and the fitness value of each sand cat is calculated, and the position of the sand cat with the highest fitness value is selected as the optimal individual position.

[0053] The position of the sand cat is updated according to its current position and random angle to obtain the corresponding updated position of the sand cat. The formula is:

[0054] Where, It is the sand cat that updates the location. is the optimal individual position, is the sensitivity range of each sand cat, is a random angle, is a random position.

[0055] The fitness value of each sand cat's updated position is calculated, and the sand cat position with the highest fitness value is selected to replace the optimal individual position until the preset number of iterations is reached. The parameters to be optimized corresponding to the optimal individual position are selected to construct a recurrent neural network to obtain a reinforced recurrent neural network.

[0056] The steps to output the propagation network features include: The key information data is converted into a time series feature matrix, where each time step corresponds to the feature vector of a node.

[0057] The input vector of each time step is used as the step vector. At each time step, the step vector and the hidden state of the previous time step are input into the hidden layer of the reinforced recurrent neural network, and the updated hidden state is calculated. The formula is expressed as:

[0058] Where, is the weight matrix from hidden state to hidden state, is the weight matrix input to the hidden state, is the bias term of the hidden layer, is the activation function, is the hidden state, is the hidden state at the previous time step, is in the time step The input vector.

[0059] The hidden state is passed to the output layer for calculation to propagate network features. The formula is expressed as:

[0060] Where, is the weight matrix from hidden state to output, is the bias term of the output layer, is the hidden state, These are propagation network characteristics. These characteristics can include node characteristics, propagation path characteristics, and propagation dynamics. Node characteristics can include node activity, node influence, node centrality, and node interest characteristics. Propagation path characteristics can include path length, path diversity, and propagation tree structure. Propagation dynamics characteristics can include propagation speed, propagation trend, propagation cycle, and propagation volatility.

[0061] The steps for analyzing and obtaining node propagation performance data include: Each node is assigned a unique identifier, and the topic text information is associated with the propagation events in the key information dissemination to generate integrated event data. These associations are stored in a relational database or graph database, and the integrated data storage is achieved through SQL (Structured Query Language) or a graph database query language (such as Cypher).

[0062] Count the total number of initial information sources and subsequent nodes involved in the dissemination, calculate the average number of disseminations for each node in different levels, draw a dissemination time series graph, and observe how the speed of information dissemination changes over time to obtain the dissemination time characteristics.

[0063] A communication network graph is constructed based on the integrated event data, where data points represent users and subject-related entities, edges represent communication relationships, and node centrality indicators are calculated from degree centrality, closeness centrality, and betweenness centrality.

[0064] Degree Centrality: Calculates each node's in-degree (the number of communication relationships received, such as the number of times it has been forwarded by other users) and out-degree (the number of communication relationships sent, such as the number of times it has forwarded to others). Nodes with high in-degree centrality are likely to be popular recipients of information, while nodes with high out-degree centrality are likely to be active disseminators of information.

[0065] Closeness centrality measures the inverse of the sum of the shortest path lengths between a node and all other nodes in the network. Nodes with high closeness centrality are able to acquire and disseminate information more quickly. For example, in an internal corporate information dissemination network, management may have high closeness centrality because they have quick access to information from various departments.

[0066] Betweenness centrality: This measures how often a node acts as an intermediary for the shortest paths in a network. Nodes with high betweenness centrality have a controlling effect on the flow of information in the network.

[0067] The relationship between keywords, sentiment tendencies and node propagation performance in the main text information is analyzed to obtain the text propagation relationship.

[0068] The user's activity index and node centrality index are correlated and analyzed, and the node propagation performance data is obtained by combining the text propagation relationship and the influence of the user's social relationship on the propagation path.

[0069] The multi-factor influence coefficient is extracted from the external factor data, and a node evaluation system is constructed according to the propagation network characteristics and node propagation performance data. The topological structure data is classified to obtain preliminary nodes, and the preliminary nodes are screened according to the multi-factor influence coefficient to obtain key nodes.

[0070] The steps for extracting multi-factor influence coefficients include: Calculate the Pearson correlation coefficient between external factor data and communication indicators. The formula is:

[0071] Where, is the Pearson correlation coefficient, It is The observed value of the external factor, It is The observed values ​​of the propagation indicators, is the index of the observation, is the sample mean of the external factor, is the sample mean of the propagation indicator The influence propagation characteristics are extracted from the external factor data based on the Pearson correlation coefficient.

[0072] A multi-factor model was constructed, and the influencing propagation characteristics and propagation indicators were used as target variables to train the multi-factor model and obtain the multi-factor influence coefficient.

[0073] like Figure 3 As shown in the figure, the steps to build a node evaluation system include: The evaluation objectives are determined based on the importance, influence and activity of the nodes in information dissemination, and target indicators are selected from the propagation network characteristics and node propagation performance data based on the evaluation objectives.

[0074] By constructing a judgment matrix to compare the relative importance of each indicator, and calculating the weight of each indicator, a node evaluation model is constructed based on the evaluation objectives and weights.

[0075] The node evaluation data is obtained by calculating the comprehensive score of each node according to the node evaluation model.

[0076] The steps of classifying and obtaining preliminary nodes include: The classification criteria are set according to the target features and weights, and the nodes in the topological structure data are classified to obtain the initial nodes.

[0077] The evaluation data corresponding to the initial node is obtained according to the node evaluation data, and the evaluation data is classified according to the hierarchical clustering method to obtain the category label of each node.

[0078] Based on multiple category labels, the distribution of classified nodes is analyzed to understand the position, quantity and role of nodes of different categories in the network, thereby determining the preliminary nodes.

[0079] The steps to obtain key nodes include: Set screening thresholds and screening targets based on target requirements and multi-factor influencing coefficients.

[0080] Arrange all preliminary nodes in ascending order according to the multi-factor influence coefficient to obtain node arrangement data, and filter nodes that meet the threshold from the node arrangement data according to the screening threshold.

[0081] The activity and credibility of the nodes that meet the threshold are further evaluated, and the nodes that meet the threshold that meet the screening target are extracted to obtain key nodes.

[0082] Based on the same general inventive concept, the present invention also protects a social network key node identification system with integrated communication features. The social network key node identification system with integrated communication features provided by the present invention is described below. The social network key node identification system with integrated communication features described below and the social network key node identification method with integrated communication features described above can be referenced to each other.

[0083] Example 1: 1. External factor data collection: Time node: Use timestamp to record the time when data is collected.

[0084] Hot topics on social platforms: Get current hot topics in real time.

[0085] News Events: Get the latest news events from news websites through the API interface.

[0086] Competitor Activities: Obtain information on competitor activities through public data.

[0087] Economic indicators: Obtain real-time economic indicators from financial data platforms.

[0088] Weather Conditions: Get real-time weather data from weather forecast websites.

[0089] Holiday: Determine whether it is a holiday based on the calendar.

[0090] Social network topology data collection: Node connection relationship: Obtain the attention, friendship, etc. between users through the API of the social network platform.

[0091] Edge weight: The edge weight is calculated based on the user's interaction frequency, interaction intensity, etc.

[0092] Network communication data collection: Published content: Obtain user published content through crawler technology or API.

[0093] Number of forwardings: Count the number of forwardings for each piece of content.

[0094] Number of comments: Count the number of comments on each piece of content.

[0095] Likes: Count the number of likes for each piece of content.

[0096] Timestamp: Records the publishing time of each piece of content.

[0097] 2. Key information data identification: text feature extraction: Keyword extraction: Use the TF-IDF algorithm to extract keywords from text content.

[0098] Word extraction: Use word frequency statistics method to extract high-frequency words.

[0099] Topic feature extraction: Use the LDA (Latent Dirichlet Allocation) algorithm to extract topic features.

[0100] User feature extraction: Activity: Counts the number of times a user posts, forwards, and comments within a unit of time.

[0101] Influence: Calculate the number of followers and followers of a user.

[0102] Interest features: Extract user interest features through text mining and topic analysis.

[0103] Propagation feature extraction: Propagation speed: Calculate the distance information travels or the number of nodes covered per unit time.

[0104] Propagation range: The number of nodes covered by the statistical information propagation.

[0105] Transmission path: records the path of information transmission.

[0106] Key information data integration: Topic text information: Identify text information related to a specific topic.

[0107] User information data: Identify key users and the information they spread.

[0108] Disseminate key messages: Identify messages that have significant dissemination effects.

[0109] Integrate key information data: Integrate subject text information, user information data and key communication information.

[0110] 3. Strengthening recurrent neural network construction and propagation network feature extraction: Defining the Sand Cat Population: Population size: Set the sand cat population size to 50.

[0111] Position vector: The position vector of each sand cat represents the parameters to be optimized in the recurrent neural network.

[0112] Initialize the population: Randomly generate the initial positions of sand cats in the search space.

[0113] Fitness calculation: Loss function: The loss function is defined as mean square error (MSE).

[0114] Fitness value: Calculate the fitness value of each sand cat and select the position of the sand cat with the highest fitness value as the optimal individual position.

[0115] Location Updates: Current position and random angle: Update the position based on the sand cat's current position and random angle.

[0116] Update position: Calculate the updated position of the sand cat and evaluate its fitness value.

[0117] Iterative optimization: Repeat the position update process until the preset number of iterations is reached.

[0118] Propagation network feature extraction: Time series feature matrix: Convert key information data into a time series feature matrix, where each time step corresponds to the feature vector of a node.

[0119] Hidden state calculation: At each time step, the step vector and the hidden state of the previous time step are input into the hidden layer of the reinforced recurrent neural network to calculate and update the hidden state.

[0120] Output propagation network features: pass the hidden state to the output layer to obtain the propagation network features.

[0121] 4. Node propagation performance data analysis: Node identification and integration event data: Unique Identifier: Assign a unique identifier to each node.

[0122] Integrate event data: associate the topic text information with the communication events in the communication key information.

[0123] Propagation time characteristic analysis: Count the number of disseminations: Count the total number of nodes that are the initial information source and subsequently participate in the dissemination.

[0124] Calculate the average number of propagations: Calculate the average number of propagations for each node in different levels.

[0125] Draw a propagation time series graph: observe how the speed of information propagation changes over time.

[0126] Construction of propagation network graph: Construct a communication network graph: data points represent users and related entities, and edges represent communication relationships.

[0127] Calculate node centrality indices: Calculate node centrality indices from degree centrality, closeness centrality, and betweenness centrality.

[0128] Analysis of text communication relationships: Keywords and sentiment tendencies: Analyze the relationship between keywords and sentiment tendencies in topic text information and node communication performance.

[0129] Node propagation performance data integration: Activity and centrality correlation analysis: Correlation analysis is performed on the user's activity index and the node centrality index.

[0130] Analysis of the impact of communication paths: combining the impact of text communication relationships and users’ social relationships on communication paths.

[0131] 5. Extraction of multi-factor influence coefficients: Correlation calculation: Pearson correlation coefficient: Calculate the Pearson correlation coefficient between external factor data and communication indicators.

[0132] Impact propagation feature extraction: Extracting influence propagation features: Extracting influence propagation features from external factor data based on the Pearson correlation coefficient.

[0133] Multifactor model construction: Model training: Use the influencing propagation characteristics and propagation indicators as target variables to train a multi-factor model.

[0134] Obtain multi-factor influence coefficients: Obtain multi-factor influence coefficients through model training.

[0135] 6. Construction of node evaluation system: Evaluation objectives and indicator selection: Evaluation objectives: Determine the evaluation objectives based on the importance, influence, and activity of the node in information dissemination.

[0136] Target indicator selection: Select target indicators from the propagation network characteristics and node propagation performance data.

[0137] Weight calculation and evaluation model construction: Judgment matrix construction: Construct a judgment matrix to compare the relative importance of each indicator.

[0138] Weight calculation: Calculate the weight of each indicator.

[0139] Node evaluation model construction: Construct a node evaluation model based on evaluation objectives and weights.

[0140] Comprehensive score calculation: Calculate comprehensive score: Calculate the comprehensive score of each node based on the node evaluation model.

[0141] Get node evaluation data: output the comprehensive score of each node.

[0142] 7. Preliminary node classification and key node screening: Classification standard setting: Target characteristics and weights: Set classification criteria based on target characteristics and weights.

[0143] Initial node classification: Classify the nodes in the topological structure data to obtain the initial nodes.

[0144] Hierarchical clustering classification: Evaluation data classification: Hierarchical clustering classification is performed based on the node evaluation data to obtain the category label of each node.

[0145] Classification result analysis: Node distribution analysis: Analyze the distribution of nodes after classification and determine the preliminary nodes.

[0146] Screening threshold and target setting: Screening thresholds and targets: Set screening thresholds and targets based on target requirements and multi-factor influencing coefficients.

[0147] Node arrangement and screening: Node arrangement: Arrange all preliminary nodes in ascending order according to the multi-factor influence coefficient.

[0148] Threshold filtering: Filter nodes that meet the threshold from the node arrangement data based on the filtering threshold.

[0149] Activity and credibility assessment: Further evaluation: The activity and credibility of nodes that meet the threshold are further evaluated.

[0150] Extract key nodes: Extract nodes that meet the screening target as key nodes.

[0151] An embodiment of the present invention provides a social network key node identification system integrating communication features, including: The multi-data collection module is used to collect external factor data and social network topology data in real time, and collect network communication data.

[0152] The network communication analysis module is used to identify network communication data according to the text-line-communication method to obtain key information data, build a reinforced recurrent neural network optimized based on the sand cat swarm optimization algorithm, input key information data, output the communication network characteristics, and analyze the key information data to obtain node communication performance data.

[0153] The information dissemination evaluation and screening module is used to extract the multi-factor influence coefficient from the external factor data, build a node evaluation system based on the propagation network characteristics and node propagation performance data, classify the topological structure data to obtain preliminary nodes, and screen the preliminary nodes according to the multi-factor influence coefficient to obtain key nodes.

[0154] This example presents a method and system for identifying key nodes in social networks that integrates communication features. By integrating external factor data, social network topology data, and network communication data, it extracts features from multiple dimensions, including text content, user behavior, and the communication process. This comprehensive consideration of multiple data and features enables more comprehensive and accurate identification of key nodes, more accurately reflecting the role and importance of nodes in network communication. Using the Sand Cat Swarm Optimization algorithm, the parameters of a reinforced recurrent neural network are optimized. By defining a sand cat population, initializing positions, calculating fitness, and updating positions, an efficient neural network model is constructed to extract communication network features. A node evaluation system is then constructed by combining these network features with node communication performance data. Evaluation objectives are determined based on importance, influence, and activity, and corresponding indicators are selected and weighted to calculate a comprehensive node score. The optimized reinforced recurrent neural network more effectively extracts communication network features, improving understanding and prediction of the communication process. The constructed node evaluation system provides a systematic approach for comprehensively evaluating the communication capabilities of nodes, contributing to a deeper understanding of their behavior and influence in social networks.

[0155] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0156] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for identifying key nodes in a social network by integrating propagation features, characterized in that: include: Collect external factor data and social network topology data in real time, and collect network communication data; The network propagation data is identified according to the text-line-transmission method to obtain key information data, a reinforced recurrent neural network optimized by the sand cat swarm optimization algorithm is constructed, the key information data is input, the propagation network characteristics are obtained as output, and the key information data is analyzed to obtain node propagation performance data; A multi-factor influence coefficient is extracted from the external factor data, a node evaluation system is constructed according to the propagation network characteristics and the node propagation performance data, and the topological structure data is classified to obtain preliminary nodes, and the preliminary nodes are screened according to the multi-factor influence coefficient to obtain key nodes.

2. The method for identifying key nodes in a social network by integrating propagation features according to claim 1, characterized in that: The steps of identifying and obtaining the key information data include: Extracting keywords, words and subject features from the text content of the network communication data to form text features, obtaining user features by analyzing user behavior features, and obtaining communication features by studying the speed, scope and path of the information dissemination process; Identifying topic text information related to the topic based on the text features, determining key users and disseminated information by analyzing the user features to obtain user information data, analyzing the dissemination features to identify key dissemination information with a satisfactory preset dissemination effect; The subject text information, the user information data and the key communication information are integrated to obtain the key information data.

3. The method for identifying key nodes in a social network by integrating propagation features according to claim 1, characterized in that: The steps of constructing the reinforced recurrent neural network include: The size of a sand cat population is defined, the position vector of each sand cat in the population represents the parameter to be optimized of the recurrent neural network, and the sand cat population is initialized, and the position of each sand cat is randomly generated in the search space; The fitness is determined according to the loss function value, and the fitness value of each sand cat is calculated, and the position of the sand cat with the highest fitness value is selected as the optimal individual position; Update the position of the sand cat according to its current position and random angle to obtain the corresponding updated position of the sand cat; The fitness value of each sand cat update position is calculated, and the sand cat position with the highest fitness value is selected to replace the optimal individual position until a preset number of iterations is reached. The parameters to be optimized corresponding to the optimal individual position are selected to construct the recurrent neural network to obtain the enhanced recurrent neural network.

4. The method for identifying key nodes in a social network by integrating propagation features according to claim 1, characterized in that: The step of outputting the propagation network characteristics includes: Convert the key information data into a time series feature matrix, where each time step corresponds to a feature vector of a node; Taking the input vector of each time step as a step vector, at each time step, inputting the step vector and the hidden state of the previous time step into the hidden layer of the reinforced recurrent neural network, and calculating and updating the hidden state; The hidden state is transferred to the output layer for calculation to obtain the propagation network feature.

5. The method for identifying key nodes in a social network by integrating propagation features according to claim 2, characterized in that: The steps for analyzing and obtaining node propagation performance data include: Assigning a unique identifier to each node, and associating the subject text information with the propagation events in the propagation key information to obtain integrated event data; Count the total number of nodes involved in the initial information source and subsequent dissemination, calculate the average number of disseminations for each node at different levels, draw a dissemination time series graph, and observe how the speed of information dissemination changes over time to obtain the dissemination time characteristics; Constructing a communication network graph based on the integrated event data, where data points represent users and subject-related entities, edges represent communication relationships, and calculating node centrality indicators from degree centrality, closeness centrality, and betweenness centrality; Analyze the relationship between keywords, sentiment tendencies and node propagation performance in the subject text information to obtain a text propagation relationship; The user's activity index is correlated with the node centrality index, and the node propagation performance data is obtained by combining the influence of the text propagation relationship and the user's social relationship on the propagation path.

6. The method for identifying key nodes in a social network by integrating propagation features according to claim 1, characterized in that: The step of extracting the multi-factor influence coefficient includes: Calculating the Pearson correlation coefficient between the external factor data and the communication indicator; extracting influence propagation features from the external factor data according to the Pearson correlation coefficient; A multi-factor model is constructed, and the multi-factor influence coefficient is obtained by using the influencing propagation characteristics and the propagation index as target variables to train the multi-factor model.

7. The method for identifying key nodes in a social network by integrating propagation features according to claim 1, characterized in that: The steps of constructing the node evaluation system include: Determine an evaluation target based on the importance, influence, and activity of the node in information dissemination, and select target indicators from the propagation network characteristics and the node propagation performance data based on the evaluation target; Compare the relative importance of each indicator by constructing a judgment matrix, calculate the weight of each indicator, and construct a node evaluation model based on the evaluation objectives and the weights; The node evaluation data is obtained by calculating the comprehensive score of each node according to the node evaluation model.

8. The method for identifying key nodes in a social network by integrating propagation features according to claim 7, characterized in that: The steps of classifying and obtaining the preliminary nodes include: According to the target index and the weight setting classification standard, the nodes in the topological structure data are classified to obtain initial nodes; Obtaining evaluation data corresponding to the initial node according to the node evaluation data, and classifying the evaluation data according to a hierarchical clustering method to obtain a category label for each node; According to multiple category labels, the distribution of nodes after classification is analyzed to understand the position, quantity and role of nodes of different categories in the network, thereby determining the preliminary nodes.

9. The method for identifying key nodes in a social network by integrating propagation features according to claim 1, characterized in that: The steps to obtain key nodes include: Setting screening thresholds and screening targets based on target requirements and the multi-factor influencing coefficients; Arrange all preliminary nodes in ascending order according to the multi-factor influence coefficient to obtain node arrangement data, and select nodes meeting the threshold from the node arrangement data according to the screening threshold; The activity and credibility of the nodes meeting the threshold are further evaluated, and the nodes meeting the threshold that meet the screening target are extracted to obtain the key nodes.

10. A social network key node identification system integrating communication features, using a social network key node identification method integrating communication features according to any one of claims 1 to 9, characterized in that: The identification system comprises: Multi-data collection module, used to collect external factor data and social network topology data in real time, and collect network communication data; A network communication analysis module is used to identify the network communication data according to the text-line-transmission combination method to obtain key information data, construct a reinforced recurrent neural network optimized by the Sand Cat Swarm Optimization algorithm, input the key information data, output the communication network characteristics, and analyze the key information data to obtain node communication performance data; An information propagation evaluation and screening module is used to extract multi-factor influence coefficients from the external factor data, construct a node evaluation system based on the propagation network characteristics and the node propagation performance data, classify the topological structure data to obtain preliminary nodes, and screen the preliminary nodes according to the multi-factor influence coefficients to obtain key nodes.

Citation Information

Patent Citations

  • Knowledge-fused social network streaming event detection system

    CN110020214A

  • Generative AI emotion propagation prediction and guidance large model construction method and system

    CN119047512A

  • Systems and Methods for Implementing Smart Assistant Systems

    US20220129556A1

Cited By

  • Advertisement putting strategy optimization method based on topological data analysis

    CN121504548A

  • An advertisement putting strategy optimization method based on topology data analysis

    CN121504548B

  • Online social network multi-element fusion propagation node identification and deduction method and system

    CN122087331A