A social network key node identification method and system fusing propagation features

By integrating propagation features into a key node identification method for social networks, and utilizing the sand cat swarm optimization algorithm and reinforced recurrent neural networks, a node evaluation system is constructed. This solves the problems of local optima and inaccurate identification in traditional methods, and achieves more accurate key node identification and propagation process understanding.

CN120632225BActive Publication Date: 2025-12-09School of Political Science, National Defense University of the Chinese People's Liberation Army
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510780707.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-12-09
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

In existing technologies, traditional key node identification methods are prone to getting stuck in local optima and fail to fully consider the multi-dimensional features in the information propagation process, resulting in inaccurate identification results.

Method used

A key node identification method for social networks using fusion propagation features is proposed. By collecting external factors and social network data in real time, a reinforced recurrent neural network based on the sand cat swarm optimization algorithm is constructed. Combining text features, user behavior features, and external factors, a node evaluation system is built to screen out key nodes.

Benefits of technology

It improves the accuracy and comprehensiveness of identifying key nodes, enhances the understanding and prediction capabilities of the propagation process, and can more realistically reflect the role and importance of nodes in the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632225B_ABST
    Figure CN120632225B_ABST
Patent Text Reader

Abstract

The application provides a kind of fusion propagation feature social network key node identification method and system, it is related to node identification technical field, method includes real-time acquisition multiple data.To network propagation data are identified to obtain key information data, construct the reinforcement recurrent neural network based on sand cat group optimization algorithm optimization, input key information data, output obtains propagation network feature, and key information data are analyzed to obtain node propagation performance data.Extract multi-factor influence coefficient from external factor data, build node evaluation system, and screen preliminary nodes according to multi-factor influence coefficient to obtain key nodes.The application extracts propagation network features and builds node evaluation system through reinforcement recurrent neural network, classifies topological structure data and determines key nodes in combination with multi-factor influence coefficient, improves the understanding and prediction ability of propagation process, and more accurately realizes the identification of key nodes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of node identification, and in particular to a social network key node identification method and system fusing propagation features. BACKGROUND

[0002] Social networks play a vital role in modern society, encompassing various online platforms where people exchange information, interact socially, and share content. Social networks have not only changed the way people communicate but also had a profound impact on information dissemination, business marketing, and many other aspects.

[0003] In social networks, information propagation is a core phenomenon. Information can spread rapidly among users, and its speed, range, and influence are subject to various factors. Understanding the rules of information propagation is of great significance for predicting public opinion trends and promoting products. For example, in the marketing field, enterprises hope to use the influence of key nodes in social networks to spread product information more widely. Key nodes play a crucial role in information propagation in social networks. These nodes may be initiators, important propagators, or have strong influence, and can drive information to spread widely in the network. Accurate identification of key nodes can help people better understand the structure and function of social networks and provide support for various applications. For example, in viral marketing, concentrating resources on key nodes can improve marketing effectiveness; in the field of network security, monitoring and protecting key nodes can enhance the security of the entire network.

[0004] Traditional key node identification methods often use conventional optimization algorithms such as gradient descent to adjust model parameters when building models. These algorithms are prone to converging to local optimal solutions rather than global optimal solutions when searching in complex high-dimensional parameter spaces, resulting in limited identification capabilities of the model for key nodes and an inability to accurately reflect the actual role of nodes in network propagation. Most existing methods only rely on partial features of nodes, such as the number of connections or activity, to identify key nodes, ignoring the comprehensive impact of multi-dimensional information such as text features, user behavior features, and external factors during information propagation. This makes the judgment of key nodes not comprehensive enough, and the identification results deviate from the actual propagation situation.

[0005] In view of the above deficiencies of the prior art, the present application aims to provide a social network key node identification method and system fusing propagation features. SUMMARY

[0006] The present application provides a social network key node identification method and system fusing propagation features to solve the defects of the prior art that parameter optimization methods are prone to local optimization and inaccurate identification of key nodes.

[0007] In one aspect, the application provides a social network key node identification method fusing propagation characteristics, comprising:

[0008] Real-time collection of external factor data and topology data of the social network, and collection of network propagation data.

[0009] Identification of the key information data from the network propagation data according to a text-line transmission method, construction of a reinforced recurrent neural network optimized based on a sand cat population optimization algorithm, input of the key information data, output of the propagation network characteristics, and analysis of the key information data to obtain node propagation performance data.

[0010] Extraction of multi-factor influence coefficients from the external factor data, construction of a node evaluation system according to the propagation network characteristics and the node propagation performance data, classification of the topology data to obtain preliminary nodes, and screening of the preliminary nodes according to the multi-factor influence coefficients to obtain key nodes.

[0011] According to the social network key node identification method fusing propagation characteristics provided by the application, the step of identifying the key information data comprises:

[0012] Extraction of keywords, words and theme features from the text content in the network propagation data to form text features, acquisition of user features by analyzing the behavior characteristics of the users, and obtaining of propagation characteristics by studying the speed, range and path of the information propagation process.

[0013] Identification of theme text information related to the theme according to the text features, determination of user information data by analyzing the user features and the information propagated, analysis of the propagation characteristics, and identification of the propagation key information with a satisfactory preset propagation effect.

[0014] Integration of the theme text information, the user information data and the propagation key information to obtain the key information data.

[0015] According to the social network key node identification method fusing propagation characteristics provided by the application, the step of constructing the reinforced recurrent neural network comprises:

[0016] Definition of the sand cat population size, definition of the position vector of each sand cat in the population as the to-be-optimized parameters of the recurrent neural network, and initialization of the sand cat population, with the position of each sand cat being randomly generated in the search space.

[0017] Determination of the fitness according to the loss function value, calculation of the fitness value of each sand cat, and selection of the position of the sand cat with the highest fitness value as the optimal individual position.

[0018] Position updating according to the current position of the sand cat and the random angle to obtain the corresponding sand cat updated position.

[0019] The fitness value of each sand cat updated position is calculated, the sand cat position with the highest fitness value is selected to replace the optimal individual position, until a preset iteration number is reached, the to-be-optimized parameter corresponding to the optimal individual position is selected to construct a recurrent neural network to obtain a reinforced recurrent neural network.

[0020] According to the social network key node identification method for fusing propagation characteristics provided by the application, the step of outputting the propagation network characteristics comprises:

[0021] The key information data is converted into a time sequence feature matrix, and each time step corresponds to a feature vector of a node.

[0022] The input vector of each time step is taken as a step vector, at each time step, the step vector and the hidden state of the previous time step are input into the hidden layer of the reinforced recurrent neural network, and the updated hidden state is calculated.

[0023] The hidden state is transmitted to the output layer for calculation to obtain the propagation network characteristics.

[0024] According to the social network key node identification method for fusing propagation characteristics provided by the application, the step of analyzing the node propagation performance data comprises:

[0025] A unique identifier is assigned to each node, and the topic text information is associated with the propagation events in the propagation key information to obtain integrated event data.

[0026] The total number of initial information sources and nodes participating in subsequent propagation is counted, the average number of propagations of each node in different levels is calculated, a propagation time sequence diagram is drawn, and the change of the information propagation speed with time is observed to obtain the propagation time characteristics.

[0027] A propagation network graph is constructed based on the integrated event data, data points represent users and subject-related entities, edges represent propagation relationships, and node centrality indexes are calculated from degree centrality, closeness centrality and betweenness centrality.

[0028] The relationship between keywords and sentiment tendencies in the subject text information and the node propagation performance is analyzed to obtain the text propagation relationship.

[0029] The activity index of the user is associated with the node centrality index for association analysis, and the influence of the text propagation relationship and the social relationship of the user on the propagation path is obtained to obtain the node propagation performance data.

[0030] According to the social network key node identification method for fusing propagation characteristics provided by the application, the step of extracting the multi-factor influence coefficient comprises:

[0031] The Pearson correlation coefficient of the external factor data and the propagation index is calculated.

[0032] Extract the influence propagation features from the external factor data according to the Pearson correlation coefficient.

[0033] Build a multi-factor model, use the influence propagation features and the propagation indicators as target variables to train the multi-factor model, and obtain multi-factor influence coefficients.

[0034] According to the social network key node identification method provided by the application, the steps of constructing the node evaluation system comprise:

[0035] According to the importance, influence and activity of the node in information propagation, determine the evaluation target, and select the target indicator from the propagation network features and node propagation performance data according to the evaluation target.

[0036] The relative importance between each index is compared by constructing a judgment matrix, and the weight of each index is calculated, and the node evaluation model is constructed according to the evaluation target and the weight.

[0037] According to the node evaluation model, the comprehensive score of each node is calculated to obtain node evaluation data.

[0038] According to the social network key node identification method provided by the application, the steps of classifying the preliminary nodes comprise:

[0039] According to the target features and the weight setting, the classification standard is set, and the nodes in the topological structure data are classified to obtain the initial nodes.

[0040] According to the node evaluation data, the evaluation data corresponding to the initial nodes is obtained, and the evaluation data is classified according to the hierarchical clustering method to obtain the category label of each node.

[0041] According to the plurality of category labels, the distribution of the classified nodes is analyzed, the position, quantity and role of the nodes of different categories in the network are understood, and the preliminary nodes are determined.

[0042] According to the social network key node identification method provided by the application, the steps of obtaining the key nodes comprise:

[0043] According to the target requirement and the multi-factor influence coefficient, the screening threshold and the screening target are set.

[0044] All the preliminary nodes are arranged in ascending order according to the multi-factor influence coefficient to obtain node arrangement data, and the nodes meeting the threshold are screened from the node arrangement data according to the screening threshold.

[0045] The active degree and the reliability of the nodes meeting the threshold are further evaluated, and the key nodes meeting the screening target are extracted from the nodes meeting the threshold.

[0046] In another aspect, the present application also provides a social network key node identification system fusing propagation characteristics, comprising:

[0047] A multi-data acquisition module is configured to acquire external factor data and topological structure data of a social network in real time, and collect network propagation data.

[0048] A network propagation analysis module is configured to identify key information data from the network propagation data according to a text line transmission method, construct a reinforced recurrent neural network optimized based on a sand cat group optimization algorithm, input the key information data, output propagation network characteristics, and analyze the key information data to obtain node propagation performance data.

[0049] An information propagation evaluation and screening module is configured to extract a multi-factor influence coefficient from the external factor data, construct a node evaluation system according to the propagation network characteristics and the node propagation performance data, classify the topological structure data to obtain preliminary nodes, and screen the preliminary nodes according to the multi-factor influence coefficient to obtain key nodes.

[0050] The social network key node identification method and system fusing propagation characteristics provided by the present application can optimize the parameters of a reinforced recurrent neural network by using a sand cat group optimization algorithm, construct an efficient neural network model, extract propagation network characteristics, solve the problems of low efficiency or easy falling into local optimization of a traditional parameter optimization method, and effectively extract propagation network characteristics, thereby improving the understanding and prediction ability of the propagation process.

[0051] The social network key node identification method and system fusing propagation characteristics provided by the present application can construct a node evaluation system by combining propagation network characteristics and node propagation performance data, determine an evaluation target, select corresponding indexes and calculate weights, calculate a node comprehensive score, classify topological structure data to obtain preliminary nodes by using hierarchical clustering, and screen the preliminary nodes according to a multi-factor influence coefficient to finally determine key nodes, thereby solving the problem that the propagation characteristics and network structure characteristics of nodes cannot be comprehensively considered, leading to inaccurate identification results, and obtaining the beneficial effect of helping to deeply understand the behavior and influence of nodes in a social network. BRIEF DESCRIPTION OF DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0053] Figure 1This is one of the flowcharts illustrating a method for identifying key nodes in a social network based on fused propagation features, provided in an embodiment of the present invention.

[0054] Figure 2 This is a second flowchart illustrating a method for identifying key nodes in a social network based on integrated propagation features, provided in an embodiment of the present invention.

[0055] Figure 3 This is one of the flowcharts illustrating a method for identifying key nodes in a social network based on fusion propagation features, provided in an embodiment of the present invention. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0057] The following is combined Figures 1-3 This invention describes a method and system for identifying key nodes in a social network based on integrated propagation characteristics.

[0058] like Figure 1 As shown, this invention provides a method, apparatus, and storage medium for identifying key nodes in a social network based on integrated propagation features. The executing entity can be a method for identifying key nodes in a social network based on integrated propagation features, including:

[0059] It collects real-time data on external factors and the topology of social networks, and gathers data on online dissemination. External factor data can include social events, etc. Online dissemination data can be obtained from social media platforms, news websites, forums, etc., collecting relevant data including text content, user information, dissemination time, dissemination path, etc.

[0060] Key information data is obtained by identifying network propagation data using the text-to-text transmission method. A reinforced recurrent neural network optimized by the sand cat swarm optimization algorithm is constructed. The key information data is input, and the propagation network characteristics are output. The node propagation performance data is obtained by analyzing the key information data.

[0061] The steps to identify and obtain key information data include:

[0062] Text features are constructed by extracting keywords, words, and thematic features from the text content of online dissemination data. User features are obtained by analyzing user behavior characteristics, and dissemination features are obtained by studying the speed, scope, and path of information dissemination.

[0063] According to the text features, theme-related text information is identified, and by analyzing user features, key users and the information they spread are determined to obtain user information data, the spread features are analyzed, and the key information with the desired spread effect is identified. According to the extracted text features, key information related to the theme is identified. For example, through keyword matching, semantic analysis and other methods, content closely related to a specific event or topic is found. By analyzing the behavior characteristics of users, key users and the information they spread are determined. For example, node users who play a key role in the spread process are identified, and the information they spread may have high importance. Analyze the spread characteristics to identify information with significant spread effect. For example, information with fast spread speed and wide range may be key information.

[0064] Integrate the theme text information, user information data and key information to obtain key information data.

[0065] As shown in Figure 2 , the steps of constructing the reinforced recurrent neural network include:

[0066] Define the size of the sand cat population, and the position vector of each sand cat in the population represents the to-be-optimized parameters of the recurrent neural network, and initialize the sand cat population. The position of each sand cat is randomly generated in the search space.

[0067] Determine the fitness according to the loss function value, calculate the fitness value of each sand cat, and select the sand cat position with the highest fitness value as the optimal individual position.

[0068] Update the position according to the current position of the sand cat and the random angle to obtain the corresponding sand cat update position, which is expressed as:

[0069]

[0070] In the formula, is the sand cat update position, is the optimal individual position, is the sensitivity range of each sand cat, is the random angle, is the random position.

[0071] Calculate the fitness value of each sand cat update position, and select the sand cat position with the highest fitness value to replace the optimal individual position until the preset iteration number is reached. The to-be-optimized parameters corresponding to the optimal individual position are selected to construct the recurrent neural network to obtain the reinforced recurrent neural network.

[0072] The output step of obtaining the spread network features includes:

[0073] Convert the key information data into a time sequence feature matrix, and each time step corresponds to a feature vector of a node.

[0074] The input vector of each time step is taken as a step vector, at each time step, the step vector and the hidden state of the previous time step are input to the hidden layer of the reinforcement recurrent neural network, and the updated hidden state is calculated, which is expressed as:

[0075]

[0076] In the formula, is the weight matrix of the hidden state to the hidden state, is the weight matrix of the input to the hidden state, is the bias term of the hidden layer, is the activation function, is the hidden state, is the hidden state of the previous time step, is the input vector at time step .

[0077] The hidden state is passed to the output layer for calculation to propagate network features, which is expressed as:

[0078]

[0079] In the formula, is the weight matrix of the hidden state to the output, is the bias term of the output layer, is the hidden state, is the propagation network feature. The propagation network feature can include node features, propagation path features, and propagation dynamic features, etc. The node features can include node activity, node influence, node centrality, and node interest features, etc. The propagation path features can include path length, path diversity, and propagation tree structure, etc. The propagation dynamic features can include propagation speed, propagation trend, propagation period, and propagation volatility, etc.

[0080] The steps of analyzing the node propagation performance data include:

[0081] A unique identifier is assigned to each node, and the topic text information is associated with the propagation events in the propagation key information to obtain integrated event data. A relational database or a graph database is used to store these association relationships, and SQL (Structured Query Language) or a graph database query language (such as Cypher) is used to realize the integrated storage of data.

[0082] The total number of initial information sources and subsequent participating nodes is counted, and the average number of propagations of each node in different levels is calculated, a propagation time sequence diagram is drawn, and the change of the propagation speed with time is observed to obtain the propagation time characteristics.

[0083] Based on the integrated event data, a propagation network graph is constructed, with data points representing users and subject-related entities, and edges representing propagation relationships. Node centrality indicators are calculated from degree centrality, closeness centrality, and betweenness centrality.

[0084] Degree centrality: The in-degree (number of received propagation relationships, such as the number of times it is forwarded by other users) and out-degree (number of sent propagation relationships, such as the number of times it forwards others) of each node are calculated. Nodes with high in-degree centrality may be popular recipients of information dissemination, and nodes with high out-degree centrality may be active disseminators of information dissemination.

[0085] Closeness centrality: Measures the reciprocal of the sum of the shortest path lengths of a node to all other nodes in the network. Nodes with high closeness centrality can quickly access and disseminate information. For example, in an internal information dissemination network of an enterprise, the management layer may have high closeness centrality because they can quickly access information from various departments.

[0086] Betweenness centrality: Calculate the frequency of a node as an intermediary in the shortest path. Nodes with high betweenness centrality have control over the flow of information in the network.

[0087] Analyze the relationship between keywords and sentiment orientation in subject text information and node propagation performance to obtain text propagation relationships.

[0088] Correlate user activity indicators with node centrality indicators, and combine text propagation relationships and user social relationships to analyze the impact on propagation paths to obtain node propagation performance data.

[0089] Extract multi-factor influence coefficients from external factor data, construct a node evaluation system based on propagation network characteristics and node propagation performance data, and classify topological structure data to obtain preliminary nodes. According to the multi-factor influence coefficients, the preliminary nodes are screened to obtain key nodes.

[0090] The steps for extracting multi-factor influence coefficients include:

[0091] Calculate the Pearson correlation coefficient between external factor data and propagation indicators, expressed as:

[0092]

[0093] In the formula, is the Pearson correlation coefficient, is the observation value of the th external factor, is the observation value of the th propagation indicator, is the index of the observation value, is the sample mean of the external factor, is a sample mean of the propagation indicator

[0094] According to the Pearson correlation coefficient, the influence propagation feature is extracted from the external factor data.

[0095] A multi-factor model is constructed, and the multi-factor model is trained using the influence propagation feature and the propagation indicator as the target variable to obtain a multi-factor influence coefficient.

[0096] As shown in Figure 3 The steps of constructing the node evaluation system include:

[0097] According to the importance, influence and activity of the node in information propagation, an evaluation target is determined, and target indicators are selected from the propagation network features and node propagation performance data according to the evaluation target.

[0098] The relative importance between each indicator is compared by constructing a judgment matrix, and the weight of each indicator is calculated, and a node evaluation model is constructed according to the evaluation target and the weight.

[0099] According to the node evaluation model, the comprehensive score of each node is calculated to obtain node evaluation data.

[0100] The steps of classifying to obtain preliminary nodes include:

[0101] According to the target feature and the weight, a classification standard is set, and the nodes in the topological structure data are classified to obtain initial nodes.

[0102] According to the node evaluation data, the evaluation data corresponding to the initial nodes is obtained, and the evaluation data is classified according to the hierarchical clustering method to obtain the class label of each node.

[0103] According to a plurality of class labels, the distribution of the classified nodes is analyzed, the position, quantity and role of the nodes of different classes in the network are understood, and preliminary nodes are determined.

[0104] The steps of obtaining key nodes include:

[0105] According to the target requirement and the multi-factor influence coefficient, a screening threshold and a screening target are set.

[0106] All preliminary nodes are arranged in ascending order according to the multi-factor influence coefficient to obtain node arrangement data, and the nodes meeting the threshold are screened from the node arrangement data according to the screening threshold.

[0107] The activity and credibility of the nodes meeting the threshold are further evaluated, and the key nodes meeting the screening target are extracted from the nodes meeting the threshold.

[0108] Based on the same overall inventive concept, the present application also protects a social network key node identification system integrating propagation characteristics, and the following describes a social network key node identification system integrating propagation characteristics provided by the present application. The social network key node identification system described below can be mutually corresponding and referenced with the social network key node identification method described above.

[0109] Example one:

[0110] 1. External factor data collection:

[0111] Time node: Use timestamp to record the time of data collection.

[0112] Social platform hot topic: Real-time access to current hot topics.

[0113] News event: Obtain the latest news event from news websites through API interface.

[0114] Competitor activity: Obtain competitor activity information through public data.

[0115] Economic indicators: Obtain real-time economic indicators from financial data platforms.

[0116] Weather conditions: Obtain real-time weather data from weather forecast websites.

[0117] Holiday: Determine whether it is a holiday according to the calendar.

[0118] Social network topology data collection:

[0119] Node connection relationship: Obtain the attention, friendship and other relationships between users through the API of the social network platform.

[0120] Edge weight: Calculate the weight of the edge according to the interaction frequency and intensity of the user.

[0121] Network propagation data collection:

[0122] Published content: Obtain the published content of the user through the crawler technology or API.

[0123] Number of reposts: Count the number of reposts of each content.

[0124] Number of comments: Count the number of comments of each content.

[0125] Number of likes: Count the number of likes of each content.

[0126] Timestamp: Record the publishing time of each content.

[0127] 2. Key information data identification: Text feature extraction:

[0128] Keyword extraction: Extract keywords from text content using TF-IDF algorithm.

[0129] Word extraction: Extract high-frequency words using word frequency statistics method.

[0130] Topic feature extraction: Extract topic features using LDA (Latent Dirichlet Allocation) algorithm.

[0131] User feature extraction:

[0132] Activity: Count the number of posts, retweets, and comments made by the user within a unit of time.

[0133] Influence: Calculate the number of followers and the number of people who follow the user.

[0134] Interest feature: Extract user interest features through text mining and topic analysis.

[0135] Propagation feature extraction:

[0136] Propagation speed: Calculate the propagation distance or the number of covered nodes within a unit of time.

[0137] Propagation range: Count the number of nodes covered by the information propagation.

[0138] Propagation path: Record the path of information propagation.

[0139] Key information data integration:

[0140] Topic text information: Identify text information related to a specific topic.

[0141] User information data: Determine key users and their propagated information.

[0142] Propagation key information: Identify information with significant propagation effect.

[0143] Integrate key information data: Integrate topic text information, user information data, and propagation key information.

[0144] 3. Reinforced recurrent neural network construction and propagation network feature extraction:

[0145] Define the sand cat population:

[0146] Population size: Set the size of the sand cat population to 50.

[0147] Position vector: The position vector of each sand cat represents the parameters to be optimized in the recurrent neural network.

[0148] Initialize the population: Randomly generate the initial positions of the sand cats within the search space.

[0149] Fitness calculation:

[0150] Loss function: Define the loss function as Mean Squared Error (MSE).

[0151] Fitness value: Calculate the fitness value of each sand cat, and select the position of the sand cat with the highest fitness value as the optimal individual position.

[0152] Position update:

[0153] Current position and random angle: Update the position based on the current position and random angle of the sand cat.

[0154] Updated position: Calculate the updated position of the sand cat and evaluate its fitness value.

[0155] Iterative optimization: Repeat the position update process until the preset number of iterations is reached.

[0156] Propagation network feature extraction:

[0157] Time series feature matrix: Convert key information data into a time series feature matrix, with each time step corresponding to a feature vector of a node.

[0158] Hidden state calculation: At each time step, input the step vector and the hidden state of the previous time step into the hidden layer of the reinforced recurrent neural network to calculate the updated hidden state.

[0159] Output propagation network features: Pass the hidden state to the output layer to obtain the propagation network features.

[0160] 4. Node propagation performance data analysis:

[0161] Node identification and integrated event data:

[0162] Unique identifier: Assign a unique identifier to each node.

[0163] Integrated event data: Associate the topic text information with the propagation events in the propagation key information.

[0164] Propagation time feature analysis:

[0165] Count the number of propagations: Count the total number of nodes involved in the initial information source and subsequent propagation.

[0166] Calculate the average number of propagations: Calculate the average number of propagations for each node in different levels.

[0167] Draw the propagation time series graph: Observe the change of information propagation speed over time.

[0168] Propagation network graph construction:

[0169] Constructing propagation network graph: data points represent users and related entities, edges represent propagation relationships.

[0170] Calculate node centrality indicators: calculate node centrality indicators from degree centrality, closeness centrality, and betweenness centrality.

[0171] Text propagation relationship analysis:

[0172] Key words and sentiment orientation: analyze the relationship between key words and sentiment orientation in topic text information and node propagation performance.

[0173] Node propagation performance data integration:

[0174] Active degree and centrality correlation analysis: correlate user active degree indicators with node centrality indicators.

[0175] Propagation path influence analysis: analyze the influence of text propagation relationship and user social relationship on propagation path.

[0176] 5. Multi-factor influence coefficient extraction:

[0177] Correlation calculation:

[0178] Pearson correlation coefficient: calculate the Pearson correlation coefficient of external factor data and propagation indicators.

[0179] Influence propagation feature extraction:

[0180] Extract influence propagation features: extract influence propagation features from external factor data based on Pearson correlation coefficient.

[0181] Multi-factor model construction:

[0182] Model training: use influence propagation features and propagation indicators as target variables to train multi-factor model.

[0183] Get multi-factor influence coefficient: get multi-factor influence coefficient through model training.

[0184] 6. Node evaluation system construction:

[0185] Evaluation target and index selection:

[0186] Evaluation target: determine evaluation target according to node importance, influence and active degree in information propagation.

[0187] Target index selection: select target index from propagation network features and node propagation performance data.

[0188] Weight calculation and evaluation model construction:

[0189] Constructing the judgment matrix: Construct a judgment matrix to compare the relative importance of each index.

[0190] Calculating weights: Calculate the weight of each index.

[0191] Building a node evaluation model: Build a node evaluation model based on the evaluation target and weight.

[0192] Comprehensive score calculation:

[0193] Calculate the comprehensive score: Calculate the comprehensive score of each node according to the node evaluation model.

[0194] Get node evaluation data: Output the comprehensive score of each node.

[0195] 7. Preliminary node classification and key node screening:

[0196] Classification criteria setting:

[0197] Target features and weights: Set the classification criteria according to the target features and weights.

[0198] Initial node classification: Classify the nodes in the topology structure data to obtain the initial nodes.

[0199] Hierarchical clustering classification:

[0200] Evaluation data classification: Classify the node evaluation data according to hierarchical clustering to obtain the class label of each node.

[0201] Classification result analysis:

[0202] Node distribution analysis: Analyze the distribution of the classified nodes to determine the preliminary nodes.

[0203] Threshold and target setting:

[0204] Threshold and target: Set the screening threshold and screening target according to the target requirements and multi-factor influence coefficient.

[0205] Node arrangement and screening:

[0206] Node arrangement: Arrange all preliminary nodes in ascending order according to the multi-factor influence coefficient.

[0207] Threshold screening: Screen the nodes that meet the threshold from the node arrangement data according to the screening threshold.

[0208] Activity and credibility evaluation:

[0209] Further evaluation: Further evaluate the activity and credibility of the nodes that meet the threshold.

[0210] Extract key nodes: Extract the nodes that meet the screening target as key nodes.

[0211] The system for identifying key nodes of a social network by fusing propagation features provided by the embodiment of the application comprises:

[0212] The multi-data acquisition module is configured to acquire external factor data and topological structure data of the social network in real time, and collect network propagation data.

[0213] The network propagation analysis module is configured to identify key information data from the network propagation data according to a text-line propagation method, construct a reinforced recurrent neural network optimized based on a sand cat group optimization algorithm, input the key information data, output propagation network features, and analyze the key information data to obtain node propagation performance data.

[0214] The information propagation evaluation and screening module is configured to extract a multi-factor influence coefficient from the external factor data, construct a node evaluation system according to the propagation network features and the node propagation performance data, classify the topological structure data to obtain preliminary nodes, and screen the preliminary nodes according to the multi-factor influence coefficient to obtain key nodes.

[0215] The method and system for identifying key nodes of a social network by fusing propagation features provided by the embodiment of the present application fuse external factor data, social network topological structure data and network propagation data, extract features from multiple dimensions such as text content, user behavior and propagation process, comprehensively consider multiple data and features, make the identification of key nodes more comprehensive and accurate, and more truly reflect the role and importance of nodes in network propagation. The sand cat group optimization algorithm is used to optimize the parameters of the reinforced recurrent neural network, and through the steps of defining a sand cat population, initializing a position, calculating fitness, and updating a position, an efficient neural network model is constructed to extract propagation network features, and a node evaluation system is constructed by combining the propagation network features and the node propagation performance data, the evaluation target is determined from the aspects of importance, influence and activity, the corresponding indicators are selected and the weights are calculated, and the comprehensive score of the node is calculated. The optimized reinforced recurrent neural network can more effectively extract the propagation network features, and improve the understanding and prediction ability of the propagation process. The constructed node evaluation system provides a systematic method for comprehensively evaluating the propagation ability of the node, and helps to deeply understand the behavior and influence of the node in the social network.

[0216] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment scheme. Those skilled in the art can understand and implement without creative labor.

[0217] Those skilled in the art can clearly understand the implementation of the embodiments by means of software and necessary general hardware platforms through the description of the above embodiments, and of course, the embodiments can also be implemented by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, and the computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0218] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacement for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for identifying key nodes in a social network with fusion of propagation features, characterized in that, The method comprises the following steps: Real-time acquisition of external factor data and topological structure data of social network, and collection of network propagation data; According to the text line transmission method, the key information data is obtained by identifying the network propagation data, a reinforced recurrent neural network based on the optimization algorithm of sand cat group is constructed, the key information data is input, and the propagation network feature is output, and the node propagation performance data is obtained by analyzing the key information data; The step of identifying the key information data comprises: Extracting keywords, words and theme features from the text content in the network propagation data to form text features, acquiring user features by analyzing user behavior characteristics, and obtaining propagation features by studying the speed, range and path of information propagation process; According to the text feature, the theme text information related to the theme is identified, the key user and the information of the transmission are determined by analyzing the user feature to obtain the user information data, the propagation key information with satisfactory preset propagation effect is identified by analyzing the propagation feature; The theme text information, the user information data and the propagation key information are integrated to obtain the key information data; The step of constructing the reinforced recurrent neural network comprises: Defining the size of the sand cat population, the position vector of each sand cat in the population representing the to-be-optimized parameters of the recurrent neural network, and initializing the sand cat population, the position of each sand cat being randomly generated in the search space; According to the loss function value, the fitness is determined, and the fitness value of each sand cat is calculated, and the sand cat position with the highest fitness value is selected as the optimal individual position; According to the current position of the sand cat and the random angle, the position is updated to obtain the corresponding sand cat updated position; The fitness value of each sand cat updated position is calculated, and the sand cat position with the highest fitness value is selected to replace the optimal individual position until a preset iteration number is reached, and the to-be-optimized parameters corresponding to the optimal individual position are selected to construct the recurrent neural network to obtain the reinforced recurrent neural network; From the external factor data, the multi-factor influence coefficient is extracted, the node evaluation system is constructed according to the propagation network feature and the node propagation performance data, and the topological structure data is classified to obtain preliminary nodes, and the preliminary nodes are screened according to the multi-factor influence coefficient to obtain key nodes.

2. The method of claim 1, wherein, The step of outputting the propagation network feature comprises: The key information data is converted into a time sequence feature matrix, and each time step corresponds to a feature vector of a node; Each time step input vector is taken as a step vector, and at each time step, the step vector and the hidden state of the previous time step are input into the hidden layer of the reinforced recurrent neural network, and the updated hidden state is calculated; The hidden state is transmitted to the output layer for calculation to obtain the propagation network feature.

3. The method of claim 1, wherein, The step of analyzing the node propagation performance data comprises: Each node is assigned a unique identifier, and the theme text information is associated with the propagation events in the propagation key information to obtain integrated event data; Count the initial information source and the total number of nodes participating in the subsequent propagation, calculate the average number of propagations of each node in different levels, draw a propagation time sequence diagram, and observe the change of the propagation speed with time to obtain the propagation time characteristics; Based on the integrated event data, a propagation network graph is constructed, data points represent users and subject-related entities, and edges represent propagation relationships. The node centrality index is calculated from the degree centrality, closeness centrality, and betweenness centrality. The relationship between keywords, sentiment orientation in the subject text information, and node propagation performance is analyzed to obtain the text propagation relationship. The user activity index is associated with the node centrality index, and the influence of the text propagation relationship and the user's social relationship on the propagation path is combined to obtain the node propagation performance data.

4. The method of claim 1, wherein, The step of extracting the multi-factor influence coefficient includes: Calculate the Pearson correlation coefficient of the external factor data and the propagation index; According to the Pearson correlation coefficient, the propagation characteristic is extracted from the external factor data; A multi-factor model is constructed, and the multi-factor model is trained using the propagation characteristic and the propagation index as target variables to obtain the multi-factor influence coefficient.

5. The method of claim 1, wherein, The step of constructing the node evaluation system includes: According to the importance, influence and activity of the node in the information propagation, determine the evaluation target, and select the target index from the propagation network characteristics and the node propagation performance data according to the evaluation target; By constructing a judgment matrix, the relative importance between each index is compared, and the weight of each index is calculated. According to the evaluation target and the weight, a node evaluation model is constructed; According to the node evaluation model, the comprehensive score of each node is calculated to obtain the node evaluation data.

6. The method of claim 5, wherein, The step of classifying the preliminary nodes includes: According to the target characteristics and weights, set the classification standard, and classify the nodes in the topological structure data to obtain the initial nodes; According to the node evaluation data, the evaluation data corresponding to the initial nodes is obtained, and the evaluation data is classified according to the hierarchical clustering method to obtain the class label of each node; According to a plurality of class labels, analyze the distribution of the classified nodes to understand the position, quantity and role of the nodes in the network of different categories, so as to determine the preliminary nodes.

7. The method of claim 1, wherein, The step of obtaining the key nodes includes: According to the target demand and the multi-factor influence coefficient, set the screening threshold and screening target; Arrange all the preliminary nodes in ascending order according to the multi-factor influence coefficient to obtain node arrangement data, and screen the threshold nodes from the node arrangement data according to the screening threshold; Further evaluate the activity and reliability of the threshold nodes, and extract the threshold nodes that meet the screening target to obtain the key nodes.

8. A system for identifying key nodes in a social network with fusion propagation features, adopting the method for identifying key nodes in a social network with fusion propagation features according to any one of claims 1 to 7, characterized in that, The recognition system includes: A multi-data acquisition module is used to collect external factor data and topological structure data of a social network in real time, and collect network propagation data; The network propagation analysis module is configured to identify key information data from the network propagation data according to a text line propagation method, construct a reinforced recurrent neural network optimized based on a sand cat group optimization algorithm, input the key information data, output propagation network features, and analyze the key information data to obtain node propagation performance data. The information propagation evaluation and screening module is configured to extract a multi-factor influence coefficient from the external factor data, construct a node evaluation system according to the propagation network features and the node propagation performance data, classify the topology structure data to obtain preliminary nodes, and screen the preliminary nodes according to the multi-factor influence coefficient to obtain key nodes.

Citation Information

Patent Citations

  • Generative AI emotion propagation prediction and guidance large model construction method and system

    CN119047512A

  • Systems and Methods for Implementing Smart Assistant Systems

    US20220129556A1