Multi-source information rapid traceability method based on user credibility

Through anti-noise clustering and graph neural network methods based on user credibility, the problem of multi-source information traceability in social networks is solved, and the rapid and accurate traceability effect is achieved, reducing the computational complexity and misjudgment rate.

CN120123601APending Publication Date: 2025-06-10UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510115874.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and accurately trace the multi-source information in social networks, especially in the presence of noisy data and abnormal nodes, which affects the accuracy of traceability.

Method used

Through a method based on user credibility, users' characteristics are counted and feature matrix is ​​constructed, the degree of similarity between users is quantified, the anti-noise clustering is performed, the user's credibility score is calculated, the social network is simplified, and the graph neural network is used to trace the multi-source information.

Benefits of technology

It realizes the rapid and accurate discovery of information sources in social networks, reduces the time and computational complexity required for traditional traceability, and improves the accuracy and robustness of traceability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123601A_ABST
    Figure CN120123601A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source information rapid traceability method based on user credibility, and aims to solve the problems that malicious false information spreading of a user is difficult to be effectively restrained and information sources are difficult to detect due to efficient and rapid information spreading characteristics in a social network. According to the method, firstly, a user credibility measuring method based on anti-noise clustering is designed, the characteristics of user nodes influencing user credibility scores in the social network are fully extracted, and on the basis, the anti-noise clustering method is used for achieving the purpose of clustering; secondly, simplifying the social network by calculating the credibility score of the user; then, on the simplified social network, a multi-source information rapid tracing method based on the graph neural network is designed, and the problem that the social network structure changes along with time is solved by introducing a circulating network architecture, so that the accuracy and efficiency of model tracing are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to network information traceability technology, and particularly to multi-source information fast traceability technology. Background Art

[0002] In the context of the current social media and Internet era, the emergence and progress of social networks have profoundly changed people's living habits, making the role of users change from passive receivers to active publishers. The widespread dissemination of false information will cause unacceptable and destructive negative impacts on individuals and society.

[0003] Information may spread false or fake information before, during, and after an event. Through the investigation of the spread of true and false news in social networks, it is found that false news spreads faster than true news, and identifying the source of information can control the spread of information in the network. Therefore, how to analyze the credibility of users on social network platforms, identify users with high credibility among them, and conduct information traceability work on this simplified social network is one of the urgent problems to be solved in building an honest, fair, and healthy social platform.

[0004] With the rapid development of social network platforms and their increasing use in spreading news and other information, the concepts of trust and trustworthiness have been given new meanings in the social environment. User credibility, as an important concept in the research of user activities in social networks, has attracted the attention of many researchers. Previous studies mainly gave methods for evaluating user credibility from three aspects: behavioral characteristics, content attributes, and text-based characteristics.

[0005] (1) Method for measuring user credibility based on behavioral characteristics: Abbasi M A, Liu H. Measuring user credibility in social media[C]. Social Computing, Behavioral-Cultural Modeling and Prediction: 6th International Conference, Washington, DC, USA, 2013: 441-448. By analyzing social network texts, constructing groups of credible, non-credible, and suspicious words, and using a weighted scoring system to evaluate text credibility, but it is no longer applicable to the current multimedia-dominated social network environment because it cannot handle image and video content.

[0006] (2) Method for measuring user credibility based on content features: Zhao L, Hua T, Lu C T, et al. A topic-focused trust model for Twitter[J]. Computer Communications, 2016, 76: 1-11. Calculate user credibility by evaluating the similarity between tweet content and authoritative sources through a model, and use a trust propagation algorithm to analyze the relationships between users (such as semantic links, interactions, friendships, etc.) to map trust transfer. However, the model was only tested in a single region, and there are limitations in regional applicability.

[0007] (3) Method for measuring user credibility based on text features: Ma Z, Gao Q. A text analysis-based method for obtaining credibility assessment of chinese microblog users[C] / / Social Computing and Social Media. Technologies and Analytics: 10th International Conference, SCSM 2018, Held as Part of HCI International 2018, Las Vegas, NV, USA, July 15-20, 2018, Proceedings, Part II 10. Springer International Publishing, 2018: 229-235. Evaluate text credibility using a weighted scoring system by analyzing social network text based on trusted, untrusted, and suspicious vocabulary groups, but it cannot handle image and video content and is no longer applicable to the multimedia-rich social network environment.

[0008] Generally speaking, the problem of information tracing in social networks can often be divided into two situations: single-source tracing and multi-source tracing. Among them, single-source tracing is applicable to scenarios where the information source is clear and the dissemination is relatively simple. Moreover, in real social networks, information often does not start spreading from a single source. Multiple users may start spreading the same or similar false information simultaneously or at different time points. Therefore, information is considered to have the property of multi-source in real social networks. Research on multi-source information tracing can be divided into methods based on ranking, methods based on network partitioning, methods based on label propagation, and methods based on deep learning, etc.

[0009] (1) Ranking-based method: First, a centrality metric is proposed to measure the suspicious levels of all infected nodes, and then the top K nodes are selected as the diffusion sources according to the ranking results.

[0010] (2) Network partitioning-based method: The network partitioning-based method transforms the multi-source problem into multiple single-source localization problems. First, different network partitioning algorithms are used to partition the infected nodes, and then a single-source localization strategy is applied in each partition to identify multiple sources.

[0011] (3) Research on multi-source localization based on label propagation: Wang Z, Wang C, Pei J, et al. Multiplesource detection without knowing the underlying propagation model[C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2017, 31(1). An attempt to identify multiple sources without knowing the underlying model. Set an original label for each node and let the infection status propagate iteratively as a label in the network. Finally, predict the source of the information according to the converged node labels, and use the local peaks of the label propagation results as the source nodes. However, nodes with overly large influence may be misjudged as the source, and it is overly dependent on the currently obtained snapshot.

[0012] (4) Deep learning-based multi-source detection method: Dong M, Zheng B, Quoc Viet Hung N, etal. Multiple rumor source detection with graph convolutional

[0013] networks[C] / / Proceedings of the 28th ACM international conference oninformation and knowledge management. 2019: 569-578. A deep learning-based model is proposed, namely the Graph Convolutional Network-based Source Identification (GCNSI) model, which accurately locates multiple information sources without prior knowledge of the underlying propagation model. However, this method has the problem of insufficient node characteristics itself.

[0014] After the applicant conducted research on existing user credibility evaluation methods and information source detection methods, it was found that most of the existing mainstream user credibility evaluation methods directly measure user credibility in social networks, without considering the existence of noisy data and abnormal nodes in social networks. Existing information tracing methods mainly focus on research in aspects such as information content, infection status, and propagation paths. However, the real social network is extremely large and complex. The dynamic changes in the social network structure may lead to the inability to accurately capture the true propagation path of information; social network data usually undergoes processing or synthesis and cannot fully reflect the true user relationships and propagation paths, thus affecting the accuracy of information tracing; the multiple forwarding and recombination of information complicate the propagation path and increase the difficulty of tracing; and the content analysis-based method lacks universality due to the diversity of information topics and is difficult to be effectively applied in all scenarios. Summary of the Invention

[0015] The problem to be solved by the present invention is to propose a way to narrow the tracing scope based on user credibility, making the tracing process more efficient, so as to be able to discover the source of information more quickly, a fast tracing method.

[0016] The technical solution adopted by the present invention to solve the above technical problems is a multi-source information fast tracing method based on user credibility, including the steps:

[0017] S1. Statistically analyze the user characteristics of all users in the network. The user characteristics include the sum of the degrees of the neighbor nodes of the user, closeness centrality, and the number of forwarding times. For each type of characteristic of all users, perform a normalization operation and assign the same weight to form a characteristic matrix for evaluating user credibility. Each column of the characteristic matrix represents a type of characteristic, and each row is the characteristic value of the three types of characteristics corresponding to a user;

[0018] S2. Quantify the similarity degree between users according to the characteristic matrix;

[0019] S3. Clustering and partitioning step:

[0020] Construct a similarity matrix using the similarity degree between users. Use the preset number of clusters as the number of clusters for spectral clustering, and construct a degree matrix with the sum of the elements in each row of the similarity matrix as the diagonal elements; subtract the degree matrix from the similarity matrix to form a difference matrix and then perform normalization;

[0021] Perform eigenvalue decomposition on the normalized difference matrix to obtain the corresponding eigenvectors; select the eigenvectors of the preset number of clusters to form a characteristic matrix and then perform normalization; each row of the characteristic matrix corresponds to a user, and each column corresponds to a clustering category;

[0022] Perform row clustering based on Kmeans on the normalized feature matrix, treating each row as a sample, and randomly select samples with a preset number of clusters as the initial cluster centers; calculate the Euclidean distance between each sample and each cluster center, assign the sample to the nearest cluster, and for each initially assigned cluster, calculate the average value of its member samples to obtain a new cluster center. Repeat the sample assignment and cluster center update until the clustering ends to obtain the final clustering result;

[0023] S4. User credibility score calculation:

[0024] Determine that the clustering weight corresponding to each clustering is the ratio of the number of sample points of this clustering type to the total number of sample points in all clustering types;

[0025] Integrate the clustering weight of the clustering where the user is located and the distance between the sample corresponding to the user and the cluster center of the clustering where the user is located to obtain the user credibility score;

[0026] S5. Network simplification steps:

[0027] Set a credibility threshold interval according to the user credibility scores of all users, regard the users falling into the credibility threshold interval as high-credibility nodes, and the remaining users as low-credibility nodes; delete the low-credibility nodes in the network structure to complete the network simplification;

[0028] S6. Tracing steps:

[0029] The tracing model EGCU based on the graph neural network generates high-order representations of the nodes in the network at each time step according to the simplified network, and inputs them into a multi-layer perceptron classifier for predicting whether the node is an information source to perform source tracing.

[0030] In addition, a computer system and a computer program product for implementing the above method are also provided.

[0031] Starting from two perspectives of user credibility measurement on the social network and multi-source information tracing, the present invention designs a method for measuring user credibility based on noise-resistant clustering, fully extracts the features of user nodes in the social network that affect the user credibility score, and on this basis, uses the noise-resistant clustering method to achieve the purpose of dividing clusters; and proposes a new formula for calculating user credibility within the cluster to calculate the user credibility score, and introduces a confidence interval to divide high- and low-credibility nodes to achieve the goal of simplifying the social network, thereby reducing the large amount of time consumed by traditional tracing.

[0032] Furthermore, on the simplified social network, a method for quickly tracing multi-source information based on graph neural network is designed. By introducing the recurrent network architecture of graph neural network, the problem of the changing social network structure over time is solved. The tracing model EGCU based on graph neural network adapts to the changes of graph structure by adjusting the parameters of the graph convolutional unit GCN. By means of the recurrent architecture, the gated recurrent unit GRU is used to update the GCN parameters; GRU and GCN are combined, and the units are connected from bottom to top to form a GCN model with double-layer convolution at each time step.

[0033] Advantages of the present invention:

[0034] 1. The present invention proposes a method for quickly tracing multi-source information based on user credibility. By combining user credibility and multi-source tracing, quick tracing is achieved.

[0035] 2. Anti-noise clustering is used to divide clusters, which can effectively discover user groups with similar behaviors or characteristics, decompose the complex network structure into multiple smaller sub-networks, thereby reducing the computational complexity. Moreover, clustering can exclude noise data and abnormal nodes, improving the accuracy and robustness of credibility calculation.

[0036] 3. Based on the fact that anti-noise clustering divides users in the social network into different clusters, the present invention proposes a calculation method for the user credibility scores between different clusters and different nodes in the cluster on the entire complete social network. This formula takes into account the hierarchical structure of the entire social network, including all sample points in the whole network, clusters, and within the clusters.

[0037] 4. A confidence interval is introduced to divide high- and low-credibility nodes, thus simplifying the social network, more comprehensively considering the volatility and uncertainty of data, reducing misjudgment and reducing the tracing time.

[0038] 5. The graph neural network model introducing the recurrent network mechanism can achieve adaptive adjustment, enabling the model to not only focus on the changes of node embeddings, but also pay more attention to its own changes and adaptability, greatly improving the accuracy and speed of model tracing. Description of the Drawings

[0039] Figure 1 It is a schematic diagram of the scenario provided by the embodiment of the present invention;

[0040] Figure 2 It is the flow of the method for quickly tracing multi-source information based on user credibility provided by the embodiment of the present invention. Detailed Embodiments

[0041] The application scenario of the present invention is as Figure 1As shown. Currently, in social networking platforms such as Weibo, WeChat, and QQ, users can freely share and spread various types of content, resulting in users having to judge the authenticity of information on the platform by themselves. Due to the characteristics of social networks, such as high efficiency, speed, and large information dissemination volume, it is difficult to effectively curb the behavior of users maliciously spreading false information, thereby increasing the difficulty for users to identify true and reliable information and providing a breeding ground for the generation and spread of information. To create a good, healthy, and trustworthy social network environment and maintain social stability and public order, the present invention proposes a multi-source information rapid traceability method based on user credibility, aiming to quickly discover and block the information dissemination path in the network.

[0042] User credibility measurement method based on noise-resistant clustering: Graph-structured data is constructed for the data in the social network. Using three features related to user credibility, namely the sum of the degrees of neighbor nodes, closeness centrality, and the number of forwarding times, a feature matrix of this data is constructed. The cosine similarity is used to calculate the similarity degree between users, and a similarity matrix is constructed. Then, the noise-resistant clustering algorithm is used to divide the social network data into different clusters. A new formula for calculating the user credibility score within the cluster is used to measure the user's credibility, and a confidence interval is introduced to determine whether the user node is a high-credibility user node or a low-credibility user node, filtering out low-credibility user nodes, thereby achieving the purpose of simplifying the social network.

[0043] Multi-source information traceability based on graph neural network: On the basis of the social network simplified by user credibility measurement, the network topology of the social network is constructed, and a model for multi-source information traceability based on graph neural network is further constructed.

[0044] Embodiment 1

[0045] As Figure 2 shown is the flow of the method of the present invention, including two main processes: user credibility measurement based on noise-resistant clustering and multi-source information traceability based on graph neural network.

[0046] S1 This invention defines user credibility as the importance or activity level of a user node in a social network. The content posted by important users in a social network is more likely to be widely spread. Considering the significant meaning of user credibility in a social network, it directly affects information dissemination, user interaction, and the overall health and sustainable development of the social network. A user with high credibility can often more easily influence others, and the information they post is also more likely to be accepted and spread by other users. On the contrary, users with low credibility may face the situation where their information is ignored. The feature matrix of the data in the social network contains various features for measuring user credibility. By analyzing these features, the similarity degree between different users can be quantified. First, a feature matrix is constructed to more intuitively present the features of the data. Normalization operations are performed on three feature vectors, namely the sum of the degrees of the neighbor nodes E(v), the closeness centrality C(v), and the number of reposts Re, to eliminate the dimensionality impact between features and assign the same weight to form a feature matrix to evaluate user credibility. Assigning the same weight to the three features assumes that each feature has the same influence on user credibility. This method of weight assignment avoids introducing biases that may be caused by overemphasizing any one feature. Each column of the feature matrix represents a feature vector, and each row is a user and their corresponding three features.

[0047] X = [x 1 , x 2 , …, x n

[0048] Where X is the feature matrix formed by n users in the social network, and x n represents the feature vector of the nth user, that is, x n = [e n , c n , r n , where e n represents the feature vector of the sum of the degrees of neighbor nodes, c n represents the feature vector of closeness centrality, and r n represents the feature vector of the number of reposts.

[0049] The sum of the degrees of neighbor nodes E(v) is the sum of the degrees of all neighbor nodes of this node. The data is processed into an undirected graph topological structure, and the degree of a node is the number of edges connecting this node to other nodes. The higher the sum of the degrees of neighbor nodes, the denser the connection between this node and other nodes, the higher the activity level and influence in the network, and the higher the credibility in the social network.

[0050]

[0051] ​The closeness centrality C(v) represents the position of a node in the entire structure. If the shortest distances between this node and other nodes in the graph structure are all small, then the closeness centrality of this node will be very high, closer to the geometric center position, and has a high credibility in the entire social network.

[0052]

[0053] Where dis(v,u) is the distance between nodes u and v.

[0054] The number of reposts Re: This feature of the number of reposts directly reflects the information dissemination behavior of user nodes in the social network. The more times a node reposts, the stronger its information dissemination ability in the network. The content it publishes may be concerned and spread by more users, and may have higher influence and credibility.

[0055] S2 can quantify the similarity degree between different users by analyzing these three features. The present invention uses the cosine similarity in the similarity algorithm to judge the similarity degree between users. The cosine similarity between user u and user v can be expressed as:

[0056]

[0057] Where e u represents the eigenvector of the sum of the degrees of the neighbor nodes of user u, and c u represents the eigenvector of the closeness centrality, and r u represents the eigenvector of the number of reposts. The subscript v represents the eigenvector corresponding to user v.

[0058] S3 performs anti-noise clustering division

[0059] (1) Since the user groups in social networks may have complex non - convex shapes, traditional distance - based clustering algorithms may not be able to divide clusters of this nature well. The present invention uses noise - resistant clustering for data partitioning. The noise - resistant clustering algorithm can identify user groups with similar characteristics or behavior patterns, and can use the similarity existing among user nodes in the social network data as the basis for clustering, rather than simply measuring based on the Euclidean distance or density between nodes, thus considering global information. At the same time, noise - resistant clustering is robust to noise data to a certain extent, and can effectively exclude noise data in the data network from the clusters, improving the accuracy and robustness of clustering. Use the cosine similarity scores obtained in the previous step to construct a similarity matrix S. Assume the initial number of clusters k, which is the target number of clusters for noise - resistant clustering partitioning. Obtain the diagonal elements of the degree matrix D by summing the elements of each row of the similarity matrix S. Then subtract the degree matrix D from the similarity matrix S to form a difference matrix L, that is, L = D - S, and normalize this difference matrix to avoid the impact on the clustering result caused by different node dimensions.

[0060] (2) Perform eigenvalue decomposition on the normalized difference matrix L to obtain eigenvalues and their corresponding eigenvectors. Select the first k eigenvectors and their corresponding eigenvalues, and form an n×k matrix U, U = [u 1 ,u 2 ,…,u k , where n is the number of user nodes in the entire network, and k is the number of clusters after clustering. And perform normalization processing on the matrix U so that the norm of each eigenvector is 1.

[0061] (3) Use the K - Means method to perform row - clustering on the normalized eigenmatrix U. Consider each row of U as a sample, and randomly select k samples as the initial cluster centers. Calculate the Euclidean distance between each sample and each cluster center, and assign the sample to the nearest cluster. For each initially assigned cluster, calculate the average value of its member samples to obtain a new cluster center. Repeat the sample assignment and cluster center update until the cluster centers no longer change significantly, and end the clustering assignment.

[0062] (4) The initially assumed number of clusters k is not necessarily the optimal value. To determine the optimal number of clusters k, the present invention uses the silhouette coefficient to determine the best_k. The silhouette coefficient can describe the clarity of the contours of each category after clustering, and is an important indicator for measuring the quality of clustering results, comprehensively considering the cohesion and separation, that is, the closeness of the sample points to other samples within the cluster and the distance between the sample points and samples in other clusters. First, perform clustering according to the initial k value to obtain the cluster to which each sample belongs. The silhouette coefficient (i) is as follows:

[0063]

[0064] where a(i) represents the cohesion of sample point i, and its calculation method is as follows:

[0065]

[0066] where j represents other sample points in the same cluster as sample i, and distance represents the distance between sample point i and sample point j. The calculation method of the separation degree b(i) is similar to that of a(i), but based on the definition of the separation degree, it is necessary to traverse other clusters to obtain multiple values {b 1 (i), b 2 (i), …, b k-1 (i)}, and select the minimum value as the final result b(i).

[0067] The smaller the value of a(i), the closer the cluster is. When a(i) < b(i), that is, when the distance within the cluster is less than the distance between clusters, the value of S(i) will tend to 1. The closer it is to 1, the more obvious the silhouette is, that is, the more compact the clustering result is. On the contrary, when a(i) > b(i), that is, when the distance within the cluster is less than the distance between clusters, the value of S(i) will tend to -1. The closer it is to -1, the more blurred the silhouette is, that is, the clustering result is more loose and the clustering effect is worse.

[0068] As the number of clusters k increases, the silhouette coefficient will show a certain change trend, and the optimal k value can be selected through analysis.

[0069] S4-1 Allocate the weights of clusters and sample points within the clusters

[0070] The user credibility is used as an indicator to measure the trustworthiness of user nodes in the network. Social networks usually have a hierarchical structure, including the entire network, clusters, and all sample points within the clusters. Distinguishing the weights of clusters at the entire network level can consider the global relationships between different clusters, help maintain the global consistency and stability of the network, ensure relatively balanced weight allocation between different clusters, and thus improve the accuracy of overall credibility calculation. In the present invention, the weights of different clusters in the entire network are processed as the ratio of the number of sample points within the current cluster to the total number of sample points in all clusters in the entire complete network. Assume that there is a cluster j in the network, and the cluster weight can be expressed as:

[0071]

[0072] m j represents the number of sample points in the jth cluster, The sum of the sample points in all clusters across the entire network, where k is the number of clusters.

[0073] Even sample points within the same cluster have different degrees of importance. Distinguishing the weights of sample points within a cluster can more finely consider local features and relationships to more accurately reflect the differences in credibility among different sample points. The distance between a sample point and the cluster center can be an important parameter for measuring the user credibility score. In the present invention, the distance from each sample point to the cluster center is calculated, and higher weights are assigned to sample points with closer distances. Assume a sample point i located in cluster j, and the cluster center vector of cluster j is C j ={c 1 ,c 2 ,…,c n}}, then the distance between sample point i and the cluster center of cluster j can be expressed as:

[0074]

[0075] Among them, the cluster center represents the position of the cluster center, and the average value of all sample points in the cluster is used as the vector of the cluster center. The vector of the sample point is represented by the elements in the similarity matrix S, and each value in the vector is the similarity score between it and all other nodes. The Euclidean distance can measure the closeness between the sample point and the cluster center. Larger weights are assigned to sample points closer to the cluster center, which can better reflect the importance and representativeness of the sample points within the cluster. Smaller weights are assigned to sample points farther from the cluster center to weaken the influence of marginal sample points on the cluster.

[0076] Considering the normalization problem of user credibility in the entire network, the present invention assumes that the sum of the user credibility scores of all sample points in all clusters is 1, and normalizes the influence factor of the distance to ensure that the sum of all weights is 1. The formula for the user credibility score of sample point i in cluster j can be expressed as:

[0077]

[0078] where j represents the jth cluster, and m j represents that there are m j sample points in the jth cluster.

[0079] Based on the obtained user credibility scores, to simplify the social network, a credibility threshold is needed to determine high-credibility nodes and low-credibility nodes. However, in a real social network, the behavior and characteristics of users may be affected by various factors, resulting in a certain degree of uncertainty in the credibility assessment. Using a single threshold is prone to misjudgment. Therefore, the present invention introduces a confidence interval to reduce the possibility of such misjudgment. The specific method is to use the obtained user credibility scores as samples, calculate their mean value and standard error. An interval is established at a 95% confidence level to ensure that the sample mean is within two standard deviations of the population mean, thereby more accurately judging the credibility of users. The user nodes falling within this credibility threshold interval are regarded as high-credibility nodes, and the remaining nodes are regarded as low-credibility nodes.

[0080] S5 The present invention constructs a traceability model EGCU based on a graph neural network in the simplified social network to predict whether a node is an information source. The main idea of the model is to continuously aggregate and spread the feature information of nodes, generate high-order representations of nodes, and perform information source prediction through a multi-layer perceptron (MLP) classifier. The specific process is as follows:

[0081] First, the initial feature matrix X of the nodes at the first time step 1 and the adjacency matrix A 1 are sent to the first layer of the first graph convolutional network GCN unit, where the GCN unit consists of multiple layers of graph convolutions, and the output of each layer is used as the input of the next layer, so as to perform inter-layer propagation in the GCN model.

[0082] To calculate the features of the nodes themselves, the adjacency matrix A 1 is added with self-loops and normalized.

[0083]

[0084] Among them, D 1 represents the degree matrix of the adjacency matrix A 1 , and I is the identity matrix. The adjacency matrix and the identity matrix are added together to make the diagonal elements 1, so as to include the features of the nodes themselves during feature aggregation. Then, by normalizing the adjacency matrix, a symmetric and normalized matrix is obtained to ensure the robustness of the model to node sets of different scales in the topological graph structure.

[0085] In inter-layer propagation, the adjacency matrix and the node feature matrix of the previous layer are used as inputs, and are converted into new node features through the weight matrix and non-linear activation function ReLU of the GCN That is:

[0086]

[0087] Among them, is the weight matrix of the l-th layer of the GCN at the first time step, is the new node feature matrix of the output at the l-th layer of the GCN at the first time step, and σ is the non-linear activation function ReLU. The GCN operation of each layer updates the node embedding based on the output of the previous layer, so as to gradually obtain the high-order representation of the node in the multi-layer convolution. Generally, That is, the initial feature matrix of the nodes at the t-th time step. At the t-th time step, the l-th layer uses the adjacency matrix A t and the node embedding matrix as inputs, and uses the parameter weight matrix of the GCN to take the node embedding matrix of the l-th layer as the output, which is also used as the input of the (l + 1)-th layer at the same time. Weighing the depth and performance to make full use of the features of the graph structure without overfitting, the present invention adopts a 2-layer GCN.

[0088] To solve the problem of the dynamics of the social network, the parameters of the GCN are adjusted in the model to adapt to the changes in the graph structure. The model uses a gated recurrent unit GRU to update the GCN parameters by means of a recurrent architecture. The present invention combines GRU with GCN and connects the units from bottom to top to form a GCN model with double-layer convolution at each time step. As time goes by, the model unfolds horizontally, and the connections between the units enable information flow, thus constructing a new traceability model EGCU based on the graph neural network. The EGCU model can realize the propagation of information of new nodes along the layers of the GCN and evolve the weight matrix of the GCN over time. The input of the GRU at the t-th step is The output of the final GRU, which is also the output of the EGCU, is x.

[0089] The EGCU model generates high-order representations of nodes at each time step and inputs them into a multi-layer perceptron (MLP) classifier to predict whether the node is the source of information.

[0090] Output = MLP(x)

[0091] where x is the node embedding representation output by the EGCU model, and output represents the prediction of whether the node is the source of information.

[0092] During the model training process, cross-entropy loss is selected to calculate the loss of the model, and the Adam optimization algorithm is used to update the relevant parameters in the model to minimize the loss function. User credibility measurement experiments are carried out using multiple data sets to realize the simplification of the social network. The traceability model EGCU based on the graph neural network is trained on the simplified social network.

[0093] Finally, the model is used for source detection of new information dissemination events.

[0094] Embodiment 2

[0095] A computer system includes a memory, a processor, and a computer program stored on the memory. It is characterized in that the processor executes the computer program to implement the steps of the above method, that is, a computer system for implementing a method for quickly tracing multi-source information based on user credibility.

[0096] Embodiment 3

[0097] A computer program product includes a computer program. It is characterized in that when the computer program is executed by a processor, it implements the steps of the method described in Embodiment 1, that is, a computer program product for implementing a method for quickly tracing multi-source information based on user credibility.

[0098] Those of ordinary skill in the art will realize that the embodiments described herein are to assist the reader in understanding the principles of the present invention and should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A multi-source information rapid tracing method based on user credibility, characterized in that: Includes steps: S1. Count the user features of all users in the network. User features include the sum of the degree of the user's neighbor nodes, proximity centrality, and forwarding times. Normalize each type of feature of all users and assign the same weight to form a feature matrix for evaluating user credibility. Each column of the feature matrix represents a type of feature, and each row is the feature value of the three types of features corresponding to a user. S2, quantify the similarity between users based on the feature matrix; S3, clustering steps: A similarity matrix is ​​constructed using the similarity between users, the preset number of clusters is used as the number of clusters for spectral clustering, and a degree matrix is ​​constructed whose diagonal elements are the sum of the elements in each row of the similarity matrix; The difference between the degree matrix and the similarity matrix is ​​used to form a discriminability matrix and then normalized; Perform eigenvalue decomposition on the normalized discriminability matrix to obtain the corresponding eigenvector; select the eigenvectors of the preset number of clusters to form a feature matrix and then normalize it; each row of the feature matrix corresponds to a user, and each column corresponds to a cluster category; Perform Kmeans-based row clustering on the normalized feature matrix, treat each row as a sample, and randomly select samples with a preset number of clusters as the initial cluster centers; Calculate the Euclidean distance between each sample and the center of each cluster, assign the sample to the nearest cluster, and for each initially assigned cluster, calculate the average value of its member samples to obtain a new cluster center. Repeat the sample assignment and cluster center update until clustering is completed to obtain the final clustering result. S4. User credibility score calculation: Determine the cluster weight corresponding to each cluster as the ratio of the number of sample points of the cluster type to the sum of the sample points in all cluster types; The user's credibility score is obtained by combining the cluster weight of the user's cluster and the distance between the user's corresponding sample and the cluster center of the cluster; S5. Network simplification steps: A credibility threshold interval is set according to the user credibility scores of all users, and users falling into the credibility threshold interval are regarded as high credibility nodes, and the remaining users are regarded as low credibility nodes; the low credibility nodes are deleted from the network structure to complete the network simplification; S6. Traceability steps: The graph neural network-based tracing model EGCU generates high-order representations of nodes in the network at each time step according to the simplified network, and inputs them into a multi-layer perceptron classifier for predicting whether the node is the source of information for source tracing.

2. The method according to claim 1, characterized in that: In S2, the cosine similarity is calculated based on the feature matrix as the similarity between users. The cosine similarity cos(u,v) between any two users, user u and user v, is expressed as: where e u The eigenvector representing the sum of degrees of neighbor nodes of user u, c u The feature vector representing the proximity centrality of user u, r u The feature vector representing the number of forwarding times of user u, and the subscript v represents the feature vector corresponding to user v.

3. The method according to claim 1, characterized in that The matrix U is an n×k matrix, U=[u1,u2,…,u k ], where n is the number of users in the entire network, k is the number of clusters after clustering, that is, the number of clustering categories; u1, u2, ..., u in the normalized feature matrix k The modulus of each eigenvector is 1.

4. The method according to claim 1, characterized in that: The optimal value of the number of clusters is determined by the silhouette coefficient, and the number of clusters corresponding to the value of the silhouette coefficient closest to 1 is selected as the optimal value.

5. The method according to claim 1, characterized in that: The cluster weight corresponding to cluster j is expressed as: Where m j It represents the number of sample points in cluster j, and k is the number of clusters after clustering.

6. The method according to claim 5, characterized in that User credibility score R of user i in cluster j ij It is expressed as: where d ij Represents the distance between the sample corresponding to user i and the cluster center of cluster j.

7. The method according to claim 1, characterized in that: A 95% confidence interval is taken from the user credibility scores of all users as the credibility threshold interval.

8. The method according to claim 1, characterized in that: The traceability model EGCU based on graph neural network adapts to changes in graph structure by adjusting the parameters of the graph convolution unit GCN. It uses the gated recurrent unit GRU to update the GCN parameters with the help of the recurrent architecture. It combines GRU with GCN and connects the units from bottom to top to form a GCN model with double-layer convolution at each time step.

9. A computer system comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method of claim 1.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claim 1 are implemented.

Citation Information

Patent Citations

  • Self-adaptive information spreading source detection method

    CN114693464A

  • Block chain-based asset circulation traceability method, apparatus and device, and storage medium

    CN117829990A