False information propagation source tracing method and device

By constructing feature representation matrix and feature vector matrix in hypergraph social network, and using hypergraph convolution and state space model, the existing false information dissemination source traceability methods are solved, and more efficient and accurate dissemination source traceability is achieved.

CN120198140APending Publication Date: 2025-06-24NORTHWESTERN POLYTECHNICAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510341622.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-24

Smart Images

  • Figure CN120198140A_ABST
    Figure CN120198140A_ABST
Patent Text Reader

Abstract

The invention provides a false information propagation source tracing method and device, relates to the technical field of propagation dynamics, and is used for solving the problems of low accuracy and relatively complex model of the existing false information propagation source tracing method. Comprising the following steps: reversely inputting a feature representation sequence corresponding to a first user node into a state space model to sequentially obtain a new intermediate state sequence and a new output sequence; updating the intermediate state of the first user node based on the state of the neighbor node and the weight of the hyperedge to obtain an updated intermediate state sequence and an updated output sequence; reversely mapping the updated output sequence into a two-dimensional vector, and performing softmax operation on the two-dimensional vector to obtain a propagation source probability or a non-propagation source probability corresponding to each updated output; and according to the user nodes included in the propagation source, the user nodes included in the non-propagation source and a loss function formula, determining training loss of a final output feature vector of the user nodes, and when the training loss converges, determining a false information propagation source.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of propagation dynamics, and more particularly to a method and device for tracing the source of false information propagation. Background Art

[0002] With the development of society and technology, the emergence of various social platforms has brought great convenience. At the same time, the spread of false information has brought great resistance to the development of society, and false information that has a major social impact is often caused by one or several sources of dissemination. Therefore, it is of great significance to timely and effectively locate the source of false information dissemination so as to control the dissemination process.

[0003] The dissemination phenomena in real life can be simulated by different dissemination models. For example, there are SI (Susceptible-Infected in English) model, SIR (Susceptible-Infected-Recovery in English) model, and SIS (Susceptible-Infected-Susceptible) model based on infection. There are also IC (Independent Cascade) model and LT (Linear Threshold) model based on influence.

[0004] For the SI model, each user in the initial social network is in an uninfected state, that is, they all do not know the false information. From a certain moment, one or more users in the social network start to spread the false information to the users with whom they have a connection. These users who first spread the false information are considered as the sources of dissemination. At each subsequent moment, the infected users will spread the false information to the users with whom they have a connection, and the probability that an uninfected user is infected is p. For the SIR model, the difference from the SI model is that users can "recover", that is, an infected user confirms that the message is false information, and from then on, they no longer believe the message and no longer participate in the dissemination process. However, in the SIS model, the recovered users can be infected again. In addition to the dissemination models based on infection, the dissemination models based on influence are also widely used. For example, the IC (Independent Cascade) model and the LT (Linear Threshold) model. In the IC model, each infected user will only spread the false information to the neighbors once. In the LT model, for a certain user, when the influence from all infected neighbors is greater than a given threshold, the user is infected, that is, the user believes the false information under the influence of the neighbors.

[0005] For non-machine learning methods, Pinto et al. proposed a method to locate the source of spread by arranging "observation points" in the social network to record the infection time. The specific approach is as follows: First, select a part of the users in the social network as observation points. After the false information starts to spread, record the infection time of these users and denote it as the observation time vector. During the location process, traverse each user in the social network, assume it is the source of spread and trigger the spread process. The time when the false information spreads to each observation point forms another theoretical time vector. Calculate the similarity of these two vectors using the probability density function of the multivariate normal distribution, and the user with the highest similarity is considered the source of spread. Wang et al. proposed a method of iterating the label value based on the Source Centrality Theory to make the label value of the user at the center of the infected area locally maximum for locating the source of spread. The specific approach is as follows: When a certain proportion of users in the social network are infected, stop the spread, obtain a snapshot of the social network at this time, and assign a label value to each user according to their status (+1 for infected users and -1 for non-infected users). Then traverse each user in the social network for the label value iteration process until the label value of each user in the social network no longer changes, and select the user with the largest local label value as the source of spread.

[0006] In recent years, machine learning methods have developed rapidly and been applied to many fields. Dong et al. first proposed a method for tracing the source of spread based on machine learning. The specific approach is as follows: When a certain proportion of users in the social network are infected, stop the spread, obtain a snapshot of the social network at this time, then calculate the user features through the LPSI algorithm, and then input the obtained feature vector into a neural network model, and select the source of spread through classification in the last layer of the model. By learning the spread process of information in the social network, Wang et al. proposed the IVGD method. The specific approach is as follows: First, learn the spread process of information in the social network, then consider the spread tracing process as the inverse process of the spread process for learning, and finally introduce an error compensation mechanism among the calculated users to select the real source of spread.

[0007] However, the existing methods have the following deficiencies:

[0008] 1. Most of the methods for tracing the source of spread are based on a single network snapshot. These methods either rely heavily on manual features or first fit the spread model based on the network snapshot and then perform reverse tracing, and there will be additional errors in the process of fitting the spread model.

[0009] 2. When considering multiple network snapshots, the existing sequence models have high complexity and do not have the ability to perceive the network topology. Summary of the Invention

[0010] An embodiment of the present invention provides a method and device for tracing the source of false information, which are used to solve the problems of low accuracy and complex models existing in the existing methods for tracing the source of false information.

[0011] An embodiment of the present invention provides a method for tracing the source of false information, including:

[0012] When it is determined that a user node in the hypergraph social network is infected with false information, obtain a network snapshot of the hypergraph social network, and based on the status information of the user node, the neighbor node information of the user node, the social network structure information of the user node, and the false information propagation information of the user node, obtain the feature vector of the user node corresponding to each user under each network snapshot and the feature vector matrix corresponding to each network snapshot;

[0013] Perform hypergraph convolution preprocessing on the feature vector matrix included in each network snapshot to obtain the feature representation matrix included in each network snapshot and the feature representation of each user node, where the feature representation is used to describe the connection relationship between the user node and the hyperedge, how many hyperedges the user node is included in, and the relationship between how many user nodes the hyperedge includes;

[0014] Reverse input the feature representation sequence corresponding to the first user node into the state space model to sequentially obtain the initial intermediate state sequence and the initial output sequence of the first user node; the initial intermediate state and the initial output are respectively based on the convolution kernel operation to obtain a new intermediate state sequence and a new output sequence;

[0015] Based on the connection relationship between the user node and the hyperedge, how many hyperedges the user node is included in, the relationship between how many user nodes the hyperedge includes, and the new intermediate state of the first user node at the previous moment, obtain the state of the neighbor node at the current moment; determine the weight of each hyperedge at the current moment according to the state of the neighbor node at the current moment, the linear convolution operation function, and the activation function; update the intermediate state of the first user node based on the state of the neighbor node and the weight of the hyperedge to obtain an updated intermediate state sequence and an updated output sequence;

[0016] Reverse map the updated output sequence into a two-dimensional vector, perform a softmax operation on the two-dimensional vector to obtain the propagation source probability or non-propagation source probability corresponding to each updated output; according to the user nodes included in the propagation source, the user nodes included in the non-propagation source, and the loss function formula, determine the training loss of the final output feature vector of the user node, and when the training loss converges, determine the source of the false information propagation.

[0017] Preferably, performing hypergraph convolution preprocessing on the feature vector matrix included in each network snapshot to obtain the feature representation matrix included in each network snapshot and the feature representation of each user node specifically includes:

[0018] Determine the first-layer feature representation matrix and the feature representation of the first-layer user nodes under the first network snapshot according to the feature vector matrix included in the first network snapshot, the connection relationship between user nodes and hyperedges in the hypergraph social network, and the degree matrix of user nodes;

[0019] Determine the feature representation matrix of the second layer and the feature representation of the second-layer user nodes under the first network snapshot according to the feature representation matrix of the first-layer user nodes under the first network snapshot, the connection relationship between user nodes and hyperedges in the hypergraph social network, and the degree matrix of user nodes;

[0020] The feature representation matrix of the first layer and the feature representation matrix of the second layer are respectively shown as follows:

[0021]

[0022] Among them, X (l+1) represents the feature representation matrix of the (l + 1)-th layer, X (l) represents the feature representation matrix of the l-th layer, H represents the connection relationship between user nodes and hyperedges in the hypergraph social network, D V represents the degree matrix of user nodes, D E represents the degree matrix of hyperedges, W (0) represents the initial trainable parameter, W (l) represents the trainable parameter of the l-th layer, σ(·) represents the activation function, X (1) represents the feature representation matrix of the first layer, represents the feature vector matrix included in the first network snapshot.

[0023] Preferably, the feature representation sequence corresponding to the first user node is reversely input into the state space model to obtain the initial intermediate state sequence and the initial output sequence of the first user node in sequence, specifically including:

[0024] Obtain the feature representations of the first user node in different network snapshots, form a feature representation sequence with the multiple feature representations corresponding to the first user node according to time, and in the reverse order of the feature representation sequence, obtain the initial intermediate state and the initial output of the first user node at different times through the following formula in sequence. The multiple initial intermediate states are arranged in chronological order to form an initial intermediate state sequence, and the multiple initial outputs are arranged in chronological order to form an initial output sequence;

[0025]

[0026] y t = Ch t + Dx t

[0027] Among them, C, D are parameter matrices, xt represents the input at time t, h t-1 represents the initial intermediate state at time t-1, h t represents the initial intermediate state at time t, y t represents the initial output at time t.

[0028] Preferably, the state of the neighbor nodes at the current moment is determined by the following formula:

[0029]

[0030] The weight of each hyperedge at the current moment is determined by the following formula:

[0031] Ω e = sigmoid(σ(MLP(h t-1 )))

[0032] Updating the intermediate state of the first user node based on the state of the neighbor nodes and the weight of the hyperedges specifically includes:

[0033] Updating the intermediate state of the first user node through the following formula to obtain the updated intermediate state:

[0034]

[0035] where h N represents the state of the neighbor nodes, H represents the connection relationship between user nodes and hyperedges in the hypergraph social network, D V is the degree matrix of the user node, indicating how many hyperedges the user node is included in, D E is the degree matrix of the hyperedge, indicating how many user nodes the hyperedge contains, h t-1 represents the initial intermediate state at time t-1, Ω e represents the weight of each hyperedge, MLP(·) represents the linear convolution operation, σ(·) represents the activation function, sigmoid(·) represents the operation that maps the data within the interval (0,1), h′ t represents the updated intermediate state at time t, x t represents the input at time t.

[0036] Preferably, the eigenvector is:

[0037]

[0038] where X i represents the eigenvector of user node i, represents the state information of user node i, represents the neighbor node information of user node i, It represents the social network structure information of user node i. It represents the false information propagation information of user node i; the ‖·‖ symbol represents vector concatenation, R is a positive integer, and 0 < R < 5.

[0039] Preferably, the loss function is:

[0040]

[0041] Among them, L represents the cross-entropy loss, V represents the set of user nodes in the hypergraph social network, v i represents user node i on the hypergraph social network, v j represents user node j of the hypergraph social network, L i represents the loss value of training user node i this time, L j represents the loss value of training user node j this time, ‖W‖2 represents the 2-norm of the parameter matrix W, λ is equal to 0.0005, |S| represents the number of propagation sources in the false information propagation, and |V| - |S| represents the number of non-propagation sources in the false information propagation.

[0042] An embodiment of the present invention provides a device for tracing the source of false information propagation, including:

[0043] A first obtaining unit, configured to obtain a network snapshot of the hypergraph social network when it is determined that there is a user node in the hypergraph social network infected with false information, and obtain the feature vector of the user node corresponding to each user and the feature vector matrix corresponding to each network snapshot according to the status information of the user node, the neighbor node information of the user node, the social network structure information of the user node, and the false information propagation information of the user node;

[0044] A second obtaining unit, configured to perform hypergraph convolution preprocessing on the feature vector matrix included in each network snapshot to obtain the feature representation matrix included in each network snapshot and the feature representation of each user node, where the feature representation is used to describe the connection relationship between the user node and the hyperedge, how many hyperedges the user node is included in, and the relationship between how many user nodes the hyperedge includes;

[0045] A third obtaining unit, configured to obtain the feature representation sequence corresponding to the first user node and input it reversely into the state space model to sequentially obtain the initial intermediate state sequence and the initial output sequence of the first user node; the initial intermediate state and the initial output are respectively based on convolution kernel operations to obtain a new intermediate state sequence and a new output sequence;

[0046] The fourth obtaining unit determines the state of neighbor nodes at the current moment according to the connection relationship between user nodes and hyperedges, the number of hyperedges containing user nodes, the number of user nodes contained in hyperedges, and the new intermediate state of the first user node at the previous moment; determines the weight of each hyperedge at the current moment according to the state of neighbor nodes at the current moment, the linear convolution operation function, and the activation function; updates the intermediate state of the first user node based on the state of neighbor nodes and the weight of hyperedges to obtain an updated intermediate state sequence and an updated output sequence;

[0047] The determination unit is used to reversely map the updated output sequence into a two-dimensional vector, perform a softmax operation on the two-dimensional vector to obtain the propagation source probability or non-propagation source probability corresponding to each updated output; determine the training loss of the final output feature vector of the user node according to the user nodes included in the propagation source, the user nodes included in the non-propagation source, and the loss function formula, and determine the false information propagation source when the training loss converges.

[0048] Preferably, according to the feature vector matrix included in the first network snapshot, the connection relationship between user nodes and hyperedges in the hypergraph social network, and the degree matrix of user nodes, determine the first-layer feature representation matrix and the feature representation of the first-layer user nodes under the first network snapshot;

[0049] According to the feature representation matrix of the first-layer user nodes under the first network snapshot, the connection relationship between user nodes and hyperedges in the hypergraph social network, and the degree matrix of user nodes, determine the feature representation matrix of the second layer and the feature representation of the second-layer user nodes under the first network snapshot;

[0050] The feature representation matrix of the first layer and the feature representation matrix of the second layer are respectively shown as follows:

[0051]

[0052] Among them, X (l+1) represents the feature representation matrix of the (l + 1)-th layer, X (l) represents the feature representation matrix of the l-th layer, H represents the connection relationship between user nodes and hyperedges in the hypergraph social network, D V represents the degree matrix of user nodes, D E represents the degree matrix of hyperedges, W (0) represents the initial trainable parameter, W (l) represents the trainable parameter of the l-th layer, σ(·) represents the activation function, X (1) represents the feature representation matrix of the first layer, represents the feature vector matrix included in the first network snapshot.

[0053] An embodiment of the present invention further provides a computer device, which includes: a processor and a memory; the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the computer device executes the above-mentioned method for tracing the source of false information dissemination.

[0054] An embodiment of the present invention further provides a computer-readable storage medium, including computer instructions, which when running on a computer device, cause the computer device to execute the above-mentioned method for tracing the source of false information dissemination.

[0055] An embodiment of the present invention provides a method and device for tracing the source of false information dissemination. This method fully considers the user node status information, neighbor node status information, social network structure information, and false information dissemination information in the hypergraph social network, constructs a feature vector matrix and a feature representation matrix represented by user nodes based on the social network snapshot; based on the feature representation matrix, the feature representation sequence of each user node is reversely input into the state space model, enabling the method to have the ability to efficiently learn the rumor propagation pattern; furthermore, a graph-aware state space model is provided, enabling the sequence model to combine network topology structure information when learning the time series forward and backward dependencies, improving the accuracy of rumor propagation source localization on the social network and the applicability of the tracing method in reality. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0057] Figure 1 It is a schematic flowchart of a method for tracing the source of false information dissemination provided by an embodiment of the present invention;

[0058] Figure 2 It is a schematic structural diagram of a device for tracing the source of false information dissemination provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0060] Figure 1Exemplarily, a schematic flowchart of a method for tracing the source of false information provided by an embodiment of the present invention is shown. This method can be applied at least in various source tracing methods, such as tracing the source of disease transmission, etc.

[0061] An embodiment of the present invention provides a method for tracing the source of false information, which proposes a heuristic false information source tracing framework based on a graph neural network on a social network, and to a certain extent solves the contradiction between the accuracy rate and the model complexity faced by existing methods. Specifically, the method includes the following steps:

[0062] Step 101, when it is determined that a user node in the hypergraph social network is infected with false information, obtain a network snapshot of the hypergraph social network, and according to the status information of the user node, the neighbor node information of the user node, the social network structure information of the user node, and the false information propagation information of the user node, obtain the feature vector of the user node corresponding to each user under each network snapshot and the feature vector matrix corresponding to each network snapshot;

[0063] Step 102, perform hypergraph convolution preprocessing on the feature vector matrix included in each network snapshot to obtain the feature representation matrix included in each network snapshot and the feature representation of each user node. The feature representation is used to describe the connection relationship between the user node and the hyperedge, how many hyperedges the user node is included in, and the relationship between how many user nodes the hyperedge includes;

[0064] Step 103, obtain the feature representation sequence corresponding to the first user node and input it reversely into the state space model to sequentially obtain the initial intermediate state sequence and the initial output sequence of the first user node; the initial intermediate state and the initial output are respectively based on the convolution kernel operation to obtain a new intermediate state sequence and a new output sequence;

[0065] Step 104, according to the connection relationship between the user node and the hyperedge, how many hyperedges the user node is included in, the relationship between how many user nodes the hyperedge includes, and the new intermediate state of the first user node at the previous moment, obtain the state of the neighbor node at the current moment; determine the weight of each hyperedge at the current moment according to the state of the neighbor node at the current moment, the linear convolution operation function, and the activation function; update the intermediate state of the first user node based on the state of the neighbor node and the weight of the hyperedge to obtain an updated intermediate state sequence and an updated output sequence;

[0066] Step 105, reversely map the updated output sequence into a two-dimensional vector, perform a softmax operation on the two-dimensional vector to obtain the propagation source probability or non-propagation source probability corresponding to each updated output; each user node included in the propagation source and the user nodes included in the non-propagation source determine the training loss of the final output feature vector of the user node based on the loss function. When the training loss converges, determine the false information propagation source.

[0067] Before introducing step 101, it is necessary to first introduce the hypergraph social network. A hypergraph is a generalized graph structure. In a hypergraph social network, different from a traditional graph (where each edge connects two nodes), a hyperedge can connect two or more nodes, and can represent the complex relationships between nodes more flexibly. The hypergraph social network is introduced in detail from the following aspects:

[0068] The hypergraph social network can be represented by G(V, E), where V represents the set of nodes in the hypergraph social network, V = {v2, v2, …, v i}, v i represents the user node i in the network, and E represents the set of hyperedges. The hyperedge e ∈ E is a non-empty subset of V. For example, suppose there is a small social platform with 5 user nodes (u1, u2, u3, u4, u5) forming the node set V. The hyperedge e1 contains the user nodes u1, u2, and u3; the hyperedge e2 contains the user nodes u3 and u4; the hyperedge e3 contains the user nodes u4 and u5, and these hyperedges form the set E.

[0069] In step 101, when false information breaks out, the number of users infected by the false information is monitored in real time, and multiple network snapshots are obtained during the spread of the false information. The network snapshots contain information such as network topology, node status, and infection time, providing data for subsequent analysis.

[0070] Specifically, the status information of the user nodes included in each network snapshot can be determined by the following formula:

[0071]

[0072] where, represents the status information of the user node i, Y represents the status of the user node i in the network snapshot, Y i = 1 indicates that the user node i is infected by the false information, and Y i = 0 indicates that the user node i is not infected by the false information.

[0073] Furthermore, the neighbor node information of the user nodes included in each network snapshot, the social network structure information of the user nodes, and the false information propagation information of the user nodes are determined in sequence by the following formula:

[0074]

[0075] where, represents the neighbor node information of the user node i, N(v i ) represents all the neighbor nodes of the user node i, |N(v i )| represents the number of all neighbors of the user node i, and v j represents the user node vi The neighbor node of, Y j Indicates the user node v j The state of, Y j = 1, indicating that the neighbor node is infected, Y j = 0, indicating that the neighbor node is not infected; Indicates the social network structure information of user node i, Indicates degree centrality, The value of is equal to the number of neighbor nodes of user node v i ; Indicates the misinformation propagation information of user node i, T i Indicates the time when user node i is infected by misinformation.

[0076] Furthermore, based on the determined status information of user nodes, neighbor node information of user nodes, social network structure information of user nodes, and misinformation propagation information of user nodes, the feature vectors of user nodes included in each network snapshot are obtained, specifically as follows:

[0077]

[0078] Among them, X i Indicates the feature vector of user node i, Indicates the status information of user node i, Indicates the neighbor node information of user node i, Indicates the social network structure information of user node i, Indicates the misinformation propagation information of user node i; The ‖·‖ symbol represents vector concatenation, R is a positive integer, and 0 < R < 5.

[0079] For example, in the above small social platform, when the scale of infected user nodes reaches 20% (i.e., 1 user), 40% (2 users), and 60% (3 users), three network snapshots are obtained respectively, and the three network snapshots are as follows:

[0080] The first network snapshot (20% infection rate)

[0081] Network topology: The hyperedge e1 connects user nodes u1, u2, and u3; The hyperedge e2 connects user nodes u3 and u4; The hyperedge e3 connects user nodes u4 and u5;

[0082] Node status: Assume that user node u1 is infected, and the remaining user nodes u2, u3, u4, and u5 are not infected;

[0083] Infection time: The infection time t1 of user node u1 = 1.

[0084] Second network snapshot (40% infection rate)

[0085] Network topology: Hyperedge e1 connects user nodes u1, u2, and u3; hyperedge e2 connects user nodes u3, u4; hyperedge e3 connects user nodes u4 and u5;

[0086] Node status: Assume that user nodes u1 and u3 are infected, and the remaining user nodes u2, u4, and u5 are not infected;

[0087] Infection time: The infection time t1 of user node u1 = 1, and the infection time t3 of user node u3 = 2.

[0088] Third network snapshot (60% infection rate),

[0089] Network topology: Hyperedge e1 connects user nodes u1, u2, and u3; hyperedge e2 connects user nodes u3, u4; hyperedge e3 connects user nodes u4 and u5;

[0090] Node status: Assume that user nodes u1, u3, and u4 are infected, and the remaining user nodes u2 and u5 are not infected;

[0091] Infection time: The infection time t1 of user node u1 = 1, the infection time t3 of user node u3 = 2, and the infection time t4 of user node u4 = 3.

[0092] In the embodiment of the present invention, for each user node in each network snapshot, its eigenvector needs to be calculated. In the first network snapshot, for each user node i, the state information of user node i can be calculated respectively through the above formulas (1), (2), (3), and (4) Neighbor node information of user node i Social network structure information of user node i And false information propagation information of the user node Then, according to formula (5), the eigenvector X of each user node i is determined i , therefore, based on the first network snapshot, 5 eigenvectors of 5 user nodes i can be determined.

[0093] In this embodiment, there are a total of three network snapshots and 5 user nodes, so a total of 15 eigenvectors of user nodes are obtained. Further, the state information, neighbor node information, social network structure information, and false information propagation information included in the 5 users in each network snapshot can be combined, that is, the 5 user node eigenvectors are composed to obtain an eigenvector matrix; in this embodiment, the three network snapshots form three eigenvector matrices.

[0094] In step 102, the feature vector matrix included in each network snapshot is preprocessed by hypergraph convolution, and the feature representation matrix included in each network snapshot and the feature representation of each user node can be obtained; the feature representation of each user node is used to describe the connection relationship between the user node and the hyperedge, how many hyperedges the user node is included in, and the relationship between how many user nodes the hyperedge includes.

[0095] In the embodiment of the present invention, for the first time, according to the feature vector matrix included in the first network snapshot, the connection relationship between the user nodes and the hyperedges in the hypergraph social network, and the degree matrix of the user nodes, the first-layer feature representation matrix under the first network snapshot is determined by the following formula:

[0096]

[0097] Further, after obtaining the first-layer feature representation matrix under the first network snapshot, convolution can be performed to obtain the first-layer feature representation matrix under the first network snapshot, specifically:

[0098]

[0099] Among them, X (l+1) represents the feature representation matrix of the (l + 1)-th layer, X (l) represents the feature representation matrix of the l-th layer, H represents the connection relationship between the user nodes and the hyperedges in the hypergraph social network, D V represents the degree matrix of the user nodes, indicating how many hyperedges the user node is included in, D E is the degree matrix of the hyperedges, indicating how many user nodes the hyperedge includes, W (0) represents the initial trainable parameter, W (l) represents the trainable parameter of the l-th layer, σ(·) represents the activation function, X (1) represents the feature representation matrix of the first layer, represents the feature vector matrix included in the first network snapshot.

[0100] It should be noted that after obtaining the first-layer feature representation matrix under the first network snapshot, the feature representation of the user nodes included in the first layer under the first network snapshot can be obtained.

[0101] In the embodiment of the present invention, hypergraph convolution preprocessing can be performed on the feature vector matrix corresponding to each network snapshot. Through hypergraph convolution preprocessing, the node features of the user nodes can include neighbor node and hyperedge information, which can enhance the model's ability to capture the hypergraph structure. It should be noted that since there are multiple network snapshots, and multiple layers of convolution may be performed in each network snapshot to obtain the feature representation matrix under that network snapshot, in order to avoid confusion caused by multiple layers of convolution in multiple network snapshots, the following is represents the feature representation matrix obtained from the first network snapshot, with indicating the feature representation matrix obtained from the first network snapshot, with representing the feature representation matrix obtained from the nth network snapshot.

[0102] For example, in the previous example, the hypergraph social network includes 5 user nodes and 3 hyperedges, and hyperedge e1 contains user nodes u1, u2, and u3; hyperedge e2 contains user nodes u3 and u4; hyperedge e3 contains user nodes u4 and u5. Then, the incidence matrix H, node degree matrix D V and hyperedge degree matrix D E can be obtained.

[0103] Specifically as follows:

[0104]

[0105] For the above incidence matrix H, which describes the connection relationship between user nodes and hyperedges, when H ij = 1, it means that hyperedge j in the hypergraph social network contains user node i; when H ij = 0, it means that hyperedge j in the hypergraph social network does not contain user node i. In this embodiment, hyperedge e1 contains user nodes u1, u2, and u3, so the first three rows in the first column are 1 and the last two rows are 0; correspondingly, hyperedge e2 contains user nodes u3 and u4, so the third and fourth rows in the second column are 1 and the rest are 0; hyperedge e3 contains user nodes u4 and u5, and the fourth and fifth rows in the fourth column are 1 and the rest are 0.

[0106] D V represents the degree matrix of user nodes, which is a diagonal matrix, and the elements on the diagonal represent how many hyperedges user node i is contained in; D E represents the hyperedge degree matrix, which is also a diagonal matrix, and the elements on the diagonal represent how many nodes hyperedge j contains.

[0107] Furthermore, if it is assumed that the feature vector matrix in the first network snapshot is:

[0108]

[0109] Among them, represents the feature vector matrix included in the first network snapshot, which is the initial user feature vector matrix and contains the original feature information of each user node.

[0110] Let W (0) be a randomly initialized matrix, which is shown as follows:

[0111]

[0112] Then, the first-layer feature representation matrix under the first network snapshot can be determined according to formula (6-1).

[0113] Furthermore, the (l + 1)-layer feature representation matrix under the first network snapshot can also be determined according to formula (6-2).

[0114] It should be noted that in this step, the preprocessing of the feature vectors of user nodes is performed for each network snapshot. In this scenario, there are three network snapshots, and each network snapshot contains the feature vectors of 5 user nodes. The user feature vectors in each network snapshot are combined into a matrix. For example, the feature vector matrix of the user nodes in the first snapshot is Each row of this matrix represents the feature vector of a user node, different rows correspond to different user nodes, and the entire vector matrix represents the feature situation of all users under this network snapshot. Therefore, the preprocessing operation takes the network snapshot as the unit and processes the feature vector matrix of all user nodes in a network snapshot as a whole.

[0115] Furthermore, in this embodiment, for the second network snapshot and the third network snapshot, the processing will also be performed according to formula (6), and the corresponding preprocessed feature representation matrices will be obtained respectively.

[0116] It should be noted that when calculating the feature representation matrix of user nodes for the second network snapshot, it can be based on "the feature vector matrix of the second network snapshot ", and then according to the above calculation of the first-layer feature representation matrix under the first network snapshot and the (l + 1)-layer feature representation matrix under the first network snapshot to obtain the (l + 1)-layer feature representation matrix under the second network snapshot That is, the feature vectors of each network snapshot are only updated layer by layer within that snapshot.

[0117] In this embodiment, three feature representation matrices can be obtained from the three network snapshots and The above and respectively represent the feature vector matrix included in the first network snapshot and the feature vector matrix included in the second network snapshot.

[0118] In step 103, the state space model is used to process sequence data, mapping a set of sequence data into another set of sequence data. It describes the state of the system at different times through the initial intermediate state, where the initial intermediate state at the current time is determined by the initial intermediate state at the previous time and the input at the previous time; the initial output represents that the initial output at the current time is obtained by a linear combination of the initial intermediate state at the current time and the input at the current time.

[0119] Obtain the feature representation sequence corresponding to the first user node, and then input it into the state space model in the reverse input manner of the representation sequence, and the initial intermediate state sequence and the initial output sequence of the first user node can be obtained in turn. Here, the first user node is one of multiple user nodes, and in this step, the processing method for each user node is the same as that of the first user node.

[0120] Specifically, the initial intermediate state and the initial output are respectively represented by the following formulas:

[0121]

[0122] y t =Ch t +Dx t (8-2)

[0123] where h1 represents the initial intermediate state at the current time, h0 represents the initial intermediate state at the previous time, x0 represents the input at the previous time, h t represents the initial intermediate state at time t, h t-1 represents the initial intermediate state at time t-1, y t represents the initial output at time t, h t represents the initial intermediate state at time t, x t represents the initial input at time t, C and D are parameter matrices.

[0124] It should be noted that the feature representation matrix is obtained in step 102. One network snapshot corresponds to one feature representation matrix, and each row of each feature representation matrix represents the feature representation of a user node after preprocessing. Therefore, before step 103, it is necessary to first summarize the feature representations corresponding to each user node according to the user nodes and the feature representation matrix, and obtain the feature representation sequence of each user node in turn. The number included in the feature representation sequence is consistent with the number of network snapshots.

[0125] Further, after obtaining the initial intermediate state and the initial output successively according to the above formulas (8-1) and (8-2), the convolution kernels with the same number as the number of feature representations can be obtained through the following formula (8-3), and then each convolution kernel and each feature representation are used to obtain the new intermediate state and the new output through the following formula (8-4).

[0126]

[0127] Among them, represents the convolution kernel, y′ represents the new output, and X′ represents the sequence of feature vectors of user nodes, that is, the feature vectors of user nodes constructed by the formula (5) for each network snapshot, and the connected sequence of feature vectors of user nodes.

[0128] For example, the previous embodiment includes three network snapshots and 5 user nodes. After step 102, three feature representation matrices are obtained and For each user node, a feature representation sequence can be obtained respectively. Here, the first user node is taken as an example for illustration:

[0129] If the feature representation sequence of the first user node is obtained as When inputting the feature representation sequence, it needs to be input in the reverse input manner:

[0130] When t = 1, Calculate and y1 = Ch1 + Dx1 to obtain the initial intermediate state h1 and the initial output y1 at t = 1.

[0131] When t = 2, Calculate and y2 = Ch2 + Dx2; to obtain the initial intermediate state h2 and the initial output y2 at t = 2.

[0132] When t = 3, Calculate and y3 = Ch3 + Dx3; to obtain the initial intermediate state h3 and the initial output y3 at t = 3.

[0133] Further, after obtaining the 3 groups of initial intermediate states h1, h2, h3 and the initial outputs y1, y2, y3 of the first user node i1, 3 elements K1, K2, K3 are determined according to the formula (8-3), and these 3 elements are used to participate in the operations of the subsequent new output and the new intermediate state;

[0134] Three groups of initial intermediate states h1, h2, h3 and three elements K1, K2, K3 of the first user node i1 are respectively input into formula (8-4), and three groups of new intermediate states h′1, h′2, h′3 are obtained in sequence, forming a new intermediate state sequence; three groups of initial outputs y1, y2, y3 of the first user node i1 and three elements K1, K2, K3 determined by formula (8-3) are respectively input into formula (8-4), and then three groups of new outputs y′1, y′2, y′3 are obtained in sequence, forming a new output sequence.

[0135] In step 104, in the hypergraph social network, the state of a user node is often affected by its neighbor nodes. The state of the neighbor nodes at the current moment can be obtained according to the connection relationship between the user node and the hyperedge, the number of hyperedges that the user node is included in, the relationship between the number of user nodes included in the hyperedge, and the new intermediate state of the first user node at the previous moment, as follows:

[0136]

[0137] Among them, h N represents the state of the neighbor nodes, H represents the connection relationship between the user nodes and the hyperedges in the hypergraph social network, D V is the degree matrix of the user nodes, h′ t-1 represents the new intermediate state at time t-1, and D E is the degree matrix of the hyperedges.

[0138] Furthermore, because different hyperedges may have different importance in the information propagation process, calculating the hyperedge weights can determine the influence of different hyperedges on the propagation. In the embodiments of the present invention, according to the state of the neighbor nodes at the current moment, the linear convolution operation function, and the activation function, the weight of each hyperedge at the current moment is determined, as follows:

[0139] Ω e = sigmoid(σ(MLP(h′ t-1 ))) (9-2)

[0140] Among them, Ω e represents the weight of each hyperedge, MLP(·) represents the linear convolution operation, σ(·) represents the activation function, and sigmoid(·) represents an operation that maps the data within the interval (0,1).

[0141] Furthermore, by combining the state of the neighbor nodes and the hyperedge weights, the new intermediate state can be updated to obtain an updated intermediate state sequence and an updated output sequence. Thus, the influence of the network topology structure and the hyperedge weights on the information propagation can be better captured. The new intermediate state update formula is as follows:

[0142]

[0143] where h′ t represents the updated intermediate state at time t, and x t represents the input at time t.

[0144] For example, in the above embodiment, through step 103 of the first user node i, three groups of new intermediate states h′1, h′2, h′3 and new outputs y′1, y′2, y′3 are output, that is, a new intermediate state sequence and a new output sequence. In this embodiment, taking t = 2 as an example, substituting h′1 into the formula Combined with the known D E , D V and H, the neighbor node state can be obtained Similarly, when t = 3, using h′2, the neighbor node state can be obtained

[0145] Furthermore, for each time t (taking values 2, 3), according to formula (9-2), the weight of each hyperedge can be calculated. When t = 2, using the new intermediate state h′1 of the previous time, through formula (9-2), the weight of each hyperedge at this time can be obtained Similarly, when t = 3, the weight of the hyperedge is calculated using h′2

[0146] Furthermore, for r = 2, substituting h′1, x2 (the input feature vector at time r = 2), (the hyperedge weight at t = 2) and the known H, D V , D E , through the formula the updated intermediate state h″2 can be obtained; similarly, when t = 3, using h′2, x3, etc. through the formula h″3 can be obtained.

[0147] Furthermore, according to the output formula, the updated output can be obtained. When t = 2, according to the formula y″2 = Ch″2 + Dx2, the updated output y″2 can be obtained; correspondingly, when t = 3, the updated output y″3 can also be obtained.

[0148] The updated intermediate state sequence and updated output sequence output by this step, which respectively incorporate the network topology structure information and hyperedge weight information, can better reflect the spread of false information in a social network with a complex topology compared to the new intermediate state sequence and new output sequence obtained in step 103, providing a more effective feature representation for subsequent model training and source of spread location.

[0149] In step 105, assuming that after step 104, the first user node obtains the updated output sequences y″1, y″2, and y″3. Since the input sequence is input in reverse, only the last element y″3 of the output sequence is concerned here.

[0150] First, map y″3 to a two-dimensional vector through a linear layer Assume y″3 = [a, b, c] (where a, b, and c are specific values obtained through previous calculations). After the linear layer transformation, we get The linear layer can be expressed as: z = W lin y″3 + b lin , where W lin is the weight matrix of the linear layer, and b lin is the bias vector. Here, the output is mapped to a two-dimensional vector, and this two-dimensional vector can correspond to two categories, such as being a propagation source and not being a propagation source, which is convenient for subsequent calculation of probability distribution and classification judgment.

[0151] Then, perform a softmax operation on and calculate according to the following formula At this time, the obtained and are normalized probability values, such that the sum of all elements is 1, that is

[0152]

[0153] It should be noted that z i represents the i-th element in the vector z, and z j represents the j-th element in the vector z. Through the softmax operation, the elements in the final output vector can be normalized and the sum of all elements is 1.

[0154] In practical applications, the number of false information propagation sources is usually only a few, while the scale of the hypergraph social network is in the thousands or tens of thousands, which leads to the problem of class imbalance. In view of this problem, in the embodiments of the present invention, the loss function is redesigned specifically as follows: In a false information propagation, the number of propagation sources and non-propagation sources are |S| and |V| - |S| respectively. Since |V| - |S| is much larger than |S|, it is necessary to balance the number of these two types of samples, as shown in the following formula:

[0155]

[0156] In formula (10 - 2), for each user node v in the propagation source set S i , and for the users not in the propagation source set, that is, the user nodes v in V - S j, calculate their cross-entropy losses \(L\) i and \(L\) j , the calculation method of the cross-entropy loss is \(L = -\log(x)\times y\), where \(x\) is the probability value predicted by the model, \(y\) is the true label of the sample (1 indicates the source of propagation, 0 indicates not the source of propagation), \(|S|\) represents the set of sources of propagation, \(|V|\) represents the set of all user nodes on the hypergraph social network, \(\|W\|_2\) represents the 2-norm of the parameter matrix \(W\), and \(\lambda\) is equal to 0.0005.

[0157] For example, if \(y''_3 = [0.5, 0.6, 0.7]\), it is mapped to a two-dimensional vector through the linear layer Let the weight matrix of the linear layer bias vector

[0158] According to the formula perform a softmax operation on \(z\):

[0159]

[0160] The probability distribution \([0.39, 0.61]\) can be obtained. Here, the vector elements are converted into probability values, which can intuitively reflect the prediction confidence of the model for each category, that is, the probabilities of the two categories. For example, 0.39 is the probability of "not the source of propagation", and 0.61 is the probability of "the source of propagation".

[0161] Assume that the known set of false information propagation sources \(S=\{u_1, u_3\}\), and the set of social network nodes \(V =\)

[0162] \(\{u_1, u_2, u_3, u_4, u_5\}\). For the user node \(u_1\) (belonging to the set of propagation sources \(S\)), its label \(y = 1\) (indicating the source of propagation), and for the user node \(u_2\) (belonging to \(V - S\)), its label \(y = 0\) (indicating not the source of propagation).

[0163] For the user node \(u_1\), according to the cross-entropy loss formula \(L = -\log(x)\times y\), where \(x\) is the probability value belonging to \(u_1\) obtained after the softmax operation (assumed to be ), then:

[0164]

[0165] For the user node \(u_2\), \(x\) is the probability value belonging to \(u_2\) obtained after the softmax operation (assumed to be ), then:

[0166]

[0167] For other user nodes u3, u4, and u5, their respective loss values L3, L4, and L5 are calculated in a similar manner.

[0168] Given that λ is equal to 0.0005, let the parameter matrix be W, calculate its 2-norm ‖W‖2, and finally, according to the loss function formula, we can obtain:

[0169] Loss = L1 + L3 + 0 + L4 + L5 + 0.0005‖W‖2

[0170] Substitute the specific values of L1, L3, L4, L5, and ‖W‖2 obtained from the calculation into the above formula to obtain the loss value Loss for this training. This loss value can be used for subsequent optimization training of the model to better identify the sources of false information dissemination.

[0171] Furthermore, perform loss training on the result obtained from formula (10 - 2). When the loss training converges or the number of training times reaches the maximum set value, the loss training can be stopped, and the loss training model can be determined as the false information dissemination source model. Based on this false information dissemination source model, for any input test sample, the corresponding dissemination source of the input test sample can be determined.

[0172] Furthermore, Table 1 shows the test results of the false information dissemination tracing method of the embodiment of the present invention on different hypergraph social networks. For model testing, first select several commonly used evaluation indicators, including Accuracy, F-Score, and AUC. Accuracy represents the proportion of correctly classified samples among all samples; F-Score is calculated from Precision and Recall. Precision represents the proportion of true dissemination sources in the calculated dissemination source set, and Recall represents the proportion of correctly located dissemination sources in the true dissemination source set; AUC tests the prediction performance of the model under different classification thresholds. By testing on public datasets, including Zoo, House, NTU2012, Mushroom, ModelNet40, 20News, PubMed, and Walmart, this method has better performance compared to the comparative methods, and the improvement range of the experimental results is approximately 15 - 25%. In addition, Table 2 provides the test results of the method provided by the embodiment of the present invention under different dissemination models. The results show that the method provided by the embodiment of the present invention can obtain good results under different models, demonstrating the robustness of the method.

[0173] Table 1 Test Results of the False Information Dissemination Tracing Method on Different Hypergraph Social Networks

[0174]

[0175] Table 2 Test Results of the False Information Propagation Tracing Method under Different Propagation Models

[0176]

[0177] Based on the same inventive concept, an embodiment of the present invention provides a false information propagation source tracing device. Since the principle of this device for solving technical problems is similar to that of a false information propagation source tracing method, the implementation of this device can refer to the implementation of the method, and the repeated parts will not be elaborated here.

[0178] As Figure 2 shown, the device includes a first obtaining unit 201, a second obtaining unit 202, a third obtaining unit 203, a fourth obtaining unit 204, and a determining unit 205.

[0179] The first obtaining unit 201 is configured to, when it is determined that a user node in the hypergraph social network is infected with false information, obtain a network snapshot of the hypergraph social network, and obtain a feature vector of the corresponding user node for each user under each network snapshot and a feature vector matrix corresponding to each network snapshot according to the status information of the user node, the neighbor node information of the user node, the social network structure information of the user node, and the false information propagation information of the user node;

[0180] The second obtaining unit 202 is configured to perform hypergraph convolution preprocessing on the feature vector matrix included in each network snapshot to obtain a feature representation matrix included in each network snapshot and a feature representation of each user node, where the feature representation is used to describe the connection relationship between the user node and the hyperedge, the relationship between how many hyperedges the user node is included in, and how many user nodes the hyperedge includes;

[0181] The third obtaining unit 203 is configured to reverse input the feature representation sequence corresponding to the first user node into the state space model to sequentially obtain an initial intermediate state sequence and an initial output sequence of the first user node; the initial intermediate state and the initial output are respectively used to obtain a new intermediate state sequence and a new output sequence based on convolution kernel operations;

[0182] The fourth obtaining unit 204 determines the state of the neighbor node at the current moment according to the connection relationship between the user node and the hyperedge, the relationship between how many hyperedges the user node is included in, and how many user nodes the hyperedge includes and the new intermediate state of the first user node at the previous moment; determines the weight of each hyperedge at the current moment according to the state of the neighbor node at the current moment, the linear convolution operation function, and the activation function; updates the intermediate state of the first user node based on the state of the neighbor node and the weight of the hyperedge to obtain an updated intermediate state sequence and an updated output sequence;

[0183] A determination unit 205 is configured to inverse-map the updated output sequence into a two-dimensional vector, perform a softmax operation on the two-dimensional vector to obtain a propagation source probability or a non-propagation source probability corresponding to each updated output; determine the training loss of the final output feature vector of the user node according to the user nodes included in the propagation source, the user nodes included in the non-propagation source, and the loss function formula, and determine the false information propagation source when the training loss converges.

[0184] It should be understood that the units included in the above false information propagation source tracing device are only logical divisions according to the functions implemented by the device. In actual applications, the above units can be superimposed or split. Moreover, the functions implemented by the false information propagation source tracing device provided in this embodiment correspond one-to-one with the false information propagation source tracing method provided in the above embodiment. For the more detailed processing procedures implemented by this device, they have been described in detail in the first method embodiment above and will not be described in detail here.

[0185] Another embodiment of the present invention further provides a computer device, which includes: a processor and a memory; the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the electronic device executes each step of the false information propagation source tracing method in the method flow shown in the above method embodiment.

[0186] Another embodiment of the present invention further provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on a computer device, the computer device is caused to execute each step of the false information propagation source tracing method in the method flow shown in the above method embodiment.

[0187] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0188] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for tracing the source of false information, characterized in that: include: When it is determined that a user node in the hypergraph social network is infected by false information, a network snapshot of the hypergraph social network is obtained, and a feature vector of the user node corresponding to each user under each network snapshot and a feature vector matrix corresponding to each network snapshot are obtained according to the status information of the user node, the neighbor node information of the user node, the social network structure information of the user node, and the false information propagation information of the user node; Perform hypergraph convolution preprocessing on the feature vector matrix included in each network snapshot to obtain the feature representation matrix included in each network snapshot and the feature representation of each user node, wherein the feature representation is used to describe the connection relationship between the user node and the hyperedge, the number of hyperedges included in the user node, and the relationship between the number of user nodes included in the hyperedge; The feature representation sequence corresponding to the first user node is inputted reversely into the state space model, and an initial intermediate state sequence and an initial output sequence of the first user node are obtained in sequence; the initial intermediate state and the initial output are respectively used to obtain a new intermediate state sequence and a new output sequence based on a convolution kernel operation; The state of the neighboring node at the current moment is obtained according to the connection relationship between the user node and the hyperedge, the number of hyperedges included in the user node, the relationship between the number of user nodes included in the hyperedge, and the new intermediate state of the first user node at the previous moment; the weight of each hyperedge at the current moment is determined according to the state of the neighboring node at the current moment, the linear convolution operation function, and the activation function; Based on the states of neighboring nodes and the weights of hyperedges, the intermediate state of the first user node is updated to obtain an updated intermediate state sequence and an updated output sequence; The updated output sequence is reversely mapped into a two-dimensional vector, and a softmax operation is performed on the two-dimensional vector to obtain the probability of a propagation source or a non-propagation source corresponding to each updated output; the training loss of the final output feature vector of the user node is determined according to the user nodes included in the propagation source, the user nodes included in the non-propagation source, and the loss function formula; when the training loss converges, the source of false information propagation is determined.

2. The method for tracing the source of false information propagation according to claim 1, characterized in that: The feature vector matrix included in each network snapshot is preprocessed by hypergraph convolution to obtain the feature representation matrix included in each network snapshot and the feature representation of each user node, including: Determine a first-layer feature representation matrix and a feature representation of first-layer user nodes under the first network snapshot according to the feature vector matrix included in the first network snapshot, the connection relationship between user nodes and hyperedges in the hypergraph social network, and the degree matrix of user nodes; Determine a feature representation matrix of the second layer under the first network snapshot and feature representations of the second layer user nodes according to a feature representation matrix of the first layer user nodes under the first network snapshot, a connection relationship between user nodes and hyperedges in the hypergraph social network, and a degree matrix of user nodes; The feature representation matrix of the first layer and the feature representation matrix of the second layer are respectively as follows: Among them, X (l+1) represents the feature representation matrix of the l+1th layer, X (l) The feature representation matrix H represents the connection relationship between user nodes and hyperedges in the hypergraph social network, D V Denotes the degree matrix of the user node, D E represents the degree matrix of the hyperedge, W (0) represents the initial trainable parameters, W (l) represents the trainable parameters of the lth layer, σ(·) represents the activation function, X (1) represents the first-layer feature representation matrix, The eigenvector matrix included in the first network snapshot is represented.

3. The method for tracing the source of false information propagation according to claim 1, characterized in that: The step of obtaining a feature representation sequence corresponding to the first user node and inputting it reversely into the state space model to sequentially obtain an initial intermediate state sequence and an initial output sequence of the first user node specifically includes: Obtain feature representations of the first user node in different network snapshots, form a feature representation sequence according to time with multiple feature representations corresponding to the first user node, and obtain the initial intermediate state and initial output of the first user node at different times in reverse order of the feature representation sequence through the following formulas, arrange the multiple initial intermediate states in time sequence to form an initial intermediate state sequence, and arrange the multiple initial outputs in time sequence to form an initial output sequence; y t =Ch t +Dx t in, C and D are parameter matrices, x t represents the input at time t, h t-1 represents the initial intermediate state at time t-1, h t represents the initial intermediate state at time t, y t Represents the initial output at time t.

4. The method for tracing the source of false information propagation according to claim 1, characterized in that: The state of the neighbor node at the current moment is determined by the following formula: The weight of each hyperedge at the current moment is determined by the following formula: Oh e =sigmoid(σ(MLP(h t-1 ))) The updating of the intermediate state of the first user node based on the state of the neighboring node and the weight of the hyperedge specifically includes: The intermediate state of the first user node is updated by the following formula to obtain an updated intermediate state: Among them, h N represents the state of neighbor nodes, H represents the connection relationship between user nodes and hyperedges in the hypergraph social network, and D V is the degree matrix of the user node, indicating how many hyperedges the user node is included in, D E is the degree matrix of the hyperedge, indicating how many user nodes the hyperedge contains, h t-1 represents the initial intermediate state at time t-1, Ω e represents the weight of each hyperedge, MLP(·) represents the linear convolution operation, σ(·) represents the activation function, sigmoid(·) represents the operation to map the data into the interval (0,1), and h′ t Indicates updating the intermediate state at time t, x t represents the input at time t, is the parameter matrix.

5. The method for tracing the source of false information propagation according to claim 1, characterized in that: The feature vector of the user node is as follows: Among them, X i represents the feature vector of user node i, Represents the status information of user node i, Represents the neighbor node information of user node i, Represents the social network structure information of user node i, represents the false information propagation information of user node i; ‖· represents vector concatenation, R is a positive integer, and 0 <R<5。 6. The method for tracing the source of false information propagation according to claim 1, characterized in that: The loss function is: Where L represents the cross entropy loss, V represents the set of user nodes in the hypergraph social network, and v i Represents user nodes i, v on the hypergraph social network j Represents user node j of the hypergraph social network, L i Indicates the loss value of user node i in this training, L j represents the loss value of user node j used in this training, ‖W‖2 represents the 2-norm of the parameter matrix W, λ is equal to 0.0005, |S| represents the number of propagation sources in the propagation of false information, and |V|-|S| represents the number of non-propagation sources in the propagation of false information.

7. A device for tracing the source of false information, characterized in that: include: A first obtaining unit is used to obtain a network snapshot of the hypergraph social network when it is determined that a user node in the hypergraph social network is infected by false information, and obtain a feature vector of the user node corresponding to each user under each network snapshot and a feature vector matrix corresponding to each network snapshot according to the state information of the user node, the neighbor node information of the user node, the social network structure information of the user node and the false information propagation information of the user node; A second obtaining unit is used to perform hypergraph convolution preprocessing on the feature vector matrix included in each network snapshot to obtain a feature representation matrix included in each network snapshot and a feature representation of each user node, wherein the feature representation is used to describe the connection relationship between the user node and the hyperedge, the number of hyperedges included in the user node, and the relationship between the number of user nodes included in the hyperedge; A third obtaining unit is used to obtain a feature representation sequence corresponding to the first user node and input it into the state space model in reverse order to obtain an initial intermediate state sequence and an initial output sequence of the first user node in sequence; the initial intermediate state and the initial output are respectively used to obtain a new intermediate state sequence and a new output sequence based on a convolution kernel operation; The fourth obtaining unit obtains the state of the neighbor node at the current moment according to the connection relationship between the user node and the hyperedge, the number of hyperedges included in the user node, the relationship between the number of user nodes included in the hyperedge, and the new intermediate state of the first user node at the previous moment; determines the weight of each hyperedge at the current moment according to the state of the neighbor node at the current moment, the linear convolution operation function, and the activation function; Based on the states of neighboring nodes and the weights of hyperedges, the intermediate state of the first user node is updated to obtain an updated intermediate state sequence and an updated output sequence; A determination unit is used to reversely map the updated output sequence into a two-dimensional vector, perform a softmax operation on the two-dimensional vector, and obtain a propagation source probability or a non-propagation source probability corresponding to each updated output; determine the training loss of the final output feature vector of the user node according to the user nodes included in the propagation source, the user nodes included in the non-propagation source, and the loss function formula; when the training loss converges, determine the source of false information propagation.

8. The false information propagation tracing device according to claim 7, characterized in that: The second obtaining unit is specifically used for: Determine a first-layer feature representation matrix and a feature representation of first-layer user nodes under the first network snapshot according to the feature vector matrix included in the first network snapshot, the connection relationship between user nodes and hyperedges in the hypergraph social network, and the degree matrix of user nodes; Determine a feature representation matrix of the second layer under the first network snapshot and feature representations of the second layer user nodes according to a feature representation matrix of the first layer user nodes under the first network snapshot, a connection relationship between user nodes and hyperedges in the hypergraph social network, and a degree matrix of user nodes; The feature representation matrix of the first layer and the feature representation matrix of the second layer are respectively as follows: Among them, X (l+1) represents the feature representation matrix of the l+1th layer, X (l) The feature representation matrix H represents the connection relationship between user nodes and hyperedges in the hypergraph social network, D V Denotes the degree matrix of the user node, D E represents the degree matrix of the hyperedge, W (0) represents the initial trainable parameters, W (l) represents the trainable parameters of the lth layer, σ(·) represents the activation function, X (1) represents the first-layer feature representation matrix, The eigenvector matrix included in the first network snapshot is represented.

9. A computer device, characterized in that: The computer device includes: a processor and a memory; the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the computer device executes the false information propagation tracing method as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that: It includes computer instructions, which, when executed on a computer device, enable the computer device to execute the false information propagation tracing method as described in any one of claims 1 to 6.