Social network-oriented text steganalysis data set construction method
By building a heterogeneous information network and a generative text steganography model of social networks, combining metapathic paths and three-dimensional dynamic regulation strategies, a steganography text data set that conforms to the real social network model is solved, and the accuracy and robustness of steganography analysis are improved.
Patent Information
- Application Number
- CN202510468647.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-29
AI Technical Summary
The existing text steganography analysis dataset cannot truly reflect the topological structure, text fragmentation and steganography text sparsity of social networks, and cannot effectively support the steganography analysis research of social networks.
A heterogeneous information network for social network platforms is built, and a generative text steganography model is used to generate multi-type steganography text. Through meta-path local group discovery and three-dimensional dynamic regulation strategies, a steganography text data set that conforms to the real social network model is generated.
The built data set can truly reflect the fragmented features and steganographic text sparsity of social networks, provide more comprehensive feature representation and model training data, improve the accuracy and robustness of steganographic analysis algorithms, and support network security and social stability.
Smart Images

Figure CN120387438A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text steganography analysis, in particular to a method for constructing a text steganography analysis data set for social networks, a device for constructing a text steganography analysis data set for social networks, an electronic device, and a computer-readable medium. Background Art
[0002] Existing text steganography analysis data sets such as T-Steg, TStego-THU, and Stego-Sandbox, although promoting the development of the text steganography analysis field to a certain extent, still have problems such as the lack of social network topologies, text fragmentation, and the mismatch between the sparsity of stego texts and real social networks, and cannot truly reflect the complexity and diversity of social network environments, thus greatly limiting the development of steganography analysis technologies for social networks. Studying a method for constructing a text steganography analysis data set for social networks and constructing a data set that conforms to the real social network model and supports text steganography analysis research in scenarios where text is fragmented and stego texts are extremely sparse is particularly important. Summary of the Invention
[0003] In view of the above problems, the present invention is proposed to provide a method for constructing a text steganography analysis data set for social networks and a corresponding device for constructing a text steganography analysis data set for social networks, an electronic device, and a computer-readable medium that overcome the above problems or at least partially solve the above problems.
[0004] The present invention discloses a method for constructing a text steganography analysis data set for social networks, the method comprising:
[0005] Constructing a heterogeneous information network of a social network platform;
[0006] Training a generative text steganography model using the data of the heterogeneous information network, and combining multiple steganography algorithms and different embedding capacity parameters to generate multiple types of stego texts to form a stego text library;
[0007] By defining multi-node association paths and combining a random walk strategy guided by transition probabilities, sampling user groups with associated interaction behavior characteristics from the heterogeneous information network;
[0008] Based on the stego text library, reconstructing the social posts published by the user groups with associated interaction behavior characteristics using a three-dimensional dynamic regulation stego text replacement strategy to generate social posts of a social network covert communication group;
[0009] After reconstruction, delete the original association relationships of the replaced social posts in the heterogeneous information network, retain the entities and association relationships in the heterogeneous information network that have not been modified, update the heterogeneous information network structure, and form a social network text steganography analysis dataset.
[0010] Optionally, construct a heterogeneous information network of a social network platform, including:
[0011] For a social network platform, based on breadth-first search, distribution diversity strategy, and numerical diversity strategy, collect and expand users to form a user network with users as nodes and specific association relationships as edges, and on the basis of the user network, collect and expand user-related entities and association relationships to form a heterogeneous information network containing multiple entities and association relationships.
[0012] Optionally, use the data of the heterogeneous information network to train a generative text steganography model, combine multiple steganography algorithms and different embedding capacity parameters to generate multiple types of steganographic texts, and form a steganographic text library, including:
[0013] Use the data of the heterogeneous information network to train a generative text steganography model;
[0014] The trained generative text steganography model uses arithmetic coding steganography algorithm, adaptive dynamic grouping steganography algorithm, and variable-length coding steganography algorithm to perform steganographic embedding on the secret information, and controls the text hiding capacity by adjusting the number of bits embedded per word, generates multiple types of steganographic texts, and forms a steganographic text library.
[0015] Optionally, by defining multi-node association paths and combining the random walk strategy guided by transition probabilities, sample user groups with associated interaction behavior characteristics from the heterogeneous information network, including:
[0016] Based on multiple types of entity nodes and the association relationships of entity nodes in the heterogeneous information network, construct a meta-path describing the indirect interaction formed by users through multi-layer interaction behaviors;
[0017] Calculate the association strength between nodes according to the meta-path, allocate transition probabilities based on the association strength and set the restart probability to constrain the walking range, randomly select unassociated seed users, collect user nodes in the path through the random walk guided by transition probabilities, stop walking after reaching the preset number, merge the sampling results of different seed users and remove duplicates to obtain user groups with associated interaction behavior characteristics.
[0018] Optionally, based on the steganographic text library, use a three-dimensional dynamic regulation steganographic text replacement strategy to reconstruct the social posts published by the user groups with associated interaction behavior characteristics to generate social posts of a social network covert communication group, including:
[0019] Based on the steganographic text library, by separately adjusting the proportion of steganographic text, the type of steganographic text, and the distribution of steganographic text, reconstruct the social posts published by the user group with associated interaction behavior characteristics to generate social posts of a social network covert communication group.
[0020] Optionally, based on the steganographic text library, by separately adjusting the proportion of steganographic text, the type of steganographic text, and the distribution of steganographic text, reconstruct the social posts published by the user group with associated interaction behavior characteristics to generate social posts of a social network covert communication group, including:
[0021] Set multiple proportions of steganographic text, select a combination of single or multiple steganographic algorithms and steganographic text types with single or multiple embedding capacities, and set the distribution mode of steganographic text in user social posts;
[0022] Calculate the number of social posts to be replaced according to the set proportion of steganographic text, generate corresponding subsets of steganographic text by combining steganographic text in the steganographic text library according to the type combination, and insert the subsets of steganographic text into the specified positions of user social posts according to the distribution mode to obtain social posts of a social network covert communication group.
[0023] Optionally, the generative text steganographic model to be trained is an RNN generative text steganographic model.
[0024] The present invention also discloses a device for constructing a text steganographic analysis dataset for a social network, and the device includes:
[0025] A heterogeneous information network construction module for constructing a heterogeneous information network of a social network platform;
[0026] A steganographic text library generation module for training a generative text steganographic model using the data of the heterogeneous information network, combining multiple steganographic algorithms and different embedding capacity parameters to generate multiple types of steganographic text to form a steganographic text library;
[0027] A local group discovery module for sampling user groups with associated interaction behavior characteristics from the heterogeneous information network by defining multi-node association paths and combining a random walk strategy guided by transition probabilities;
[0028] A social post reconstruction module for reconstructing the social posts published by the user group with associated interaction behavior characteristics based on the steganographic text library by using a three-dimensional dynamic regulation steganographic text replacement strategy to generate social posts of a social network covert communication group;
[0029] The social network text steganography analysis dataset integration module is used to reconstruct and then delete the original association relationships of the replaced social posts in the heterogeneous information network, retain the entities and association relationships in the heterogeneous information network that have not been modified, update the heterogeneous information network structure, and form a social network text steganography analysis dataset.
[0030] The present invention also discloses an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0031] The memory is used to store a computer program;
[0032] When the processor is used to execute the program stored on the memory, it realizes the method for constructing a social network-oriented text steganography analysis dataset as described in the present invention.
[0033] The present invention also discloses one or more computer-readable media, on which instructions are stored. When executed by one or more processors, the processors are caused to execute the method for constructing a social network-oriented text steganography analysis dataset as described in the present invention.
[0034] The present invention has the following advantages:
[0035] The method for constructing a social network-oriented text steganography analysis dataset of the present invention constructs a heterogeneous information network for a social network platform, and uses the heterogeneous information network data to train a generative text steganography model. Through the model, secret information is processed to generate multiple types of steganographic texts, forming a social network steganographic text library. The local community discovery technology based on meta-paths is used to sample special user groups to ensure that the proportion of this group in social network users is sparse. Then, a three-dimensional dynamic regulation user social post reconstruction strategy is designed to reconstruct the social posts of the sampled users from three dimensions: the proportion, type, and distribution of steganographic texts, so as to dynamically regulate the complexity and sparsity of steganographic texts in the dataset. This method successfully constructs a dataset that conforms to the real social network model to support text steganography detection research in scenarios where text is fragmented and steganographic texts are extremely sparse, which is of great significance for promoting the development of text steganography analysis technology, maintaining network security and social stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of the steps of the method for constructing a social network-oriented text steganography analysis dataset provided by an embodiment of the present invention;
[0037] Figure 2 is a schematic diagram of a twitter heterogeneous information network provided by an embodiment of the present invention;
[0038] Figure 3It is the framework diagram of the RNN generative text steganography model provided by the embodiment of the present invention;
[0039] Figure 4 It is the flowchart for constructing the social network text steganography analysis dataset provided by the embodiment of the present invention;
[0040] Figure 5 It is the structural block diagram of the device for constructing the social network-oriented text steganography analysis dataset provided by the embodiment of the present invention;
[0041] Figure 6 It is the block diagram of an electronic device provided by the embodiment of the present invention;
[0042] Figure 7 It is the schematic diagram of a computer-readable medium provided by the embodiment of the present invention. Detailed implementation manners
[0043] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0044] From a theoretical perspective, the text steganography analysis dataset constructed by the present invention can truly reflect the fragmented characteristics of texts in social networks and the sparse distribution of steganographic texts, and has rich information such as context information and entity association relationships in social networks, providing more comprehensive feature representations and model training data for steganography analysis algorithms for social networks, as well as a more accurate and reliable experimental verification platform. This will help improve the accuracy and robustness of steganography analysis algorithms and promote the continuous progress of text steganography analysis technology. From a practical application perspective, this dataset can provide strong technical support for network security supervision agencies, which will help protect users' personal privacy and information security, prevent the spread of malicious information, and improve the governance level of the cyberspace and maintain social stability. This research not only has theoretical significance but also has important practical application value.
[0045] The following will take the Twitter online social network as a research example to introduce the method for constructing the social network-oriented text steganography analysis dataset of the present invention. This method includes the following five parts: constructing a heterogeneous information network, generating a steganographic text library, local community discovery, user tweet reconstruction, and dataset construction.
[0046] Refer to Figure 1 , which shows the step flowchart of the method for constructing the social network-oriented text steganography analysis dataset provided by the embodiment of the present invention, and specifically may include the following steps:
[0047] Step 101, construct a heterogeneous information network of the social network platform;
[0048] In an alternative embodiment of the present invention, a heterogeneous information network of a social network platform is constructed, including:
[0049] For the social network platform, users are collected and expanded based on breadth-first search, distribution diversity strategy, and numerical diversity strategy to form a user network with users as nodes and specific association relationships as edges. And based on the user network, a heterogeneous information network containing various entities and association relationships is formed by collecting and expanding user-related entities and association relationships.
[0050] The process of constructing the heterogeneous information network is divided into two stages:
[0051] Stage 1: User network collection. Starting from the seed users, a preset number of their followers and followees are obtained through breadth-first search to form a basic user set. The distribution diversity strategy is used to stratify and sample numerical metadata into high, medium, and low intervals, and boolean metadata is sampled according to true and false classifications. And the numerical diversity strategy is used to preferentially extract neighbor users with the greatest difference from the current user's metadata to form a user network with users as nodes and follow relationships as edges.
[0052] Specifically, this stage mainly focuses on the construction of the user network. First, in a breadth-first search (BFS) manner, starting from the selected seed users, 1000 followers and 1000 followees are retrieved using the Twitter API as the basic user set. To ensure the diversity of users in the dataset, two diversity-aware strategies, distribution diversity and numerical diversity, are adopted to optimize the breadth-first retrieval for user expansion, aiming to more comprehensively cover different types of users to ensure that the collected users are widely representative. This stage constructs a homogeneous graph with users as nodes and follow relationships as edges, where are user nodes and are follow relationship edges. Table 1 describes the user metadata used in the diversity-aware sampling.
[0053] Table 1 User metadata used in the diversity-aware sampling
[0054]
[0055] Distribution diversity: Given user metadata, different types of users will fall into the metadata distribution in different ways. The goal of distribution diversity is to sample users at the top, middle, and bottom of the distribution. For numerical metadata, among the neighbors of the user and their metadata values, k users with the highest values, k users with the lowest values, and k users randomly selected from the remaining users are selected. For boolean type metadata, k users with the value true and k users with the value false are selected.
[0056] Value Diversity: Given a user and their metadata, the metadata values of different neighbor users are different. The neighbor users with larger differences from the current user's metadata values are preferentially sampled to ensure the diversity of the collected users. For numerical metadata, the probability that an extended neighbor user v ∈ N(u) of user u is sampled is denoted as p(v) ∝ |u num - v num |, where u num is the metadata value of user u. For boolean-type metadata, k users are selected from the opposite values.
[0057] Phase II: Heterogeneous Graph Construction. Based on user nodes, tweets, user lists, hashtag entities, and their associated relationships are extended, including interaction behaviors such as retweets, quotes, replies, and mentions. Relevant tweets are dynamically extended through topic associations, and finally, a heterogeneous information network containing four types of entities: users, tweets, lists, and hashtags, and multiple associated relationships is formed.
[0058] Specifically, in this phase, these users' tweets, associated lists, hashtags, and 12 other relationships between users and these new entities are mainly collected based on the user network constructed in Phase I. Among them, for user entities, their metadata information, tweets, lists, and following relationships are collected. For tweet entities, detailed information of each tweet is collected, including retweeted, quoted, and replied tweets, as well as mentioned users. In addition, all hashtags in list tweets are collected, and more tweets related to the topic are searched using the Twitter API. After the extension of entities and relationships in this phase, finally, a Twitter heterogeneous information network containing 4 types of entities (92,932,326 nodes) and 14 types of relationships (170,185,937 edges) is presented. As Figure 2 shown, the left side shows an HIN instance modeling the twitter social network, and the right side is the heterogeneous information network schema (HIN Scheme), which describes the relationships between nodes. The detailed information of nodes and edges is shown in Tables 2 and 3.
[0059] Table 2 Entities in the Heterogeneous Information Network
[0060]
[0061] Table 3 Relationships in the Heterogeneous Information Network
[0062]
[0063] Step 102: Use the data of the heterogeneous information network to train a generative text steganography model, combine multiple steganography algorithms and different embedding capacity parameters to generate multiple types of stego texts, and form a stego text library;
[0064] The generative text steganography method can automatically generate a natural text according to the secret information without the original carrier, has strong anti-detection ability and high embedding efficiency, and is the most widely used text steganography technology at present. In order to ensure the diversity of steganographic texts in the constructed steganographic text library, the present invention uses the steganography algorithm and capacity as parameter variables, and embeds five types of secret information bitstreams using three advanced generative text steganography algorithms to generate multiple types of steganographic texts. The present invention uses a generative text steganography model (RNN-Stego) to generate steganographic texts.
[0065] In an alternative embodiment of the present invention, the generative text steganography model is trained using the data of the heterogeneous information network, combined with multiple steganography algorithms and different embedding capacity parameters, to generate multiple types of steganographic texts, forming a steganographic text library, including:
[0066] Training the generative text steganography model using the data of the heterogeneous information network;
[0067] The trained generative text steganography model uses arithmetic coding steganography algorithm, adaptive dynamic grouping steganography algorithm and variable length coding steganography algorithm to perform steganographic embedding on the secret information, and controls the text hiding capacity by adjusting the number of bits embedded per word, generating multiple types of steganographic texts, forming a steganographic text library.
[0068] In the process of generating steganographic texts, first preprocess the Twitter texts collected in the previous section and use them as a corpus to train the generative text steganography model RNN-Stego (the model framework is as Figure 3 shown), so that the model learns the statistical distribution characteristics of the language. When generating steganographic texts, three steganography algorithms, namely arithmetic coding (AC), adaptive dynamic grouping (ADG) and variable length coding (VLC), are used to encode the secret information. Among them, VLC maps the secret information to conditional probabilities through Huffman coding, effectively reducing the distribution difference between the steganographic text and the carrier text. AC uses the inverse process of arithmetic coding for information hiding. This data compression method is suitable for encoding element sequences with known probability distributions. It first generates (uniform sampling) information, and then maps the information to (word) sequences, minimizing the statistical feature distribution difference between the steganographic text and the carrier text while achieving information hiding. ADG divides the conditional probabilities into several "buckets" as equal-sum as possible, and it can be mathematically proven to achieve the theoretical minimum difference. The text hiding capacity is adjusted by adjusting the number of bits embedded per word (Bit Per Word, BPW) of the steganographic text generation model. Finally, a steganographic text library is constructed:
[0069] D stego =Ua∈{AC,ADG,VLC} , c∈[1,5] S(a, c) (1)
[0070] Among them, S(a, c) represents generating stego-text using steganography algorithms a ∈ {AC, ADG, VLC} and steganography hiding capacity c ∈ [1, 5]. Table 4 shows the average length of stego-text generated under different steganography algorithms and embedding payloads in the stego-text library.
[0071] Table 4D stego The average length of stego-text generated under different steganography algorithms and embedding payloads
[0072]
[0073]
[0074] Step 103: By defining multi-node association paths and combining the random walk strategy guided by transition probabilities, sample user groups with associated interaction behavior characteristics from the heterogeneous information network;
[0075] In social network analysis, a meta-path is an important analysis tool that can reveal the complex association relationships between different entities. The present invention proposes a local group discovery method based on meta-path constraints to sample special users. These user groups may have similar social behavior patterns or social relationships and are important carriers for the spread of stego-text. By sampling these users, the spread of stego-text is concentrated within the scope of such special groups, simulating the phenomenon of group aggregation in social networks.
[0076] In an optional embodiment of the present invention, sampling user groups with associated interaction behavior characteristics from the heterogeneous information network by defining multi-node association paths and combining the random walk strategy guided by transition probabilities includes:
[0077] Based on multiple types of entity nodes and the association relationships of entity nodes in the heterogeneous information network, construct a meta-path that describes the indirect interaction formed by users through multi-layer interaction behaviors;
[0078] Calculate the association strength between nodes according to the meta-path, allocate transition probabilities based on the association strength, set the restart probability to constrain the walking range, randomly select unassociated seed users, collect user nodes in the path through the random walk guided by the transition probability, stop walking after reaching the preset number, and merge and deduplicate the sampling results of different seed users to obtain user groups with associated interaction behavior characteristics.
[0079] First, introduce several basic concepts:
[0080] Definition 1: Heterogeneous Information Network (HIN). A HIN is represented as a directed graph \(G=(V, E, \varphi, \psi)\), where \(V\) is the set of nodes, \(E\) is the set of edges. \(\varphi: V\rightarrow N\) is the node type mapping function, \(\psi: E\rightarrow R\) is the relationship type mapping function, and \(|N| + |R|>2\). Then each node \(v\in V\) belongs to a node type \(\varphi(v)\in N\), and each edge \(e\in E\) belongs to a relationship type \(\psi(e)\in R\).
[0081] Definition 2: Network Schema. Given a HIN \(G=(V, E, \varphi, \psi)\) with a node type mapping function \(\varphi: V\rightarrow N\) and a relationship type mapping function \(\psi: E\rightarrow R\), its network schema is denoted as \(S\) G \(=(N, R)\). The network schema emphasizes how the node types in \(N\) are associated through the relationships in \(R\). For example, Figure 2 The left side gives an HIN instance constructed from the Twitter social network, Figure 2 and the right side describes the network schema of four node types and the relationship types between them in the Twitter heterogeneous information network.
[0082] Definition 3: Meta-path. A meta-path \(P\) is a path defined on \(S\) G and is denoted as where \(L\) is the length of the meta-path \(P\), \(N\) i \(\in N(1\leq i\leq L + 1)\), \(R\) j \(\in R(1\leq j\leq L)\). For convenience, the meta-path is usually represented as a sequence of node type names, i.e., \(P=(N_1, N_2, \ldots, N\) L+1 ). If there exists a path \(p=(u_1, \ldots, u\) G ) in \(S\) L and \(p\) satisfies \(\varphi(u\) i ) = \(N\) i \((1\leq i\leq L)\), then \(p\) is called a path instance of \(P\) and is denoted as \(p\in P\). Different meta-paths imply different semantics. For example, the meta-path \(P_1=(User, Tweet, User)\) means that users like the same tweet, while the meta-path \(P_2=(User, Tweet, Tag, Tweet, User)\) means that users post / retweet / like tweets with the same topic. Among them, the path \(p=(u_1, t_5, u_2)\) is a path instance of \(P_1\).
[0083] Definition 4: P-connected and P-neighbor. If node \(u\) i can be linked to node \(u\) j through a path instance of the meta-path \(P\), then \(u\) j is called a P-connected node of \(u\) i . The P-connected nodes of \(u\) i are all P-neighbors of \(u\) i .
[0084] Based on the above definitions, first construct a meta-path to guide the random walk. Given that the covert communication between text steganographers needs to be hidden, they tend to interact indirectly. The present invention constructs the following meta-path based on the indirect interaction links formed by users posting or liking tweets with the same #hashtag:
[0085]
[0086] Where U, T, H ∈ N respectively represent the three types of nodes of "user", "tweet", and "topic" in the heterogeneous information network. post, love, hashtag ∈ R respectively represent the three types of edges of "post", "like", and "hashtag". This path captures the behavioral characteristics of users interacting through tweets and topics.
[0087] Next, define the correlation degree between nodes based on the total number of path instances between nodes:
[0088]
[0089] Where represents a path instance starting from node u i and ending at node u j . This correlation degree reflects the interaction intensity between nodes.
[0090] Based on the correlation degree between nodes, define the transition probability of node u i to u j based on the meta-path P as:
[0091]
[0092] Where N(u i ) represents the set of P-neighbors of node u i . To avoid excessive deviation from the target area, set the restart probability ɑ = 0.15, that is, there is a 15% probability of returning to the initial user node each time of the walk.
[0093] When collecting the candidate user set, in order to increase the diversity of the user group, randomly select multiple seed users and ensure that the seed users are not P-neighbors of each other. Starting from each seed user, according to the guidance of the meta-path, select the P-neighbor node with a greater transition probability to move, and record the user nodes in the walk path. When the number of recorded user nodes reaches 5000, stop the random walk. Combine the user groups sampled by random walk starting from different seeds, and remove the duplicate user nodes to ensure that the users in the final candidate user set are unique. Algorithm 1 in Table 5 describes the detailed process of local group discovery based on meta-path constraints.
[0094] Table 5 Detailed process of local group discovery based on meta-path constraints
[0095]
[0096]
[0097] Step 104: Based on the stego-text library, use a stego-text replacement strategy with three-dimensional dynamic regulation to reconstruct the social posts published by the user group with associated interaction behavior characteristics, generating social posts of a social network covert communication group.
[0098] In an alternative embodiment of the present invention, based on the stego-text library, using a stego-text replacement strategy with three-dimensional dynamic regulation to reconstruct the social posts published by the user group with associated interaction behavior characteristics, generating social posts of a social network covert communication group, includes:
[0099] Based on the stego-text library, by separately adjusting the stego-text ratio, stego-text type, and stego-text distribution, reconstruct the social posts published by the user group with associated interaction behavior characteristics, generating social posts of a social network covert communication group.
[0100] In the process of constructing the dataset, the behavior of users publishing stego-text for covert communication in the social network is simulated by replacing the tweets of some users with stego-text. In order to flexibly control the distribution and sparsity of stego-text in the dataset and ensure the authenticity and reliability of the dataset, the present invention designs a three-dimensional dynamic regulation strategy (S-RTD), by separately adjusting the three dimensions of stego-text ratio (Stego Ratio), stego-text type (Stego Type), and stego-text distribution (Stego Distribution), to simulate covert communication scenarios with different complexities and sparsities.
[0101] In an alternative embodiment of the present invention, based on the stego-text library, by separately adjusting the stego-text ratio, stego-text type, and stego-text distribution, reconstruct the social posts published by the user group with associated interaction behavior characteristics, generating social posts of a social network covert communication group, includes:
[0102] Set multiple stego-text ratios, select a combination of single or multiple steganography algorithms and single or multiple embedding capacities of stego-text types, and set the distribution pattern of stego-text in the user's social posts;
[0103] Calculate the number of social posts to be replaced according to the set stego-text ratio, generate corresponding stego-text subsets for the stego-text in the stego-text library according to the type combination, and insert the stego-text subsets into the specified positions of the user's social posts according to the distribution pattern, obtaining the social posts of a social network covert communication group.
[0104] The following is a detailed elaboration of the details of the three-dimensional dynamic regulation strategy:
[0105] Let the set of target users be U = {u1, u2, …, u 5000}, and its tweet set be The stego text library is D stego = {S a,c}, where a ∈ SA = {AC, ADG, VLC} and c ∈ SC = [1, 5]. Define the three-dimensional regulation strategy as follows:
[0106] (1) Stego Ratio (SR). When a steganographic communication user publishes stego text, they may also publish some normal text to conceal the existence of the stego text. The present invention designs different stego ratios ρ ∈ [0.1, 0.3, 0.5, 0.7, 0.9, 1.0] to replace the tweet set T u of user u. Then the number of tweets that user u needs to replace with stego text is:
[0107]
[0108] (2) Stego Type (ST). When a user publishes stego text, there are multiple batch steganography strategies. For example, for a fixed payload, a steganographic communication user may use different steganographic algorithms to concentrate the embedding of secret information in a few tweets; it is also possible to disperse the payload into multiple texts to reduce the amount of information embedded in each text, thereby reducing the possibility of being discovered. Taking the steganographic algorithm and capacity as parameters, different types of stego text subsets are formed, uniformly labeled as S T . Both the steganographic algorithm and the capacity have two regulation methods: "single" or "multiple", thus deriving the following four stego text subsets:
[0109] a) When using a single steganographic algorithm and a single capacity, the formed subset is S(a, c), where a represents a specific steganographic algorithm and c represents a specific capacity:
[0110] S T = S(a, c) ∈ {S a,c |a × c} (6)
[0111] where a × c represents all possible (a, c) sets obtained by combining one element from each of the sets SA and SC.
[0112] b) When using multiple steganographic algorithms and a single capacity, the formed subset is denoted as S(~, c), where the ~ symbol represents the diversity of the steganographic algorithm:
[0113] S T = S(~, c) = {S a,c | a ∈ SA} (7)
[0114] c) When a single steganography algorithm and multiple capacities are adopted, the formed subset is denoted as S(a, ~), where the ~ symbol represents the diversity of capacities:
[0115] S T = S(a, ~) = {S a,c | c ∈ SC} (8)
[0116] d) When multiple steganography algorithms and multiple capacities are adopted simultaneously, the formed subset is denoted as S(~, ~):
[0117] S T = S(~, ~) = {S a,c | a ∈ SA, c ∈ SC} (9)
[0118] (3) Stego Distribution (SD). When a large number of secret messages need to be urgently released, steganography users need to continuously release multiple stego texts to complete the steganography communication task. If there is enough time, the stego texts can be released in small amounts and multiple times, and the steganography users can intersperse the stego texts with normal texts and share them on social platforms. Therefore, stego texts may appear densely in a certain time period and be relatively sparse at other times. In the present invention, when replacing the user's tweets, the distribution mode of the stego text sequence is set to be continuously and batch - distributed in the front, middle, and back parts of the user's tweet sequence, or randomly and dispersedly distributed in the user's tweet sequence. The function for reconstructing the user's original tweet sequence under the regulation of the stego text distribution is as follows:
[0119]
[0120] where S T represents the stego text subset, and the Index(n u , C T ) function realizes randomly selecting C u indices from n T positions.
[0121] To better understand the S - RTD strategy, an example is given for illustration. Suppose the original tweet sequence of a user U ∈ U sample is T U = [t1, t2, …, t n . It should be emphasized that T U is an ordered sequence arranged in the order of release time, where t iDenote the i-th tweet. When the three-dimensional dynamic regulation strategy is: ρ = 50% in SR, ST adopts a single algorithm and capacity (assuming a = AC, c = 1), and m = previous in SD. At this time, the stego text subset is Subsequently, use S T1 to replace the first U tweets of T The reconstructed tweet sequence of user U
[0122] Step 105, after reconstruction, delete the original association relationship of the replaced social post in the heterogeneous information network, retain the entities and association relationships in the heterogeneous information network that have not been modified, update the heterogeneous information network structure, and form a social network text steganalysis dataset.
[0123] After using the local group discovery algorithm to collect users with indirect interactions in a small range and reconstructing the tweets of these users using the stego text replacement strategy of S-RTD three-dimensional dynamic regulation, delete the association relationship of the replaced tweets in the heterogeneous information network, while keeping other unchanged entities and relationships. Finally, construct the SN-Stego dataset that simulates the text steganalysis environment with different degrees of fragmentation and sparsity. Figure 4 Describes the detailed process of constructing the SN-Stego dataset.
[0124] Evaluate the social network text steganalysis dataset constructed in this research example:
[0125] 1. Statistical analysis
[0126] Through the above steps, a complex and diverse steganalysis environment is simulated. SN-Stego is compared with the existing three mainstream text steganalysis datasets T-Steg, TStego-THU, and Stego-Sandbox. The statistical results are shown in Table 6. It can be seen that SN-Stego has a much larger data scale, which is 100 times that of TStego-THU. It is worth noting that SN-Stego contains rich entities and association relationships. Compared with other datasets that only contain isolated text data or simple reply relationships, the heterogeneous information network structure of SN-Stego can reveal more deep-level and potential stego features and information, providing a more comprehensive and reliable research and testing platform for researchers.
[0127] Table 6 Statistical comparison of SN-Stego and three mainstream text steganalysis datasets
[0128]
[0129] 2. Experimental analysis
[0130] Through experiments, the limitations of existing text steganalysis methods when applied to real social network scenarios with fragmented texts and sparse stego texts are revealed, highlighting the urgency and necessity of developing new text steganalysis methods for social networks, and demonstrating the important value of the method for constructing a text steganalysis dataset for social networks proposed in this invention and the constructed SN-Stego dataset in supporting text steganalysis research for social networks.
[0131] (1) Baseline models and experimental settings
[0132] The following five mainstream deep learning-based text steganalysis models are used as baseline models: FCN identifies the semantic association relationships between words in a text based on a single-layer fully connected network, and uses the fact that the embedding of secret information will destroy the statistical correlation between words for steganalysis; RNN uses a bidirectional recurrent neural network (BiRNN) to extract the conditional distribution features of each word in the text; CSW refines the word correlations in the text into continuous word correlations, cross-word correlations, and cross-sentence correlations, and uses convolutional sliding windows of various sizes (CSW) to extract these correlation features for text steganalysis; ATT uses an LSTM module to obtain the sequential features of words, and then uses an attention mechanism to find the inconsistent local features in the text; GNN is the first language steganalysis method that attempts GNN and can obtain the global information between words. They all use BERT to extract text features.
[0133] Given that most experiments are carried out in steganalysis environments with a low proportion of stego texts, the F1 score can more accurately reflect the performance of the model when dealing with imbalanced datasets. Therefore, the F1 score is selected as the key metric to measure the detection performance of the model.
[0134] All experimental codes are written based on PyTorch and executed on a GeForce RTX3080 GPU with 10Gb of graphics memory. Other parameters are shown in Table 7.
[0135] Table 7 Experimental related parameter settings
[0136]
[0137] (2) Experimental results and analysis
[0138] The F1 scores of them under different Sparsity Ratios of Stegos (SRS) were tested. The experimental results are shown in Table 8. It can be seen that in the case of more stego texts (SRS is 50%), these existing methods are relatively easy to correctly detect stego texts, and some methods can obtain F1 scores above 50 points. However, this situation is quite idealized because in the real social network scenario, the proportion of stego texts is usually extremely sparse or even non-existent. As the SRS decreases, the steganalysis environment gets closer to the real scenario, and as can be seen from Table 8, the detection performance of these baseline models has also significantly decreased as expected. When the proportion of stego texts is only 10%, the detection performance of these baseline models is very poor.
[0139] Table 8 Detection F1 Results of Baseline Models in Different Stego Text Sparsity Scenarios
[0140]
[0141] In summary, the text steganalysis model developed under the existing text steganalysis dataset is difficult to apply in the real social network environment. The dataset construction method proposed in the present invention can construct a dataset that conforms to the structural characteristics, text fragmentation characteristics, and stego sparsity characteristics of the real social network, which is of great significance for improving the accuracy and robustness of steganalysis technology and expanding and deepening the research on text steganalysis for social networks.
[0142] The present invention has the following advantages:
[0143] 1. When constructing a heterogeneous information network, two diversity perceptions are used to improve the breadth-first algorithm, so that the collected source data has diversity.
[0144] 2. When generating a stego text library, taking steganography algorithms and embedding payloads as parameters, different types of stego texts are generated to simulate the diversity and fragmentation characteristics of social network texts.
[0145] 3. When sampling local groups, meta-paths are used for constraint, so that the sampled users have similar social behavior patterns or social relationships and only contain a small number of users. Using this special group as a candidate set to simulate the hidden communication group in the social network, ensuring that the spread of stego texts is concentrated within a small range of special groups, which not only conforms to the group aggregation characteristics of the social network but also meets its extremely sparse characteristics in the social network.
[0146] 4. When simulating the distribution of steganographic texts, a three-dimensional dynamic steganographic text replacement strategy is designed to achieve dynamic regulation and replacement from three dimensions: the proportion, type, and distribution of steganographic texts, so as to flexibly control the distribution and sparsity of steganographic texts in the dataset, thereby ensuring the authenticity and reliability of the dataset.
[0147] 5. A dataset that conforms to the real social network model and can support text steganography detection research in the scenario of fragmented texts and extremely sparse steganographic texts is constructed. This dataset has a significant improvement in terms of data scale, network model, and application scenario compared with the datasets commonly used in existing text steganography analysis research, and is of great significance for improving the accuracy and robustness of steganography analysis technology and expanding and deepening text steganography analysis research for social networks.
[0148] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0149] Refer to Figure 5 , which shows the structural block diagram of the text steganography analysis dataset construction device for social networks provided in the embodiments of the present invention, and specifically may include the following modules:
[0150] The heterogeneous information network construction module 501 is used to construct the heterogeneous information network of the social network platform;
[0151] The steganographic text library generation module 502 is used to train a generative text steganography model using the data of the heterogeneous information network, combine multiple steganography algorithms and different embedding capacity parameters to generate multiple types of steganographic texts, and form a steganographic text library;
[0152] The local group discovery module 503 is used to sample user groups with associated interaction behavior characteristics from the heterogeneous information network by defining multi-node association paths and combining the random walk strategy guided by the transition probability;
[0153] The social post reconstruction module 504 is used to reconstruct the social posts published by the user groups with associated interaction behavior characteristics based on the steganographic text library by using the three-dimensional dynamic regulation steganographic text replacement strategy to generate the social posts of the social network covert communication group;
[0154] The social network text steganography analysis dataset integration module 505 is used to reconstruct and then delete the original association relationships of the replaced social posts in the heterogeneous information network, retain the entities and association relationships in the heterogeneous information network that have not been modified, update the heterogeneous information network structure, and form a social network text steganography analysis dataset.
[0155] In an embodiment of the present invention, the heterogeneous information network construction module includes:
[0156] The heterogeneous information network construction sub-module is used to collect and expand users on the social network platform based on breadth-first search, distribution diversity strategy, and numerical diversity strategy, form a user network with users as nodes and specific association relationships as edges, and on the basis of the user network, collect and expand user-related entities and association relationships to form a heterogeneous information network containing multiple entities and association relationships.
[0157] In an embodiment of the present invention, the steganographic text library generation module includes:
[0158] The steganographic model training sub-module is used to train a generative text steganographic model using the data of the heterogeneous information network;
[0159] The steganographic text generation sub-module is used to use the trained generative text steganographic model to perform steganographic embedding on the secret information using arithmetic coding steganography algorithm, adaptive dynamic grouping steganography algorithm, and variable-length coding steganography algorithm, and control the text hiding capacity by adjusting the number of bits embedded in a single character, generate various types of steganographic texts, and form a steganographic text library.
[0160] In an embodiment of the present invention, the local group discovery module includes:
[0161] The meta-path construction sub-module is used to construct a meta-path that describes the indirect interaction formed by users through multi-layer interaction behaviors based on multiple types of entity nodes and the association relationships of entity nodes in the heterogeneous information network;
[0162] The local user group determination sub-module is used to calculate the association strength between nodes according to the meta-path, allocate transition probabilities based on the association strength and set the restart probability to constrain the wandering range, randomly select uncorrelated seed users, collect user nodes in the path through random wandering guided by the transition probability, stop wandering after reaching the preset number, merge the sampling results of different seed users and remove duplicates to obtain a user group with correlated interaction behavior characteristics.
[0163] In an embodiment of the present invention, the social post reconstruction module includes:
[0164] A social post reconstruction sub-module, which is used to reconstruct the social posts published by the user group with associated interaction behavior characteristics based on a steganographic text library, by respectively adjusting the proportion of steganographic text, the type of steganographic text, and the distribution of steganographic text, so as to generate social posts of a social network covert communication group.
[0165] In an embodiment of the present invention, the social post reconstruction sub-module includes:
[0166] A three-dimensional steganographic text replacement strategy selection unit, which is used to set various proportions of steganographic text, select a combination of single or multiple steganographic algorithms and steganographic text types with single or multiple embedding capacities, and set the distribution mode of steganographic text in the user's social posts;
[0167] A social post reconstruction unit, which is used to calculate the number of social posts to be replaced according to the set proportion of steganographic text, generate corresponding steganographic text subsets for the steganographic text in the steganographic text library according to the type combination, and insert the steganographic text subsets into the specified positions of the user's social posts according to the distribution mode, so as to obtain the social posts of the social network covert communication group.
[0168] In an embodiment of the present invention, the generative text steganographic model to be trained is an RNN generative text steganographic model.
[0169] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiment.
[0170] In addition, an embodiment of the present invention also provides an electronic device, as Figure 6 shown, which includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604. Among them, the processor 601, the communication interface 602, and the memory 603 complete the communication with each other through the communication bus 604.
[0171] The memory 603 is used to store a computer program.
[0172] The processor 601 is used to implement the method for constructing a text steganographic analysis data set for a social network as described in the above embodiment when executing the program stored in the memory 603.
[0173] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0174] The communication interface is used for communication between the above terminal and other devices.
[0175] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0176] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0177] As Figure 7 shown, in another embodiment provided by the present invention, there is also provided a computer-readable storage medium 701, in which instructions are stored. When it runs on a computer, it causes the computer to execute the method for constructing a text steganography analysis data set for a social network described in the above embodiment.
[0178] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the method for constructing a text steganography analysis data set for a social network described in the above embodiment.
[0179] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in the form of a computer program product in whole or in part. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)).
[0180] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0181] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.
[0182] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.
Claims
1. A method for constructing a text steganography analysis dataset for social networks, characterized in that The method includes: Constructing a heterogeneous information network of a social network platform; Using the data of the heterogeneous information network to train a generative text steganography model, combining multiple steganography algorithms and different embedding capacity parameters to generate multiple types of stego texts, and forming a stego text library; By defining multi-node association paths and combining a random walk strategy guided by transition probabilities, sampling user groups with associated interaction behavior characteristics from the heterogeneous information network; Based on the stego text library, using a three-dimensional dynamic regulation stego text replacement strategy to reconstruct the social posts published by the user groups with associated interaction behavior characteristics, and generating social posts of a social network covert communication group; After reconstruction, delete the original association relationship of the replaced social posts in the heterogeneous information network, retain the entities and association relationships in the heterogeneous information network that have not been modified, update the structure of the heterogeneous information network, and form a social network text steganography analysis data set.
2. The method according to claim 1, wherein Constructing a heterogeneous information network of a social network platform includes: For a social network platform, collecting and expanding users based on breadth-first search, distribution diversity strategy, and numerical diversity strategy to form a user network with users as nodes and specific association relationships as edges, and on the basis of the user network, collecting and expanding user-related entities and association relationships to form a heterogeneous information network containing multiple entities and association relationships.
3. The method according to claim 1, wherein Using the data of the heterogeneous information network to train a generative text steganography model, combining multiple steganography algorithms and different embedding capacity parameters to generate multiple types of stego texts, and forming a stego text library, including: Using the data of the heterogeneous information network to train a generative text steganography model; The trained generative text steganography model uses arithmetic coding steganography algorithm, adaptive dynamic grouping steganography algorithm, and variable-length coding steganography algorithm to perform stego embedding on secret information, and controls the text hiding capacity by adjusting the number of bits embedded in a single character, generating multiple types of stego texts, and forming a stego text library.
4. The method according to claim 1, characterized in that, By defining multi-node association paths and combining a random walk strategy guided by transition probabilities, sampling user groups with associated interaction behavior characteristics from the heterogeneous information network, including: Based on multiple types of entity nodes and the association relationships of entity nodes in the heterogeneous information network, constructing a meta-path that describes the indirect interaction formed by users through multi-layer interaction behaviors; Calculating the association strength between nodes according to the meta-path, assigning transition probabilities based on the association strength and setting a restart probability to constrain the walking range, randomly selecting unassociated seed users, collecting user nodes in the path through the random walk guided by the transition probability, stopping walking after reaching the preset number, and merging and de-duplicating the sampling results of different seed users to obtain user groups with associated interaction behavior characteristics.
5. The method according to claim 1, wherein Based on the stego text library, using a three-dimensional dynamic regulation stego text replacement strategy to reconstruct the social posts published by the user groups with associated interaction behavior characteristics, and generating social posts of a social network covert communication group, including: Based on the steganographic text library, by separately adjusting the proportion of steganographic text, the type of steganographic text, and the distribution of steganographic text, reconstruct the social posts published by the user group with associated interaction behavior characteristics to generate social posts of the social network covert communication group.
6. The method according to claim 5, characterized in that, Based on the steganographic text library, by separately adjusting the proportion of steganographic text, the type of steganographic text, and the distribution of steganographic text, reconstruct the social posts published by the user group with associated interaction behavior characteristics to generate social posts of the social network covert communication group, including: Set multiple proportions of steganographic text, select a combination of single or multiple steganographic algorithms and steganographic text types with single or multiple embedding capacities, and set the distribution pattern of steganographic text in user social posts; Calculate the number of social posts to be replaced according to the set proportion of steganographic text, generate corresponding subsets of steganographic text by combining steganographic text in the steganographic text library according to type, and insert the subsets of steganographic text into the specified positions of user social posts according to the distribution pattern to obtain social posts of the social network covert communication group.
7. The method according to claim 1, characterized in that The generative text steganography model to be trained is an RNN generative text steganography model.
8. A device for constructing a text steganography analysis dataset for a social network, characterized in that, The device includes: A heterogeneous information network construction module for constructing a heterogeneous information network of the social network platform; A steganographic text library generation module for training a generative text steganography model using the data of the heterogeneous information network, combining multiple steganographic algorithms and different embedding capacity parameters to generate multiple types of steganographic text and form a steganographic text library; A local group discovery module for sampling user groups with associated interaction behavior characteristics from the heterogeneous information network by defining multi-node association paths and combining a random walk strategy guided by transition probabilities; A social post sequence reconstruction module for reconstructing the social posts published by the user group with associated interaction behavior characteristics based on the steganographic text library using a three-dimensional dynamic regulation steganographic text replacement strategy to generate social posts of the social network covert communication group; A social network text steganography analysis dataset integration module for deleting the original association relationships of the replaced social posts in the heterogeneous information network after reconstruction, retaining the unchanged entities and association relationships in the heterogeneous information network, and updating the heterogeneous information network structure to form a social network text steganography analysis dataset.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store computer programs; When the processor is used to execute the program stored on the memory, it implements the method for constructing a social network-oriented text steganography analysis dataset as described in any one of claims 1-7.
10. One or more computer-readable media, on which instructions are stored, which when executed by one or more processors cause the processors to execute the method for constructing a social network-oriented text steganography analysis dataset as described in any one of claims 1-7.