Social robot detection method and system based on tweet filter
By adopting a detection method based on tweet filters in social networks, combining user description and attribute information, and integrating multiple information encodings, the problem of limited detection capabilities of identifying social network robots in the prior art is solved, and more efficient and accurate robot detection is achieved.
Patent Information
- Application Number
- CN202311587387.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art, when identifying robots in social networks, is limited by the simple encoding of semantic information by pre-trained language models, resulting in limited detection capabilities, and new robots can circumvent existing detection systems through design.
Using a detection method based on tweet filters, we build a Twitter user text corpus, pre-trained scoring device and tweet filter, combining user description and attribute information, and fuse multiple information encodings to obtain user representations, thereby improving the accuracy and robustness of robot detection.
It effectively improves the efficiency of semantic representation, reduces the number of tweets to the most critical content, enhances the recognition ability of various robots, can respond to new robot avoidance strategies in a timely manner, and improves the accuracy and robustness of detection.
Smart Images

Figure CN120045967A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of social robot detection, and specifically, to a method and system for detecting social robots based on a tweet filter. Background Art
[0002] With the rapid development of the Internet, online social networks have experienced remarkable booming growth. Platforms such as Twitter, Facebook, and Weibo enable individuals to easily share and spread new information and their own opinions. However, social networks are not only a gathering place for real users but also a habitat for a large number of social robots, which are automated programs that simulate normal human behavior on the network and often have other motives. Some social robots actively participate in online discussions of important events. In addition, they also spread low-credibility information. The emergence of social robots has seriously affected the order of social networks and seriously threatened the security of cyberspace. Therefore, it is of great practical significance to detect bot accounts on social networks.
[0003] Social network robots disguise their automated nature by simulating the behavior of real users. Identifying robots in social networks is crucial for maintaining the integrity of online discourse, so many research efforts have been dedicated to identifying robots active in social networks. However, due to the diversity and dynamic behavior of social robots, there are still many challenges. On the one hand, active users generate a large number of tweets. However, previous methods have mainly been limited to using pre-trained language models (PLMs) to encode semantic information, thus limiting the final detection ability of the model. On the other hand, most current methods attempt to identify robots through the attribute information of users, which means that new robots can be specifically designed to evade existing detection systems.
[0004] However, social robots often need to send tweets to achieve their purposes, such as promoting advertisements. The tweets they post are different from those of real users, which makes it possible to detect robots through tweets. At the same time, various large-scale pre-trained language models (PLMs) based on Transformer have emerged and shown strong performance in various natural language processing tasks. This is mainly because large-scale PLMs improve the semantic information encoding performance and enrich the semantic representation input. However, most of these methods do not conduct in-depth research on how to more effectively extract semantic information and simply use PLMs to encode semantic information, thus limiting the final detection ability of the model.
[0005] Therefore, it is necessary to propose a new technical solution to improve the above technical problems. Summary of the Invention
[0006] Aiming at the defects in the prior art, the purpose of the present invention is to provide a social bot detection method and system based on a tweet filter.
[0007] A social bot detection method based on a tweet filter provided by the present invention, the method includes the following steps:
[0008] Step S1: Construct a Twitter user text corpus using the user data in the dataset;
[0009] Step S2: The pre-trained scorer calculates a score for each tweet in the Twitter user text corpus;
[0010] Step S3: Use the rules applicable to tweet screening to screen the tweets through a tweet filter;
[0011] Step S4: Integrate multiple information encodings to obtain the representation of the user and detect social bots.
[0012] Preferably, the user data in step S1 is (D, T, P), where D is the user description, T is the user tweet, and P is the user attribute.
[0013] Preferably, step S1 includes the following steps:
[0014] Step S1.1: Extract the tweet, description, and label information of user U from the source dataset;
[0015] Step S1.2: Select L tweets from the M tweets of each user. If M is less than L, then all M tweets are selected. The i-th randomly selected tweet is represented by c, and then combined with the user description D and the label y to form min{L, M} data; c represents a combination of a tweet and a description; each data is composed of the combination c and the label y.
[0016] Preferably, step S2 includes the following steps:
[0017] Step S2.1: Combine the tweet and the user description, and use the PLM model BERT for feature extraction;
[0018] Step S2.2: Train the tweet scorer, and then associate the corresponding user description with all the tweets of the user to obtain the corresponding score;
[0019] Step S2.3: Sort the scores of each tweet from low to high to obtain a ranking.
[0020] Preferably, the tweet filter in step S3 includes a Twitter user text corpus, a tweet scorer, and tweet screening rules.
[0021] The present invention also provides a social robot detection system based on a tweet filter, and the system includes the following modules:
[0022] Module M1: Construct a Twitter user text corpus using the user data in the dataset;
[0023] Module M2: The pre-trained scorer calculates a score for each tweet in the Twitter user text corpus;
[0024] Module M3: Use the rules applicable to tweet screening to screen the tweets through a tweet filter;
[0025] Module M4: Integrate multiple information encodings to obtain the representation of the user and detect social robots.
[0026] Preferably, the user data in the module M1 is (D, T, P), where D is the user description, T is the user's tweet, and P is the user attribute.
[0027] Preferably, the module M1 includes the following modules:
[0028] Module M1.1: Extract the tweet, description, and label information of user U from the source dataset;
[0029] Module M1.2: Select L tweets from the M tweets of each user. If M is less than L, then all M tweets are selected. The i-th randomly selected tweet is represented by c, and then combined with the user description D and the label y to form min{L, M} pieces of data; c represents a combination of a tweet and a description; each piece of data is composed of the combination c and the label y.
[0030] Preferably, the module M2 includes the following modules:
[0031] Module M2.1: Combine the tweet and the user description, and use the PLM model BERT for feature extraction;
[0032] Module M2.2: Train the tweet scorer, and then associate the corresponding user description with all the user's tweets to obtain the corresponding score;
[0033] Module M2.3: Sort the scores of each tweet from low to high to obtain a ranking.
[0034] Preferably, the tweet filter in the module M3 includes a Twitter user text corpus, a tweet scorer, and tweet screening rules.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1. The present invention proposes an innovative tweet filter mechanism aimed at improving the efficiency of semantic representation. By applying specific rules and ranking techniques of a pre-trained scorer, a large number of tweets are successfully reduced to the most critical content, focusing on the most informative tweets.
[0037] 2. The present invention can efficiently integrate a robot detection model with various user information. The robot detection model proposed by the present invention can efficiently integrate various user information. Compared with traditional methods, this model not only relies on user attribute information but also combines tweet content and description information, making robot detection more robust and accurate. In this way, even when faced with new evasion strategies adopted by robots, timely and accurate judgments can be made. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Other features, objectives, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0039] Figure 1 It is the overall framework diagram of the social robot detection system based on the tweet filter of the present invention;
[0040] Figure 2 It is the process diagram of constructing the tweet user text corpus of the present invention;
[0041] Figure 3 It is the process schematic diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0042] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.
[0043] Example 1:
[0044] According to a social robot detection method based on a tweet filter provided by the present invention, the method includes the following steps, referring to Figure 3 :
[0045] Step S1: Construct a tweet user text corpus using the user data in the dataset; the user data is (D, T, P), where D is the user description, T is the user tweet, and P is the user attribute.
[0046] Step S1.1: Extract the tweet, description, and label information of user U from the source dataset;
[0047] Step S1.2: Select L tweets from each user's M tweets. If M is less than L, then select all M tweets. The i-th randomly selected tweet is represented by c, and then combine it with the user description D and the label y to form min{L, M} pieces of data; c represents a combination of a tweet and a description; each piece of data consists of the combination c and the label y.
[0048] Step S2: The pre-trained scorer calculates a score for each tweet in the Twitter user text corpus.
[0049] Step S2.1: Combine the tweet and the user description, and use the PLM model BERT for feature extraction.
[0050] Step S2.2: Train the tweet scorer, and then associate the corresponding user description with all the user's tweets to obtain the corresponding scores.
[0051] Step S2.3: Sort the scores of each tweet from low to high to obtain a ranking.
[0052] Step S3: Use the rules applicable to tweet screening to screen the tweets through a tweet screener; the tweet screener includes a Twitter user text corpus, a tweet scorer, and tweet screening rules.
[0053] Step S4: Integrate multiple information encodings to obtain the user's representation and detect social robots.
[0054] The present invention also provides a social robot detection system based on a tweet screener. The social robot detection system based on a tweet screener can be implemented by executing the process steps of the social robot detection method based on a tweet screener. That is, those skilled in the art can understand the social robot detection method based on a tweet screener as a preferred implementation manner of the social robot detection system based on a tweet screener.
[0055] Example 2:
[0056] The present invention also provides a social robot detection system based on a tweet screener. The system includes the following modules:
[0057] Module M1: Construct a Twitter user text corpus using the user data in the dataset; the user data is (D, T, P), where D is the user description, T is the user's tweet, and P is the user attribute.
[0058] Module M1.1: Extract the tweet, description, and label information of user U from the source dataset.
[0059] Module M1.2: Select L tweets from M tweets of each user. If M is less than L, then select all M tweets. The i-th randomly selected tweet is represented by c, and then combined with the user description D and the label y to form min{L, M} pieces of data; c represents a combination of a tweet and a description; each piece of data consists of the combination c and the label y.
[0060] Module M2: The pre-trained scorer calculates a score for each tweet in the Twitter user text corpus.
[0061] Module M2.1: Combine the tweet and the user description, and use the PLM model BERT for feature extraction.
[0062] Module M2.2: Train the tweet scorer, and then associate the corresponding user description with all the user's tweets to obtain the corresponding scores.
[0063] Module M2.3: Sort the tweets in ascending order according to the scores of each tweet to obtain a ranking.
[0064] Module M3: Use the rules applicable to tweet screening to screen the tweets through a tweet screener; the tweet screener includes the Twitter user text corpus, the tweet scorer, and the tweet screening rules.
[0065] Module M4: Integrate multiple information encodings to obtain the user's representation and detect social bots.
[0066] Example 3:
[0067] With the rapid development of the Internet, online social networks have experienced remarkable booming growth. Platforms such as Twitter, Facebook, and Weibo enable individuals to easily share and spread new information and their own opinions. However, social networks are not only a gathering place for real users but also a habitat for a large number of social bots. These automated programs simulate normal human behavior on the Internet and often have other motives. Some social bots actively participate in online discussions of important events. In addition, they also spread low-credibility information. The emergence of social bots has seriously affected the order of social networks and posed a serious threat to network space security. Therefore, detecting bot accounts on social networks has important practical significance.
[0068] Social network bots mask their automated nature by mimicking the behavior of real users. Identifying bots in social networks is crucial for maintaining the integrity of online discourse, so many research efforts have been dedicated to identifying bots active in social networks. However, due to the diversity and dynamic behavior of social bots, there are still many challenges. On the one hand, active users generate a large number of tweets. However, previous methods have mainly been limited to using pre-trained language models (PLMs) to encode semantic information, thus limiting the final detection ability of the model. On the other hand, most current methods attempt to identify bots through the attribute information of users, which means that new bots can be specifically designed to evade existing detection systems.
[0069] However, social bots often need to send tweets to achieve their purposes, such as promoting advertisements. The tweets they post are different from those of real users, which makes it possible to detect bots through tweets. At the same time, various large-scale pre-trained language models (PLMs) based on Transformer have emerged and demonstrated strong performance in various natural language processing tasks. This is mainly because large-scale PLMs improve the semantic information encoding performance and enrich the semantic representation input. However, most of these methods have not conducted in-depth research on how to more effectively extract semantic information and simply use PLMs to encode semantic information, thus limiting the final detection ability of the model.
[0070] The practical problem to be solved by the present invention is how to process richer semantic information and avoid missing bots that cannot be detected by existing systems. It becomes crucial to screen out the key parts from a large number of user tweets. This inspires us to design a tweet filter to assist in bot detection. Based on this, we propose a tweet filter. Specifically, the present invention uses a pre-trained scorer to rank the tweets of a large number of users and screens them according to specific rules. In addition, the present invention preprocesses the tweets and fuses multiple information encodings to obtain the representation of users, thereby enhancing its ability to identify various bots.
[0071] The architecture of the entire system is as Figure 1 shown. This detection system uses a pre-trained scorer to rank the tweets of a large number of users and screens them according to specific rules. In addition, this detection system preprocesses the tweets and fuses multiple information encodings to obtain the representation of users, thereby strengthening the detection ability for various bots. As Figure 1As shown, the orange background represents the descriptions of Twitter users, the yellow background represents the tweets of Twitter users, and the green background shows the attributes of Twitter users, such as the number of followers and following. In short, the bot detection task relies on the following information to identify the label y: user description D, user tweets T, and user attributes P. In the tweet filter, we first constructed a Twitter user text corpus using the user data (D, T, P) in the dataset, which will be used to train the scorer. Then, the scorer calculates a score for each tweet, and this score is used for subsequent tweet filtering. Subsequently, we introduced some rules applicable to tweet filtering. In the bot detection model, we integrated multiple information encodings to obtain the user representation, further enhancing the detection ability for various fake bots. Next, I will introduce the model in detail for each module.
[0072] Module 1: Tweet Filter. This module consists of components such as the Twitter user text corpus, tweet scorer, and tweet filtering rules.
[0073] Figure 2 Details of the steps to construct the Twitter user text corpus are described. First, we extract the tweet, description, and label information of user U from the source dataset. Then, we select L tweets from the M tweets of each user. If M is less than L, we select all M tweets. The i-th randomly selected tweet is represented by c, and then combined with the user description D and label y to form min{L, M} pieces of data. Here, c represents the combination of a tweet and a description. Therefore, each piece of data consists of the combination c and the label y. Since a single user may generate multiple data chunks in the Twitter user text corpus, we can easily create a larger corpus by including the data of all users in the source dataset.
[0074] Based on the constructed Twitter user text corpus, we will train a scorer to score and filter the user's tweets. First, we define a feature extraction method that combines the tweet and user description. We use the large-scale PLM model BERT for feature extraction. Once the tweet scorer is fully trained, we associate the corresponding user description with all the user's tweets to obtain the corresponding scores. Then, we can sort the scores of each tweet from low to high to obtain a ranking.
[0075] Considering the large number of tweets posted by users, we try to select K tweets through the ranking r. We hope to find out which parts of the top-ranked tweets can best distinguish real users from bots. In this process, we found the most effective selection rule through a large number of experiments, that is, to select K tweets with scores in the middle range.
[0076] Module 2: Robot Detection Model. After selecting the tweets, we classify them by comprehensively considering the user's attribute information and descriptions. In addition, we introduce tweet-derived features to assist in robot detection.
[0077] For the description vector, we use BERT to extract the description features. The description information is input into the BERT model for processing.
[0078] For the tweet vector, we propose two tweet encoding methods to obtain the corresponding vectors. One is to concatenate K tweets and use BERT for encoding at once; the other method is to encode K tweets separately and then take the average.
[0079] For the tweet-derived vectors, their definitions include the cleaning rate (rate_clean), the number of "@", "#", "http", and the ratios of "@", "#", "http".
[0080] For the attribute vector P, we only utilize the user's "verified" attribute to assist our detection. This "verified" attribute indicates to other users that the account has been verified by Twitter.
[0081] Now we can construct the user's representation vector. We can flexibly select the combination of the above vectors, connect them together, and form the user's representation vector r:
[0082]
[0083] In the classifier layer, we use a single-layer neural network with an activation function for prediction:
[0084]
[0085] The loss function consists of the supervised label and the regularization term:
[0086]
[0087] The present invention proposes an innovative tweet filter mechanism. Traditional methods simply use pre-trained language models to encode semantic information, which limits the ultimate detection ability of the model. However, the present invention proposes an innovative tweet filter mechanism aimed at improving the efficiency of semantic representation. By applying specific rules and ranking techniques of a pre-trained scorer, we have successfully reduced a large number of tweets to the most critical content, focusing on the most informative tweets. A robot detection model that can efficiently integrate various user information. Nowadays, most methods attempt to identify robots through user attribute information, which enables new types of robots to be specifically designed to circumvent existing detection systems. The robot detection model proposed by the present invention can efficiently integrate various user information. Compared with traditional methods, this model not only relies on user attribute information but also combines tweet content and description information, making robot detection more robust and accurate. In this way, even when faced with new evasion strategies adopted by robots, our model can make timely and accurate judgments.
[0088] Those skilled in the art can understand this embodiment as a more specific illustration of Embodiment 1 and Embodiment 2.
[0089] Those skilled in the art know that in addition to implementing the systems, various devices, modules, and units provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the systems, various devices, modules, and units provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same functions. Therefore, the systems, various devices, modules, and units provided by the present invention can be regarded as a kind of hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structures within the hardware component; the devices, modules, and units for implementing various functions can also be regarded as either software modules for implementing the method or structures within the hardware component.
[0090] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be arbitrarily combined with each other.
Claims
1. A method for detecting social bots based on a tweet filter, characterized in that, the method comprises the following steps: Step S1: Construct a Twitter user text corpus using the user data in the dataset; Step S2: A pre-trained scorer calculates a score for each tweet in the Twitter user text corpus; Step S3: Use rules applicable to tweet filtering to filter tweets through a tweet filter; Step S4: Integrate multiple information encodings to obtain a representation of the user and detect social bots.
2. The method for detecting social bots based on a tweet filter according to claim 1, characterized in that, the user data in step S1 is (D, T, P), where D is the user description, T is the user's tweet, and P is the user attribute.
3. The method for detecting social bots based on a tweet filter according to claim 1, characterized in that, step S1 comprises the following steps: Step S1.1: Extract the tweet, description, and label information of user U from the source dataset; Step S1.2: Select L tweets from the M tweets of each user. If M is less than L, then all M tweets are selected. The i-th randomly selected tweet is represented by c, and then combined with the user description D and the label y to form min{L, M} pieces of data; c represents a combination of a tweet and a description; each piece of data consists of the combination c and the label y.
4. The method for detecting social bots based on a tweet filter according to claim 1, characterized in that, step S2 comprises the following steps: Step S2.1: Combine the tweet and the user description, and use the PLM model BERT for feature extraction; Step S2.2: Train the tweet scorer, and then associate the corresponding user description with all the tweets of the user to obtain the corresponding score; Step S2.3: Sort the scores of each tweet from low to high to obtain a ranking.
5. The method for detecting social bots based on a tweet filter according to claim 1, characterized in that, the tweet filter in step S3 includes a Twitter user text corpus, a tweet scorer, and tweet filtering rules.
6. A social bot detection system based on a tweet filter, characterized in that, the system comprises the following modules: Module M1: Construct a Twitter user text corpus using the user data in the dataset; Module M2: A pre-trained scorer calculates a score for each tweet in the Twitter user text corpus; Module M3: Use rules applicable to tweet filtering to filter tweets through a tweet filter; Module M4: Integrate multiple information encodings to obtain a representation of the user and detect social bots.
7. The social bot detection system based on a tweet filter according to claim 6, characterized in that, the user data in module M1 is (D, T, P), where D is the user description, T is the user's tweet, and P is the user attribute.
8. The social bot detection system based on a tweet filter according to claim 6, characterized in that, module M1 comprises the following modules: Module M1.1: Extract the tweets, descriptions, and tag information of user U from the source dataset; Module M1.2: Select L tweets from M tweets of each user. If M is less than L, then select all M tweets. The i-th randomly selected tweet is represented by c, and then combined with the user description D and the tag y to form min{L, M} pieces of data; c represents a combination of a tweet and a description; each piece of data consists of the combination c and the tag y.
9. The social bot detection system based on a tweet filter according to claim 6, characterized in that, the module M2 includes the following modules: Module M2.1: Combine the tweets and the user description, and use the PLM model BERT for feature extraction; Module M2.2: Train the tweet scorer, and then associate the corresponding user description with all the tweets of the user to obtain the corresponding score; Module M2.3: Sort the scores of each tweet from low to high to obtain a ranking.
10. The social bot detection system based on a tweet filter according to claim 6, characterized in that, the tweet filter in the module M3 includes a Twitter user text corpus, a tweet scorer, and tweet filtering rules.
Citation Information
Cited By
Social platform robot detection method based on multiple agents
CN122114166A