A social robot detection system and method based on micro-blog platform text features
By integrating explicit, implicit, and deep textual semantic features on the Weibo platform, a social robot detection system was developed, addressing the shortcomings of existing methods in terms of detection accuracy and versatility. This system enables effective identification and efficient detection of generative content generated by large models.
Patent Information
- Application Number
- CN202311091925.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-29
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2043-08-29
AI Technical Summary
Existing social robot detection methods lack a unified fusion approach, making it difficult to effectively detect generative content from large models. Furthermore, graph models rely on network construction, resulting in poor transferability. Text detection methods are fragmented and lack in-depth mining of new features, leading to insufficient detection accuracy and versatility.
A text feature detection system based on the Weibo platform is adopted. It combines explicit text feature extraction module, implicit text feature extraction module and deep text semantic feature extraction module, and combines sentiment detection, stance detection, spam content detection and nickname detection. The XGBoost classifier is used to judge social robots. Multiple text features are integrated to improve detection accuracy and versatility.
It improves the accuracy and versatility of social robot detection, reduces the dependence on graph model data requirements, can effectively identify large model-generated text, and improves the overall efficiency and transferability of the detection algorithm.
Smart Images

Figure CN116991973B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of social networking, specifically to social bot detection technology, and more specifically, to a social bot detection system and method based on text features of the Weibo platform. Background Technology
[0002] Social bots are accounts active on social media that possess machine-controlled attributes. Their scale is expanding rapidly, and their influence is gradually increasing. Social bot groups operated by group control software have already exhibited many problems such as guiding public opinion, spreading rumors, and manipulating traffic. As a result, the issue of social bot detection is gradually gaining attention.
[0003] Social bots have evolved through three generations: from isolated accounts exhibiting obvious bot behavior, to group accounts with numerous network links, and finally to highly realistic accounts that utilize new generative synthesis techniques to create and publish highly credible content. As the influence of social bots grows, their impact on society also increases, and the potential risks are receiving increasing attention. The detection and regulation of social bots have become a major focus for scholars and industry professionals.
[0004] As robot camouflage technology continues to advance, social robot detection technology is also constantly evolving. It can be broadly categorized into human-based methods, traditional machine learning-based methods, and deep learning-based methods. Among these, the human-based method, also known as the crowdsourcing method, involves outsourcing the detection of social robots to humans. This method has become less common as the scale of social robots has increased, due to the high cost of training human personnel. Traditional machine learning methods and deep learning are involved in the following areas: First, individual social bot detection methods, evolving from support vector machines, Naive Bayes, and random forests to deep neural network methods, with continuous advancements in feature engineering and the use of new features; second, social bot group detection methods, including clustering algorithms, community detection algorithms, and seed account-based expansion methods, which are mainly based on unsupervised training and are also combined with individual detection methods for in-depth analysis of the discovered groups; third, graph network detection methods, which mainly utilize social network relationships to construct graph networks and use the graph network structure and features for computation, further developing into methods such as graph convolutional neural networks; and finally, text-based social bot detection methods, which initially extracted features by statistically analyzing word frequencies in text content, later evolving to use BERT models for representation learning to represent text semantics for recognition, and with the development of large models such as ChatGPT, methods for detecting generative text have emerged. There are many branches of text-based social robot detection methods, but these branches are relatively independent and lack a unified method to integrate them. Furthermore, no new text-based features have been added to the current detection model to improve the accuracy of social robot detection.
[0005] Most existing social robot detection methods focus on individual robot detection. As features become richer and deep models are utilized, these methods are gradually shifting towards graph depth models. Meanwhile, text detection based on natural language processing is gradually becoming the main tool for extracting text features.
[0006] However, several problems exist:
[0007] The first issue is that current text detection methods in the field of social robots rarely involve the detection of large model-generated content. A few applications in the field of social robots are still at the detection level of GPT2, which is lagging behind the rapidly developing technology.
[0008] The second is that pure natural language processing based on text information has better transferability detection, while graph models rely on network construction;
[0009] The third is that text-based social robot detection methods involve many sub-fields, such as spam detection, sentiment detection, stance detection, and even nickname detection, all of which reflect the characteristics of robots from different perspectives. Currently, there is a lack of a systematic method to summarize and integrate existing text-based detection methods and features to form a set of social robot detection methods, and there is also no solution to further improve the accuracy of detection by deeply mining the text. Summary of the Invention
[0010] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a new social robot detection system and method based on text features of the Weibo platform.
[0011] According to a first aspect of the present invention, a social bot detection system based on text features of the Weibo platform is provided. The system includes: an explicit text feature extraction module for extracting explicit text features corresponding to the account metadata text and the original comment / forward text of the Weibo platform account; an implicit text feature extraction module for extracting implicit text features corresponding to the account metadata text and the original comment / forward text of the Weibo platform account; a deep text semantic feature extraction module for performing sentiment detection, stance detection, spam content detection, nickname detection, and text generation detection on the account metadata text and the original comment / forward text of the Weibo platform account to obtain corresponding deep text semantic features; and a social bot determination module for concatenating the explicit text features, implicit text features, and deep text semantic features corresponding to the account metadata text and the original comment / forward text of the Weibo platform account to obtain a fusion feature corresponding to the account metadata text and the original comment / forward text of the Weibo platform account, and determining whether the Weibo platform account is a social bot based on the fusion feature.
[0012] Optionally, the explicit text features include the following 20 categories: maximum original post length, minimum original post length, median original post length, average original post length, maximum comment length, minimum comment length, median comment length, average comment length; frequency of @ usage in original posts, comments, and reposts; frequency of hashtag usage, and frequency of different hashtag themes after hashtag theme classification; frequency of various parts of speech usage in original posts, comments, and reposts; statistics on punctuation usage; whether the post content contains traditional Chinese characters or English content; whether the post content, nickname, or introduction involves specific entities; total word count, total word count after deduplication, total word count after deduplication; frequency of emoticons used in the post; reposts... Whether the post is empty, and the frequency of such empty posts; whether the post content begins or ends with numbers, punctuation, links, or hashtags, and the corresponding content percentage; the number of Weibo domains covered in the post content, and the distribution percentage of the top 5 most frequent domains; the number of posts with links, the percentage of posts with links, the maximum number of duplicate posts with links, and the number of links after deduplication; the frequency and distribution of IP addresses from provinces and cities; the content percentage of long posts, short posts, and very short posts; the frequency of specific words used in the post; the frequency and number of derogatory words, positive words, and degree adverbs used; the frequency and number of domain-specific terms used in comments or posts; the frequency, percentage, and proportion of posts, videos, and images published or forwarded; whether other account names are mentioned in the text.
[0013] Optionally, the implicit text features include: the similarity between account nicknames and bios based on vector calculations; the clustering degree of comments under a single original post; the degree of repetition and frequency of repetition of published content, comments, and reposts, and content homogenization measured by averaging the similarity of pairwise content vectors; the continuous similarity between two consecutive posts based on the posting order; the correlation between the post content and the central sentence vector of the topic; the number of posts with a similarity exceeding 0.9; external link content sampling, including the similarity between the external link content and the central sentence vector of the post, and the similarity with the central sentence vector of the participating topic; the calculation of account vectors by average pooling of account text vectors, and the calculation of the average similarity between the account and its nearest neighbor account vectors; the similarity between text content and known spam content and bot language; the number of clusters after the topic clustering of published content and the extraction of language represented by the central vector within each cluster, and the results of sentiment detection.
[0014] Optionally, the explicit text feature extraction module is a natural language processing model obtained by training the corresponding explicit text features of the Weibo platform account's meta-information text and the original comment / forward text as input.
[0015] Optionally, the implicit text feature extraction module uses a fastText network trained with the account metadata text corresponding to the Weibo platform account and the original comment / forward text as input, and the corresponding implicit text features.
[0016] Optionally, the deep text semantic feature extraction module includes: a sentiment detection network, used to perform sentiment detection on the account metadata text and the original comment / forward text corresponding to the Weibo platform account, configured to train a BERT-like model with the account metadata text and the original comment / forward text as input and the sentiment detection result as output; a stance detection network, used to perform stance detection on the account metadata text and the original comment / forward text corresponding to the Weibo platform account, configured to train a BERT-like model with the account metadata text and the original comment / forward text corresponding to the Weibo platform account as input and the stance detection result as output; and a spam detection network, used to detect spam content on the account metadata text and the original comment / forward text corresponding to the Weibo platform account. The system is configured to train a BERT-like model using the account metadata text corresponding to the Weibo platform account and the original comment / forward text as input, and the spam content detection result as output; a nickname detection network is used to detect nicknames in the account metadata text corresponding to the Weibo platform account and the original comment / forward text, and is configured to train a BERT-like model using the account metadata text corresponding to the Weibo platform account and the original comment / forward text as input; a text generation detection network is used to perform text generation detection in the account metadata text corresponding to the Weibo platform account and the original comment / forward text, and is configured to train a BERT-like model using the account metadata text corresponding to the Weibo platform account and the original comment / forward text as input, and the text generation detection result as output.
[0017] Optionally, the social robot judgment module uses the fusion features of the account metadata text corresponding to the Weibo platform account and the original comment repost text as input, and the classification result of whether the Weibo platform account is a social robot as output, and trains a classifier using the XGBoost method.
[0018] According to a second aspect of the present invention, a method for detecting social bots based on text features of a Weibo platform is provided. The method includes: S1, obtaining the account metadata text and the original comment / forward text corresponding to the Weibo platform to be detected; S2, using the system described in the first aspect of the present invention to detect whether the account corresponding to the Weibo platform is a social bot based on the account metadata text and the original comment / forward text corresponding to the Weibo platform to be detected.
[0019] Compared with existing technologies, the advantages of this invention are as follows: This invention summarizes the features of text-based social bot detection methods and adds some new engineering features, which can improve the accuracy of social bot detection while enhancing the versatility and transferability of the detection algorithm, and is less limited by the data requirements of methods such as graph model algorithms. This invention uses a more comprehensive set of explicit text features, performs feature processing using implicit vector representations, and integrates deep semantic methods for text from multiple perspectives, especially incorporating large-model generative text detection into social bot detection to address the increasingly sophisticated forged social bot accounts. The method proposed in this invention can also be applied to the text feature extraction process of some deep learning models. Compared with complex algorithms that utilize more graph network content such as adjacency matrices for calculation, this invention is simpler in its content and has improved overall efficiency. Attached Figure Description
[0020] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:
[0021] Figure 1 This is a social robot detection system based on text features of the Weibo platform according to an embodiment of the present invention;
[0022] Figure 2 This is a schematic diagram of the explicit text feature extraction process according to an embodiment of the present invention;
[0023] Figure 3 This is a schematic diagram of the implicit text feature extraction process according to an embodiment of the present invention;
[0024] Figure 4 This is a schematic diagram of the deep semantic feature extraction process according to an embodiment of the present invention;
[0025] Figure 5 A schematic diagram of the social robot's judgment process. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention is further described in detail below through specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0027] As described in the background section, existing social robot detection schemes are inefficient and inaccurate, and cannot effectively adapt to the upgrading of social robot camouflage techniques. During their research on social robot detection, the inventors discovered that graph network methods rely heavily on network data construction, have poor transferability compared to text detection methods, and existing research on social robot text detection is fragmented across specific domains, lacking comprehensive fusion methods for feature aggregation. Furthermore, there are few algorithms specifically designed for text detection of large-scale generative content. Moreover, there are still new features to be explored and utilized in text detection, indicating room for further improvement.
[0028] Through in-depth research on social robot text detection, the inventors discovered that the above problems can be solved in the following ways: First, existing explicit features related to text content are summarized, and some new effective features are proposed to form a social robot text detection method that integrates various text features, improving the algorithm's transferability and versatility while ensuring accuracy. Second, implicit vectors are processed to calculate more interpretable features that also cover implicit feature content. Third, deep semantic features are further extracted through spam detection, sentiment detection, stance detection, nickname detection, and large-model generative text detection, especially large-model generative text detection. Finally, the first three are integrated to make the final social robot judgment, achieving the integration and optimization of social robot text detection methods.
[0029] In summary, this invention addresses the shortcomings of existing social robot detection algorithms, which are heavily reliant on data and whose text detection methods, while possessing good transferability, are fragmented and lack sufficient support for the latest large-model generative text detection. Furthermore, there is still room for optimization in text feature construction. This invention proposes a social robot detection scheme based on Weibo platform text features. This scheme integrates existing text features and proposes new ones, optimizing previous algorithms in explicit feature extraction, implicit feature processing, and deep semantic mining. In particular, it integrates and extends spam detection, sentiment detection, stance detection, nickname detection, and large-model generative text detection into the social robot text detection model, ultimately using the XGBoost method to determine the social robot.
[0030] Specifically, the present invention provides a social robot text monitoring solution that integrates 20 categories of explicit features, 10 categories of implicit processing features, and 5 categories of deep text semantic features. In particular, it proposes new and effective engineering features in terms of feature construction, including whether the content of the post starts or ends with a number / punctuation / link / hashtag and the proportion of similar content. It also calculates the continuous similarity between two consecutive posts based on the posting order, the similarity between external link content and the post center vector, the similarity between external link content and the central sentence vector of the topic, and the similarity between the text content and known spam content and robot language, thereby improving the effectiveness of detection.
[0031] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0032] According to one embodiment of the present invention, such as Figure 1 The present invention provides a social bot detection system capable of multi-directional fusion and in-depth mining of usable features in text. The system includes: an explicit text feature extraction module for extracting explicit text features from the account metadata text corresponding to a Weibo platform account and the original comment / forward text; an implicit text feature extraction module for extracting implicit text features from the account metadata text corresponding to a Weibo platform account and the original comment / forward text; a deep text semantic feature extraction module for performing sentiment detection, stance detection, spam content detection, nickname detection, and text generation detection on the account metadata text corresponding to the Weibo platform account and the original comment / forward text to obtain corresponding deep text semantic features; and a social bot determination module for concatenating the explicit text features, implicit text features, and deep text semantic features of the account metadata text corresponding to the Weibo platform account and the original comment / forward text to obtain a fusion feature, and determining whether the Weibo platform account is a social bot based on the fusion feature. The functions of each module are described in detail below.
[0033] I. Explicit Text Feature Extraction Module
[0034] The main function of this module is to extract statistical features from the text content of Weibo platform accounts, that is, to perform preliminary feature engineering to obtain explicit text features. Feature extraction is a common, or rather, a general, method in algorithm engineering. The difference between this invention and existing technologies in feature extraction lies mainly in the content of the constructed features and the thresholds for specific applications. As mentioned in the background section, existing technologies suffer from a loose technical structure. This invention summarizes the main aspects of text detection, integrating features into a single model rather than focusing on a single aspect. Furthermore, addressing the issue of insufficient feature construction in existing technologies, this invention proposes some new features, adding text features that are clearly useful for identifying bots. These features have proven effective in actual detection processes.
[0035] According to an embodiment of the present invention, the explicit text features constructed in the present invention include the following twenty categories: 1) Maximum original blog post length, minimum original blog post length, median original blog post length, average original blog post length, maximum comment length, minimum comment length, median comment length, and average comment length; 2) Frequency and recurrence of the use of @ in original blog posts, comments, and reposts; 3) Frequency and recurrence of hashtag usage, and frequency and recurrence of different hashtag themes after hashtag theme classification; 4) Frequency and recurrence of various parts of speech in original blog posts, comments, and reposts, such as the usage of verbs, nouns, and adjectives; 5) Statistics on the use of punctuation marks, such as the use of parentheses and whether it ends with a period; 6) Whether the post content contains traditional Chinese characters or English content; 7) Whether the post content, nickname, and introduction involve entities such as celebrities, companies, and brands; 8) Total number of characters in the post, total number of characters after deduplication, total number of words, and total number of words after deduplication; 9) Frequency and recurrence of the use of emoticons in the post. 10) Whether the forwarded text is empty, and the frequency of empty text; 11) Whether the content of the post starts or ends with numbers / punctuation / links / hashtags, and the proportion of similar content; 12) The number of Weibo fields involved in the content of the post, and the distribution of the top 5 fields in terms of frequency; 13) The number of posts with links, the proportion of posts with links, the maximum number of repeated posts with links, and the number of links after deduplication of posts with links; 14) The frequency and distribution of IP addresses from provinces and cities; 15) The proportion of long posts (more than 140 characters), the proportion of short posts, and the proportion of very short posts (less than 10 characters); 16) The frequency of use of specific words in the post, such as the use of some special abbreviations; 17) The frequency and number of use of derogatory words, commendatory words, and degree adverbs; 18) The frequency and number of use of professional terms in the comments or posts; 19) The frequency, frequency, and proportion of posting or forwarding text, video, and image content; 20) Whether other account names are mentioned in the text.
[0036] In summary, the specific explicit feature extraction process is as follows: Figure 2 As shown, the process first obtains the account metadata text and the original comment / forward text, then performs explicit text feature engineering extraction, and finally obtains twenty categories of explicit text content features, namely statistical features.
[0037] According to one embodiment of the present invention, the explicit feature extraction module is generally obtained by pre-training a natural language processing model. Since model training is generally a technique known to those skilled in the art, the training process of the model is not described in detail in the embodiments of the present invention.
[0038] II. Implicit Text Feature Extraction Module
[0039] The main function of this module is to transform text vectors to extract text features, i.e., implicit text content features, from the vector strength. In the embodiments of this invention, it is mainly implemented based on the FastText network. FastText is a pre-trained shallow network for word vector calculation and text classification, which has better training efficiency than deep networks while maintaining comparable accuracy.
[0040] According to one embodiment of the present invention, the implicit text content features constructed in the present invention include the following ten categories: 1) similarity between account nicknames and profiles calculated based on vectors; 2) comment clustering degree under a single original post; 3) content repetition degree and repetition frequency of published content, comments, and reposts, and the average similarity of content vectors is used to measure content homogenization; 4) continuous similarity between two consecutive posts based on the posting order; 5) correlation between the post content under a topic and the central sentence vector of the topic; 6) number of posts with a similarity exceeding 0.9; 7) sampling of external link content, the similarity between the external link content and the central vector of the post, and the similarity between the external link content and the central sentence vector of the participating topic; 8) mean-pooling of account text vectors to calculate account vectors, and the average similarity between the account and its nearest neighbor account vectors; 9) similarity between text content and known spam content and robot rhetoric, and a knowledge base of robot-filled rhetoric extracted from foreign literary classics such as "Stray Birds" and "Gitanjali" was discovered through rhetoric extraction for detection. 10) Extract the number of clusters and the discourse represented by the in-cluster center vectors after the content topics are clustered, and further sentiment detection is to be carried out.
[0041] In summary, the specific implicit feature extraction process is as follows: Figure 3 As shown, similar to the explicit feature extraction process, the process first obtains the account metadata text and the original comment forwarding text, then generates implicit feature vectors for the text, processes the implicit vectors to extract features, and finally obtains ten categories of implicit text content features.
[0042] III. Deep Text Semantic Feature Extraction Module
[0043] This module primarily performs sentiment detection, stance detection, spam detection, nickname detection, and text generation detection on the text content of Weibo platform accounts to form deep text semantic features. Existing technologies only perform single-function detection for deep semantic features, such as stance detection, without first determining whether the stance is fabricated by a bot, or performing other detections. In contrast, the solution of this invention performs multi-faceted detection, generating more comprehensive deep semantic features and improving detection accuracy.
[0044] According to one embodiment of the present invention, sentiment detection, stance detection, spam content detection, nickname detection, and text generation detection are all performed using BERT-like models after corresponding training. Sentiment detection refers to detecting the sentiment of text content based on part-of-speech analysis such as negation words and the use of degree adverbs, combined with the BERT model. The output value is between -1 and 1, and from this, sentiment analysis features are derived: whether the sentiment of the account text changes over time or is only unidirectionally output; whether the sentiment of the account text continuously reverses; the average positive / negative sentiment of the account text on a single topic; the average positive / negative sentiment of all topics of the account; the degree of deviation from the average sentiment of other users under the topic; and the degree of deviation from the average sentiment of the nearest neighbor accounts. Stance detection refers to performing stance detection based on the pre-trained BERT-SPC model, combined with the specific Weibo topics involved, to perform three-class classification detection, dividing stances into support, opposition, and neutrality. Similar to sentiment detection, it can be extended to stance features such as all topics of the user and the user's stances with neighbors. Spam content detection refers to judging spam content based on factors including, but not limited to, text length, spelling errors, content information content calculation, and the detection of reasonable sentence dependency relationships in the posted content using a sentence dependency tree. Nickname detection uses rules such as nickname length, whether the nickname contains the word "bot," and whether the nickname contains more than five uncommon characters. Text generation detection mainly uses a classifier trained on a BERT model. Through a black-box approach, it obtains real human-edited data from Chinese Wikipedia as positive samples and obtains negative samples from a large model API with input prompts. A general classifier is trained without knowing the structural parameters of the large model. The trained model is then used to detect text content on Weibo accounts to determine whether it is generated by the large model. After these five detections, multi-dimensional deep semantic features are generated based on the detection results.
[0045] In summary, the specific deep semantic feature extraction process is as follows: Figure 4 As shown, the method first obtains the account metadata text and the original comment forwarding text, then performs sentiment detection, stance detection, spam content detection, nickname detection, and text generation detection respectively, and finally concatenates the five detection results to generate multi-dimensional deep semantic features.
[0046] IV. Social Robot Detection Module
[0047] This module's function is to concatenate the extraction results from the previous three modules and then determine whether an account is a social bot. According to one embodiment of the present invention, the invention combines the XGBoost algorithm to form a text-based social bot classifier and determines whether an account is a social bot based on the concatenation result.
[0048] It should be noted that XGBoost is a tool based on the Gradient Boosting Tree (GBDT) algorithm. Essentially, it uses ensemble learning principles to fit residuals and classify samples through multiple decision trees. In other words, XGBoost itself is a tool for training a model, requiring robot and non-robot samples, along with their features, to train and generate the model. Human-edited data from the Chinese Wikipedia is used as positive samples, and negative samples are obtained from large model APIs with input prompts. In this invention, the input to the XGBoost tool is the concatenated features. The decision tree formed after XGBoost training is the classification model, which determines the probability of a sample being a robot based on different feature values, identifying samples with higher probabilities as robots. Since this training process is known to those skilled in the art, it will not be described in detail in this embodiment.
[0049] In summary, the specific process for determining a social robot is as follows: Figure 5 As shown, firstly, all text-related features involved in the first three steps are concatenated to organically combine the twenty explicit text features, ten implicit text features, and five deep semantic detection features from the previous modules. Then, based on the constructed classification model, social robot text detection and judgment are performed to determine whether a Weibo account is a social robot.
[0050] As can be seen from the preceding embodiments, the present invention integrates 20 categories of explicit features, 10 categories of implicit processing features, and 5 categories of deep text semantic features. These features have been screened through engineering practice to form a multi-angle, multi-granularity social robot text detection solution, namely a text feature detection classifier based on ensemble learning. This classifier has good versatility, high accuracy, and strong interpretability, and can detect various problems in the deep semantics of social robot text, thus addressing the challenges of social robot text detection brought about by generative large models. Furthermore, this invention proposes new and effective engineering features for feature construction, including whether the content of a post begins or ends with a number / punctuation / link / hashtag and the proportion of similar content; statistical analysis of the continuous similarity between two consecutive posts based on the posting order; similarity between external link content and the post's central vector; similarity between external link content and the central sentence vector of the participating topic; and the similarity between the text content and known spam content and bot dialogue (a knowledge base of bot dialogue extracted from foreign literary classics such as *Stray Birds* and *Gitanjali* was discovered through dialogue extraction and used for detection). These new features are effective features discovered in practical engineering. Adding these features has a good effect on detecting specific types of social bots, and can identify some social bots that publish content copied from certain content pools, effectively supplementing existing algorithms. In addition, in terms of deep semantics, this invention applies a black-box model generative text detection sample construction method to the field of social bot text detection. Specifically, it obtains real human-edited data from Chinese Wikipedia as positive samples and obtains negative samples from large model APIs for prompt input to train an effective classifier. Furthermore, this invention proposes effective nickname detection rules, such as whether the nickname contains the word "bot" or whether the nickname contains more than 5 uncommon characters. This enables the solution of this invention to cope with the ever-evolving large-scale model-generated text, and can distinguish between real people and large models from the perspectives of text style and common usage. At the same time, it also improves the detection effect for extremely short nickname texts.
[0051] Compared to existing technologies, this invention summarizes the features of text-based social bot detection methods and adds some new engineering features. This improves the accuracy of social bot detection while enhancing the versatility and transferability of the detection algorithm, and is less limited by the data requirements of methods such as graph model algorithms. This invention uses comprehensive explicit text features, performs implicit vector representation feature processing, and integrates deep semantic methods from multiple perspectives, particularly incorporating large-model generative text detection into social bot detection to address increasingly sophisticated fake social bot accounts. The method proposed in this invention can also be applied to text feature extraction in some deep learning models. Compared to complex algorithms that utilize more graph network content such as adjacency matrices for computation, this invention is simpler in its approach and offers improved overall efficiency.
[0052] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.
[0053] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.
[0054] Computer-readable storage media can be tangible devices that hold and store instructions for use by an instruction execution device. Computer-readable storage media can be, for example, including but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.
[0055] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A social robot detection system based on text features of the Weibo platform, characterized in that, The system includes: The explicit text feature extraction module uses a natural language processing model to extract statistical features from the account metadata text and the original comment / forward text corresponding to the Weibo platform account to obtain the corresponding explicit text features. The implicit text feature extraction module uses the fastText network to transform the account metadata text corresponding to the Weibo platform account and the original comment and forward text into text vectors to extract the text features of the vector strength and obtain the corresponding implicit text features. The deep text semantic feature extraction module uses the BERT model to perform sentiment detection, stance detection, spam content detection, nickname detection, and text generation detection on the account metadata text corresponding to the Weibo platform account and the original comment and forwarded text, respectively. The module then concatenates the five detection results to obtain the corresponding deep text semantic features. The social bot determination module is used to concatenate the account metadata text corresponding to the Weibo platform account with the explicit text features, implicit text features, and deep text semantic features corresponding to the original comment and forwarded text to obtain the fusion feature of the account metadata text corresponding to the Weibo platform account and the original comment and forwarded text, and determine whether the Weibo platform account is a social bot based on the fusion feature.
2. The system according to claim 1, characterized in that, The explicit text features include the following 20 categories: Maximum original blog post length, minimum original blog post length, median original blog post length, average original blog post length, maximum comment length, minimum comment length, median comment length, average comment length; Frequency and range of use of @ in original blog posts, comments, and reposts; Hashtag usage frequency and frequency, and the frequency and frequency of different themes after hashtag topic categorization; Frequency of use of various parts of speech in original blog posts, comments, and reposts; Statistics on punctuation usage; Does the post contain traditional Chinese characters or English content? Does the content of the post, nickname, or bio involve any specific entity? Total word count, total word count after deduplication, total word count, total word count after deduplication; The frequency and frequency of using emojis in posts; Whether the forwarded text is empty, and the frequency and duration of such empty text; Whether the content of the post begins or ends with numbers, punctuation, links, or hashtags, and the corresponding content percentage; The number of Weibo topics covered in the posts, and the distribution percentage of the top 5 most frequent topics; Number of posts with links, percentage of posts with links, maximum number of duplicate posts with links, and number of links in posts with links after deduplication. IP addresses originating from provinces and cities, and their distribution frequency and frequency. The content percentage of long articles, the content percentage of short articles, and the content percentage of very short articles; Frequency of use of specific words in the article; Frequency and number of uses of derogatory words, commendatory words, and adverbs of degree; The frequency and range of use of domain-specific terminology in comments or posts; The frequency, range, and percentage of text, video, and image content published or forwarded; Whether to mention other account names in the text.
3. The system according to claim 1, characterized in that, The implicit text features include: Similarity between account nicknames and profiles based on vector calculations; Comment aggregation rate under a single original post; The degree and frequency of repetition in published content, comments, and reposts are measured by averaging the similarity scores of each pair of content vectors to determine content homogenization. The consecutive similarity of content between two consecutive posts is calculated based on the order of posting. The degree of correlation between the content of posts under a topic and the central sentence vector of the topic; The number of posts with a similarity score exceeding 0.9; External link content sampling includes the similarity between the external link content and the central vector of the post, as well as the similarity between the external link content and the central sentence vector of the topic. The account text vector is averaged to calculate the account vector, and the average similarity between the account and its nearest neighbor account vector is calculated. The degree of similarity between the text content and known spam content and robot language; The results of content topic clustering, cluster number, and discourse extraction represented by the in-cluster center vectors, followed by sentiment detection.
4. The system according to claim 1, characterized in that, The explicit text feature extraction module is a natural language processing model that takes the account metadata text corresponding to the Weibo platform account and the original comment repost text as input and the corresponding explicit text features as training.
5. The system according to claim 1, characterized in that, The implicit text feature extraction module uses the account metadata text corresponding to the Weibo platform account and the original comment / forward text as input, and the corresponding implicit text features are trained to obtain a fastText network.
6. The system according to claim 1, characterized in that, The deep text semantic feature extraction module includes: A sentiment detection network is used to perform sentiment detection on the account metadata text and the original comment / forward text corresponding to the Weibo platform account. It is configured to train a BERT-like model with the account metadata text and the original comment / forward text corresponding to the Weibo platform account as input and the sentiment detection result as output. A stance detection network is used to detect the stance of the account metadata text and the original comment repost text corresponding to the Weibo platform account. It is configured to train a BERT-like model with the account metadata text and the original comment repost text corresponding to the Weibo platform account as input and the stance detection result as output. The spam content detection network is used to detect spam content in the account metadata text and original comment / forward text corresponding to the Weibo platform account. It is configured to train a BERT-like model with the account metadata text and original comment / forward text corresponding to the Weibo platform account as input and the spam content detection result as output. The nickname detection network is used to detect nicknames in the account metadata text and original comment / forward text corresponding to the Weibo platform account. It is configured to train a BERT-like model with the account metadata text and original comment / forward text corresponding to the Weibo platform account as input and the nickname detection result as output. The text generation and detection network is used to perform text generation and detection on the account metadata text and the original comment and forwarded text corresponding to the Weibo platform account. It is configured to train a BERT-like model with the account metadata text and the original comment and forwarded text corresponding to the Weibo platform account as input and the text generation and detection result as output.
7. The system according to claim 1, characterized in that, The social robot judgment module takes the fusion features of the account metadata text corresponding to the Weibo platform account and the original comment repost text as input, and the classification result of whether the Weibo platform account is a social robot as output, and obtains a classifier trained using the XGBoosst method.
8. A method for detecting social bots based on text features from the Weibo platform, characterized in that, The method includes: S1. Obtain the account metadata text and original comment / forward text corresponding to the Weibo platform to be detected; S2. Using the system described in any one of claims 1-7, detect whether the account corresponding to the Weibo platform is a social bot based on the account metadata text and the original comment / forward text corresponding to the Weibo platform to be detected.
9. A computer-readable storage medium, characterized in that, It stores a computer program that can be executed by a processor to implement the steps of the method of claim 8.
10. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to perform the steps of the method as described in claim 8.
Citation Information
Patent Citations
Method for identifying robot user on micro-blog platform
CN102571485A
Social robot detection method and system, storage medium and electronic equipment
CN112487176A