Social robot identification method and system based on deep learning

Through preprocessing based on the expression library and large language model, emotion feature extraction and joint encoding of graph convolutional networks, the problem of unutilized differences in the semantics and emotional expressions of emoticons in social robot detection is solved, and a more efficient social robot recognition effect is achieved.

CN120670984AInactive Publication Date: 2025-09-19曾卡芊
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510756535.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing social bot detection methods ignore the semantic value and emotional expression differences of emoticons, resulting in the loss of key features, and fail to fully explore the distinguishing features of emotional expression patterns between bot accounts and real users.

Method used

By converting tweet emoticons into text descriptions based on an emoticon library, combining a large language model to generate tweets that incorporate emoticon semantics, extracting emotional features, and using a graph convolutional network model for joint encoding and feature weight adjustment, a global context-aware comprehensive feature vector is generated. The social relationship topology is then integrated through a relational graph convolutional network, and finally the model parameters are optimized using the cross-entropy loss function.

Benefits of technology

It effectively alleviates the problems of sparse text expression and missing emotional information, enhances the separability of emotional behavior patterns between social robots and real users, and improves the classification performance and robustness of social robot recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670984A_ABST
    Figure CN120670984A_ABST
Patent Text Reader

Abstract

The invention discloses a social robot identification method and system based on deep learning, and relates to the technical field of robots, and the method comprises the steps: carrying out the smooth processing of a tweet through a large language model, and generating a tweet fused with expression semantics in combination with a natural language processing model; capturing an emotion expression difference between a robot account and a real user; generating a comprehensive feature vector of global context sensing; outputting fusion features; the output fusion features are mapped to a high-dimensional space through a linear layer, and the detection probability of a robot account is output through an activation function; and based on adaptive moment estimation, gradient back propagation is carried out by using a cross entropy loss function, and parameters of the graph convolutional network model are optimized to obtain a final classification result. According to the method, the problems of sparse text expression and emotion information loss are effectively relieved, and the separability of the social robot and the real user in emotion behavior modes is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robotics technology, and in particular to a social robot identification method and system based on deep learning. Background Art

[0002] Research on social robot detection has undergone years of development, and traditional methods are mainly divided into the following four categories: content feature-based methods build detection models by analyzing the vocabulary distribution, metadata information (such as URL ratio) and basic statistical laws of user-generated texts; behavior feature-based methods focus on identifying anomalies from user activity timing patterns (such as posting frequency and active time periods) or interactive behaviors (such as forwarding and comment chains); deep learning-based methods use models such as convolutional neural networks (CNN) and recurrent neural networks (RNN) to automatically extract deep features of user content, significantly improving detection performance in complex scenarios; graph neural network-based methods capture potential correlation features by modeling user social relationship networks, such as the follow-up and forwarding relationships in the graph structure, using graph convolutional networks or graph attention networks.

[0003] To overcome the traditional methods' reliance on feature engineering, deep learning technology has achieved a quantum leap in detection capabilities through end-to-end learning. Deep learning-based social bot detection methods automatically extract deep features of user content and behavior through neural networks, significantly improving detection performance in complex scenarios. The core advantage of these methods lies in their ability to bypass the limitations of traditional manual feature engineering and learn discriminative patterns directly from raw data.

[0004] Existing social robot detection methods have almost two major limitations: First, the rich semantics and emotional expression value of emoticons are ignored. Existing studies often regard them as noise and directly eliminate them in the preprocessing stage, resulting in the loss of key features; second, there are significant differences in the emotional expression patterns between robot accounts and real users, including statistical deviations in emotional complexity, extreme value span, dynamic fluctuations and expression consistency. However, this discriminative feature has not yet been fully explored.

[0005] Currently, no effective solutions have been proposed for the problems in related technologies. Summary of the Invention

[0006] In response to the problems in the related technology, the present invention proposes a social robot identification method and system based on deep learning to overcome the above-mentioned technical problems existing in the existing related technology.

[0007] To this end, the specific technical solutions adopted in the present invention are as follows:

[0008] According to one aspect of the present invention, a method for identifying a social robot based on deep learning is provided, comprising:

[0009] Based on the emoji library, the emojis in the tweets are converted into corresponding text descriptions. The tweets are then fluently processed using a large language model, and combined with a natural language processing model to generate tweets that incorporate emoji semantics.

[0010] Extracting several emotional features from tweets that incorporate emoji semantics to capture the differences in emotional expression between bot accounts and real users;

[0011] A graph convolutional network model is used to jointly encode several features, and the weights of various features are adaptively adjusted through an attention gating mechanism to generate a comprehensive feature vector that is globally context-aware.

[0012] The comprehensive feature vector is input into the relationship graph convolution network, and the user social relationship topology is fused through two layers of graph convolution operations, and the fused features are output;

[0013] The output fusion features are mapped to a high-dimensional space through a linear layer, and the detection probability of the robot account is output through an activation function;

[0014] Based on adaptive moment estimation and using the cross entropy loss function for gradient back propagation, the graph convolutional network model parameters are optimized to obtain the final classification results.

[0015] Furthermore, based on the emoticon library, the emoticons in the tweets are converted to corresponding text descriptions. The tweets are then fluently processed using a large language model, and combined with the natural language processing model to generate tweets that incorporate emoticon semantics.

[0016] Defining the vocabulary and emoji structure of tweets based on the joint set form;

[0017] Use the emoji library to build an emoji-text mapping table, and detect and convert emojis in tweets;

[0018] Embed the mapped emoji text into the original tweet position to generate a new tweet containing the descriptive emoji text;

[0019] A large language model is used to smoothen the mapped new tweets. In case rare or custom emojis are not included in the emoji library, the original character encoding is retained and preset tags are added as placeholders to avoid information loss.

[0020] Furthermore, we extract several emotional features from tweets that incorporate emoji semantics to capture the differences in emotional expression between bot accounts and real users, including:

[0021] Based on the sentiment dictionary matching of tweets that integrate expression semantics, the positive sentiment value, negative sentiment value, neutral sentiment value and sentiment complexity of each tweet are calculated respectively in combination with context adjustment rules;

[0022] The emotional complexity of each tweet is calculated using the entropy formula to quantify the diversity and uncertainty characteristics of emotional expression;

[0023] Take the average of the sentiment features of all the user's tweets to obtain the user-level sentiment feature vector;

[0024] The extreme span of sentiment is defined as the absolute difference between the maximum positive sentiment and the maximum negative sentiment in the tweets of the same user, in order to capture the emotional fluctuation range of sentiment expression;

[0025] Based on the volatility of emotions, the standard deviation of the emotional complexity of users’ tweets is calculated to measure the stability of emotional expressions.

[0026] Emotional consistency is defined as the mean cosine similarity of the emotional complexity between adjacent tweets to capture the continuity of user emotional expression;

[0027] The user-level emotional feature vector, emotional extreme span, emotional volatility and emotional consistency are concatenated to obtain the final emotional feature vector.

[0028] Furthermore, the user-level emotional feature vector includes emotional positivity, emotional negativity, emotional neutrality, and emotional complexity;

[0029] Among them, the expression of the positivity of emotion is:

[0030]

[0031] The expression of negativity of emotion is:

[0032]

[0033] The expression of emotional neutrality is:

[0034]

[0035] The complexity of emotion is expressed as:

[0036]

[0037] Where i represents the i tweet; N represents the total number of tweets of the user; Expresses positivity of emotion; Indicates the negativity of emotion; Indicates the degree of emotional neutrality; To express the complexity of emotions; Indicates positive sentiment value; Indicates negative sentiment value; Indicates neutral sentiment value; Indicates emotional complexity.

[0038] Furthermore, a graph convolutional network model is used to jointly encode several features, and the weights of various features are adaptively adjusted through the attention gating mechanism to generate a comprehensive feature vector with global context awareness, including:

[0039] A natural language processing model and a multi-layer perceptron are used to encode several features and mapped to a unified dimension through a fully connected layer. A leaky corrected linear unit function is introduced to improve nonlinear expression capabilities.

[0040] The aligned features are fed into a graph convolutional network to fuse user social structure and attribute information, capture high-order correlations between multiple modalities, and generate structure-aware embeddings.

[0041] Set up a single-head attention mechanism and dynamically evaluate feature importance through query-key interaction to achieve context-aware weight adjustment;

[0042] The weighted sum of each feature is used to generate a global context-aware comprehensive feature vector.

[0043] Furthermore, the expression of the single-head attention mechanism is:

[0044]

[0045] Where, α k W represents the attention weight of the k-th feature; a represents the parameter matrix; q represents the learnable global query vector; q T represents the transpose of q; Represents the vector of the k-th feature after a specific transformation; Represents the feature vectors of different types of features after specific transformation.

[0046] Furthermore, the output fusion features are mapped to a high-dimensional space through a linear layer, and the detection probability of the robot account is output through the activation function, including:

[0047] The fused user feature vector is input into the linear layer and the initial hidden representation is generated by the leaky rectified linear unit function;

[0048] Utilize a multi-layer graph convolutional network to aggregate neighbor relationships, gradually expand the receptive field, and extract context-aware representations of user nodes;

[0049] The final node representation is input to the multi-layer perceptron and mapped to a high-dimensional space, and an activation function is used to output the detection probability of the robot account.

[0050] Furthermore, based on adaptive moment estimation and using the cross-entropy loss function for gradient backpropagation, the graph convolutional network model parameters are optimized, and the final classification results are obtained, including:

[0051] A fully connected layer is used to map the context-aware representation of the extracted user node into a two-dimensional space, and an activation function is used to calculate the category probability distribution;

[0052] The total loss function of the graph convolutional network model is used to measure the deviation between the predicted probability and the true label, and to balance the classification accuracy and the complexity of the graph convolutional network model;

[0053] Based on adaptive moment estimation, all learnable parameters are summed up and the graph convolutional network model parameters are updated to obtain the final classification result.

[0054] Furthermore, the expression of the total loss function is:

[0055]

[0056] Where L represents the total loss function of the graph convolutional network model; y i represents the true label; represents the predicted probability; Indicates summing the marked users; represents the sum of all learnable parameters in the graph convolutional network model framework; λ represents the decay coefficient; w represents the learnable parameters in the graph convolutional network model framework; θ represents the set of all learnable parameters in the graph convolutional network model framework; Y represents the set of labeled users.

[0057] According to another aspect of the present invention, a social robot recognition system based on deep learning is also provided, the system comprising:

[0058] The emoji preprocessing module converts emojis in tweets into corresponding text descriptions based on the emoji library. It also uses a large language model to streamline tweets and combines it with a natural language processing model to generate tweets that incorporate emoji semantics.

[0059] The emotional feature processing module is used to extract several emotional features from tweets that incorporate expression semantics, capturing the differences in emotional expression between robot accounts and real users;

[0060] The user node feature processing module is used to jointly encode several features using a graph convolutional network model and adaptively adjust the weights of various features through an attention gating mechanism to generate a comprehensive feature vector that is aware of the global context;

[0061] The feature fusion module is used to input the comprehensive feature vector into the relationship graph convolution network, fuse the user social relationship topology through two layers of graph convolution operations, and output the fused features;

[0062] The heterogeneous social graph modeling module is used to map the output fusion features to a high-dimensional space through a linear layer and output the detection probability of the robot account through an activation function;

[0063] The learning and optimization module is used to optimize the graph convolutional network model parameters based on adaptive moment estimation and gradient backpropagation using the cross-entropy loss function to obtain the final classification results.

[0064] The beneficial effects of the present invention are:

[0065] 1. This paper uses the Emoji library and the large language model GPT-4 to perform sentiment semantic enhancement on tweets, and combines it with RoBERTa to generate high-quality semantic embeddings, effectively alleviating the problems of sparse text expression and missing sentiment information. It further proposes a seven-dimensional sentiment feature quantification framework to characterize user expression differences from three dimensions: sentiment amplitude, volatility, and consistency, and enhance the separability of emotional behavior patterns between social robots and real users.

[0066] 2. This invention constructs a five-dimensional feature fusion mechanism, dynamically weighted and integrating heterogeneous features from multiple sources based on an attention-gated network, generating a context-aware global user representation. Furthermore, the Relationship-Aware Graph Convolutional Network (RGCN) introduces the user attention network structure, integrating graph topology information with user behavioral characteristics to achieve deep recognition and global behavioral modeling of social robots. Furthermore, the model further enhances the expressive power of features through the use of multi-layer perceptrons and can effectively disseminate information in large-scale, sparse social graphs, thereby improving classification performance and robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0068] Figure 1 is a flow chart of a social robot identification method based on deep learning according to an embodiment of the present invention;

[0069] Figure 2 This is a principle block diagram of a social robot recognition system based on deep learning according to an embodiment of the present invention.

[0070] In the picture:

[0071] 1. Emoji preprocessing module; 2. Emotional feature processing module; 3. User node feature processing module; 4. Feature fusion module; 5. Heterogeneous social graph modeling module; 6. Learning and optimization module. DETAILED DESCRIPTION

[0072] To further illustrate each embodiment, the present invention provides drawings, which are part of the disclosure of the present invention. They are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. By referring to these contents, ordinary technicians in this field should be able to understand other possible implementation methods and the advantages of the present invention.

[0073] According to an embodiment of the present invention, a social robot identification method and system based on deep learning are provided.

[0074] The present invention will now be further described with reference to the accompanying drawings and specific embodiments. Figure 1 As shown, the social robot identification method based on deep learning according to an embodiment of the present invention includes:

[0075] S1. Convert tweet emojis based on the emoji library to obtain corresponding text descriptions; use a large language model to streamline the tweets and combine them with a natural language processing model to generate tweets that incorporate emoji semantics;

[0076] S2. Extract several emotional features from tweets that incorporate expression semantics to capture the differences in emotional expression between robot accounts and real users;

[0077] S3. Utilize the graph convolutional network model to jointly encode several features and adaptively adjust the weights of various features through the attention gating mechanism to generate a comprehensive feature vector that is globally context-aware.

[0078] S4. Input the comprehensive feature vector into the relationship graph convolution network, fuse the user social relationship topology through two layers of graph convolution operations, and output the fused features;

[0079] S5. Map the output fusion features to a high-dimensional space through a linear layer, and output the detection probability of the robot account through an activation function;

[0080] S6. Based on adaptive moment estimation and using the cross-entropy loss function for gradient backpropagation, the graph convolutional network model parameters are optimized to obtain the final classification results.

[0081] Specifically, the present invention realizes end-to-end modeling through multi-stage collaborative processing. First, the tweets are preprocessed: the Emoji library (i.e., expression library) is used to convert emoticons into corresponding text descriptions, and then the GPT-4 model (i.e., large language model) is used to smooth the tweets. The RoBERTa model (i.e., natural language processing model) is combined to generate tweet embedding representations that integrate expression semantics, solving the problem that traditional word vectors are insufficient in representing the semantic information contained in emoticons, and enhancing the expression ability of emotional features. Furthermore, seven emotional features are extracted from the preprocessed tweets to capture the statistical differences in emotional expression between robot accounts and real users. Then, ESA-BotRGCN (i.e., graph convolutional network model) jointly encodes tweet features, emotional features, user description features, numerical features, and category features, and adaptively adjusts the weights of various features through the attention gating mechanism to generate a comprehensive feature vector with global context awareness. The comprehensive features are then input into the relational graph convolutional network (RGCN), and the user social relationship topology is integrated through a two-layer graph convolution operation, and Dropout is used to suppress overfitting. Finally, the above output features are mapped to a high-dimensional space through a linear layer, and then the detection probability of the robot account is output through the Softmax function (i.e., activation function). The Adam optimizer (i.e., adaptive moment estimation) is used and gradient backpropagation is performed based on the cross-entropy loss function to optimize the model parameters, thereby obtaining the final classification result.

[0082] In this optional embodiment, the emoticons in the tweet are converted based on the emoticon library to obtain corresponding text descriptions; the tweets are fluently processed using a large language model, and a tweet integrating emoticon semantics is generated in combination with a natural language processing model, including:

[0083] Defining the vocabulary and emoji structure of tweets based on the joint set form;

[0084] Use the emoji library to build an emoji-text mapping table, and detect and convert emojis in tweets;

[0085] Embed the mapped emoji text into the original tweet position to generate a new tweet containing the descriptive emoji text;

[0086] A large language model is used to smoothen the mapped new tweets. In case rare or custom emojis are not included in the emoji library, the original character encoding is retained and preset tags are added as placeholders to avoid information loss.

[0087] Specifically, through emoji-text mapping and language model optimization, the semantic and emotional value of emojis is fully preserved. First, each tweet t is defined as the union of text words and emojis:

[0088] t={w,e}。

[0089] Where w={w1,w2,...,w M} represents the text word sequence in the tweet, M is the total number of words; e={e1,e2,...,e K} represents the set of emojis in the tweet, and K is the total number of emojis. The tweets are then preprocessed, and the specific process includes:

[0090] 1) Emoji mapping: First, the tweet is detected for emojis. If no emojis are present, no processing is done. If emojis are present, the emojis are mapped to the tweet using the mapping table M. k Convert to the corresponding text description text k :

[0091] f(e k )=text k (k=1,2,...,K).

[0092] After comparing the features of the Emoji, Emojiswitch, and EmojiBase libraries, this paper selected the Emoji library to construct an emoji-to-text mapping table M. Table 1 lists 10 typical mapping examples. The Emoji library covers all Unicode emojis and provides long text descriptions in multiple languages, which more completely preserves the semantics and emotions of emojis.

[0093] Table 1. Examples of commonly used emoticons and their text mappings

[0094]

[0095] The mapped tweet t′ is expressed as:

[0096] t′=w∪{f(e1),f(e2),...,f(e K )}.

[0097] 2) Text coherence optimization: Tweets converted from emojis often suffer from poor text coherence and semantic ambiguity, such as the tweet "Let's go to the beach!" After conversion, it becomes "Let's go to the beach! :beach_with_umbrella:", which is semantically incoherent. Therefore, the present invention introduces the GPT-4 model to semantically reconstruct the text, limiting the length of the generated text to no more than 200% of the original tweet, so that the tweet can fully express its semantic and emotional information. For example, the above tweet is converted to "Let's go to the beach and enjoy the sun under anumbrella!"

[0098] 3) Abnormal emoji processing: For rare or custom emojis not included in the emoji library, this invention retains their original Unicode encoding (i.e., character encoding) and adds a special marker "[UNK]" as a placeholder to avoid information loss. Subsequently, all new tweets are fed into the model for training instead of the original tweets.

[0099] In this optional embodiment, several emotional features are extracted from tweets that incorporate expression semantics to capture the differences in emotional expression between robot accounts and real users, including:

[0100] Based on the sentiment dictionary matching of tweets that integrate expression semantics, the positive sentiment value, negative sentiment value, neutral sentiment value and sentiment complexity of each tweet are calculated respectively in combination with context adjustment rules;

[0101] The emotional complexity of each tweet is calculated using the entropy formula to quantify the diversity and uncertainty characteristics of emotional expression;

[0102] Take the average of the sentiment features of all the user's tweets to obtain the user-level sentiment feature vector;

[0103] The extreme span of sentiment is defined as the absolute difference between the maximum positive sentiment and the maximum negative sentiment in the tweets of the same user, in order to capture the emotional fluctuation range of sentiment expression;

[0104] Based on the volatility of emotions, the standard deviation of the emotional complexity of users’ tweets is calculated to measure the stability of emotional expressions.

[0105] Emotional consistency is defined as the mean cosine similarity of the emotional complexity between adjacent tweets to capture the continuity of user emotional expression;

[0106] The user-level emotional feature vector, emotional extreme span, emotional volatility and emotional consistency are concatenated to obtain the final emotional feature vector.

[0107] Specific, emotional characteristics By quantifying the emotional expression patterns of user tweets, we capture the statistical differences between bot accounts and real users in different emotional characteristics. This paper proposes seven emotional features, including basic emotional features (positive, negative, neutral sentiment values, and sentiment complexity), and three advanced features based on these basic emotional features (sentiment extreme span, volatility, and consistency), which significantly enhance the distinguishability of features. The specific definitions are as follows:

[0108] 1) Basic Sentiment and Complexity: We perform sentiment analysis on each user's tweet, extracting positive sentiment values ​​based on a predefined sentiment dictionary that contains positive, negative, and neutral sentiment intensity values ​​for each word, combined with contextual adjustment rules (such as negation, emphasis, and punctuation). Negative sentiment value Neutral sentiment value and emotional complexity

[0109]

[0110] Where i represents the i-th tweet, j represents the j-th word in the i-th tweet, M is the total number of words in the i-th tweet, and I Pos (w j ) and I Neg (w j ) represents each word w j The positive and negative emotional intensity, C(w j ) represents the context adjustment factor, and the neutral sentiment value is the remaining part not covered by positive or negative sentiment. The diversity and volatility of emotional expression can be measured by entropy:

[0111]

[0112] in, is the normalized ratio of each sentiment category. The sentiment features of all tweets of a user are averaged to obtain the user-level sentiment feature vector. The user-level sentiment feature vector includes sentiment positivity, sentiment negativity, sentiment neutrality, and sentiment complexity.

[0113] Among them, the expression of the positivity of emotion is:

[0114]

[0115] The expression of negativity of emotion is:

[0116]

[0117] The expression of emotional neutrality is:

[0118]

[0119] The complexity of emotion is expressed as:

[0120]

[0121] Where i represents the i tweet; N represents the total number of tweets of the user; Expresses positivity of emotion; Indicates the negativity of emotion; Indicates the degree of emotional neutrality; To express the complexity of emotions; Indicates positive sentiment value; Indicates negative sentiment value; Indicates neutral sentiment value; Indicates emotional complexity.

[0122] 2) Emotional extreme span: To capture the emotional fluctuation range of emotional expression, define the emotional extreme span S Span is the absolute difference between the maximum positive sentiment and the maximum negative sentiment in tweets of the same user:

[0123]

[0124] in, and are all positive, and this indicator quantifies the degree of extremeness of user emotional expression. Span Usually higher; while robot accounts are task-driven (such as concentrated propaganda or attack), with single emotional expression, S Span This feature can significantly improve the robot detection performance.

[0125] 3) Emotional volatility: Further introduce emotional volatility S Var , by calculating the standard deviation of the emotional complexity of user tweets, we can measure the stability of emotional expression:

[0126]

[0127] The emotional complexity of real users fluctuates greatly (S Var High), while robots generate content programmatically, so the emotional complexity distribution is more concentrated (S Var Low).

[0128] 4) Emotional consistency: Define emotional consistency S Consist is the mean cosine similarity of the sentiment complexity between adjacent tweets, which is used to capture the continuity of user sentiment expression:

[0129]

[0130] in, Represents the sentiment vector of the i-th tweet. Since robot accounts generate content in batches, the sentiment similarity of adjacent tweets is high (S Consist approaches 1), while real user emotional expressions are more random (S Consist lower).

[0131] 5) The final sentiment feature vector is formed by concatenating the above indicators:

[0132]

[0133] in, Indicates that the dimension is seven.

[0134] In this optional embodiment, a graph convolutional network model is used to jointly encode several features, and the weights of various features are adaptively adjusted through an attention gating mechanism to generate a global context-aware comprehensive feature vector including:

[0135] A natural language processing model and a multi-layer perceptron are used to encode several features and mapped to a unified dimension through a fully connected layer. A leaky corrected linear unit function is introduced to improve nonlinear expression capabilities.

[0136] The aligned features are fed into a graph convolutional network to fuse user social structure and attribute information, capture high-order correlations between multiple modalities, and generate structure-aware embeddings.

[0137] Set up a single-head attention mechanism and dynamically evaluate feature importance through query-key interaction to achieve context-aware weight adjustment;

[0138] The weighted sum of each feature is used to generate a global context-aware comprehensive feature vector.

[0139] Specifically, ESA-BotRGCN (i.e., graph convolutional network model) uses multimodal user information to address the challenge of bot accounts disguised as accounts, accurately identifying malicious bot accounts and preventing them from spreading false information. Specifically, ESA-BotRGCN jointly encodes five types of features, including sentiment features, user description features, tweet content features, numerical features, and categorical features. This paper will explain the encoding process of the latter four types of features:

[0140] 1) User description features: User description features are extracted from the personal profile of the account homepage. To capture its semantic information, pre-trained RoBERTa is used for encoding. First, the user description text is input into RoBERTa (i.e., a natural language processing model) to generate a word-level embedding vector:

[0141]

[0142] in, is the representation of the user description; m represents the mth word in the user description, Q is the total number of words; D s is the RoBERTa embedding dimension. Then the representation vector of the user description is derived:

[0143]

[0144] Among them, W B and b B is a learnable parameter, φ is the activation function, and D is the embedding dimension of the Twitter user. In the following content of this invention, Leaky-ReLU (i.e., leaky rectified linear unit function) is used as the activation function φ.

[0145] 2) Tweet content features: After emoticon preprocessing, the semantic and emotional information of tweets are enhanced. The processed tweets are used as input and encoded using RoBERTa in a similar way to the user description features. First, RoBERTa (i.e., natural language processing model) is used to generate an embedding vector for each tweet. Indicates that the generated embedding vector is D-dimensional, that is, the vector has D dimensions and the value of each dimension is a real number. Then take the average of the embedding vectors of all the user's tweets To obtain the representation r of the user's tweet t .

[0146] 3) Numerical features Numerical attributes refer to measurable indicators used to quantify user behaviors and attributes. As shown in Table 2, these features can provide numerical descriptions of user behavior patterns and preferences. The numerical attributes are processed by multi-layer perceptrons (MLPs) and graph neural networks (GNNs). Specifically, 10 numerical features are first obtained directly from the Twitter API, of which the first three are related to Emojis. Z-score normalization is then used to eliminate dimensional differences. Finally, a fully connected layer is used to obtain the representation of the user's numerical features.

[0147] Table 2. User's Numerical Features

[0148]

[0149]

[0150] 4) Categorical features Categorical features are qualitative indicators used to describe user attributes. As shown in Table 3, these features are usually expressed in binary form (yes / no). Similar to numerical attributes, feature engineering is avoided and instead they are encoded using multi-layer perceptrons (MLPs) and graph neural networks (GNNs). Specifically, 11 categorical features are directly obtained from the Twitter API, then converted into sparse vectors using one-hot encoding. Finally, they are concatenated and transformed through a fully connected layer and Leaky-ReLU to obtain the representation of user classification features.

[0151] Table 3. User's Categorical Features

[0152]

[0153]

[0154] In this optional embodiment, the output fusion features are mapped to a high-dimensional space through a linear layer, and the detection probability of the robot account is output through an activation function, including:

[0155] The fused user feature vector is input into the linear layer and the initial hidden representation is generated by the leaky rectified linear unit function;

[0156] Utilize a multi-layer graph convolutional network to aggregate neighbor relationships, gradually expand the receptive field, and extract context-aware representations of user nodes;

[0157] The final node representation is input to the multi-layer perceptron and mapped to a high-dimensional space, and an activation function is used to output the detection probability of the robot account.

[0158] Specifically, the model (i.e., graph convolutional network model) realizes multimodal feature fusion through feature encoding and attention mechanism. Given a user node feature set They represent sentiment features, user descriptions, tweet content, numerical and categorical features respectively. The fusion process is defined as follows:

[0159] 1) Feature alignment and nonlinear transformation: The fully connected layer projects heterogeneous features into a unified dimensional space, eliminating dimensionality differences and introducing nonlinear expression capabilities.

[0160]

[0161] in, is the learnable weight matrix; is the bias term; D = 256 is the unified embedding dimension, and Leaky-ReLU (i.e., leaky rectified linear unit function) alleviates the gradient vanishing problem.

[0162] 2) Single-head attention weight allocation: Design a single-head attention mechanism to dynamically evaluate feature importance through query-key interaction:

[0163]

[0164] Among them, α k Represents the attention weight of the k-th category feature; is the parameter matrix; is a learnable global query vector used to capture global dependencies across features; T is the transpose of q; The k-th feature undergoes a specific transformation (such as with the weight matrix W a The vector after the feature representation before the operation is used to participate in the attention weight calculation; and Similarly, the feature vectors of different types of features (here j represents different categories, j∈{senti,b,i,num,cat}) are processed and participate in the normalized calculation of attention weights; j∈{senti,b,i,num,cat} is a set representation, j is an element in the set, and each element in {senti,b,i,num,cat} represents features of different categories, such as "senti" represents emotional features; "num" represents numerical features; "cat" represents categorical features. These categories are traversed through j to calculate the attention weights of features of different categories.

[0165] 3) Context-aware feature fusion: weighted summation generates a global comprehensive feature vector Input the subsequent RGCN module for social graph modeling:

[0166]

[0167] In this way, the model (graph convolutional network model) can dynamically adjust the influence of each feature during training, thereby improving the ability to identify feature importance. This fusion method enables the model (graph convolutional network model) to more effectively integrate information from different sources and provide better performance in complex tasks.

[0168] In this optional embodiment, based on adaptive moment estimation and using the cross entropy loss function for gradient back propagation, the graph convolutional network model parameters are optimized, and the final classification results include:

[0169] A fully connected layer is used to map the context-aware representation of the extracted user node into a two-dimensional space, and an activation function is used to calculate the category probability distribution;

[0170] The total loss function of the graph convolutional network model is used to measure the deviation between the predicted probability and the true label, and to balance the classification accuracy and the complexity of the graph convolutional network model;

[0171] Based on adaptive moment estimation, all learnable parameters are summed up and the graph convolutional network model parameters are updated to obtain the final classification result.

[0172] Specifically, in order to solve the limitations of traditional graph neural networks in modeling multiple relationships, the present invention uses a relational graph convolutional network (RGCN) to achieve deep feature extraction of heterogeneous social graphs. The present invention constructs a heterogeneous information network G = (V, E, R), where the node set V represents the user account and the edge set E contains two types of relationships R = {r1, r2} = {"following", "follower"}, which respectively represent the active attention and passive attention behaviors between users. Each user node v i The eigenvector r∈V i, which is the comprehensive vector after fusion by the attention mechanism. RGCN can integrate the rich information contained in the node context by designing an efficient information transmission and aggregation mechanism. In this process, RGCN not only considers the node (i.e. user node) v i Its own eigenvector r i , it also fully learns the node's neighbor information to generate context-aware node representations. This process gradually expands the receptive field through multi-layer stacking, thereby capturing global topological patterns. Furthermore, social networks typically exhibit large-scale and sparse graph structures, with relatively few connections between nodes, which leads to certain obstacles in the process of information dissemination. RGCN's advantage in processing large-scale sparse graphs lies in: it updates the relationship weights between nodes and the embedding vectors representing the relationships through an edge weight learning mechanism based on dynamic relationship attention, effectively filling in any missing information in the graph.

[0173] Specifically, the present invention uses the graph convolution operation of RGCN as follows: first, the nodes are initialized, and the user feature vector r i The initial hidden representation is generated by linear transformation and activation function:

[0174]

[0175] Among them, the weight matrix W1 and the bias b1 are learnable parameters; φ is the Leaky-ReLU activation function (i.e., leaky rectified linear unit function). Then, in the l-th layer convolution, the present invention transforms the node (i.e., user node) v i The representation is updated to:

[0176]

[0177] Among them, N r (i) represents the node v i The set of neighbors with relation r; Θ self and Θ r are the weight matrices for self-connections and relation r, respectively. Multiple layers of RGCN are then stacked, and multi-layer perceptrons (MLPs) are inserted between layers to further transform the user representation, gradually expanding the receptive field of the nodes to capture the global topological pattern:

[0178]

[0179] Among them, the weight matrix and bias is a learnable parameter, and the final node represents h i Mapped to high-dimensional discriminant space through MLP.

[0180] Specifically, ESA-BotRGCN models social robot detection as a binary classification task, where the label y of the user accounti ∈{0,1} represents the real user (0) and the social robot (1). Here, the present invention will explain the classification mechanism and loss function design of the model (i.e., the graph convolutional network model).

[0181] 1) Classifier design: The model (i.e., graph convolutional network model) extracts context-aware representations of user nodes through the relational graph convolutional network (RGCN) This means that the node represents h i Is a 128-dimensional vector, that is, the vector has 128 dimensions, the value of each dimension is a real number and the classifier is built based on this. First, the node is represented by h i Mapped to two-dimensional space through a fully connected layer:

[0182] z i =W O ·h i +b O .

[0183] in, is the weight matrix; is the bias term; z i Is the node representation h i The output obtained by the fully connected layer is a transition value from the original node feature representation to the final classification probability output. The Softmax function (i.e. activation function) is then used to calculate the category probability distribution:

[0184]

[0185] Among them, z i0 It is z i The first element in the vector (corresponding to one dimension in two-dimensional space), It means that the base is the natural constant e, z i0 The result of the exponential operation of the exponential. In the Softmax function, through such an exponential operation, z i Each element in is converted into a numerical value for calculating the category probability, and then normalized (divided by the sum of the exponential values ​​of all categories) to obtain the final probability of each category.

[0186] 2) Loss Function: The total loss (total loss function) of the model (i.e., graph convolutional network model) consists of cross-entropy loss and weight decay term, aiming to balance classification accuracy and model complexity. The expression of the total loss function is:

[0187]

[0188] Where L represents the total loss function of the graph convolutional network model; y i represents the true label; represents the predicted probability; Indicates summing the marked users; represents the sum of all learnable parameters in the graph convolutional network model framework; λ represents the attenuation coefficient, which is used to control model parameters and suppress overfitting; w represents the learnable parameters in the graph convolutional network model framework; θ represents the set of all learnable parameters in the graph convolutional network model framework; Y is the set of labeled users. The first part is the cross entropy loss, which measures the predicted probability. True label y i The second part is the weight decay term.

[0189] According to another embodiment of the present invention, Figure 2 As shown, a social robot recognition system based on deep learning is also provided, which includes:

[0190] Emoji preprocessing module 1 is used to convert emojis in tweets based on the emoji library to obtain corresponding text descriptions; it uses a large language model to smooth the tweets and combines it with a natural language processing model to generate tweets that incorporate emoji semantics;

[0191] Emotional feature processing module 2 is used to extract several emotional features from tweets that incorporate expression semantics, capturing the differences in emotional expression between robot accounts and real users;

[0192] User node feature processing module 3 is used to jointly encode several features using a graph convolutional network model, and adaptively adjust the weights of various features through an attention gating mechanism to generate a comprehensive feature vector with global context awareness;

[0193] Feature fusion module 4 is used to input the comprehensive feature vector into the relationship graph convolution network, fuse the user social relationship topology through two layers of graph convolution operations, and output the fused features;

[0194] Heterogeneous social graph modeling module 5, used to map the output fusion features to a high-dimensional space through a linear layer, and output the detection probability of the robot account through an activation function;

[0195] The learning and optimization module 6 is used to optimize the graph convolutional network model parameters based on adaptive moment estimation and gradient backpropagation using the cross-entropy loss function to obtain the final classification results.

[0196] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A social robot identification method based on deep learning, characterized in that: include: Convert tweet emojis based on the emoji library to obtain corresponding text descriptions; Use a large language model to streamline tweets and combine it with a natural language processing model to generate tweets that incorporate emoji semantics. Extracting several emotional features from tweets that incorporate emoji semantics to capture the differences in emotional expression between bot accounts and real users; A graph convolutional network model is used to jointly encode several features, and the weights of various features are adaptively adjusted through an attention gating mechanism to generate a comprehensive feature vector that is globally context-aware. The comprehensive feature vector is input into the relationship graph convolution network, and the user social relationship topology is fused through two layers of graph convolution operations, and the fused features are output; The output fusion features are mapped to a high-dimensional space through a linear layer, and the detection probability of the robot account is output through an activation function; Based on adaptive moment estimation and using the cross entropy loss function for gradient back propagation, the graph convolutional network model parameters are optimized to obtain the final classification results.

2. A social robot identification method based on deep learning according to claim 1, characterized in that: The emoticons in the tweet are converted based on the emoticon library to obtain corresponding text descriptions; The tweets are fluently processed using a large language model and combined with a natural language processing model to generate tweets that incorporate emoji semantics. Defining the vocabulary and emoji structure of tweets based on the joint set form; Use the emoji library to build an emoji-text mapping table, and detect and convert emojis in tweets; Embed the mapped emoji text into the original tweet position to generate a new tweet containing the descriptive emoji text; A large language model is used to smoothen the mapped new tweets. In case rare or custom emojis are not included in the emoji library, the original character encoding is retained and preset tags are added as placeholders to avoid information loss.

3. The social robot identification method based on deep learning according to claim 1, characterized in that: The extraction of several emotional features from tweets that incorporate expression semantics to capture the differences in emotional expression between robot accounts and real users includes: Based on the sentiment dictionary matching of tweets that integrate expression semantics, the positive sentiment value, negative sentiment value, neutral sentiment value and sentiment complexity of each tweet are calculated respectively in combination with context adjustment rules; The emotional complexity of each tweet is calculated using the entropy formula to quantify the diversity and uncertainty characteristics of emotional expression; Take the average of the sentiment features of all the user's tweets to obtain the user-level sentiment feature vector; The extreme span of sentiment is defined as the absolute difference between the maximum positive sentiment and the maximum negative sentiment in the tweets of the same user, in order to capture the emotional fluctuation range of sentiment expression; Based on the volatility of emotions, the standard deviation of the emotional complexity of users’ tweets is calculated to measure the stability of emotional expressions. Emotional consistency is defined as the mean cosine similarity of the emotional complexity between adjacent tweets to capture the continuity of user emotional expression; The user-level emotional feature vector, emotional extreme span, emotional volatility and emotional consistency are concatenated to obtain the final emotional feature vector.

4. The social robot identification method based on deep learning according to claim 3, characterized in that: The user-level emotional feature vector includes emotional positivity, emotional negativity, emotional neutrality, and emotional complexity; The expression of the positivity of the emotion is: The expression of the negativity of the emotion is: The expression of the neutrality of the emotion is: The expression of the complexity of the emotion is: Where i represents the i tweet; N represents the total number of tweets of the user; Expresses positivity of emotion; Indicates the negativity of emotion; Indicates the degree of emotional neutrality; To express the complexity of emotions; Indicates positive sentiment value; Indicates negative sentiment value; Indicates neutral sentiment value; Indicates emotional complexity.

5. The social robot identification method based on deep learning according to claim 1, characterized in that: The graph convolutional network model is used to jointly encode several features, and the weights of various features are adaptively adjusted through the attention gating mechanism to generate a comprehensive feature vector with global context awareness, including: A natural language processing model and a multi-layer perceptron are used to encode several features and map them to a unified dimension through a fully connected layer. A leaky corrected linear unit function is introduced to improve nonlinear expression capabilities. The aligned features are fed into a graph convolutional network to fuse user social structure and attribute information, capture high-order correlations between multiple modalities, and generate structure-aware embeddings. Set up a single-head attention mechanism and dynamically evaluate feature importance through query-key interaction to achieve context-aware weight adjustment; The weighted sum of each feature is used to generate a global context-aware comprehensive feature vector.

6. The social robot identification method based on deep learning according to claim 5, characterized in that: The expression of the single-head attention mechanism is: Where, α k Represents the attention weight of the k-th category feature; W a represents the parameter matrix; q represents the learnable global query vector; q T represents the transpose of q; Represents the vector of the k-th feature after a specific transformation; Represents the feature vectors of different types of features after specific transformation.

7. The social robot identification method based on deep learning according to claim 1, characterized in that: The output fusion features are mapped to a high-dimensional space through a linear layer, and the detection probability of the robot account is output through an activation function, including: The fused user feature vector is input into the linear layer and the initial hidden representation is generated by the leaky rectified linear unit function; Utilize a multi-layer graph convolutional network to aggregate neighbor relationships, gradually expand the receptive field, and extract context-aware representations of user nodes; The final node representation is input to the multi-layer perceptron and mapped to a high-dimensional space, and an activation function is used to output the detection probability of the robot account.

8. The social robot identification method based on deep learning according to claim 1, characterized in that: The adaptive moment estimation is used and the cross entropy loss function is used for gradient back propagation to optimize the graph convolutional network model parameters. The final classification results include: A fully connected layer is used to map the context-aware representation of the extracted user node into a two-dimensional space, and an activation function is used to calculate the category probability distribution; The total loss function of the graph convolutional network model is used to measure the deviation between the predicted probability and the true label, and to balance the classification accuracy and the complexity of the graph convolutional network model; Based on adaptive moment estimation, all learnable parameters are summed up and the graph convolutional network model parameters are updated to obtain the final classification result.

9. The social robot identification method based on deep learning according to claim 8, characterized in that: The expression of the total loss function is: Where L represents the total loss function of the graph convolutional network model; y i represents the true label; represents the predicted probability; Indicates summing the marked users; represents the sum of all learnable parameters in the graph convolutional network model framework; λ represents the decay coefficient; w represents the learnable parameters in the graph convolutional network model framework; θ represents the set of all learnable parameters in the graph convolutional network model framework; Y represents the set of labeled users.

10. A social robot recognition system based on deep learning, used to implement the social robot recognition method based on deep learning according to any one of claims 1 to 9, characterized in that: The system includes: The emoji preprocessing module converts emojis in tweets into corresponding text descriptions based on the emoji library. It also uses a large language model to streamline tweets and combines it with a natural language processing model to generate tweets that incorporate emoji semantics. The emotional feature processing module is used to extract several emotional features from tweets that incorporate expression semantics, capturing the differences in emotional expression between robot accounts and real users; The user node feature processing module is used to jointly encode several features using a graph convolutional network model and adaptively adjust the weights of various features through an attention gating mechanism to generate a comprehensive feature vector that is aware of the global context; The feature fusion module is used to input the comprehensive feature vector into the relationship graph convolution network, fuse the user social relationship topology through two layers of graph convolution operations, and output the fused features; The heterogeneous social graph modeling module is used to map the output fusion features to a high-dimensional space through a linear layer and output the detection probability of the robot account through an activation function; The learning and optimization module is used to optimize the graph convolutional network model parameters based on adaptive moment estimation and gradient backpropagation using the cross-entropy loss function to obtain the final classification results.

Citation Information

Cited By

  • Social robot detection method based on BERT and bidirectional dynamic feature fusion, electronic equipment and storage medium

    CN121051661A

  • Social robot detection method based on bert and bidirectional dynamic feature fusion, electronic device and storage medium

    CN121051661B

  • Social network account classification method based on sparse generative graph model

    CN121881055A

  • Relationship enhancement graph-based social robot detection method, apparatus and device, and storage medium

    CN121935735A

  • A method, apparatus, device, and storage medium for detecting social robots based on relationship enhancement graphs.

    CN121935735B