Social media robot detection method, system and device and storage medium
By constructing a multi-source teacher model and combining it with the positive correlation contrastive learning method, the social media robot detection model is distilled and trained, which solves the problems of low robot detection accuracy and poor adaptability in existing technologies and achieves efficient and robust robot detection effects.
Patent Information
- Application Number
- CN202510999578.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-09-12
AI Technical Summary
Existing social media bot detection methods are unable to effectively deal with the disguise and diversity of bot accounts, resulting in low accuracy, high false positive rate and poor adaptability to new disguised behaviors.
A joint optimization strategy of multi-source supervision and contrastive alignment is adopted. By constructing a text teacher model and a graph structure teacher model, combined with a positive correlation contrastive learning mechanism, the student model is distilled and trained to improve its detection performance and robustness.
The detection performance, robustness and adaptability of social media robot detection to multimodal loss situations have been significantly improved, the model deployment cost has been reduced, and the adaptability to new disguised behaviors has been improved.
Smart Images

Figure CN120633715A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology and relates to a social media robot detection method and system, in particular to a social media robot detection method, system, device and storage medium based on multi-teacher distillation and positive correlation contrast learning. Background Art
[0002] With the rapid development of social media, a large number of bot accounts have emerged, frequently engaging in activities such as public opinion manipulation, false propaganda, and inflated traffic, posing a serious threat to platform ecosystems and information security. These bots mimic human behavior through automated scripts and are highly camouflaged and diverse. Traditional detection methods that rely on manual rules or single modal features struggle to cope, suffering from low accuracy, high false positive rates, and limited adaptability to new types of camouflaged behaviors.
[0003] Currently, mainstream social media bot detection methods primarily focus on multimodal modeling and deep learning. Some methods utilize features such as the language style, sentiment, and posting frequency of text content to build models. These methods are suitable for detecting content anomalies, but struggle to capture interactions between users. Another area of research employs social graph modeling, building a graph structure based on user attention, interaction, and forwarding relationships, and introducing graph neural networks to classify nodes. This approach can effectively identify structural anomalies or group-based bot accounts, but relies on complete graph data and is limited by data loss or privacy protection issues. To improve overall performance, researchers have proposed multimodal fusion strategies that integrate multi-dimensional information such as text, behavior, and social graphs, leveraging attention mechanisms or multi-branch networks to achieve joint modeling of multi-source features. While this improves performance, the training and deployment costs are high, and practical applications remain challenging.
[0004] To reduce reliance on complex modalities and improve model deployment efficiency, some recent research has introduced knowledge distillation mechanisms. By constructing a teacher-student architecture, these methods transfer knowledge from a high-performance teacher model to a lightweight student model, enabling effective detection under weak or no-graph conditions. However, most of these methods still perform distillation based on a single teacher model and lack the ability to collaboratively model multimodal teacher information. Furthermore, the distillation process typically only aligns the teacher's output, neglecting to model the consistency of the teacher and student representation space structures. This results in limited student models' ability to integrate multi-source knowledge, and their generalization capabilities still need to be improved.
[0005] Based on the above-mentioned defects of the existing technology, there is an urgent need to study a new social media robot detection method and system. Summary of the Invention
[0006] In order to overcome the shortcomings of the existing technology, the present invention proposes a social media robot detection method, system, device and storage medium, which significantly improves the detection performance, robustness and adaptability of the student model in social media robot detection tasks through a joint optimization strategy of multi-source supervision and comparative alignment.
[0007] In order to achieve the above object, the present invention provides the following technical solutions: A social media robot detection method, characterized by comprising the following steps: 1) Collecting multi-dimensional basic attribute information of users on social media platforms and converting the multi-dimensional basic attribute information into pseudo-text with semantic expression capabilities based on rule templates; 2) Build a text teacher model and a graph structure teacher model, and train them based on open source datasets; 3) constructing a student model and performing distillation training on the student model based on the text teacher model and the graph structure teacher model; 4) Using the outputs of the text teacher model and the graph structure teacher model as positive samples and the output of the student model as anchor samples, the student model is aligned and optimized using a positive correlation contrast learning strategy; 5) Use the aligned and optimized student model to detect social media robots.
[0008] Preferably, in step 1), the multi-dimensional basic attribute information of social media platform users collected includes registration time, number of fans, number of followings, content publishing frequency, account authentication status, device type used, and IP active area, and before being converted into pseudo-text with semantic expression capabilities, it is first standardized and cleaned, including removing missing or invalid records, unifying the time format, and normalizing continuous values.
[0009] Preferably, in step 2), the text teacher model adopts a pre-trained language model, whose input is user tweets and personal introduction text, and whose output includes category probability distribution and intermediate layer representation.
[0010] Preferably, in step 2), the graph structure teacher model adopts a heterogeneous graph neural network, whose input is the user's social relationship graph, and the output includes category probability distribution and node embedding vector.
[0011] Preferably, in step 3), the student model adopts the LoRA-BERT model to efficiently process the pseudo text and is connected to a multi-layer perceptron as a classification head, whose input is the pseudo text and output includes category probability distribution and [CLS] vector.
[0012] Preferably, in step 3), the distillation training of the student model includes output layer distillation, wherein the output layer distillation is to use the category probability distribution output by the text teacher model and the graph structure teacher model as the supervision target, so that the student model jointly minimizes the KL loss of the category probability distribution output by the student model and the category probability distribution output by the text teacher model and the graph structure teacher model during the training process:
[0013] in, and Represent the category probability distributions output by the text teacher model and the graph structure teacher model, respectively, represents the category probability distribution output by the student model, 、 is an adjustable weight coefficient used to balance the influence of the text teacher model and the graph structure teacher model on the training of the student model.
[0014] Preferably, in step 3), performing distillation training on the student model further includes intermediate layer distillation, and the intermediate layer distillation specifically includes: 31) Using mean square error or cosine similarity as the alignment loss function, the [CLS] vector output by the student model is aligned with the intermediate layer representation output by the text teacher model; 32) Mapping the node embedding vector output by the graph structure teacher model to the same dimension as the [CLS] vector output by the student model through a learnable mapping layer, and then using mean square error or cosine similarity as the alignment loss function to align the [CLS] vector output by the student model with the node embedding vector output by the graph structure teacher model.
[0015] In addition, the present invention also provides a social media robot detection system, which is characterized by comprising: A user data collection and pseudo-text construction module, which is used to collect multi-dimensional basic attribute information of social media platform users and convert the multi-dimensional basic attribute information into pseudo-text with semantic expression capabilities according to a rule template; A teacher model construction and training module, which is used to construct a text teacher model and a graph structure teacher model, and train the text teacher model and the graph structure teacher model based on an open source dataset; A student model construction and distillation module, which is used to construct a student model and perform distillation training on the student model based on the text teacher model and the graph structure teacher model; A positive correlation comparison learning module is used to use the outputs of the text teacher model and the graph structure teacher model as positive samples, the output of the student model as anchor samples, and align and optimize the student model through a positive correlation comparison learning strategy; The social media robot detection module is used to detect social media robots using the aligned and optimized student model.
[0016] Furthermore, the present invention also provides a social media robot detection device, characterized by comprising: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the social media robot detection method as described above. Finally, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the steps of the social media robot detection method described above.
[0017] Compared with the prior art, the social media robot detection method and system of the present invention has one or more of the following beneficial technical effects: 1. Integrating multi-source knowledge to overcome the limitations of a single teacher: This paper introduces multiple teacher models from different modalities (e.g., text and graph structures) to fully exploit the complementary information in multimodal data. This addresses the problem of one-sided knowledge representation and insufficient generalization capabilities in traditional methods due to reliance on a single teacher model. Furthermore, in real-world deployments where raw user data (e.g., tweets and graph structures) is difficult to obtain or subject to privacy restrictions, this paper designs a metadata-based pseudo-text construction strategy. This strategy converts structured information such as user registration time, posting frequency, number of followers, device type, and IP activity into natural language, generating pseudo-text as the sole input to the student model. This enables efficient social media bot detection under low-resource conditions.
[0018] 2. Improve the representation alignment ability of the student model and the multi-teacher model: This invention constructs positive sample pairs between students and multiple teachers through a positive correlation contrast learning mechanism, aligns the multi-source semantic structure from the representation space level, effectively alleviates the interference caused by semantic inconsistency between teachers on the student model training, and enhances the student model's ability to uniformly understand multimodal semantics.
[0019] 3. Balancing detection performance and model deployment efficiency: While maintaining high detection accuracy, the student model generated by the proposed strategy is a lightweight structure that can be deployed in environments with no or weak graph conditions. It has strong practicality and adaptability, solving the problems of high deployment cost and strong dependence of existing graph neural network models.
[0020] 4. Significantly improved robustness and generalization capabilities: Thanks to the redundant supervision of multiple teachers and the alignment mechanism reinforced by contrastive learning, the student model can maintain stable performance in the face of data missing, modality inconsistency, or unknown robot behavior, and has stronger generalization and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 4 is a flow chart of the social media robot detection method of the present invention.
[0022] Figure 2 Schematic diagram of the social media robot detection system of the present invention. DETAILED DESCRIPTION
[0023] Before describing in detail any embodiment of the present invention, it should be understood that the present invention is not limited in its application to the construction and arrangement details of the components set forth in the following description or illustrated in the following figures. The present invention is capable of other embodiments and can be practiced or carried out in various ways. In addition, it should be understood that the words and terms used herein are for descriptive purposes and should not be considered restrictive. The use of "including" or "having" and their variations herein is intended to cover the items and their equivalents set forth below and additional items. Unless otherwise specified or limited, the terms "mounted", "connected", "supported" and "coupled" and their variations are used broadly and cover direct mounting and indirect mounting, connection, support and coupling. In addition, "connected" and "coupled" are not limited to physical or mechanical connections or couplings. Furthermore, on the first hand, in the disclosure of the present invention, the terms "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like to indicate orientations or positional relationships are based on the orientations or positional relationships shown in the accompanying drawings, which are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore the above terms cannot be understood as limitations on the present invention; on the second hand, the term "one" should be understood as "at least one" or "one or more", that is, in one embodiment, the number of an element may be one, while in another embodiment, the number of the element may be multiple, and the term "one" cannot be understood as a limitation on the quantity.
[0024] Current social media platforms are highly active, with information disseminating rapidly. Numerous bot accounts controlled by automated programs remain active across various platforms. These accounts, through mass content generation, manipulation of public opinion, and fraudulent marketing, pose a serious threat to platform ecosystem security, user trust, and the overall public opinion environment. To improve the accuracy and robustness of social media bot detection, recent research has gradually introduced multimodal modeling strategies, combining textual information, user behavior characteristics, and social graph data to enhance the model's ability to recognize complex semantic and behavioral patterns.
[0025] At the same time, knowledge distillation techniques are widely used in this field to transfer knowledge from a high-performance but computationally expensive teacher model to a lightweight student model, adapting to resource-constrained or data-restricted deployments. However, existing methods generally suffer from the following limitations: First, most use a single teacher model for distillation, which fails to fully exploit the complementarity of multimodal information. Second, the representation spaces between the teacher and student models differ significantly, lacking an effective semantic alignment mechanism, which impacts distillation effectiveness and model generalization.
[0026] To address the above technical bottlenecks, the present invention proposes a student model collaborative optimization method that integrates multi-teacher knowledge distillation and positive correlation contrastive learning. The core ideas include: (1) constructing a multi-teacher system consisting of a text teacher model and a graph structure teacher model to achieve joint modeling and guided distillation of knowledge in different modalities; (2) designing a student model that takes metadata + pseudo-text as input, combining lightweight structure and easy deployment; (3) introducing a positive correlation contrastive learning mechanism, taking multiple teacher outputs as positive samples, and enhancing the representation consistency between the student model and the teacher model through the semantic alignment goal. Therefore, the present invention significantly improves the detection performance, robustness, and adaptability of the student model to multi-modal loss situations in the social media robot detection task through the joint optimization strategy of multi-source supervision and contrastive alignment, and has good practical application value and promotion prospects.
[0027] Figure 1 Flowchart showing the social media robot detection method of the present invention. Figure 1 As shown, the social media robot detection method of the present invention includes the following steps: 1. User data collection and pseudo-text construction.
[0028] The multi-dimensional basic attribute information of social media platform users is collected and converted into pseudo-text with semantic expression capabilities according to a rule template.
[0029] Specifically, first, we collect multi-dimensional basic attribute information of social media platform users, covering key metadata such as registration time, number of fans, number of followers, content publishing frequency, account authentication status, device type used, and IP active areas.
[0030] Next, the user's multi-dimensional basic attribute information is standardized and cleaned, including removing missing or invalid records, unifying time formats, and normalizing continuous values. Numerical fields can be directly used as model input features; categorical fields are vectorized using one-hot encoding or embedding.
[0031] Finally, in order to give full play to the advantages of the pre-trained language model in natural language semantic understanding, the present invention further converts the above-mentioned structured multi-dimensional basic attribute information into a pseudo-text form with semantic expression capabilities according to a customized rule template.
[0032] In the present invention, the rule template may be: This user registered in {registration time} year, account {description of whether it is authenticated}, currently has {number of followers} followers, and follows {number of followers} users. The cumulative number of posts is {total number of posts}, with an average of {average posting frequency} posts per day. This user primarily logs in using {device type} devices, and their IP address active region is {description of changes in IP active region}.
[0033] With the above rule templates, the collected structured, multi-dimensional basic attribute information can be converted into a pseudo-text form with semantic expression capabilities. For example, the converted pseudo-text is "This user registered in 2018, the account is in a verified state, currently has 2,300 followers, and follows 80 users. A total of 300 pieces of content have been distributed, with an average of three posts per day. This user primarily logs in using an Android device, and their IP address activity area is relatively stable." The use of pseudo-text not only enhances the semantic readability of structured information, but also facilitates subsequent processing and integration with the language model (student model).
[0034] 2. Teacher model construction and training.
[0035] A text teacher model and a graph structure teacher model are constructed, and the text teacher model and the graph structure teacher model are trained based on an open source dataset.
[0036] In this paper, multiple teacher models are trained offline to extract user representations from two dimensions: text and graph structure. The two types of teacher models jointly provide rich soft-label supervision signals for the subsequent student model distillation.
[0037] Among them, the text teacher model uses the RoBERTa pre-trained language model. The input is user tweets and personal introduction text after data cleaning. The output includes soft labels (category probability distribution) and intermediate layer representations (such as [CLS] vectors and feature representations of multiple hidden layers).
[0038] The graph structure teacher model uses the RGCN heterogeneous graph neural network, with the input being the user's social relationship graph, and the output including soft labels (category probability distribution) and corresponding node embedding vectors.
[0039] In the present invention, the data for training the text teacher model and the graph structure teacher model can come from open source datasets, such as the Cresci dataset, the Twibot20 dataset or the TwiBot22 dataset. These datasets have text sequence data, graph relationship data and metadata, which can meet the training requirements of the present invention for the text teacher model and the graph structure teacher model.
[0040] 3. Student model construction and distillation.
[0041] A student model is constructed and distillation training is performed on the student model based on the text teacher model and the graph structure teacher model.
[0042] In the present invention, the structure of the student model adopts LoRA-BERT to efficiently process pseudo-text generated by metadata, and accesses a multi-layer perceptron (MLP) as a classification head, whose input is the pseudo-text and output includes category probability distribution and [CLS] vector.
[0043] In the present invention, distilling the student model specifically includes: 1. Output layer distillation - KL loss.
[0044] During the output layer distillation phase, the primary goal is to guide the student model in learning preliminary classification capabilities using the category probability distributions (i.e., soft labels) output by the teacher model. Specifically, the text teacher model and the graph teacher model each output a category probability distribution for determining whether a user is a bot account. For example, the probability distribution for an account being a "bot" is [0.92, 0.08]. These soft labels provide richer category boundary information than hard labels, helping the student model learn smoother and more generalizable decision boundaries. The student model's input is pseudo-text generated based on user metadata, encoded using LoRA-BERT, and then fed into a multi-layer perceptron (MLP) classifier to output predictions (category probability distributions). To align the student model's output with the soft labels of the two teacher models, Kullback–Leibler divergence (KL divergence) is used as the primary loss function for distillation training. KL divergence measures the difference between two probability distributions. The distillation objective is to minimize the KL loss between the category probability distribution output by the student model and the category probability distributions output by the two teacher models, thereby ensuring that the student model closely matches the teacher model's decision behavior in the output space. In the case of multiple teacher models, the distillation process not only considers the category probability distribution output by the text teacher model, but also simultaneously introduces the category probability distribution output by the graph structure teacher model, forming a dual supervision signal. The soft labels of both serve as supervision targets, and the student model jointly minimizes the KL loss function between itself and the outputs of the two teacher models during training, which can be expressed as follows: , Where, and Represent the category probability distributions output by the text teacher model and the graph structure teacher model, respectively, represents the category probability output by the student model, 、 It is an adjustable weight coefficient used to balance the influence of the two types of teacher models on the student model training, and can be set and adjusted as needed.
[0045] 2. Intermediate layer distillation-intermediate representation alignment.
[0046] In the middle-layer distillation stage, the present invention designs differentiated distillation strategies for different types of teacher models.
[0047] For the text teacher model, since both the teacher and student models use pre-trained language models (such as RoBERTa and LoRA-BERT), their structures are highly similar, and the hidden representation dimensions of the intermediate layers are consistent, intermediate representation alignment distillation can be performed at multiple levels. Specifically, the [CLS] vectors output by multiple hidden layers of the student model (such as the 6th and 12th layers) can be aligned with the corresponding hidden layers of the text teacher model. Using mean squared error (MSE) or cosine similarity as the loss function, deep semantic alignment between the layers is achieved.
[0048] However, for graph-structured teacher models (such as RGCN), their network structure is completely different from that of the student model, and their intermediate representation is the embedding vector of the nodes in the graph, which does not have hierarchical structure or contextual sequence features. Therefore, intermediate layer distillation can only be performed at the final user representation layer. In order to achieve alignment with the student model in the semantic space, a learnable mapping layer (such as linear projection or a small MLP) is first used to map the node embedding vector output by the graph-structured teacher model to the same dimension as the [CLS] vector output by the student model. Subsequently, MSE or cosine similarity is used as the alignment loss function to enable the student model to learn the social topological semantic information contained in the graph-structured teacher model.
[0049] By introducing multiple teacher models from different modalities (such as text and graph structures), the present invention fully mines the complementary information in multimodal data, solving the problems of one-sided knowledge expression and insufficient generalization ability caused by relying on a single teacher model in traditional methods.
[0050] 4. Positive correlation comparative learning.
[0051] In this paper, a positive correlation contrastive learning strategy is introduced, and the outputs of the text teacher model and the graph structure teacher model are regarded as positive samples, and the output of the student model is used as the anchor sample. The positive correlation contrastive learning strategy is used to enhance the student model's ability to model cross-modal consistency.
[0052] The present invention constructs positive sample pairs between students and multiple teachers through a positive correlation contrast learning mechanism, aligns multi-source semantic structures at the representation space level, effectively alleviates the interference caused by semantic inconsistency between teachers on student model training, and enhances the student model's ability to uniformly understand multimodal semantics.
[0053] Therefore, the present invention's distillation training of the student model includes three parts: KL loss, intermediate representation alignment, and positive correlation contrastive learning. During specific training, the student model can be trained using a joint loss function. The total loss is a weighted combination of the output layer's distillation loss (KL loss), the intermediate representation layer alignment loss, and the positive correlation contrastive loss of the intermediate layers: Loss = α·output layer distillation loss + β·intermediate representation layer alignment loss + γ·positive correlation contrastive loss, thereby achieving effective fusion and compression of multimodal knowledge. α, β, and γ are adjustable weight coefficients used to balance the impact of the three losses on student model training, and can be set and adjusted as needed.
[0054] 5. Social media robot detection.
[0055] In this paper, for social media bot detection, only the student model, trained through positive correlation contrast learning, needs to be deployed. This student model relies solely on pseudo-text generated from metadata as input, without requiring access to the user's actual text content or social graph structure information. This significantly reduces system deployment costs and reliance on data resources. Furthermore, this approach avoids direct processing of sensitive user data, effectively improving the student model's privacy protection capabilities in real-world scenarios and achieving a lightweight and scalable social media bot detection system.
[0056] Figure 2 The figure shows the structure of the social media robot detection system of the present invention. Figure 2 As shown, the social media robot detection system of the present invention includes: 1. User data collection and pseudo-text construction module.
[0057] The user data collection and pseudo-text construction module is used to collect multi-dimensional basic attribute information of users on the social media platform and convert the multi-dimensional basic attribute information into pseudo-text with semantic expression capabilities according to a rule template.
[0058] 2. Teacher model construction and training module.
[0059] The teacher model construction and training module is used to construct a text teacher model and a graph structure teacher model, and train the text teacher model and the graph structure teacher model based on an open source dataset.
[0060] 3. Student model construction and distillation module.
[0061] The student model construction and distillation module is used to construct a student model and perform distillation training on the student model based on the text teacher model and the graph structure teacher model.
[0062] 4. Positive correlation comparison learning module.
[0063] The positive correlation comparison learning module is used to use the outputs of the text teacher model and the graph structure teacher model as positive samples, and the output of the student model as anchor samples, and align and optimize the student model through a positive correlation comparison learning strategy; 5. Social media robot detection module.
[0064] The social media robot detection module is used to detect social media robots using the aligned and optimized student model.
[0065] In addition, the present invention also provides a social media robot detection device, which includes: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the social media robot detection method as described above. Finally, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is configured to implement the steps of the social media robot detection method described above when the program is executed by a processor.
[0066] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Those skilled in the art may, based on the principles of the present invention, modify or replace the technical solutions of the present invention with equivalents without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A social media robot detection method, characterized in that: The following steps are involved: 1) Collecting multi-dimensional basic attribute information of users on social media platforms and converting the multi-dimensional basic attribute information into pseudo-text with semantic expression capabilities based on rule templates; 2) Build a text teacher model and a graph structure teacher model, and train them based on open source datasets; 3) constructing a student model and performing distillation training on the student model based on the text teacher model and the graph structure teacher model; 4) Using the outputs of the text teacher model and the graph structure teacher model as positive samples and the output of the student model as anchor samples, the student model is aligned and optimized using a positive correlation contrast learning strategy; 5) Use the aligned and optimized student model to detect social media robots.
2. The social media robot detection method according to claim 1, characterized in that: In step 1), the multi-dimensional basic attribute information of social media platform users collected includes registration time, number of fans, number of followings, content publishing frequency, account authentication status, device type used, and IP active area. In addition, before being converted into pseudo-text with semantic expression capabilities, it is first standardized and cleaned, including removing missing or invalid records, unifying the time format, and normalizing continuous values.
3. The social media robot detection method according to claim 1, characterized in that: In step 2), the text teacher model adopts a pre-trained language model, whose input is user tweets and personal introduction text, and whose output includes category probability distribution and intermediate layer representation.
4. The social media robot detection method according to claim 3, characterized in that: In step 2), the graph structure teacher model adopts a heterogeneous graph neural network, whose input is the user's social relationship graph, and the output includes category probability distribution and node embedding vector.
5. The social media robot detection method according to claim 1, characterized in that: In step 3), the student model adopts the LoRA-BERT model to efficiently process the pseudo text and is connected to a multi-layer perceptron as a classification head, whose input is the pseudo text and output includes category probability distribution and [CLS] vector.
6. The social media robot detection method according to claim 1, characterized in that: In step 3), the distillation training of the student model includes output layer distillation, wherein the output layer distillation uses the category probability distribution output by the text teacher model and the graph structure teacher model as the supervision target, so that the student model jointly minimizes the KL loss of the category probability distribution output by the student model and the category probability distribution output by the text teacher model and the graph structure teacher model during the training process: in, and Represent the category probability distributions output by the text teacher model and the graph structure teacher model, respectively, represents the category probability distribution output by the student model, 、 is an adjustable weight coefficient used to balance the influence of the text teacher model and the graph structure teacher model on the training of the student model.
7. The social media robot detection method according to claim 6, characterized in that: In step 3), performing distillation training on the student model further includes intermediate layer distillation, and the intermediate layer distillation specifically includes: 31) Using mean square error or cosine similarity as the alignment loss function, the [CLS] vector output by the student model is aligned with the intermediate layer representation output by the text teacher model; 32) Mapping the node embedding vector output by the graph structure teacher model to the same dimension as the [CLS] vector output by the student model through a learnable mapping layer, and then using mean square error or cosine similarity as the alignment loss function to align the [CLS] vector output by the student model with the node embedding vector output by the graph structure teacher model.
8. A social media robot detection system, characterized in that: include: A user data collection and pseudo-text construction module, which is used to collect multi-dimensional basic attribute information of social media platform users and convert the multi-dimensional basic attribute information into pseudo-text with semantic expression capabilities according to a rule template; A teacher model construction and training module, which is used to construct a text teacher model and a graph structure teacher model, and train the text teacher model and the graph structure teacher model based on an open source dataset; A student model construction and distillation module, which is used to construct a student model and perform distillation training on the student model based on the text teacher model and the graph structure teacher model; A positive correlation comparison learning module is used to use the outputs of the text teacher model and the graph structure teacher model as positive samples, the output of the student model as anchor samples, and align and optimize the student model through a positive correlation comparison learning strategy; The social media robot detection module is used to detect social media robots using the aligned and optimized student model.
9. A social media robot detection device, characterized in that: include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the social media robot detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the social media robot detection method according to any one of claims 1 to 7 are implemented.