Fine-grained network spoofing detection method based on multi-modal large model
Through multimodal large model and data augmentation technology, the problem of modal singleness and coarse detection of cyberbullying detection on short video platforms is solved, and fine-grained detection of cyberbullying attributes, roles, topics and types is realized, which improves detection accuracy and generates applicable Chinese data sets.
Patent Information
- Application Number
- CN202510408405.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-25
AI Technical Summary
The existing cyberbullying detection technology has problems such as single modality, coarse detection granularity, low detection accuracy and neglect of repetition of cyberbullying on short video social platforms. In addition, there is a lack of data sets of Chinese multi-modal cyberbullying, making it difficult to adapt to the complexity of multi-round multi-role cyberbullying conversations.
A multi-modal large model-based cyberbullying detection method is adopted, including multi-modal coding module, cross-modal alignment module and multi-task learning module. Combined with thinking chain prompts and data enhancement technology, multiple rounds of multi-role Chinese cyberbullying data sets are generated to achieve fine-grained detection of cyberbullying attributes, roles, topics and types.
It improves the accuracy of cyberbullying detection, especially on short video platforms, which can simultaneously detect bullying attributes, roles, types and topics, solves the problem of data scarcity and supports more accurate cyberbullying detection and data synthesis.
Smart Images

Figure CN120372530A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer - supported multimodal processing. Specifically, it relates to a fine - grained cyberbullying detection method based on a multimodal large model. Background Art
[0002] (1) Cyberbullying Detection Technology
[0003] The popularization of the Internet and the wide application of smart phones have led to a rapid growth in social media users. Although information technology has brought convenience to people, it has also given rise to cyberbullying.
[0004] Cyberbullying refers to bullying behavior carried out using digital technology, usually through tools such as social media, instant messaging platforms, gaming platforms, and mobile phones, with the purpose of intimidating, irritating, or humiliating others. Cyberbullying has the following characteristics: using text, images, videos, etc. as media on social media platforms; usually occurring in post comments and chat sessions, and requiring an understanding from the expressed intention and context; its behavior patterns are closely related to the regional cultural background; it contains intentionality, repetition, and power imbalance; usually involves multiple parties such as victims, bullies, and bystanders; bullying topics include age, gender, occupation, behavioral activities, etc.; bullying types include slander, harassment, defamation, etc.
[0005] Existing cyberbullying detection research has shifted from initially being based on machine - learning models to mainly using deep - learning models and pre - trained large models at the model level, and from binary classification of cyberbullying attributes to more fine - grained detection at the detection granularity level, such as cyberbullying roles, types, and themes. From the perspective of cyberbullying behavior patterns, most research focuses on specific words and emotional features in text, and some research also endeavors to extract cyberbullying features from user profiles, sentence intentions, emojis, and images. For example, some research has modeled text data hierarchically at the word - level and comment - level, and first trains and then fuses the text, visual, and user - profile modalities separately, and verifies the effectiveness of the proposed multimodal cyberbullying detection framework through experiments. In addition, some research has analyzed the cyberbullying behavior patterns in the Indian social environment, modeled two - modality data of text and emojis (related to emotions and feelings), and improved the detection performance by emphasizing emotional features. Then, some research has used image - recognition APIs (such as Oxford and GPT - 4V) to convert images with bullying elements into text descriptions, and then used these texts to train a cyberbullying classifier, verifying the advancement of using multimodal information for cyberbullying detection and the feasibility of classifying bullying images through image descriptions.
[0006] Although the existing studies above have shown certain effectiveness and good detection performance in multi-modal cyberbullying detection, there is still room for improvement in their detection accuracy. The reason is that they often only focus on a single sentence and do not consider the context of multi-turn conversations. In addition, the current cyberbullying detection algorithms for short video platforms (including videos, pictures, and texts) are still blank, and the detection granularity can only focus on one or two aspects of cyberbullying, which is not conducive to the detection of cyberbullying and the adoption of further intervention measures.
[0007] (2) Multi-turn multi-role cyberbullying data synthesis technology
[0008] Under the upsurge of short video social platforms, the new cyberbullying behavior patterns not only require automated detection algorithms to follow up and innovate, but the cyberbullying data that can adapt to the times is also an important foundation for supporting the implementation of the algorithms. The current situation of the existing cyberbullying datasets is as follows: most of the existing cyberbullying datasets are in English or Hindi in terms of language type; in terms of annotation granularity, most of them only perform binary annotation on the cyberbullying attributes, that is, whether it is cyberbullying; in addition, the sample length often only contains a single sentence; in terms of data modality, most of them are text modality. The problems they have are insufficient language diversity, making it difficult to apply to the Chinese cultural environment; lack of finer-grained annotation of cyberbullying, which is not conducive to deeper analysis and detection; incomplete context, lack of multi-turn multi-role interaction data, ignoring the repeatability of cyberbullying; only focusing on text modality, with a single modality type. In addition, because the original cyberbullying data may lead to data privacy leakage and they are subject to the control of the platform's content review department, cyberbullying data is usually more scarce and difficult to obtain. Therefore, synthesizing cyberbullying data through advanced large language models has become one of the trends to solve the data problems in this field. Currently, some studies have synthesized cyberbullying data through generative models, and then used the synthetic data to train classifiers. The effectiveness of the synthesized data has been verified by evaluating the classification performance of the classifiers on real data. However, the synthesis methods proposed in these studies are still simple prompt engineering and cannot adapt to the complexity of synthesizing multi-turn multi-role cyberbullying conversation data. The synthesized data still only focuses on the prediction of the cyberbullying attributes of a single sentence, and there are still large deficiencies in terms of language type, annotation granularity, modality type, and context integrity, and cannot support the implementation of automated detection algorithms that adapt to the development of the Internet, especially short video platforms. Summary of the Invention
[0009] Aiming at the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to provide a fine-grained cyberbullying detection method based on a multimodal large model. The present invention is applicable to fine-grained cyberbullying detection in short video social platforms, including four aspects: bullying attributes, themes, types, and roles. By deeply grasping the characteristics of cyberbullying's intentionality and repeatability and making up for the gap in the cyberbullying detection framework for video stream information, more accurate and fine-grained detection of cyberbullying is achieved, promoting the construction and healthy development of a clean online space environment. At the same time, the present invention proposes to semi-automatically generate cyberbullying conversations based on chain-of-thought prompting and data augmentation techniques to synthesize a multi-round and multi-role Chinese cyberbullying dataset.
[0010] The technical solution of the present invention is specifically introduced as follows.
[0011] The present invention provides a cyberbullying detection method based on a multimodal large model, including the following steps:
[0012] (1) Training a cyberbullying detection model based on a multimodal large model using a cyberbullying dataset; wherein:
[0013] The cyberbullying dataset is cyberbullying posts and their comments containing multi-round and multi-role interactions in an online social platform, and each user comment will be labeled with bullying attributes, bullying roles, bullying themes, and bullying types;
[0014] The cyberbullying detection model includes a multimodal encoding module, a cross-modal alignment module, and a multi-task learning module:
[0015] The multimodal encoding module uses a pre-trained large model to encode multimodal information to obtain visual feature representations and text feature representations; visual features include scene and human action information;
[0016] The cross-modal alignment module synchronizes representations of different modalities into a unified space to obtain aligned multimodal representations for the fine-grained cyberbullying detection task;
[0017] The multi-task learning module, based on the aligned multimodal representations, takes multi-task learning as the goal to detect the bullying attributes, bullying roles, bullying themes, and bullying types of cyberbullying;
[0018] (2) Based on the trained cyberbullying detection model, perform cyberbullying detection on the video-text or image-text data to be tested.
[0019] In the present invention, in step (1), the cyberbullying dataset is constructed by using a generative large model to imitate and expand context-rich cyberbullying post conversations on a real dataset; the constructed dataset samples are based on a conversation, that is, a post and its comments below, as a unit, including three modalities: video, image, and text. For video-text data, it includes video fragments of human conversations and subtitle text extracted from the video; for image-text data, it includes photos posted in the post and feelings or descriptions about the photos; the tags include four major categories: cyberbullying attributes, roles, themes, and categories, and the language type is Chinese;
[0020] In the present invention, using the real dataset and the seed dataset as the input of the generative large model, after the synthetic dataset is obtained by the generative large model semi-automatically generating data under the guidance of the chain-of-thought prompt, it is then processed through data governance and data evaluation, and data governance and data evaluation are carried out crosswise; among them: during the chain-of-thought prompt process, three prompts are designed for each stage of the generation task: task description, generation conditions, and context examples, and the prompts are iteratively optimized through the generation results of each stage.
[0021] In the present invention, in the data governance stage, for the synthetic dataset, first, a heuristic method is used to select high-quality samples, that is, synthetic data with low fidelity is filtered out. Secondly, for the original data with relatively high quality but incorrect labels, label enhancement is carried out, that is, both manual annotation and a third-party auxiliary model voting annotation mechanism are used to correct the labels.
[0022] In the present invention, in the data evaluation stage, the synthetic dataset after data governance is directly evaluated for fidelity, and its evaluation dimensions are divided into three aspects: factual accuracy, logical consistency, and label accuracy; in the factual accuracy evaluation, each sample is checked for rule constraints based on a pre-constructed knowledge base in the cyberbullying field; in the logical consistency evaluation, focus is on the context self-consistency of multi-round conversations, and three types of problems are mainly identified: sudden changes in the behavior of cyberbullying roles, unreasonable jumps in topics, and mismatches in victim feedback; in the label accuracy evaluation, four scoring models are used to calculate the accuracy scores of various labels of the sample, and the labels are four major categories: cyberbullying attributes, roles, themes, and categories.
[0023] In the present invention, in step (1), in the multi-modal encoding module, based on the VideoSwin Transformer pre-trained model, the visual modality information is encoded to obtain visual feature representations, and based on Qwen2.5-7B-Instruct, the text modality information is encoded to obtain text feature representations.
[0024] In the present invention, in step (1), in the cross-modal alignment module, first, for visual feature representation, one-dimensional convolution is used to limit the number of tokens in the prefix and reduce the computational overhead, and a linear layer is used to reduce the dimension of the visual features to align them with the dimension of the text embedding space; then, in order to establish a unified representation space dominated by text, multi-head cross-attention is used to associate features of different modalities to obtain soft visual token inputs acceptable to the language model; then, special tokens are used to connect the visual input and the text token input, thereby obtaining the aligned multi-modal unified representation.
[0025] In the present invention, in step (1), in the multi-task learning module, first, the aligned multi-modal representation is input into Qwen2.5-7B-Instruct, and then the sequence output of the language model is averaged, and then passed through four task-specific fully connected layers, and finally passed through the softmax layer to output the predicted classification labels for the bullying attributes, bullying roles, bullying themes, and bullying types of cyberbullying.
[0026] In the present invention, in step (1), the bullying attribute is whether the sample to be detected belongs to cyberbullying; the bullying roles include victims, bullies, neutral bystanders, and just bystanders; the bullying themes include age, gender, appearance, background, and behavioral activities; the bullying types include slander, harassment, defamation, exclusion, and exposure.
[0027] In the present invention, in step (1), the loss function for training the cyberbullying detection model is the weighted sum of individual task losses, and the loss function for the task is categorical cross-entropy.
[0028] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0029] 1) It solves the problems that the detection algorithms for short video social platforms still have single modality, coarse detection granularity, low detection accuracy, and neglect of the repeatability of cyberbullying. By designing a fine-grained cyberbullying detection model based on a multi-modal large model (denoted as VedioSwin+Qwen2.5-7B-Instruct), and using the best model (VedioMAE+Roberta) on the public dataset ToxCMM as a baseline for comparative experiments, as Figure 6 shown, the experimental results show that the designed model has excellent cyberbullying detection performance and is superior to the baseline. The F1 value for detecting whether it is cyberbullying is 94.38, and the Accuracy is 94.21. In addition, multi-task ablation experiments and multi-modal ablation experiments are also carried out, and the experimental results are respectively as Figure 7 and Figure 8As shown, the results indicate that the designed model can effectively detect cyberbullying attributes, roles, types, and themes simultaneously, and its supplementary image and video modalities can further improve the accuracy of cyberbullying detection.
[0030] 2) Solved the problem of the lack of a multi-modal cyberbullying dataset with Chinese and fine-grained annotations. In the method of using few-shot chain-of-thought prompting technology to synthesize cyberbullying data, the present invention incorporates the steps of iterative optimization prompting and screening mechanisms to cope with the complexity of synthesizing multi-modal multi-role cyberbullying conversation data; introduces text-to-image technology in the chain-of-thought prompting to solve the problem of the lack of visual modality data in the cyberbullying field; for the annotation of synthetic data, adopts an artificial annotation and third-party auxiliary model voting annotation mechanism for label enhancement to improve the accuracy of annotation; in terms of enhancing data generalization, for the first time, proposes to enhance the simulated real data distribution by performing conversation rewriting on a real dataset and mixing it with synthetic data. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a structural diagram of the fine-grained cyberbullying detection model based on a multi-modal large model of the present invention.
[0032] Figure 2 It is a structural diagram of the multi-round multi-role cyberbullying data synthesis technology of the present invention.
[0033] Figure 3 It is an illustration of multi-round multi-role cyberbullying data samples of the present invention.
[0034] Figure 4 It is an illustration of the fact accuracy check prompt template of the present invention.
[0035] Figure 5 It is an illustration of the logical consistency check prompt template of the present invention.
[0036] Figure 6 It is an illustration of the baseline model comparison experiment results of the present invention.
[0037] Figure 7 It is an illustration of the multi-task ablation experiment results of the present invention.
[0038] Figure 8 It is an illustration of the multi-modal ablation experiment results of the present invention.
[0039] Figure 9 It is an illustration of the cyberbullying detection system interface of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The technical solutions of the present invention will be described in detail below with reference to the drawings and embodiments.
[0041] The present invention proposes a fine-grained cyberbullying detection method based on a multimodal large model, which specifically includes:
[0042] (1) A fine-grained cyberbullying detection method based on a multimodal large model is designed. The pre-trained large model is used to encode multimodal information and perform cross-modal spatio-temporal alignment to obtain the aligned representation of multimodality. Then, in the decoder part, multi-task learning is used as the goal to detect the attributes, roles, themes, and types of cyberbullying.
[0043] (2) A multi-round and multi-role cyberbullying data synthesis method is designed. The advanced generative large model (such as Doubao-1.5-pro) is used to rewrite and imitate cyberbullying post conversations with context on the real dataset through chain-of-thought prompting and data augmentation techniques, obtaining the cyberbullying data to support the implementation of the proposed fine-grained cyberbullying detection algorithm based on the multimodal large model.
[0044] Through the above two steps, the detection of the cyberbullying attributes, roles, themes, and types of short video posts can be finally achieved.
[0045] 1. Fine-grained Cyberbullying Detection Model Based on Multimodal Large Model
[0046] The model structure is as Figure 1 shown. The overall model can be divided into three parts: a multimodal encoding module, a cross-modal alignment module, and a multi-task learning module.
[0047] a) Multimodal Encoding Module
[0048] First, use VideoSwin Transformer to encode visual modality information. VideoSwin Transformer is a pure Transformer architecture designed specifically for video recognition tasks. Its core idea is to capture spatio-temporal dependencies in videos while maintaining computational efficiency by introducing spatio-temporal locality and hierarchical modeling. When modeling spatio-temporal locality, VideoSwin Transformer divides the video into local spatio-temporal windows and calculates self-attention only within the windows, significantly reducing the computational load. At the same time, through the shifted window strategy, adjacent windows can interact, thereby capturing spatio-temporal dependencies across windows. When extracting hierarchical features, it follows the hierarchical structure of Swin Transformer and gradually extracts spatio-temporal features of different granularities through multi-stage downsampling. In each stage, the resolution of the feature map is halved and the number of channels is doubled through the Patch Merging operation (similar to pooling in CNN), forming a pyramid structure. This design supports multi-scale feature fusion and is suitable for tasks such as action recognition and spatio-temporal detection. In practical applications, for the input of cyberbullying video clips, VideoSwin demonstrates powerful processing capabilities. Taking a cyberbullying video containing a person arguing and accompanied by verbal attacks as an example, VideoSwin first divides the video clip into multiple non-overlapping 3D blocks (T×H×W), each block containing the local spatio-temporal region of consecutive frames, and the input size is assumed to be T×H×W×3 (time×height×width×channels). These small blocks are embedded and mapped to a low-dimensional feature space to generate feature vectors, just like assigning a unique "identity label" to each small block for subsequent model understanding and processing. Then, the self-attention mechanism is used to capture the long-term dependencies in the video, calculate the attention weights between the small blocks, determine their importance levels, and then fuse the information to extract representative features. This process is like precisely screening out key information in the "information ocean" of the video, taking key elements such as the person's arguing actions and facial expression changes as the focus of attention. In the multi-layer Transformer encoder, the features processed by the self-attention mechanism are further refined and fused, and techniques such as layer normalization are used to improve the training efficiency and stability. Finally, the model outputs the deeply processed video feature representation, providing strong visual information support for subsequent cyberbullying detection.
[0049] Secondly, the Qianwen pre-trained large language model Qwen2.5-7B is used to encode text modality information. Qwen2.5-7B uses a byte-level byte-pair encoding (BPE) tokenizer, and its architecture includes multiple Transformer layers, each layer equipped with a causal attention mechanism and a feed-forward neural network. Secondly, it uses grouped query attention to replace multi-head attention, optimizes the KV cache during inference, and significantly improves the inference efficiency of the model. Qwen2.5-7B uses a large-scale pre-training dataset of over 2.2 trillion tokens, covering various data types such as text and code. On multiple evaluation datasets, Qwen2.5-7B has shown excellent performance, especially in the field of natural language understanding and generation. In the cyberbullying detection scenario, cyberbullying detection is essentially a natural language understanding task, which classifies cyberbullying attributes in the context of conversation level, sentence level, and word level by focusing on specific words, sentiment tendencies, intention viewpoints, etc. Taking a multi-turn cyberbullying conversation as a data sample, special tokens are added in the multi-turn conversation To distinguish the comment speeches of different roles, a specific BPE tokenizer matching Qwen2.5-7B is used to divide tokens and map them into vector representations in a high-dimensional space, thereby achieving the encoding of cyberbullying texts.
[0050] In summary, the visual encoder VideoSwin and the text encoder Qwen2.5-7B together constitute the multi-modal encoder of the cyberbullying detection framework. VideoSwin focuses on extracting visual features in videos, including information such as scenes and human actions, while Qwen2.5-7B is responsible for deeply understanding and encoding the text content. In actual cyberbullying scenarios, the actions and expressions of people in the video may imply the occurrence of bullying behavior, while the text comments clearly express the intentions and verbal attack content of the bullies. Through the collaborative work of these two encoders, the model can obtain cyberbullying information from multiple perspectives, making up for the limitations of single-modal encoders.
[0051] b) Cross-modal alignment module
[0052] Modal encoders are usually trained independently, which leads to differences between the generated representations. In cyberbullying detection, the visual modality and the text modality each contain unique information. If they cannot be effectively synchronized, it will affect the model's overall understanding of multi-modal data and the detection accuracy. Therefore, it is crucial to synchronize the representations of different modalities into a unified space to enhance the overall coherence and effectiveness of multi-modal processing.
[0053] For the visual high-dimensional feature Z encoded by the VideoSwin model v = VideoSwin(V), where V is the video input. To limit the number of tokens in the prefix and reduce the computational overhead, we use one-dimensional convolution to compress the length of the multi-modal feature into a compact and consistent value. The one-dimensional convolution operation can be regarded as sliding a convolution kernel over the feature vector and performing weighted summation on local features, thereby achieving the adjustment of the feature length. Suppose the visual high-dimensional feature is Z v , whose shape is (N, C, L), where N represents the number of samples, C represents the number of channels, and L represents the feature length. Through the one-dimensional convolutional layer Conv, with a convolution kernel size of k and a stride of s, the convolution operation can be expressed as:
[0054]
[0055] In the formula, W conv is the weight matrix of the convolution kernel, and b conv is the bias term. After the convolution operation, the obtained feature is then passed through a fully connected layer FC to reduce the hidden size of the visual feature and align it with the token embedding dimension of the language model. The transformation formula of the linear layer is:
[0056] C v = FC(Conv(Z v )) = W fc ·Conv(Z v ) + b fc
[0057] where W fc is the weight matrix of the linear layer, and b fc is the bias term. The finally obtained C v is the fixed-length visual abstract feature after convolution and the linear layer, and this fixed length is the same as the text embedding dimension. Then, to establish a text-dominated unified representation space, multi-head cross-attention (MHCA) is used to associate the features of the visual-text modality. Multi-head cross-attention involves applying scaled dot-product attention to three inputs, and the formula is as follows:
[0058] Attention
[0059]
[0060] where Q is the query vector from one modality, and K and V are the key and value from another modality, and d k is the dimension of the key and query vectors. In this study, for the text representation E t and the visual representation C v , the alignment of the video and image representations with the text embedding space is achieved through the following formula:
[0061]
[0062] In the above formula, E t is the text representation, and C s v is the aligned visual soft representation for inputting into the language model. The multi-head cross-attention mechanism can capture the inter-modal correlation information from different subspaces, enabling the visual features to be better aligned with the text features and integrated into the text-dominated representation space.
[0063] Finally, the special token <sep>to connect the visual representation C s v and the text representation E t to obtain an aligned multimodal unified representation for the fine-grained cyberbullying detection task. The connection operation can be simply expressed as:
[0064]
[0065] where M is the final multimodal unified representation, which integrates visual and text information. Through such a modality synchronization process, it provides rich and aligned feature inputs for the subsequent multi-task learning module.
[0066] c) Multi-task learning module
[0067] In this module, the aligned multimodal representation obtained from the previous module is first input into the language model. Subsequently, the sequence output of the language model will be averaged and passed through four task-specific fully connected layers, and finally, a softmax layer outputs the predicted classification labels for cyberbullying attributes, roles, themes, and types. In terms of the loss function, the loss function for all tasks is categorical cross-entropy, and the final loss function Loss f is the weighted sum of the individual task-specific losses Loss s . For M tasks, the contribution of the loss of each task to the overall loss is determined by the loss weight β. The loss function formula for multi-task learning is as follows:
[0068]
[0069] where the parameter β i represents end-to-end learning and indicates the contribution of task i to the multi-task loss, making the parameter updates between different tasks have different importance.
[0070] 2. Multi-round multi-role cyberbullying data synthesis technology
[0071] The synthesis goal of the cyberbullying dataset is to generate cyberbullying posts and their comments containing multi-round multi-role interactions in online social platforms, and each user comment will be labeled with bullying attributes, bullying roles, bullying themes, and bullying types. The original post can be in the form of text and images, or in the form of a video with subtitles. In addition, the synthetic data should be as close to the real data as possible, which is beneficial for the model to generalize to unseen real cyberbullying data. Therefore, the present invention proposes a multi-round multi-role Chinese cyberbullying data synthesis technology to achieve the synthesis goal. This technology is a method for semi-automatically synthesizing cyberbullying data through a generative large model using the chain of thought prompting technology and data augmentation technology.
[0072] The innovation of this method lies in the following aspects: First, in the method of using few-shot chain-of-thought prompting technology to synthesize cyberbullying data, a step of incorporating iterative optimization prompting and screening mechanisms is introduced to address the complexity of synthesizing multi-modal and multi-role cyberbullying conversation data. Second, text-to-image technology is introduced in the chain-of-thought prompting to solve the problem of the lack of visual-modal data in the field of cyberbullying. Third, for the annotation of synthetic data, an artificial annotation and a third-party auxiliary model voting annotation mechanism are adopted for label enhancement to improve the accuracy of annotation. Fourth, in terms of enhancing data generalization, for the first time, it is proposed to enhance the simulation of real data distribution by performing conversation rewriting on a real dataset and mixing it with synthetic data.
[0073] As Figure 2 shown, the multi-round and multi-role cyberbullying data synthesis technology is sequentially divided into four stages, namely real dataset collection, data generation, data governance, and data evaluation. Its core route is as follows:
[0074] In the real dataset collection stage, first, a publicly available real cyberbullying dataset A (consisting of a pure-text Chinese cyberbullying dataset COLDataset and a multi-modal cyberbullying dataset ToxCMM) is collected. The sample form is a single sentence or an image-text pair, and the latter is semantically aligned by humans. Second, an artificially designed seed dataset B is constructed. The samples of this dataset correspond to the synthesis target of the cyberbullying data described at the beginning of this section and will be used as context examples in subsequent prompt engineering.
[0075] In the data generation stage, first, design the task description, generation conditions, and context examples in prompt engineering. It should be noted that during the chain of thought prompting process, specific prompts for the above three items should be designed for each stage of the generation task, and the prompts should be iteratively optimized based on the stage generation results. The task description specifically refers to the instruction information given to the large model, including the role the large model needs to play, format instructions, task description, and necessary knowledge expansion (such as the characteristics and harms of cyberbullying). For example, the samples in the real dataset A are single comments or sentences, corresponding to the task of expanding and restoring the comments in a multi-round dialogue scenario (the samples in A are real); the samples in dataset B are subsets of the target dataset C, corresponding to the task of imitating comments and simulating scenarios. The generation conditions refer to restrictions on the overall structure, language style, and word count of the synthetic data, and also limit the types of labels of the synthetic data and their candidates. For example, the candidates for the bullying attribute label are "yes" and "no"; the candidates for the bullying role label are bully, victim, neutral bystander, just bystander (mediator), negative bystander (bully); the candidates for the bullying theme label are behavior activities, age, gender, appearance, background; the types of bullying are harassment, slander, libel, social exclusion, exposure. The context examples refer to the reference standards for the large model to generate data, which will enable the large model to better understand the goals of the synthetic task in in-context learning. Secondly, use the real dataset A and the seed dataset B as the input of the generative large model, and use the designed chain of thought prompts to guide the large model to semi-automatically generate data in multiple steps. That is, for the labeled samples in the real dataset A, they will go through the processes of data preprocessing, label mapping and relabeling, conversation expansion, image generation, and data integration to obtain the synthetic dataset A*. It should be noted that in the image generation stage, first use the cyberbullying conversation to generate a text description of the image, and then generate a realistic style image based on this description; for the seed dataset B, only need to go through the processes of conversation imitation, image generation, and data integration to obtain the synthetic dataset B*.
[0076] Figure 3 This is a diagram of the multi-round and multi-role cyberbullying data samples of the present invention.
[0077] In the data governance stage, for the samples in A* and B*, first, use a heuristic method to select high-quality samples, that is, filter out the synthetic data with low fidelity (the content of the synthetic data has factual errors, logical inconsistencies, label errors, or irrelevant content). Secondly, for the original data with relatively high quality but incorrect labels, label enhancement should be carried out, that is, simultaneously adopt the mechanisms of manual annotation and third-party auxiliary model voting annotation to correct the labels. After this stage is completed, the final dataset C for model training is obtained.
[0078] In the data evaluation stage, it is necessary to directly evaluate the fidelity of the synthetic dataset C. The evaluation dimensions can be divided into three aspects: factual accuracy, logical consistency, and label accuracy, so as to verify the usability of the synthesized data. It should be noted that data evaluation and data governance are carried out crosswise because governing data depends on the results of data evaluation.
[0079] In the factual accuracy evaluation, each sample is checked for rule compliance based on a pre-constructed knowledge base in the field of cyberbullying. The knowledge base covers common sense rules (such as "×× accusations need to be associated with verifiable ×× backgrounds") and temporal logic constraints (such as the mutual exclusivity of "not exercising for three months" and "daily check-in records"). The model makes a knowledge compliance judgment on the input sample through a specific prompt template (such as Figure 4 ). If any violation is detected, the corresponding error type is marked and the factual accuracy score is calculated. For example, for a dialogue sample containing "It is recommended to check ××, only those who ×××× like to show off their figures", if the victim has not publicly disclosed the ×× background information, the model will trigger a "Rule 2 violation" mark and reduce the factual score accordingly. The factual accuracy score Score fact is defined as:
[0080]
[0081] In the logical consistency evaluation, focus on the context self-consistency of multi-turn conversations, and mainly identify three types of problems: sudden changes in the behavior of cyberbullying characters, unreasonable jumps in topics, and mismatches in victim feedback. By designing inference prompt words as shown in Figure 5 , the model needs to analyze the consistency of the roles' positions in the conversation (such as whether the bully suddenly turns into a supporter), the logic of topic evolution (such as the rationality of the transition from "fitness effect" to "×× accusation"), and the alignment degree of attack-feedback (such as whether the victim's response matches the bullying content). The logic score Score logic is calculated through a weighted penalty mechanism, where Ⅱ is an indicator function that takes 1 when the condition is true and 0 otherwise. Feedback mismatch is given the highest weight because it directly affects the authenticity of the conversation and the annotation accuracy of multiple subsequent labels. The logic score Score logic is defined as:
[0082]
[0083] In the label accuracy evaluation, first, the synthetic dataset C is randomly divided into two sub-datasets with a ratio of 3:7 according to the number of samples. The division process must ensure that there is no data leakage between the two datasets. Then, the sample labels in the smaller sub-dataset are manually checked and corrected twice to ensure high annotation accuracy. Then, four Roberta-base models are used to train a multi-classification task for single-class labels on the smaller sub-dataset to obtain a label scoring model. The training of the four models is used for bullying attribute classification, role classification, theme classification, and type classification respectively. The reason for using multiple models for single-task training instead of a single model for multi-task training here is to ensure the high accuracy of the label scoring model itself as much as possible. For each sample in the larger sub-dataset, four Roberta scoring models are used to calculate the accuracy scores of various labels of the sample (1 point is obtained if the annotated label is consistent with the label output by the scoring model, otherwise 0 points). The label accuracy score Score label of each sample is also calculated through a weighted penalty mechanism:
[0084] Score lable = 0.4 × Score 属性 + 0.2 × Score 角色 + 0.2 × Score 主题 + 0.2 × Score 类型
[0085] After the faithfulness evaluation of each sample in the synthetic dataset, the following statistical data are obtained: 98.2% of the samples in the synthetic dataset C have a factual accuracy score Score fact ≥ 0.85; the proportion of samples with a logical consistency score Score logic ≥ 0.8 is 95.3%; the proportion of samples with a label accuracy score Score label ≥ 0.9 in the larger (7 / 10) sub-dataset of dataset C is 97.1%. For the samples that do not meet the standards, the data governance module will handle them by deletion or manual correction. The evaluation results of the above three evaluation dimensions verify the usability and effectiveness of the cyberbullying data synthesized by the present invention.
[0086] Implementation example:
[0087] In this example, a fine-grained cyberbullying detection model based on a multi-modal large model (denoted as VedioSwin+Qwen2.5-7B-Instruct) is designed, and a comparative experiment is carried out with the best model (VedioMAE+Roberta) on the public dataset ToxCMM as the baseline, as Figure 6 As shown, the experimental results show that the designed model has excellent cyberbullying detection performance and is superior to the baseline. The F1 value for detecting whether it is cyberbullying is 94.38, and the Accuracy is 94.21. In addition, multitask ablation experiments and multimodal ablation experiments were also carried out. The experimental results are shown in Figure 7 and Figure 8 respectively. The results show that the designed model can effectively detect cyberbullying attributes, roles, types, and themes simultaneously, and its supplementary image and video modalities can further improve the accuracy of cyberbullying detection.
[0088] The present invention also developed a cyberbullying detection system (the system interface is shown in Figure 9 ), aiming at the problem that cyberbullying incidents occur frequently in short video social platforms and the cyberbullying detection algorithms for video streams and posts are still blank, to provide an automated fine-grained cyberbullying detection system for platform content reviewers, relevant online community managers, and researchers in this field. The system is built using Python and Flask and deployed on a linux server. The system includes a file upload module and a cyberbullying detection module. After uploading video-text or image-text data, the system will detect the cyberbullying attributes, roles, types, and themes in the input data. It is worth noting that for the convenience of users, the system also supports inputting text by parsing the text in pictures (such as screenshots of multi-round conversations).< / sep>
Claims
1. A method for detecting cyberbullying based on a multimodal large model, characterized in that It includes the following steps: (1) Training a cyberbullying detection model based on a multimodal large model using a cyberbullying dataset; where: The cyberbullying dataset is cyberbullying posts and their comments containing multi-round multi-role interactions in an online social platform, and each user comment will be labeled with bullying attributes, bullying roles, bullying themes, and bullying types; The cyberbullying detection model includes a multimodal encoding module, a cross-modal alignment module, and a multi-task learning module; The multimodal encoding module encodes multimodal information using a pre-trained large model to obtain visual feature representations and text feature representations; visual features include scene and human action information; The cross-modal alignment module synchronizes representations of different modalities into a unified space to obtain aligned multimodal representations for fine-grained cyberbullying detection tasks; The multi-task learning module detects the bullying attributes, bullying roles, bullying themes, and bullying types of cyberbullying based on the aligned multimodal representations with multi-task learning as the goal; (2) Based on the trained cyberbullying detection model, implementing cyberbullying detection on the video-text or image-text data to be measured.
2. The network bullying detection method according to claim 1, wherein In step (1), the cyberbullying dataset is constructed by using a generative large model to imitate and expand context-rich cyberbullying post conversations on a real dataset; the constructed dataset samples include video-text data and image-text data; taking a conversation, that is, a post and its comments below, as a unit, it includes three modalities: video, image, and text, and the labels include four major categories: cyberbullying attributes, roles, themes, and categories, and the language type is Chinese.
3. The network bullying detection method according to claim 1, wherein In step (1), using the real dataset and the seed dataset as the input of the generative large model, after the synthetic dataset is obtained by guiding the generative large model to semi-automatically generate data through chain-of-thought prompting, it is then processed through data governance and data evaluation; where: during the chain-of-thought prompting process, three prompts are designed for each stage of the generation task: task description, generation conditions, and context examples, and the prompts are iteratively optimized through the generation results of each stage.
4. The network bullying detection method according to claim 3, wherein In the data governance stage, for the synthetic dataset, first, a heuristic method is used to select high-quality samples, that is, filtering out synthetic data with low fidelity. Secondly, for the original data with relatively high quality but incorrect labels, label enhancement is performed, that is, both manual annotation and a third-party auxiliary model voting annotation mechanism are used to correct the labels.
5. The network bullying detection method according to claim 3, wherein In the data evaluation stage, the fidelity of the synthetic dataset after data governance is directly evaluated. The evaluation dimensions include three aspects: factual accuracy, logical consistency, and label accuracy. In the factual accuracy evaluation, each sample is checked for rule compliance based on the pre-constructed knowledge base in the cyberbullying domain. In the logical consistency evaluation, focus is on the context self-consistency of multi-turn conversations, and three types of problems are mainly identified: sudden changes in the behavior of cyberbullying characters, unreasonable jumps in topics, and mismatches in victim feedback. In the label accuracy evaluation, four scoring models are used to calculate the accuracy scores of various labels of the sample. The labels are four major categories: cyberbullying attributes, roles, themes, and categories.
6. The network bullying detection method according to claim 1, wherein In step (1), in the multi-modal encoding module, the visual modal information is encoded based on the VideoSwin Transformer pre-trained model to obtain visual feature representations, and the text modal information is encoded based on the large language model Qwen2.5-7B-Instruct to obtain text feature representations.
7. The network bullying detection method according to claim 1, characterized in that In step (1), in the cross-modal alignment module, first, for the visual feature representations, one-dimensional convolution is used to limit the number of tokens in the prefix and reduce the computational overhead, and a linear layer is used to reduce the dimension of the visual features to align them with the dimension of the text embedding space. Then, to establish a unified representation space dominated by text, multi-head cross-attention is used to associate features of different modalities to obtain soft visual token inputs that can be accepted by the language model. Then, special tokens are used to connect the visual input and the text token input to obtain the aligned multi-modal unified representation.
8. The network bullying detection method according to claim 1, wherein In step (1), in the multi-task learning module, first, the aligned multi-modal representations are input into the large language model Qwen2.5-7B-Instruct. Subsequently, the sequence output of the large language model is averaged, then passed through four task-specific fully connected layers, and finally, the predicted classification labels for the cyberbullying attributes, cyberbullying roles, cyberbullying themes, and cyberbullying types are output through the softmax layer.
9. The network bullying detection method according to claim 1, wherein In step (1), the cyberbullying attribute refers to whether the sample to be detected belongs to cyberbullying; the cyberbullying roles include victims, bullies, neutral bystanders, and just bystanders; the cyberbullying themes include age, gender, appearance, background, and behavioral activities; the cyberbullying types include slander, harassment, vilification, exclusion, and exposure.
10. The network bullying detection method according to claim 1, wherein, In step (1), the loss function for training the cyberbullying detection model is the weighted sum of individual task losses, and the loss function for each task is categorical cross-entropy.
Citation Information
Cited By
Data processing method, electronic device, storage medium and computer program product
CN121279458A