Network spoofing detection method based on OpenAI embedding and RNN
By combining OpenAI embedding and RNN and LSTM network models, the adaptability and accuracy of the cyberbullying detection methods in the prior art are solved, and efficient identification and accurate detection of cyberbullying behaviors are achieved.
Patent Information
- Application Number
- CN202510247444.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
Existing cyberbullying detection methods are difficult to adapt to changing cyberbullying behaviors. Machine learning-based methods require a lot of labeled data and are difficult to capture the complexity and nuance of cyberbullying languages. Manual reviews are time-consuming and labor-intensive and error-prone.
Using the cyberbullying detection method based on OpenAI embedding and RNN, the user post data is collected from social media platforms, cleaned and text classification, and RNN network and LSTM network models are constructed, combined with OpenAI embedding model for training and testing, and the context semantics of the text are captured.
It improves the accuracy and robustness of cyberbullying detection, can more effectively identify complex cyberbullying behaviors and potential insulting expressions, and reduces the rate of misjudgment.
Smart Images

Figure CN120180184A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cyberbullying detection, and particularly to a cyberbullying detection method based on OpenAI embedding and RNN. Background Art
[0002] With the popularization of social media, the problem of cyberbullying has become increasingly serious. Cyberbullying not only causes psychological and emotional harm to victims, but may also lead to serious consequences such as suicide. Therefore, developing effective cyberbullying detection methods has important social significance.
[0003] Existing cyberbullying detection methods are mainly divided into two categories: rule-based methods and machine learning-based methods. Rule-based methods rely on predefined rules and patterns and are difficult to adapt to the ever-changing cyberbullying behaviors; machine learning-based methods (such as support vector machine SVM and naive Bayes) require a large amount of labeled data for training, but it is difficult to capture the complexity and nuances of cyberbullying language. In addition, the manual review method is time-consuming and laborious and prone to errors, and the concealment and indefinability of cyberbullying behaviors make the traditional detection methods have limited effects. Therefore, how to automatically detect cyberbullying has become an urgent problem to be solved. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a cyberbullying detection method based on OpenAI embedding and RNN.
[0005] To achieve the above purpose, the present invention provides the following solutions:
[0006] A cyberbullying detection method based on OpenAI embedding and RNN, comprising:
[0007] Collecting user post data from a social media platform; the user post data is labeled with whether it contains a cyberbullying behavior label;
[0008] Cleaning the collected user post data to remove irrelevant information and obtaining cleaned data;
[0009] Performing text classification on the cleaned data based on the context relationship to obtain classified data;
[0010] Dividing the classified data into a training set and a test set, and annotating the training set, marking the data with a cyberbullying behavior label as the negative class, and the data without a cyberbullying behavior label as the positive class;
[0011] Constructing an initial detection model; the initial detection model includes an RNN network, an LSTM network, and an OpenAI embedding model;
[0012] Input the text in the training set into the RNN network, and let the OpenAI embedding model convert the text into a vector representation. Then input the vector representation into the LSTM network for training to obtain a trained model, and use the test set to test the trained model to obtain a well-tested cyberbullying detection model;
[0013] Input the data of the post to be tested into the cyberbullying detection model to obtain a detection result.
[0014] Preferably, the user post data includes the user's questions and answers.
[0015] Preferably, clean the collected user post data to remove irrelevant information to obtain clean data, including:
[0016] Use regular expressions to match HTML tags, non-alphabetic, non-numeric characters, and URLs in the user post data, and remove the HTML tags, non-alphabetic, non-numeric characters, and URLs;
[0017] Replace consecutive repeated characters in the removed text with a single character;
[0018] Convert all uppercase characters in the replaced text to lowercase characters;
[0019] Delete the extra spaces in the converted text to obtain the clean data.
[0020] Preferably, perform text classification on the clean data based on the context relationship to obtain classification data, including:
[0021] Input the clean data into the text context relationship feature extraction layer to extract context relationship feature information;
[0022] Input the clean data into the global feature extraction layer to extract global feature information;
[0023] Fuse the context relationship feature information and the global feature information to obtain fused text features;
[0024] Input the fused text features into the trained text classification model to obtain the classification data.
[0025] Preferably, input the clean data into the text context relationship feature extraction layer to extract context relationship feature information, including:
[0026] Use a pre-trained language model to extract the initial feature information of the clean data;
[0027] Input the initial feature information into the forward gated recurrent unit and the backward gated recurrent unit;
[0028] Concatenate the outputs of the forward gated recurrent unit and the backward gated recurrent unit to obtain the context relationship feature information.
[0029] Preferably, concatenating the outputs of the forward gated recurrent unit and the backward gated recurrent unit to obtain the context relationship feature information includes:
[0030] Using the formula:
[0031]
[0032] Concatenate the outputs of the forward gated recurrent unit and the backward gated recurrent unit to obtain the context relationship feature information; where, H t ={H1, H2, … H l} represents the initial feature information input at time t, represents the output of the forward gated recurrent unit at time t, represents the output of the backward gated recurrent unit at time t, GRU represents the gated recurrent unit, and G represents the context relationship feature information.
[0033] Preferably, inputting the cleaned data into the global feature extraction layer to extract the global feature information includes:
[0034] Input the initial feature information of the cleaned data into the convolutional layer and the pooling layer in sequence to obtain the global feature information; where, the global feature information extraction formula is:
[0035] c i = f(ω · H + b)
[0036]
[0037] In the formula, f is the activation function, ω is the convolution kernel, h is the convolution kernel size, b is the bias, and c i is the i-th extracted feature vector of the convolutional layer, represents the value after the features extracted by 3 different convolution kernels pass through the max pooling layer, and C represents the extracted global feature information.
[0038] Preferably, the text classification model includes a fully connected layer and an output layer connected in sequence; the loss function of the text classification model is: Where, p i is the predicted probability of the positive class, p i = σ(z i ), where σ(·) represents the Sigmoid function; y i∈{0,1} indicates that the sample is a positive or negative class; γ is the focusing coefficient in FocalLoss; m is the Margin hyperparameter; w0 and w1 are the positive and negative class weight coefficients respectively, used to balance class imbalance; y i is the true label, y i = 1 represents the positive class, y i = 0 represents the negative class; N is the total number of training samples, z i represents the result output by the fully connected layer for the i-th sample.
[0039] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:
[0040] The present invention provides a cyberbullying detection method based on OpenAI embedding and RNN, including: collecting user post data from a social media platform; the user post data is labeled with whether it contains cyberbullying behavior; cleaning the collected user post data to remove irrelevant information to obtain cleaned data; performing text classification on the cleaned data based on context relationships to obtain classified data; dividing the classified data into a training set and a test set, and annotating the training set, marking the data with the label of containing cyberbullying behavior as the negative class, and marking the data without the label of containing cyberbullying behavior as the positive class; constructing an initial detection model; the initial detection model includes an RNN network, an LSTM network, and an OpenAI embedding model; inputting the text in the training set into the RNN network, and enabling the OpenAI embedding model to convert the text into a vector representation, and inputting the vector representation into the LSTM network for training to obtain a trained model, and using the test set to test the trained model to obtain a well-tested cyberbullying detection model; inputting the to-be-tested post data into the cyberbullying detection model to obtain a detection result. By combining the pre-trained OpenAI embedding with RNN and LSTM to capture the context semantics of the text, and combining the data cleaning and classification processes, the present invention can efficiently improve the accuracy and robustness of cyberbullying detection. Description of the Drawings
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0042] Figure 1 is the flowchart of the method provided by the embodiment of the present invention. Detailed Embodiments
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0044] The object of the present invention is to provide a cyberbullying detection method based on OpenAI embedding and RNN. By combining the pre-trained OpenAI embedding with RNN and LSTM to capture the context semantics of the text, and combining the data cleaning and classification processes, the accuracy and robustness of cyberbullying detection can be effectively improved.
[0045] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0046] Figure 1 The flowchart of the method provided for the embodiments of the present invention is as Figure 1 shown. The present invention provides a cyberbullying detection method based on OpenAI embedding and RNN, including:
[0047] Step 100: Collect user post data from a social media platform; the user post data is labeled with a tag indicating whether it contains cyberbullying behavior.
[0048] Step 200: Clean the collected user post data to remove irrelevant information and obtain cleaned data.
[0049] Step 300: Perform text classification on the cleaned data based on the context relationship to obtain classified data.
[0050] Step 400: Divide the classified data into a training set and a test set, and label the training set. The data with a tag indicating cyberbullying behavior is marked as the negative class, and the data without a tag indicating cyberbullying behavior is marked as the positive class.
[0051] Step 500: Build an initial detection model; the initial detection model includes an RNN network, an LSTM network, and an OpenAI embedding model.
[0052] Step 600: Input the text in the training set into the RNN network, and let the OpenAI embedding model convert the text into a vector representation, and input the vector representation into the LSTM network for training to obtain a trained model, and use the test set to test the trained model to obtain a well-tested cyberbullying detection model.
[0053] Step 700: Input the data of the post to be tested into the cyberbullying detection model to obtain the detection result.
[0054] Preferably, the user post data includes the user's questions and answers.
[0055] Specifically, step 100 of this embodiment mainly collects user post data by calling the official API provided by the social media platform or using an authorized crawler system. To ensure the legality and accuracy of the data, it is usually necessary to first obtain the permission of the platform party or relevant users and comply with the platform's terms of use. The collected data covers the original text content posted by users, the basic information of the publisher (if compliance permits), and metadata such as timestamps to ensure sufficient context support in subsequent processing and analysis.
[0056] In specific implementation, this embodiment screens the required languages, keywords, or time ranges according to research needs. For example, to better focus on cyberbullying-related content, keywords similar to abusive and offensive language can be key-captured or filtered. At the same time, combined with the platform's tagging function or potential text classification mechanism, information such as "containing bullying" and "not containing bullying" can be obtained together during the acquisition stage to provide a preliminary annotation basis for subsequent classification and training.
[0057] After completing the preliminary data acquisition, this embodiment stores it structurally in a database or a distributed data processing system. The storage format can include JSON, CSV, or a NoSQL table structure more suitable for large-scale processing, so that it can be efficiently indexed and extracted during subsequent data cleaning, classification, and model training. At the same time, by recording the question-and-answer form in the user post at the data level, structural information such as interview-style conversations and Q&A scenarios can be incorporated into subsequent text analysis, providing richer context features for the cyberbullying detection model.
[0058] Preferably, the collected user post data is cleaned to remove irrelevant information to obtain cleaned data, including:
[0059] Use regular expressions to match HTML tags, non-alphabetic, non-numeric characters, and URLs in the user post data, and remove the HTML tags, non-alphabetic, non-numeric characters, and URLs;
[0060] Replace consecutive repeated characters in the removed text with a single character;
[0061] Convert all uppercase characters in the replaced text to lowercase characters;
[0062] Delete the extra spaces in the converted text to obtain the cleaned data.
[0063] Specifically, step 200 of this embodiment includes:
[0064] Step 201: In the present invention, first, by using regular expressions to match in the user post data, HTML tags, non-alphabetic characters, non-numeric characters, and possibly included URL links are located and deleted. This can remove such as " Tags such as "<a href=””" or link formats such as "http: / / ”"www.” are also removed, as well as noise elements that affect model understanding, such as "@#$” and emojis, to ensure that subsequent processing only focuses on the effective text content.
[0065] Step 202: After the above cleaning is completed, it is necessary to further process consecutive repeated characters in the text (such as "!!!” and "hahaha”). By detecting such characters and merging them into a single character, the text length can be greatly reduced, noise caused by users frequently repeating certain letters or symbols can be avoided, and it also helps the subsequent model to more accurately extract semantic features.
[0066] Step 203: For all uppercase characters in the processed text, they should be uniformly converted to their corresponding lowercase forms. Case handling allows the model to obtain consistent feature vectors when processing different cases of the same word, and also ensures that there is less ambiguity during pre-training word vector or word frequency statistics operations, avoiding duplicate calculations of near-synonymous uppercase and lowercase words during model training and subsequent analysis.
[0067] Step 204: In the last step, it is necessary to delete the extra spaces in the text to prevent the remaining extra spaces after deleting special characters or merging repeated characters from affecting word segmentation or subsequent feature extraction. After removing the extra spaces, cleaner data that is more concise and convenient for subsequent processing is obtained, laying a good foundation for subsequent text classification and training of the cyberbullying detection model.
[0068] Preferably, based on the context relationship, the cleaned data is classified to obtain classification data, including:
[0069] Input the cleaned data into the text context relationship feature extraction layer to extract context relationship feature information;
[0070] Input the cleaned data into the global feature extraction layer to extract global feature information;
[0071] Fuse the context relationship feature information and the global feature information to obtain fused text features;
[0072] Input the fused text features into the trained text classification model to obtain the classification data.
[0073] Preferably, inputting the cleaned data into the text context relationship feature extraction layer to extract context relationship feature information includes:
[0074] Use a pre-trained language model to extract the initial feature information of the cleaned data;
[0075] Input the initial feature information into the forward gated recurrent unit and the backward gated recurrent unit;
[0076] Concatenate the outputs of the forward gated recurrent unit and the backward gated recurrent unit to obtain the context relationship feature information.
[0077] Preferably, concatenating the outputs of the forward gated recurrent unit and the backward gated recurrent unit to obtain the context relationship feature information includes:
[0078] Using the formula:
[0079]
[0080] Concatenate the outputs of the forward gated recurrent unit and the backward gated recurrent unit to obtain the context relationship feature information; where, H t ={H1, H2, … H l} represents the initial feature information input at time t, represents the output of the forward gated recurrent unit at time t, represents the output of the backward gated recurrent unit at time t, GRU represents the gated recurrent unit, and G represents the context relationship feature information.
[0081] Specifically, by concatenating the outputs of the forward gated recurrent unit and the backward gated recurrent unit, the present invention can more fully learn the text context relationship and obtain the context information.
[0082] Preferably, inputting the cleaned data into the global feature extraction layer to extract the global feature information includes:
[0083] Input the initial feature information of the cleaned data into the convolutional layer and the pooling layer in sequence to obtain the global feature information; where, the global feature information extraction formula is:
[0084] c i =f(ω·H + b)
[0085]
[0086] In the formula, f is the activation function, ω is the convolution kernel, h is the convolution kernel size, b is the bias, and c i is the i-th extracted feature vector of the convolutional layer, represents the value after the features extracted by 3 different convolution kernels pass through the max pooling layer, and C represents the extracted global feature information.
[0087] Specifically, in the text context relationship feature extraction layer, first, the cleaned text is passed through a pre-trained language model (such as BERT, GPT, etc.) to obtain the corresponding initial feature vectors, enabling each word or token to have a relatively rich semantic representation. Next, these initial features are sequentially input into a forward gated recurrent unit (GRU) and a backward gated recurrent unit (GRU). The former processes the sequence data in the order from left to right of the text, while the latter processes the data in reverse from right to left. By this bidirectional way, the contextual correlation features of the text are captured.
[0088] After completing the extraction of the above forward and backward information, the outputs of the two are concatenated to obtain a feature representation containing a more comprehensive context relationship. In other words, the words at the same moment will retain their association with the previous text in the forward GRU, and their association with the subsequent text in the backward GRU. By concatenating these two results, the bidirectional dependence of the context before and after can be centrally reflected in a comprehensive vector representation. This bidirectional concatenation method allows the model to more fully learn the internal logical structure and semantic context of the text, thus achieving a more accurate recognition performance in the subsequent cyberbullying detection process.
[0089] In the global feature extraction layer, convolution and pooling are used to capture higher-level features of the text. Specifically, the initial feature information of the cleaned data can be first input into several convolutional kernels of different sizes. Different convolutional kernels can capture n-gram features of different granularities and patterns, and then the maximum pooling (or average pooling) operation is used to obtain the most significant features of each convolutional output. Finally, the pooling results obtained from multiple convolutional kernels are combined to obtain a relatively complete global feature representation, providing richer and multi-dimensional information support for subsequent tasks such as discriminating cyberbullying.
[0090] Optionally, the model construction process in step 500 of this embodiment includes:
[0091] 1. Select an RNN model: Construct a recurrent neural network (RNN) model and use long short-term memory (LSTM) units because LSTM units can effectively capture the sequential dependence of the text.
[0092] 2. Select an OpenAI embedding model: Use the pre-trained embedding model provided by OpenAI, such as text-embedding-ada-002 or text-embedding-ada-001. This model can convert the text into a high-dimensional vector representation and capture the semantic information of the text.
[0093] 3. Model training:
[0094] Data Input: Input the text in the training set into the RNN model, and use the OpenAI embedding model to convert the text into a vector representation. For example, input x = {x1, x2, ···, x n} into the model, where x i represents the i-th word in the text.
[0095] 4. Embedding Representation: Use the OpenAI embedding model to convert the text x into a vector representation h = {h1, h2, ···, h n}, where h i represents the embedding vector of the i-th word in the text.
[0096] LSTM Calculation: Input the embedding vector h into the LSTM cell, and calculate the output y of the LSTM cell, where y represents the final representation of the text.
[0097] Loss Function: Use the loss function to evaluate the difference between the model prediction result and the true label.
[0098] Model Optimization: Use the Adam optimizer to minimize the loss function and update the model parameters so that the model can more accurately predict cyberbullying.
[0099] Preferably, the text classification model includes a fully connected layer and an output layer connected in sequence; the loss function of the text classification model is: where p i is the prediction probability for the positive class, p i = σ(z i ), where σ(·) represents the Sigmoid function; y i ∈{0, 1} represents that the sample is a positive class or a negative class; γ is the focusing coefficient in FocalLoss; m is the Margin hyperparameter; w0 and w1 are the positive and negative class weight coefficients respectively, used to balance the class imbalance; y i is the true label, y i = 1 represents the positive class, y i = 0 represents the negative class; N is the total number of training samples, and z i represents the result output by the i-th sample after passing through the fully connected layer.
[0100] Specifically, when constructing the text classification task, the model in this embodiment combines the ideas of focal loss and adding an offset. By further introducing an adjustment term in the loss function, it improves the attention to difficult-to-classify samples. Compared with common models that only use focal loss or only rely on simple classification boundaries, it can better distinguish subtle positive and negative class differences during the training process. This combination helps to address the problems of sample imbalance and high difficulty in recognizing some bullying texts, thus showing stronger adaptability and robustness in the real scenario of cyberbullying detection.
[0101] When designing the model, the weight - boosting effect of focal loss on difficult - to - classify samples is fully utilized, enabling the network to gradually focus on those subtle difference scenarios during training that have potential offensive or bullying intentions but are difficult to distinguish from ordinary text. Meanwhile, the introduced additional offset enables the classification boundary to have better discrimination ability for the context, thus achieving a smoother transition between high - confidence and low - confidence text recognition and reducing the misjudgment rate.
[0102] Regarding the determination method of positive and negative class weight coefficients, it can usually be set according to the quantity ratio of positive and negative samples in the training dataset: if the quantity of negative samples is much larger than that of positive samples, the positive class weight is appropriately increased to prevent the model from "favoring" predicting as the negative class; vice versa. It is also possible to test different weight combinations based on cross - validation or search strategies in practical applications, so as to select the weight scheme that best meets the business requirements (such as recall rate or accuracy) indicators. Through this appropriate weight configuration, the negative impact brought by the imbalance of construction data can be offset to the greatest extent.
[0103] As an alternative implementation manner, this embodiment provides a specific application process:
[0104] S1. First, pre - process the social media posts (such as word - segmenting, removing stop words, removing HTML tags, etc.), and then input the pre - processed text into the OpenAI API to obtain the embedded representation of the text.
[0105] Its essence is to send a text string in JSON format to the OpenAI API and receive the text feature vector returned by the API as the output.
[0106] These feature vectors contain the vector representation of each word in the text, thus retaining the semantic information and context relationship of the text.
[0107] The obtained embedded representation is input into a recurrent neural network (RNN) for classification to determine whether the post contains cyberbullying.
[0108] Finally, output the result to inform the user whether the text belongs to cyberbullying and the corresponding classification result.
[0109] For example, for the sentence "The shoes you wear today are so ugly.", the pre - processing may include word - segmenting, removing stop words (such as "very", "ne", etc.), removing HTML tags, etc.
[0110] In addition, this embodiment also performs context classification on the classification result of whether it belongs to cyberbullying.
[0111] S2. Call the OpenAI API to convert the text into an embedded representation.
[0112] Input the preprocessed text sentence by sentence into the OpenAI API to obtain the word vector representation of the text.
[0113] S3. Input the embedded representation of the text into the RNN for classification.
[0114] Use a pre-trained RNN model to classify the text to be detected. The RNN can learn from and remember text patterns and sequential relationships from the previously processed results, enabling it to effectively identify cyberbullying.
[0115] S4. Output the detection result according to the classification result of the RNN.
[0116] If the classification score is higher than 0.9, it is considered that the text contains cyberbullying; otherwise, it is considered that the text does not contain cyberbullying.
[0117] Optionally, this embodiment can also collect the reporting feedback of users and count the number and frequency of reports of unfriendly behaviors to determine which words or behavior patterns are more likely to be identified as cyber violence, thereby dynamically adjusting the classification criteria.
[0118] The beneficial effects of the present invention are as follows:
[0119] (1) The OpenAI embedding model can map the text to a vector space in a high-dimensional manner through rich pre-trained semantic knowledge, and combine the ability of the RNN and LSTM to capture sequential information to achieve more accurate text feature extraction and context relationship description.
[0120] (2) When processing sequential data, the RNN / LSTM can retain the context information, making the network more sensitive to the associated words, metaphorical expressions, and context clues in cyberbullying texts, thereby improving the ability to identify complex sentence patterns and potential insults.
[0121] (3) The data cleaning process removes the noise irrelevant to the detection; at the same time, since the pre-trained OpenAI model itself covers diverse corpora, combined with the robustness of the RNN / LSTM in feature extraction, it can better accommodate different language styles such as abbreviations, slang, and typos.
[0122] (4) By comprehensively modeling the context features and global semantics, the trained cyberbullying detection model can not only make accurate judgments on known forms of bullying texts, but also be more capable of dealing with variants with novel words or similar semantics.
[0123] (5) The multi-layer network structure (RNN / LSTM + OpenAI Embeddings) improves the recall ability for true bullying texts while ensuring accuracy, reduces the missed detection rate, and has higher practical value for actual application scenarios (such as real-time monitoring of social media platforms).
[0124] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.
[0125] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A cyberbullying detection method based on OpenAI embedding and RNN, characterized in that, include: Collect user post data from social media platforms; A label indicating whether the user post data contains cyberbullying behavior; Cleaning the collected user post data to remove irrelevant information and obtain cleaned data; Based on the contextual relationship, text classification is performed on the cleaned data to obtain classified data; Dividing the classified data into a training set and a test set, and annotating the training set, marking the data containing labels of cyberbullying behavior as a negative class, and marking the data not containing labels of cyberbullying behavior as a positive class; Constructing an initial detection model; the initial detection model includes an RNN network, an LSTM network and an OpenAI embedding model; Inputting the text in the training set into the RNN network, and using the OpenAI embedding model to convert the text into a vector representation, and inputting the vector representation into the LSTM network for training to obtain a trained model, and using the test set to test the trained model to obtain a tested cyberbullying detection model; The post data to be tested is input into the cyberbullying detection model to obtain the detection result.
2. The cyberbullying detection method based on OpenAI embedding and RNN according to claim 1, characterized in that The user post data includes user's questions and answers.
3. The cyberbullying detection method based on OpenAI embedding and RNN according to claim 1, characterized in that The collected user post data is cleaned to remove irrelevant information to obtain cleaned data, including: Use regular expressions to match HTML tags, non-alphabetic, non-numeric characters, and URLs in the user post data, and remove the HTML tags, non-alphabetic, non-numeric characters, and URLs; Replace consecutive repeated characters in the removed text with a single character; Convert all uppercase characters in the replaced text to lowercase characters; Redundant spaces in the converted text are deleted to obtain the cleaned data.
4. The cyberbullying detection method based on OpenAI embedding and RNN according to claim 1, characterized in that Based on the contextual relationship, the cleaned data is subjected to text classification to obtain classified data, including: Inputting the cleaned data into a text context feature extraction layer to extract context feature information; Inputting the cleaned data into a global feature extraction layer to extract global feature information; Fusing the contextual feature information and the global feature information to obtain a fused text feature; The fused text features are input into a trained text classification model to obtain the classification data.
5. The cyberbullying detection method based on OpenAI embedding and RNN according to claim 4, characterized in that Inputting the cleaned data into the text context feature extraction layer to extract context feature information includes: Extracting initial feature information of the cleaned data using a pre-trained language model; Inputting the initial feature information into a forward gated recurrent unit and a reverse gated recurrent unit; The outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual relationship feature information.
6. The method for detecting cyberbullying based on OpenAI embedding and RNN according to claim 5, characterized in that: The outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual feature information, including: Using the formula: The outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual feature information; wherein, H t ={H1,H2,…H l } represents the initial feature information input at time t, represents the output of the positive gated recurrent unit at time t, Represents the output of the reverse gated recurrent unit at time t, GRU represents the gated recurrent unit, and G represents the contextual relationship feature information.
7. The method for detecting cyberbullying based on OpenAI embedding and RNN according to claim 6, characterized in that: Inputting the cleaned data into the global feature extraction layer to extract global feature information includes: The initial feature information of the cleaned data is sequentially input into the convolution layer and the pooling layer to obtain the global feature information; wherein the global feature information extraction formula is: c i =f(ω·H+b) In the formula, f is the activation function, ω is the convolution kernel, h is the convolution kernel size, b is the bias value, c is i is the feature vector extracted by the i-th convolutional layer, It represents the value of the features extracted by three different convolution kernels after the maximum pooling layer, and C represents the extracted global feature information.
8. The method for cyberbullying detection based on OpenAI embedding and RNN according to claim 4, characterized in that: The text classification model includes a fully connected layer and an output layer connected in sequence; the loss function of the text classification model is: Among them, p i is the predicted probability of the positive class, p i =σ(z i ), where σ(·) represents the Sigmoid function; y i ∈{0,1} indicates that the sample is positive or negative; γ is the focusing coefficient in FocalLoss; m is the Margin hyperparameter; w0, w1 are the weight coefficients of the positive and negative classes, respectively, which are used to balance the class imbalance; yi is the true label, yi=1 represents the positive class, yi=0 represents the negative class; N is the total number of training samples, and zi represents the result of the i-th sample output through the fully connected layer.