Abnormal language detection method and system based on social media short text
By applying an abnormal language detection model of two-way LSTM and self-attention mechanism in social media, the problem of difficult to identify harmful speech in the prior art is solved, and more efficient text semantic understanding and abnormal language detection are achieved.
Patent Information
- Application Number
- CN202510281883.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively identify and detect harmful speech in social media, especially in the absence of strong management measures, where manual testing is not sufficient to combat toxic speech.
Using an abnormal language detection model based on bidirectional LSTM and self-attention mechanism, a BLAM model is constructed to perform abnormal language detection by extracting the context information of words in the dataset and generating different connection weights.
This method can more accurately capture the deep semantic features of the text, effectively handle long-distance dependencies, and improve the accuracy and recognition ability of abnormal language detection.
Smart Images

Figure CN120216637A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of social media and deep learning, and particularly relates to an abnormal language detection method and system based on short texts of social media. Background Art
[0002] With the progress of society, people's lives have been significantly improved. The rise of the Internet has not only met people's needs, but also provided countless platforms for people's daily life, social interaction and entertainment. In order to enhance user engagement and promote closer interaction between applications and users, these platforms have added comment and communication functions. Therefore, every Internet user can freely express their opinions. Essentially, everyone is a media. An individual's voice can reach a wide audience through Internet platforms, ensuring that no one's voice is ignored.
[0003] Therefore, the identification and detection of harmful remarks play a crucial role in maintaining social harmony, creating a positive online environment, ensuring a pleasant user experience, and safeguarding the well-being of individuals. However, in the absence of strong management measures, manual detection alone is not sufficient to effectively combat toxic remarks. Therefore, there is an urgent need to explore the use of deep learning methods for abnormal language detection. Summary of the Invention
[0004] The present invention provides an abnormal language detection method based on short texts of social media, and the method includes:
[0005] Step S1, collecting comment data to generate a data set;
[0006] Step S2, constructing an abnormal language detection model based on a bidirectional LSTM and a self-attention mechanism;
[0007] Step S3, inputting the data set into the abnormal language detection model to extract feature information, and classifying the data set based on the feature information to obtain an abnormal language detection result.
[0008] Optionally, in the step S2, the content of constructing an abnormal language detection model based on a bidirectional LSTM and a self-attention mechanism specifically includes:
[0009] Using a bidirectional LSTM to extract the context information of words in the data set and obtain semantic meanings;
[0010] Based on the semantic meanings, using a self-attention mechanism to generate different connection weights.
[0011] Optionally, the calculation method of the bidirectional LSTM includes:
[0012] f t =σ(W f1 ht-1 +W f2 X t +b f )
[0013] i t = σ(W i1 h t-1 +W i2 X t +b i )
[0014] O t = σ(W o1 h t-1 +W o2 X t +b o )
[0015] where f t , i t , o t represent the forget gate, input gate, output gate respectively, and σ(·) is the Logistic function with an output range of (0, 1). W f1 , W i1 , W o1 , W f2 , W i2 , W o2 are the weight matrices of the forget gate, and X t is the input at the current moment, and h t-1 is the external state at the previous moment.
[0016] Optionally, the output of the abnormal language detection model is:
[0017] c t = f t · c t-1 + i t · c t
[0018] o t = σ(W o1 C t-1 + W o2 f t + X t )
[0019] where c t is used to update the memory cell state: adding the results of the forget gate and the input gate to update the memory cell state, and o t is to calculate the output gate, and calculate the value of the output gate according to the current input and the memory cell state.
[0020] The present invention also discloses an abnormal language detection system based on short social media texts, and the system includes:
[0021] A data acquisition module, configured to acquire comment data and generate a data set;
[0022] A model construction module, configured to construct an abnormal language detection model based on a bidirectional LSTM and a self-attention mechanism;
[0023] An abnormal detection module, configured to input the data set into the abnormal language detection model to extract feature information, and classify the data set based on the feature information to obtain an abnormal language detection result.
[0024] Optionally, in the model construction module, the content of constructing an abnormal language detection model based on a bidirectional LSTM and a self-attention mechanism specifically includes:
[0025] Using a bidirectional LSTM to extract the context information of words in the data set and obtain semantic meanings;
[0026] Based on the semantic meanings, using a self-attention mechanism to generate different connection weights.
[0027] Optionally, the calculation method of the bidirectional LSTM includes:
[0028] f t = σ(W f1 h t-1 + W f2 X t + b f )
[0029] i t = σ(W i1 h t-1 + W i2 X t + b i )
[0030] O t = σ(W o1 h t-1 + W o2 X t + b o )
[0031] where f t , i t , o t represent the forget gate, input gate, output gate respectively, and σ(·) is a Logistic function with an output range of (0, 1), and W f1 , W i1 , W o1 , W f2 , W i2 , W o2 are the weight matrices of the forget gate, and X tInput at the current moment, h t-1 is the external state at the previous moment.
[0032] Optionally, the output of the abnormal language detection model is:
[0033] c t = f t · c t-1 + i t · c t
[0034] o t = σ(W o1 C t-1 + W o2 f t + X t )
[0035] where c t is used to update the memory cell state: adding the results of the forget gate and the input gate to update the memory cell state, and o t is to calculate the output gate, and calculate the value of the output gate according to the current input and the memory cell state.
[0036] Compared with the prior art, the beneficial effects of the present invention are:
[0037] 1. This method uses publicly available data from part of Wikipedia, organizes and processes it to construct a new dataset, and the classification prediction results are closer to the actual situation. Compared with traditional machine learning methods, such as logistic regression and XGBoost, although they have a solid mathematical foundation and relatively simple algorithms, they have limitations in text feature extraction. They often require manual design of feature engineering and cannot automatically and effectively capture the deep semantic information in the text. Especially when dealing with short texts, it is easy to ignore the context information, resulting in a decrease in the recognition accuracy. The BLAM model can convert the text into a vector representation rich in semantic information by introducing the TEDA model for word vector embedding. The TEDA model can learn the complex relationships between words in the text and consider the context information, so as to more accurately capture the deep semantic features of the text. In addition, the BLAM model also combines a bidirectional LSTM network, which can effectively extract the long-term dependencies of the text and further enhance the model's understanding and recognition ability of text features.
[0038] 2. It has the ability to handle long - distance dependencies. Although traditional LSTM networks can handle the long - term dependencies of text, in practical applications, due to problems such as information transmission ability and gradient vanishing, they often can only establish short - distance dependencies, resulting in insufficient ability to capture long - distance dependencies. The BLAM model effectively solves this problem by introducing the self - attention mechanism. The self - attention mechanism can calculate the attention weights for each word in the text and assign different information importance according to the weights, thus better capturing the long - distance dependencies in the text. Regardless of the distance between words, the self - attention mechanism can effectively establish the association between them, so as to more accurately identify abnormal language in the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] To more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments are briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0040] Figure 1 It is a method step diagram of the abnormal language detection method based on short social media texts according to the embodiment of the present invention;
[0041] Figure 2 It is a distribution diagram of classification target values of the abnormal language detection method based on short social media texts according to the embodiment of the present invention;
[0042] Figure 3 It is an experimental result loss diagram of the abnormal language detection method based on short social media texts according to the embodiment of the present invention;
[0043] Figure 4 It is an experimental result AUC diagram of the abnormal language detection method based on short social media texts according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0045] To make the above - mentioned objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the drawings and specific embodiments.
[0046] Embodiment 1
[0047] An abnormal language detection method based on short texts of social media, such as Figure 1 shown, the method includes:
[0048] Step S1, collect comment data and generate a data set.
[0049] The data mainly comes from the comment data publicly available on Wikipedia. The collected data is cleaned, mainly removing redundant data and irrelevant data, solving the problem of data imbalance, and setting the abnormal harmful data in the data above 1 / 10 to avoid data imbalance.
[0050] Further process relevant public data. Since CivilComments public comments have been transferred from its platform to a persistent open archive, researchers will be able to understand and improve the civility of online conversations in the coming years. It expands the data annotation of various harmful dialogue attributes by human raters, where the text of each comment is in the comment text column. Each piece of data in Train (training set) comes with a toxicity label (target). For evaluation purposes, a target value target≥0.5 will be considered abnormal speech (harmful).
[0051] Step S2, build an abnormal language detection model based on a bidirectional LSTM and a self-attention mechanism.
[0052] The detection model of this method is established based on the multi-task detection model TEDA. TEDA is a language pre-training model released by Google. It emphasizes no longer using the traditional unidirectional language model or the shallow splicing method of two unidirectional language models for pre-training, but adopting a bidirectional Transformer structure. By using a deep neural network with a Transformer architecture, it provides a dense vector representation for natural language. TEDA is a multi-task model that utilizes the self-supervised characteristics of large-scale text data to construct self-supervised tasks of next sentence prediction (NSP) and masked language model (MLM). The idea of MLM is to randomly mask some words from the input corpus and then predict the words from the context. However, MLM cannot understand the relationship between sentences. To obtain sentence-level representation and enable the model to judge whether sentence B is the context of sentence A, TEDA uses the NSP task for pre-training to generate a deep bidirectional language representation that can integrate left and right context information.
[0053] The detection model BLAM is constructed by combining bidirectional LSTM and self-attention mechanism to obtain feature information, including: To effectively solve the problem of gradient explosion or disappearance, the Recurrent Neural Network (RNN) constructs the Long Short-Term Memory Network (LSTM) by introducing a threshold mechanism. The LSTM network can be regarded as a special case of the RNN. It only passes the main part of the data to the next layer instead of all the data. The main improvements are in the following two aspects:
[0054] New internal state: h t ∈h d Circular information storage, hidden layer h t ∈RD, LSTM consists of a forget gate, an input gate, and an output gate. The forget gate f t determines how much information from the t-1 moment is discarded. The input gate i t determines how much information will be stored in the t slot. (3) The output gate O t determines how much information needs to be given to h t The calculations of the forget gate, input gate, and output gate are as follows:
[0055] f t =σ(W f1 h t-1 +W f2 X t +b f )
[0056] i t =σ(W i1 h t-1 +W i2 X t +b i )
[0057] O t =σ(W o1 h t-1 +W o2 X t +b o )
[0058] The BLAM model uses a two - layer Bi - LSTM as the shared layer for multi - task learning. On the one hand, Bi - LSTM has been recognized for its entity recognition performance. On the other hand, Bi - LSTM is used to extract the semantic information features of speech, which can better extract the context information of words, thereby obtaining the semantic meaning of speech and enhancing the ability of the network. In Bi - LSTM, the direction of one LSTM is the forward direction of the input sequence, and the direction of the other LSTM is the reverse direction of the input sequence. When extracting features from speech, the two - direction LSTMs do not share states. The state of the forward - sequence - direction LSTM is only transmitted in the forward - sequence direction, and the state of the reverse - sequence - direction LSTM is only transmitted in the reverse direction, but they will be concatenated as the output of the entire Bi - LSTM later.
[0059] Due to the information transmission ability and the vanishing of gradients, the upper - layer Bi - LSTM layer of the model can only establish short - distance dependencies. If we want to establish long - distance dependencies between input words, there are two ways to achieve this. One is to build a deep network, but the computational cost will increase. However, the length of the speech text in this paper is uncertain, and the fully - connected layer cannot handle the variable length of the input sequence. In practical applications, for different input lengths, the size of the connection weights is also different. In this case, the attention mechanism can be used to "dynamically" generate the weights of different connections. Finally, the overall output of the model is as follows:
[0060] c t =f t ·c t-1 +i t ·c t
[0061] o t =σ(W o1 C t-1 +W o2 f t +X t )
[0062] Step S3: Input the dataset into the abnormal language detection model to extract feature information, and classify the dataset based on the feature information to obtain the abnormal language detection result.
[0063] Each piece of data in Train (training set) has a toxicity label (target). For evaluation purposes, a target value target≥0.5 will be considered as abnormal speech (harmful).
[0064] The update direction of the Adam parameters can calculate the exponentially weighted average t of the gradient g, and the exponentially weighted average of the squared gradient g adaptively adjusts the learning rate. The Adam adaptive momentum estimation algorithm can be regarded as a combination of the momentum method and the RMSprop algorithm. It not only uses momentum as the parameter update direction but also adaptively adjusts the learning rate. On the one hand, the Adam algorithm calculates the exponentially weighted average of the squared gradient g2, and on the other hand, it calculates the exponentially weighted average of the gradient g. The calculation formula is as follows:
[0065] M t = β1M t-1 +(1 - β1)gt
[0066] G t = β2G t-1 +(1 - β2)gt☉gt
[0067] According to the training results, compared with the weighted and sampled data, if the initial loss value of the original data is small, the convergence speed is slow. The recall value of the model trained with downsampled data shows a linear increase, while the recall values of the original data and the weighted data are low and fluctuate greatly. The AUC value of the downsampled data performs better than the other two and has a good trend. Then, for the model trained with the downsampled data, it is inferred that if the Epoch value increases, the values of each index will be better.
[0068] According to Figure 3 、 Figure 4 's experimental loss graph and AUC curve graph to analyze the experimental results and avoid data imbalance. In order to explore the performance of the model under different data distributions, corresponding experiments are carried out using the original dataset, the dataset with different weights for positive and negative samples, and the balanced dataset after downsampling. The Batch size is 512, and the Epoch is 8 rounds. From the above training results, compared with the weighted and sampled data, the initial loss value of the original data is small, but the convergence speed is slow. The recall value of the model trained with downsampled data shows a linear increase, while the recall values of the original data and the weighted data are low and fluctuate greatly. The AUC value of the downsampled data performs better than the other two and has a good trend. The final experimental comparison results are shown in Table 1:
[0069] Table 1
[0070] Model Accuracy Precision Recall Auc F1 LR 0.75 0.79 0.50 0.85 0.61 XGBoost 0.80 0.77 0.76 0.90 0.73 LSTM 0.81 0.86 0.78 0.92 0.83 OUR 0.89 0.90 0.88 0.98 0.89
[0071] As can be seen from the performance comparison table in step S3, the performance of the BLAM model proposed by this method is superior to other models. Although machine learning methods such as Logistic Regression (LR) and XGBoost have a solid mathematical foundation and relatively simple algorithms, they cannot automatically extract text features and have certain limitations. Therefore, the experimental results cannot reach the best. The LSTM network not only has the memory ability to receive its own information but also has the memory ability of historical information and has long-term dependence. Therefore, it is more suitable for short text tasks than the convolutional network. Although LSTM can theoretically establish long-distance dependencies, in practice, it can only establish short-distance dependencies due to the information transmission ability and the disappearance of gradients. To better establish long-distance dependencies between input sequences and obtain long-distance information interaction, BLAM establishes long-distance dependencies by adding a self-attention mechanism. The self-attention mechanism calculates the attention of each word and all words, so the maximum path length is only 1, capturing long-distance dependencies regardless of the distance between them.
[0072] Some users take advantage of the cross-timeliness, transparency and other characteristics of the Internet platform to make abnormal remarks, which seriously endanger the network environment and social security. Based on this, a new attention-based BLAM model based on the long short-term memory network combined with the TEDA model is proposed by this method, and the reliability of the model is proved by experiments.
[0073] Example Two
[0074] An abnormal language detection system based on short texts of social media, the system includes:
[0075] A data collection module, used to collect comment data and generate a data set.
[0076] The data mainly comes from the comment data publicly available on Wikipedia. The collected data is cleaned, mainly removing redundant data and irrelevant data, solving the problem of data imbalance, and setting the abnormal data in the data above 1 / 10 to avoid data imbalance.
[0077] Further process relevant public data. Since the CivilComments public comments have been transferred from their platform to a persistent open archive, researchers will be able to understand and improve the civility of online conversations in the coming years. The data annotation of various harmful conversation attributes by human raters is extended, where the text of each comment is in the comment text column. Each piece of data in Train (training set) has a toxicity label (target). For evaluation purposes, a target value target≥0.5 will be considered an abnormal remark (harmful). The target values of the final experiment are as Figure 2 shown.
[0078] The model construction module is used to construct an abnormal language detection model based on bidirectional LSTM and self-attention mechanism.
[0079] The detection model of this method is established based on the multi-task detection model TEDA. TEDA is a language pre-training model released by Google. It emphasizes that instead of using the traditional unidirectional language model or the shallow splicing method of two unidirectional language models for pre-training, it adopts a bidirectional Transformer structure. By using a deep neural network with a Transformer architecture, it provides a dense vector representation for natural language. TEDA is a multi-task model that utilizes the self-supervised characteristics of large-scale text data to construct self-supervised tasks of next sentence prediction (NSP) and masked language model (MLM). The idea of MLM is to randomly mask some words from the input corpus and then predict the words from the context. However, MLM cannot understand the relationship between sentences. In order to obtain sentence-level representation and enable the model to judge whether sentence B is the context of sentence A, TEDA uses the NSP task for pre-training to generate a deep bidirectional language representation that can integrate left and right context information.
[0080] By combining bidirectional LSTM and self-attention mechanism, the detection model BLAM is constructed to obtain feature information, including: To effectively solve the problem of gradient explosion or disappearance, the recurrent neural network (RNN) constructs the long short-term memory network (LSTM) by introducing a threshold mechanism. The LSTM network can be regarded as a special case of the RNN. It only passes the main part of the data to the next layer instead of all the data. The main improvements are in the following two aspects:
[0081] New internal state: h t ∈h d Recurrent information storage, hidden layer h t ∈RD, LSTM consists of a forget gate, an input gate, and an output gate. The forget gate f t determines how much information from the t-1 moment is discarded. The input gate i t determines how much information will be stored in the t slot. (3) The output gate O t determines how much information needs to be given to h t The calculations of the forget gate, input gate, and output gate are as follows:
[0082] f t =σ(W f1 h t-1 +W f2 X t +b f )
[0083] i t =σ(W i1h t-1 +W i2 X t +b i )
[0084] O t =σ(W o1 h t-1 +W o2 X t +b o )
[0085] The BLAM model uses two layers of Bi-LSTM as the shared layer for multi-task learning. On the one hand, Bi-LSTM has been recognized for its entity recognition performance; on the other hand, Bi-LSTM is used to extract the semantic information features of speech, which can better extract the context information of words, thereby obtaining the semantic meaning of speech and enhancing the ability of the network. In Bi-LSTM, the direction of one LSTM is the forward direction of the input sequence, and the direction of the other LSTM is the reverse direction of the input sequence. When extracting features from speech, the LSTMs in the two directions do not share states. The state of the LSTM in the forward sequence direction is only transmitted in the forward sequence direction, and the state of the LSTM in the reverse sequence direction is only transmitted in the reverse direction, but they will be concatenated as the output of the entire Bi-LSTM at the same time.
[0086] Due to the information transmission ability and the disappearance of gradients, the upper Bi-LSTM layer of the model can only establish short-distance dependencies. If you want to establish long-distance dependencies between input words, there are two ways to achieve this. One is to build a deep network, but the computational cost will increase. However, the length of the speech text in this article is uncertain, and the fully connected layer cannot handle the variable length of the input sequence. In practical applications, for different input lengths, the size of the connection weights is also different. In this case, the attention mechanism can be used to "dynamically" generate the weights of different connections. Finally, the overall output of the model is as follows:
[0087] c t =f t ·c t-1 +i t ·c t
[0088] o t =σ(W o1 C t-1 +W o2 f t +X t )
[0089] Anomaly detection module, used to input the dataset into the anomaly language detection model to extract feature information, classify the dataset based on the feature information, and obtain the anomaly language detection result.
[0090] Each piece of data in the Train (training set) is associated with a toxicity label (target). For evaluation purposes, a target value target ≥ 0.5 is considered an abnormal speech (harmful).
[0091] The update direction of the Adam parameters can calculate the exponentially weighted average t of the gradient g, and the exponentially weighted average of the squared gradient g adaptively adjusts the learning rate. The Adam adaptive momentum estimation algorithm can be regarded as a combination of the momentum method and the RMSprop algorithm. It not only uses momentum as the parameter update direction but also adaptively adjusts the learning rate. On the one hand, the Adam algorithm calculates the exponentially weighted average of the squared gradient g 2 and on the other hand, calculates the exponentially weighted average of the gradient g. The calculation formula is as follows:
[0092] M t = β1M t-1 +(1 - β1)gt
[0093] G t = β2G t-1 +(1 - β2)gt ⊙ gt
[0094] According to the training results, compared with the weighted and sampled data, if the initial loss value of the original data is smaller, the convergence speed is slower. The recall value of the model trained with downsampled data shows a linear increase, while the recall values of the original data and the weighted data are lower and fluctuate more. The AUC value of the downsampled data performs better than the other two and has a good trend. It is inferred that for the model trained with the downsampled data, if the Epoch value increases, the values of each index will be better.
[0095] According to Figure 3 、 Figure 4 's experimental loss graph and AUC curve graph to analyze the experimental results and avoid data imbalance. To explore the performance capabilities of the model under different data distributions, corresponding experiments are conducted using the original dataset, datasets with different weights for positive and negative samples, and the balanced dataset after downsampling. The Batch size is 512 and the Epoch is 8 rounds. From the above training results, compared with the weighted and sampled data, the initial loss value of the original data is smaller, but the convergence speed is slower. The recall value of the model trained with downsampled data shows a linear increase, while the recall values of the original data and the weighted data are lower and fluctuate more. The AUC value of the downsampled data performs better than the other two and has a good trend. The final experimental comparison results are shown in Table 2:
[0096] Table 2
[0097] Model Accuracy Precision Recall Auc F1 LR 0.75 0.79 0.50 0.85 0.61 XGBoost 0.80 0.77 0.76 0.90 0.73 LSTM 0.81 0.86 0.78 0.92 0.83 OUR 0.89 0.90 0.88 0.98 0.89
[0098] As can be seen from the performance comparison table in step S3, the performance of the BLAM model proposed by this method is better than that of other models. Machine learning methods such as Logistic Regression (LR) and XGBoost have solid mathematical foundations and relatively simple algorithms, but they cannot automatically extract text features and have certain limitations. Therefore, the experimental results cannot reach the best. The LSTM network not only has the memory ability to receive its own information but also has the memory ability of historical information and has long-term dependence. Therefore, it is more suitable for short text tasks than the convolutional network. Although LSTM can theoretically establish long-distance dependencies, due to the information transmission ability and the disappearance of gradients, it can only establish short-distance dependencies in practice. To better establish the long-distance dependencies between input sequences and obtain long-distance information interaction, BLAM establishes long-distance dependencies by adding a self-attention mechanism. The self-attention mechanism calculates the attention of each word to all words, so the maximum path length is only 1, capturing long-distance dependencies regardless of the distance between them.
[0099] Some users take advantage of the characteristics of the Internet platform such as cross-timeliness and transparency to make abnormal remarks, seriously endangering the network environment and social security. Based on this, a new attention-based BLAM model based on the long short-term memory network combined with the TEDA model is proposed by this method, and the reliability of the model is proved through experiments.
[0100] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for detecting abnormal language based on short texts in social media, characterized in that: The method comprises: Step S1, collect comment data and generate a data set; Step S2: construct an abnormal language detection model based on bidirectional LSTM and self-attention mechanism; Step S3: input the data set into the abnormal language detection model to extract feature information, classify the data set based on the feature information, and obtain an abnormal language detection result.
2. The abnormal language detection method based on social media short text according to claim 1 is characterized in that: In step S2, the content of constructing an abnormal language detection model based on the bidirectional LSTM and self-attention mechanism specifically includes: Use bidirectional LSTM to extract contextual information of words in the dataset and obtain semantic meaning; Based on the semantic meaning, a self-attention mechanism is used to generate different connection weights.
3. The abnormal language detection method based on social media short text according to claim 2 is characterized in that: The calculation method of the bidirectional LSTM includes: f t =σ(W f1 h t-1 +W f2 X t +b f ) i t =σ(W i1 h t-1 +W i2 X t +b i ) O t =σ(W o1 h t-1 +W o2 X t +b o ) Among them, f t 、i t , o t Respectively represent the forget gate, input gate, output gate and σ(·) is the Logistic function with an output interval of (0,1), W f1 , W i1 , W o1 , W f2 , W i2 , W o2 is the weight matrix of the forget gate, X t The current input, h t-1 is the external state at the previous moment.
4. The abnormal language detection method based on social media short text according to claim 2 is characterized in that: The output of the abnormal language detection model is: c t =f t ·c t-1 +i t ·c t o t =σ(W o1 C t-1 +W o2 f t +X t ) Among them, c t It is used to update the state of the memory cell: add the results of the forget gate and the input gate to update the state of the memory cell. t It is to calculate the output gate, and calculate the value of the output gate according to the current input and memory unit state.
5. An abnormal language detection system based on social media short text, the system using the abnormal language detection method according to any one of claims 1 to 4, characterized in that: The system comprises: Data collection module, used to collect comment data and generate data sets; Model building module, used to build an abnormal language detection model based on bidirectional LSTM and self-attention mechanism; The abnormality detection module is used to input the data set into the abnormal language detection model to extract feature information, classify the data set based on the feature information, and obtain an abnormal language detection result.
6. The abnormal language detection system based on social media short text according to claim 5 is characterized in that: In the model construction module, the content of constructing an abnormal language detection model based on bidirectional LSTM and self-attention mechanism specifically includes: Use bidirectional LSTM to extract contextual information of words in the dataset and obtain semantic meaning; Based on the semantic meaning, a self-attention mechanism is used to generate different connection weights.
7. The abnormal language detection system based on social media short text according to claim 5 is characterized in that: The calculation method of the bidirectional LSTM includes: f t =σ(W f1 h t-1 +W f2 X t +b f ) i t =σ(W i1 h t-1 +W i2 X t +b i ) O t =σ(W o1 h t-1 +W o2 X t +b o ) Among them, f t 、i t , o t Respectively represent the forget gate, input gate, output gate and σ(·) is the Logistic function with an output interval of (0,1), W f1 , W i1 , W o1 , W f2 , W i2 , W o2 is the weight matrix of the forget gate, X t The current input, h t-1 is the external state at the previous moment.
8. The abnormal language detection system based on social media short text according to claim 5 is characterized in that: The output of the abnormal language detection model is: c t =f t ·c t-1 +i t ·c t o t =σ(W o1 C t-1 +W o2 f t +X t ) Among them, c t It is used to update the state of the memory cell: add the results of the forget gate and the input gate to update the state of the memory cell. t It is to calculate the output gate, and calculate the value of the output gate according to the current input and memory unit state.