A method for detecting aggressive speech with fusion of post attributes

By constructing a heterogeneous graph to fuse post attributes, extracting text features using ERNIE and BiGRU, and capturing the relationship features between post attributes through a heterogeneous graph attention network, the problem of insufficient detection accuracy and generalization ability in existing technologies is solved, and more efficient offensive speech detection is achieved.

CN117056514BActive Publication Date: 2026-03-17ANHUI UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311024914.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-12
Publication Date
2026-03-17
Estimated Expiration
2043-08-12

AI Technical Summary

Technical Problem

Existing methods for detecting offensive speech struggle to cover all forms and variations of offensive speech. Machine learning-based methods have poor generalization ability, while deep learning-based methods, although highly accurate, fail to fully utilize the relationship information between post attributes.

Method used

By constructing a heterogeneous graph, fusing post attributes, extracting text features using ERNIE and BiGRU, capturing the relationship features between post attributes through a heterogeneous graph attention network, and combining it with Softmax classification, the detection of whether post text contains offensive content is achieved.

Benefits of technology

It improves the accuracy and generalization of offensive speech detection, enabling better understanding and analysis of post information in social media data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117056514B_ABST
    Figure CN117056514B_ABST
Patent Text Reader

Abstract

This invention discloses a method for detecting offensive speech by fusing post attributes, belonging to the text processing technology of the natural language processing field. The method includes the following steps: S1: Preprocessing the dataset; S2: Processing the processed posts using ERNIE and BiGRU to obtain the text features of the posts; S3: Constructing a heterogeneous graph using post attributes and embedding nodes into the heterogeneous graph; S4: Updating nodes through a two-layer attention mechanism of a heterogeneous graph attention network, with the last updated node containing relational representations of post attributes; S5: Concatenating the text features of the posts with the relational features containing post attributes and inputting them into a Softmax layer to obtain the offensive speech detection result. This invention, by employing a heterogeneous graph attention network, obtains relational features containing post attributes, which, after fusing with the original post text features, yield a semantically richer feature vector, making detection more accurate and providing strong technical support for purifying the online environment and creating a harmonious online atmosphere.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing, specifically a method for detecting offensive speech that integrates post attributes. Background Technology

[0002] With the widespread use of the internet, social media has become a crucial channel for people to communicate, obtain information, and express opinions. While people can freely express their views, this freedom also makes offensive language more likely. Offensive language can influence public judgment of facts, incite violence, and exacerbate opposing viewpoints. Therefore, detecting offensive language on social media is crucial. From a personal perspective, it benefits user experience and mental health; from a business perspective, it helps website administrators control posts, protect platform interests, and improve platform reputation; from a social perspective, it helps relevant departments monitor offensive language and create a harmonious online environment.

[0003] Existing methods for detecting offensive speech can be broadly categorized into three types: rule-based methods, machine learning-based methods, and deep learning-based methods. Rule-based methods identify specific keywords, phrases, or sentence structures to determine the presence of offensive intent, but they struggle to cover various forms and variations of offensive speech. Machine learning-based methods train models using machine learning algorithms to automatically detect offensive speech, typically requiring large amounts of labeled data for training. These models may overfit the training data, resulting in poor generalization ability on unseen data. Deep learning-based methods can learn feature representations directly from raw text data through end-to-end learning, eliminating the need for manual feature design and reducing the workload of feature engineering. Compared to the higher classification accuracy of rule-based and machine learning methods, current deep learning methods rely on the semantic information of post context for offensive speech detection.

[0004] The difference in the offensive speech detection method proposed in this invention lies in that, while considering the contextual text features of the text itself, it also incorporates the relationship between post attributes (post, topic, author), and captures more complex structural information and richer semantic information through heterogeneous graphs, making the detection more accurate. Summary of the Invention

[0005] The purpose of this invention is to provide a method for detecting offensive speech by integrating post attributes. This method takes post attributes into account through a heterogeneous graph, obtains the relationship features between post attributes through a heterogeneous graph attention network, and then concatenates them with the text features of the post itself before inputting them into a classification layer to determine whether the post text contains offensive content.

[0006] To achieve its objectives, the present invention employs the following technical solution:

[0007] An offensive speech detection method integrating post attributes, comprising the following steps:

[0008] Step 1: First, perform preprocessing operations on the text data, such as word segmentation, removing irrelevant characters, emojis, etc.;

[0009] Step 2: After processing the posts in the processed data through ERNIE and BiGRU, obtain the text features of the posts;

[0010] Step 3: Use the post attributes to construct a heterogeneous graph and perform node embedding on the heterogeneous graph;

[0011] Step 4: After two layers of attention mechanisms of the heterogeneous graph attention network, obtain a relationship representation containing post attributes;

[0012] Step 5: Concatenate the text features of the post and the relationship features containing post attributes and send them into Softmax to obtain the detection result.

[0013] Among them, in the said Step 1, the specific steps of the text preprocessing are: <00​​​​​​​​​​​​​​​​​​​​​​​​Step 2.3: Compressing the obtained post text vectors using both average pooling and max pooling can reduce information loss. Average pooling preserves the overall trend and average information in the features to some extent. Max pooling can retain more obvious features, which helps improve the model's invariance and robustness.

[0021] Step 3, which involves obtaining the heterogeneous graph and embedding its nodes, includes the following steps:

[0022] Step 3.1: Construct a heterogeneous graph of post attributes based on meta-paths. The heterogeneous graph contains three types of nodes: post, topic, and author. Define two meta-paths: post-author-post and post-topic-post.

[0023] Step 3.2: Vectorize the nodes in the heterogeneous graph. Since the length of post text is generally long, the ERNIE model is used to embed the post into vectors, and then the resulting matrix vector is flattened into a one-dimensional vector through a fully connected layer.

[0024] Step 3.3: Author and topic nodes in heterogeneous graphs are generally only a few characters long, so Word2Vec is used directly for word embedding, and then the word vectors are transformed into one-dimensional vectors through a fully connected layer.

[0025] Step 3.4: Based on the metapath post-author-post, obtain the neighboring nodes of the post node. At this point, all the post neighboring nodes come from the same author. Based on the metapath post-topic-post, obtain the neighboring nodes of the post node. At this point, all the post neighboring nodes have the same topic.

[0026] Step 4, which involves obtaining a relational representation containing post attributes, includes the following steps:

[0027] Step 4.1: First, map each node to the same dimension using a transformation matrix.

[0028] Step 4.2: Learn the attention value of each post node through node-level attention, obtain the weight coefficient of each feature vector through the Softmax function, and update the node vector.

[0029] Step 4.3: Learn the attention value for each meta-path through semantic attention, obtain the importance of each meta-path through the normalization function, and update the feature vector of the node again using the obtained weights.

[0030] The vector fusion steps in step 5 are as follows:

[0031] Step 5.1: Concatenate the post nodes in the three-way relationship representation obtained through the heterogeneous graph attention mechanism with the text sentence representation obtained after BiGRU and fully connected layers to obtain features that have both post text semantics and post relationships.

[0032] Step 5.2: Input the concatenated vector into the Softmax function to obtain the detection result.

[0033] The offensive speech detection method that integrates post attributes provided by this invention has the following advantages:

[0034] (1) A heterogeneous graph-based method for mining post attribute relationships was constructed, taking into account factors other than the post itself. First, a heterogeneous graph was constructed using post attributes, and appropriate meta-paths were defined. Then, based on the meta-paths, isomorphic subgraphs under the meta-paths were found, and the relationships between nodes and their neighbors were learned through the node attention of the heterogeneous graph attention network. Finally, the weights based on each meta-path were learned through semantic attention, thereby learning a post content relationship representation that includes the relationships between the three.

[0035] (2) Text semantic features are obtained through ERNIE and BiGRU, and the post attribute relationship representation is obtained by using a heterogeneous graph attention network. The two are concatenated to fuse semantic features and post attribute relationship representation, resulting in a richer and more comprehensive feature vector, which can better understand and analyze post information in social media data. Attached Figure Description

[0036] Figure 1 Flowchart of a method for detecting offensive speech that integrates post attributes;

[0037] Figure 2 A schematic diagram illustrating the process of obtaining text features from a post;

[0038] Figure 3 A schematic diagram illustrating the transformation process from text data to heterogeneous graph data;

[0039] Figure 4 A schematic diagram illustrating the process of obtaining post attribute relationship features based on heterogeneous graph attention networks. Detailed Implementation

[0040] The present invention will be further explained and illustrated below through specific embodiments.

[0041] Example 1: This invention provides a method for detecting offensive speech by integrating post attributes, such as... Figure 1 As shown. The specific steps are as follows:

[0042] S1: First, preprocess the text data, including word segmentation, removal of irrelevant characters and emojis, etc.

[0043] S1.1: Word segmentation is performed using the jieba library to split the text into a sequence of words.

[0044] S1.2: Stop word filtering is carried out to filter out stop words such as "de", "shi", "le", etc. to reduce the size of the feature space.

[0045] S1.3: Special symbols are removed, such as punctuation marks, special symbols, web link tags, etc.

[0046] S2: Obtain the text features of the post. The following is a detailed description in combination with Figure 2 as follows:

[0047] S2.1: The preprocessed post text X = {x1, x2,..., x n}, where x i represents the i-th symbol or word in the text. It is input into the ERNIE model to map the text to a high-dimensional vector space to obtain the vector embedding of the text, resulting in the input vector sequence V = {v1, v2,..., v n}, where v i represents the vector representation of x i . Capture the semantic information of the text, learn the potential semantic dependencies to make it more generalizable, and enhance the learning effect of the model.

[0048] S2.2: Input the text encoded by ERNIE into the BiGRU model. Through a series of GRU units for information transfer and feature extraction, the feature vector H = {h1, h2,..., h n} of the text is obtained. The formula is as follows:

[0049] f t = σ(W f [h t-1 , x t )

[0050] [[ID=!44]]r t = σ(W r [h t-1 , x t )

[0051] h t ′ = tanh(W[r[[ID=!58]] t h t-1 , x t )

[0052] h t = (1 - f t ) × h t-1 + f t × h t ′ It should be noted that there seems to be a formatting issue in the original text where some of the tags are not in a standard form (e.g., the repeated use of "!" in the translation of formulas). If possible, it would be beneficial to correct the original text for a more accurate translation. Also, the specific meaning of the model formulas and operations might require more context and knowledge of the relevant fields for a more comprehensive understanding.

[0053] Parameter description: f t and r t These represent updating the door and resetting the door, respectively; h t ' represents the vector representation of the hidden layer at time t; h t-1 This represents the output of the hidden layer at time t-1; x t W represents the word vector input at time t; W represents the weight matrix; W f and W r It is the weight matrix for the update gate and the reset gate.

[0054] S2.3: Compression of the obtained post text vectors using both average pooling and max pooling can reduce information loss. Average pooling preserves the overall trend and average information in the features to some extent. Max pooling can retain more obvious features, which helps improve the model's invariance and robustness.

[0055] S3: Obtain the heterogeneous graph and its node embeddings. The following section combines... Figure 3 A detailed explanation is provided below:

[0056] S3.1: Construct a heterogeneous graph about post attributes based on meta-paths. The heterogeneous graph has three types of nodes: post, topic, and author. Define two meta-paths: post-author-post and post-topic-post.

[0057] S3.2: Vectorize the nodes in the heterogeneous graph. Since the length of post text is generally long, the ERNIE model is used to embed the post into vectors, and then the resulting matrix vector is flattened into a one-dimensional vector through a fully connected layer.

[0058] S3.3: Author and topic nodes in heterogeneous graphs are generally only a few characters long, so Word2Vec is used directly for word embedding, and then the word vectors are transformed into one-dimensional vectors through a fully connected layer.

[0059] S3.4: Based on the metapath post-author-post, obtain the neighboring nodes of the post node. At this time, all the neighboring nodes of the post node come from the same author. Based on the metapath post-topic-post, obtain the neighboring nodes of the post node. At this time, all the neighboring nodes of the post node have the same topic.

[0060] S4: Obtain the relation representation containing post attributes, as shown below. Figure 4 A detailed explanation is provided below:

[0061] S4.1: First, each node is mapped to the same dimension using a transformation matrix, calculated as follows:

[0062]

[0063] Where, ξ i Let ξ be the post node vector before transformation. i ′ represents the converted post node vector.

[0064] S4.2: Through node-level attention, the weight of each post node is learned, calculated as follows:

[0065]

[0066] Where σ is the Sigmoid activation function, || is the concatenation function, and att node It is node-level attention, a Φ It is a node-level attention vector.

[0067] S4.2.1: The weight coefficients of each feature vector are obtained through the Softmax function, as shown in the following formula:

[0068]

[0069] in, It is a matrix vector used to adjust the shape.

[0070] S4.2.2: Using the attention values ​​as weights, perform a weighted summation calculation to obtain the vector embedding of post node i under this metapath after one round of message passing:

[0071]

[0072] Where, ξ i ' is the transformed post node vector, These are the weight coefficients of the eigenvectors.

[0073] S4.2.3: After applying node-level attention to all post nodes, we can obtain the post node embedding representations for specific semantics under these two meta-paths as follows: This represents information about posts published by the same user. Represents posts that fall under the same topic.

[0074] S4.3: The importance of each meta-path is learned through semantic-level attention, calculated as follows:

[0075]

[0076] in, q is the weight of each meta-path, and V represents all nodes. q is the semantic-level attention vector, and W is the weight matrix. It is the metapath Φ PThe vector embedding of post i is given below, where b is the bias vector, and it is shared across all metapaths.

[0077] S4.3.1: After obtaining the importance of each meta-path, the above results are Softmax normalized for semantic-level attention and final root node embedding calculation. The calculation formula is as follows:

[0078]

[0079]

[0080] in, It is a semantic-level attention value. It is the weight of each meta-path. It is the embedding of post nodes under the metapath, where Z represents the set of node feature vectors Z = {z1, z2, ..., z} after passing through the HAN network layer once. i}

[0081] S4.4: A two-layer attention mechanism using a Heterogeneous Graph Attention Network (HAN) is used to obtain post nodes containing post attribute relationships. Node-level attention learns the importance of a node to its neighbors based on a specific meta-path, while semantic-level attention learns the weight information of each meta-path. Different isomorphic subgraphs are generated through different meta-paths, and message passing and information aggregation are performed in these subgraphs. Finally, the weighted aggregation of the vector attention from the subgraphs under each meta-path is performed for back-end propagation. Through this two-layer attention learning, the final node can capture the complex structural and rich semantic information of the heterogeneous graph.

[0082] S5: Fusion of text feature vectors and post attribute relationship vectors.

[0083] S5.1: The post nodes in the three-way relationship representation obtained through the heterogeneous graph attention mechanism are concatenated with the text sentence representation obtained after BiGRU and fully connected layers to obtain features that have both post text semantics and post relationships. The calculation formula is as follows:

[0084] T = concat(T) text T rep )

[0085] Where concat() is the concatenation function, T text It is a textual feature of the post, T rep It is a characteristic of the relationship between posts.

[0086] S5.2: Input the concatenated vector into the Softmax function to obtain the classification result. The calculation formula is as follows:

[0087] Q = softmax(u t T+d t )

[0088] Among them, u t It is the weight, d t It is the bias, and T is the concatenated vector.

[0089] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and not restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

[0090] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A method for detecting aggressive speech with fusion of post attributes, characterized in that Comprise the following steps: Step 1: first, the text data preprocessing operation, word segmentation, remove punctuation, space, emoticons; Step 2: the post text after processing data through ERNIE and BiGRU in the text features of post text are obtained; Step 3: the use of post text attribute construction of heterogeneous graph, and carry out heterogeneous graph node embedding; Step 4: after two layers of attention mechanism of heterogeneous graph attention network update node, get the relationship representation containing post text attribute; Step 5: the text features of post text and the relationship features containing post text attribute are spliced into the Softmax layer to obtain the detection result; Among them, step 4 includes: Step 4.1: through the node level attention, learn the importance of each neighbor node to the node, and update the feature vector of the node; Step 4.1.1: the attention value of each post text is calculated as follows: where σ is a sigmoid activation function, || is a concatenation function, is the node-level attention, is the node-level attention vector; Step 4.1.2: the weight coefficient of each feature vector is obtained by the softmax function, and the formula is as follows: wherein is a matrix vector used to adjust the shape; Step 4.1.3: the attention value is used as the weight for weighted summation calculation, and the vector embedding of post text node i under the meta path is obtained after one round of message passing under the meta path: wherein, is the converted post node vector, is the weight coefficient of the feature vector; Step 4.1.4: After all the post nodes go through the node-level attention, we can get the post node embedding representation about the specific semantics under the two meta-paths as , represent the post information published by the same user, represent the posts containing the same topic; Step 4.2: learn the importance of each meta path through semantic attention, and update the feature vector of the node again; Step 4.2.1: learn the attention value of each meta path, calculated as follows: wherein, is the weight of each meta-path, V represents all nodes, q is the attention vector of semantic level, and W is the weight matrix, is the meta-path the vector embedding of the post i, b is the bias vector, and is shared for all meta-paths; Step 4.2.2: after obtaining the importance of each meta path, the above result is normalized by Softmax as the semantic level attention and the final root node embedding calculation, the calculation formula is as follows: wherein, is a semantic-level attention value, is a weight of each meta-path, is a post node embedding under the meta-path, and Z represents a node feature vector set after one round of transmission through the HAN network layer ; Step 3 includes: Step 3.1: based on the meta path, the heterogeneous graph about the post text attribute is constructed, and there are three types of nodes in the heterogeneous graph, namely post text, theme and author, and two meta paths are defined: post text-author-post text, post text-theme-post text; Step 3.2: the nodes in the heterogeneous graph are vectorized, because the length of the post text is longer, the ERNIE model is used for vector embedding of the post text, and then the obtained matrix vector is flattened into a one-dimensional vector through the full connection layer; the author and theme nodes in the heterogeneous graph have only a few characters, so the word2vec is used for word embedding, and then the word vector is changed into a one-dimensional vector through the full connection layer; Step 3.3: according to the meta path post text-author-post text, the neighbor nodes of the post text node are obtained, at this time the post text neighbor nodes are from the same author; according to the meta path post text-theme-post text, the neighbor nodes of the post text node are obtained, at this time the post text neighbor nodes have the same theme.

2. The method of claim 1, wherein the detecting of the aggressive speech in the fusion post attribute comprises: Step 1 includes: Step 1.1: word segmentation, use jieba library to cut the text into a sequence of words; Step 1.2: stop word filtering, filter out "of", "is", "of" stop words to reduce the size of the feature space; Step 1.3: remove special symbols, marks, remove punctuation, emoticons, web link tags.

3. The method of claim 1, wherein the detecting of the aggressive speech in the fusion post attribute comprises: determining a sentiment of the fusion post attribute; and determining a sentiment of a comment of the fusion post attribute. Step 2 includes: Step 2.1: the preprocessed post text is processed through ERNIE and BiGRU model to extract the context features of the text; Step 2.2: By compressing the obtained post text vector in two ways of average pooling and maximum pooling, the loss of information can be reduced. The average pooling retains the overall trend and average information in the features to some extent, and the maximum pooling can retain the more obvious features, which helps to improve the invariance and robustness of the model.

4. The method of claim 1, wherein the detecting of the aggressive speech in the fusion post attribute comprises: Step 5 includes: Step 5.1: The post node in the relationship representation obtained by the heterogeneous graph attention mechanism is spliced with the text sentence representation obtained after the BiGRU and fully connected layer, to obtain features with both post text semantics and post relationship, and the calculation formula is as follows: wherein concat() is a concatenation function, is a text feature of the post, is a relationship feature of the post. Step 5.2: The spliced vector is input into the Softmax function to obtain the result, and the calculation formula is as follows: wherein, is a weight, is a bias, concat is a vector concatenation function, and T is the concatenated vector.

Citation Information

Patent Citations

  • Rumor detection method and system based on dynamic heterogeneous graph and multi-level attention

    CN115659966A

  • Social network rumor detection method based on emotion perception and graph convolutional network

    CN116431760A