News agitation propaganda intention detection method and system

Through comprehensive text feature extraction, syntactic semantic similarity separation and dual-pooling feature enhancement technical means, the problem of lack of high-level information, feature interweaving and sparse features in the detection of news text agitation propaganda intention in the existing technology is solved, and the accuracy of news agitation propaganda intention detection and model performance are improved.

CN120104802APending Publication Date: 2025-06-06UNIV OF JINAN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510184838.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When dealing with the agitated publicity text, the existing news text agitated publicity intention detection method faces the problems of lack of advanced information acquisition ability, difficulty in distinguishing feature interweaving, and sparse features, resulting in low model learning effect and generalization ability.

Method used

Comprehensive technical means such as text feature extraction, syntactic semantic similarity separation, dual-pooling feature enhancement, etc., are used to construct syntactic occlusion matrix, calculate KL divergence and feature difference norms, to explore high-order syntactic features, separate interleaving features, solve feature sparse problems, and provide high-quality feature vectors to classifiers to realize accurate detection of news and publicity intentions.

Benefits of technology

It improves the accuracy of the detection of news and publicity intentions, enhances the model's ability to capture text syntactic structure information, reduces the difficulty of feature learning, and improves the model's learning effect and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104802A_ABST
    Figure CN120104802A_ABST
Patent Text Reader

Abstract

The invention provides a news agitation propaganda intention detection method and system. The method comprises the following steps: inputting a preprocessed to-be-detected text into a trained agitation propaganda intention prediction model to obtain a prediction result; the model comprises a text feature extraction module, a syntax and semantic similarity separation module, a double-pooling feature enhancement module and a classifier. The text feature extraction module constructs syntactic, semantic and sequence diagrams; the syntax and semantic similarity separation module separates the syntax graph and the semantic graph to obtain a corrected graph; a double-pooling feature enhancement module performs feature enhancement on the corrected syntactic graph, the corrected semantic graph and the sequence graph, and performs fusion to obtain a final vector; and the classifier obtains an agitating propaganda intention prediction result based on the final vector. Through comprehensive text feature extraction, syntactic and semantic similarity separation, double-pooling feature enhancement and the like, high-order syntactic features are mined, interlaced features are separated, the problem of feature sparsity is solved, high-quality feature vectors are provided for a classifier, and accurate detection of news agitation propaganda intentions is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semantic inspection, and in particular to a method and system for detecting news propaganda intentions. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] In the Internet age, every netizen can publish news content on social media platforms such as blogs, Weibo, and WeChat, and spread it quickly with the help of social networks. The news content published by netizens usually has the purpose of agitation and propaganda, and because the content has not been strictly reviewed, it is very easy to contain subjective bias and mislead readers' judgment. If people with ulterior motives publish such agitation and propaganda news and spread it widely, it is easy to incite public sentiment and even disrupt social order.

[0004] Existing news text propaganda intention detection mainly includes source-based classification methods and content-based classification methods. Among them:

[0005] (1) Source-based classification method: By analyzing the publisher, dissemination channel and other attributes of the news text, it is possible to detect whether the news has propaganda intent. This method relies on external data and is suitable for quickly screening news content. However, propaganda news media may publish non-propaganda news to enhance credibility; non-propaganda news media may also publish propaganda news to increase traffic. Therefore, propaganda news cannot be accurately detected based on the source of the news alone.

[0006] (2) Content-based classification methods directly analyze the content of the news text itself to detect whether the news has propaganda intent. This type of method analyzes from multiple dimensions such as syntax, semantics, and context. Although it has high accuracy, it has a long training cycle and high training cost due to its high computational complexity.

[0007] In recent years, content-based news propaganda intent classification methods often use deep learning methods such as graph convolutional neural networks. By constructing a text graph and using the adjacency information of nodes in the graph, global text features are captured. However, this type of method still faces some challenges when dealing with propaganda text. For example, the syntactic representation lacks the ability to obtain non-adjacent but indirectly related high-order information, which makes it difficult for the model to fully capture the syntactic structure information of the text and affects the understanding and judgment of the text semantics; the feature representation distribution of the syntactic graph and the semantic graph is intertwined, which makes it difficult to accurately distinguish syntactic and semantic features during feature extraction, increasing the difficulty of feature learning; the sparse features of the text graph will reduce the effective information available to the model, reducing the learning effect and generalization ability of the model. Summary of the invention

[0008] In order to solve the above problems, the present invention proposes a method and system for detecting news propaganda intentions. Through comprehensive text feature extraction, syntactic and semantic similarity separation, double-pooling feature enhancement, etc., high-order syntactic features are mined, intertwined features are separated, and the feature sparsity problem is solved. High-quality feature vectors are provided for the classifier, thereby realizing accurate detection of news propaganda intentions.

[0009] In order to achieve the above object, the present invention adopts the following technical solution:

[0010] In a first aspect, the present invention provides a method for detecting news propaganda intentions, comprising:

[0011] Obtain the text to be detected and preprocess the text to be detected;

[0012] The pre-processed text to be detected is input into the trained propaganda intention prediction model to obtain the propaganda intention prediction result;

[0013] Among them, the propaganda intention prediction model includes a text feature extraction module, a syntactic and semantic similarity separation module, a dual-pooling feature enhancement module and a classifier; the text feature extraction module is used to extract text features of the input text and construct a text graph, including a syntactic graph, a semantic graph and a sequence graph; the syntactic and semantic similarity separation module is used to perform similarity separation based on the syntactic graph and the semantic graph to obtain a modified syntactic graph and a modified semantic graph; the dual-pooling feature enhancement module is used to perform feature enhancement on the modified syntactic graph, the modified semantic graph and the sequence graph respectively, and fuse the enhanced features to obtain a final feature vector; the classifier is used to obtain the propaganda intention prediction result based on the final feature vector.

[0014] In a second aspect, the present invention provides a news propaganda intention detection system, comprising:

[0015] A data acquisition unit is configured to acquire the text to be detected and pre-process the text to be detected;

[0016] The intention detection unit is configured to input the preprocessed text to be detected into the trained propaganda intention prediction model to obtain the propaganda intention prediction result;

[0017] Among them, the propaganda intention prediction model includes a text feature extraction module, a syntactic and semantic similarity separation module, a dual-pooling feature enhancement module and a classifier; the text feature extraction module is used to extract text features of the input text and construct a text graph, including a syntactic graph, a semantic graph and a sequence graph; the syntactic and semantic similarity separation module is used to perform similarity separation based on the syntactic graph and the semantic graph to obtain a modified syntactic graph and a modified semantic graph; the dual-pooling feature enhancement module is used to perform feature enhancement on the modified syntactic graph, the modified semantic graph and the sequence graph respectively, and fuse the enhanced features to obtain a final feature vector; the classifier is used to obtain the propaganda intention prediction result based on the final feature vector.

[0018] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for detecting news propaganda intentions described in the first aspect.

[0019] In a fourth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for detecting news propaganda intentions described in the first aspect are implemented.

[0020] Compared with the prior art, the present invention has the following beneficial effects:

[0021] (1) In terms of feature extraction, the present invention constructs syntactic graph, semantic graph and sequence graph respectively through syntactic, semantic and sequence feature extraction submodules to comprehensively capture text features, and syntactic feature extraction can mine non-adjacent but indirectly related high-order information; the syntactic and semantic similarity separation module uses KL regularizer and differential regularizer to separate the syntactic and semantic feature distributions, solves the problem of feature interweaving, and makes feature extraction more accurate; the dual-pooling feature enhancement module extracts smooth and significant features through mean pooling and maximum pooling, and then enhances them through feedforward neural network and gating mechanism, expands the node information update range, and solves the problem of sparse text graph features; finally, the classifier makes predictions based on the final feature vector. The modules work together to optimize the model performance from many aspects, effectively improving the accuracy of news propaganda intention detection.

[0022] (2) The present invention constructs a syntactic masking matrix to prune the multi-head attention embedded in a specific syntactic structure, mines non-adjacent but indirectly related high-order syntactic features, enhances the ability of syntactic representation to obtain high-order information, overcomes the problem of traditional syntactic representation that is difficult to obtain high-order information, helps the model comprehensively capture the syntactic structure information of the text, and promotes the understanding and judgment of the text semantics.

[0023] (3) The present invention separates syntactic and semantic similarities by calculating the KL divergence and feature difference norm of the syntactic graph and the semantic graph, thereby solving the problem of interweaving distribution of feature representations of the syntactic graph and the semantic graph, making it possible to more accurately distinguish syntactic and semantic features during the feature extraction process and reducing the difficulty of feature learning.

[0024] (4) The present invention extracts the smooth features and significant features of the text as a feature enhancement layer to enhance the original features, expand the node information update range, solve the problem of sparse text graph features, increase the effective information available to the model, and improve the model learning effect and generalization ability.

[0025] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their description are used to explain the present invention but do not constitute a limitation of the present invention.

[0027] Figure 1 A main flow chart of a method for detecting news propaganda intention provided by an embodiment of the present invention;

[0028] Figure 2 An overall framework diagram of a method for detecting news propaganda intentions provided by an embodiment of the present invention;

[0029] Figure 3 A schematic diagram of extracting syntactic dependency provided by an embodiment of the present invention;

[0030] Figure 4 A schematic diagram of a syntactic feature extraction process provided by an embodiment of the present invention;

[0031] Figure 5 A flow chart of regularization for separation of syntactic and semantic similarity provided by an embodiment of the present invention;

[0032] Figure 6 A flowchart of a dual-pooling feature enhancement operation provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0033] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0034] Terminology explanation:

[0035] 1. Agitation: Agitation is a form of language expression with clear intentions and strategies aimed at influencing the opinions or behaviors of others.

[0036] 2. Distributed interweaving: a feature redundancy problem caused by node overlap during the construction of a text graph.

[0037] 3. Syntactic structure embedding: a method that represents the syntactic structure of a text as a vector representation that is easy to operate.

[0038] 4. Text graph: A graphical structure with words as vertices, used to represent the vocabulary in a text and their relationships.

[0039] 5. Sequence graph: A text graph in which nodes are words and edges represent sequence relationships. It is used to represent the order relationship and dependency structure of word pairs in the text.

[0040] 6. Syntactic graph: A text graph in which nodes are words and edges represent syntactic relationships, used to represent the syntactic dependencies and grammatical structure of word pairs in the text.

[0041] 7. Semantic graph: A text graph whose nodes are words and edges represent semantic relationships, used to represent the contextual information and potential semantic associations of word pairs in the text.

[0042] 8. Stanford NLy: A natural language processing tool that includes a part-of-speech tagger, a named entity recognizer, a parser, a bootstrap pattern learning, and an open information extraction tool. This patent uses it to extract syntactic dependencies between words in a text.

[0043] 9. BERT model: A pre-trained language model based on the Transformer architecture for extracting features from text.

[0044] 10. KL Regularization: A regularization method that constrains the model by minimizing the KL divergence.

[0045] 11. Differential Regularization: A method to prevent overfitting by limiting or regularizing the gradients of model parameters.

[0046] Embodiment 1

[0047] like Figure 1 As shown, this embodiment discloses a method for detecting news propaganda intention, comprising the following steps:

[0048] S1: Obtain the text to be detected and preprocess the text to be detected;

[0049] S2: Input the preprocessed text to be detected into the trained propaganda intention prediction model to obtain the propaganda intention prediction result.

[0050] Next, combine Figure 2 , a news propaganda intention detection method disclosed in this embodiment is described in detail.

[0051] 1. Data acquisition and preprocessing

[0052] First, the news text to be recognized is obtained and preprocessed. The specific operations include: removing non-text data, stop words, case conversion, stem extraction, and removing low-frequency words. Through these data cleaning steps, characters irrelevant to the propaganda recognition task are removed, thus obtaining a neat text.

[0053] 2. Prediction of propaganda intentions

[0054] The preprocessed text to be detected is input into the trained propaganda intention prediction model to perform feature extraction and propaganda intention detection to obtain the propaganda intention prediction result.

[0055] Among them, the propaganda intention prediction model includes a text feature extraction module, a syntactic and semantic similarity separation module, a dual-pooling feature enhancement module and a classifier.

[0056] 1. Text feature extraction module

[0057] The text feature extraction module includes a syntactic feature extraction submodule, a semantic feature extraction submodule and a sequence feature extraction submodule, which are respectively used to extract the syntactic feature vector, the semantic feature vector and the sequence feature vector of the text to be detected;

[0058] Based on the syntactic feature vector, semantic feature vector and sequence feature vector, corresponding syntactic graph, semantic graph and sequence graph are generated respectively.

[0059] (1) Syntactic feature extraction submodule

[0060] The syntactic feature extraction submodule is used to extract syntactic feature vectors. The syntactic dependency between word pairs is extracted as an additional feature and embedded into the multi-head attention mechanism. The syntactic masking matrix is ​​constructed based on the syntactic distance. The syntactic masking matrix is ​​used to adjust the learning range of the attention head. By fusing the features of each attention head, the syntactic representation of nodes that are not directly adjacent but indirectly related is enhanced, and then these comprehensive information is converted into vector form, and finally the syntactic feature vector is obtained.

[0061] First, use Stanford NLP to extract the syntactic dependencies between word pairs, such as Figure 3 As shown in Figure 2, syntactic dependencies are embedded into a multi-head attention mechanism with syntactic masking to extract syntactic features, such as Figure 4 As shown in Figure 2, each syntactic dependency connects two words and is used to describe the specific structural type of the syntax, which is embedded as an additional feature into each head of the multi-head attention.

[0062] Specifically, a multi-head attention mechanism is used to construct an adjacency matrix, and the number of adjacency matrices is the same as the number of attention heads. A syntactic mask is constructed for each attention head according to the syntactic distance, the connection weight of nodes beyond the syntactic distance is reduced, and the attention range of the attention head is controlled to prune the syntactic features. Syntactic masking matrix with syntactic distance k The formula for calculation is:

[0063]

[0064] Among them, L(·) represents the shortest distance between two nodes, i and j are two different nodes, i is the horizontal coordinate of the adjacency matrix, j is the vertical coordinate of the adjacency matrix, k represents the preset syntactic distance, and p represents the number of attention heads.

[0065] Finally, the features learned by each attention head are fused to enhance the syntactic representation of nodes that are not directly adjacent but indirectly related. The matrix after masking is:

[0066]

[0067] Among them, A k Represents the adjacency matrix learned by the k-th head in the multi-head attention.

[0068] (2) Semantic feature extraction submodule

[0069] The semantic feature extraction submodule is used to extract semantic feature vectors. With the help of the calculated semantic weights of word pairs, the semantic associations of each word pair in the corpus are comprehensively considered, the word pair information that meets the similarity threshold conditions is integrated and vectorized, and the semantic weights and other related information are converted into vector representations to obtain semantic feature vectors.

[0070] Specifically, in order to further enrich the representation of the relationship between words in the text, semantic features are extracted through cosine similarity. BERT is used to capture long-distance dependencies and contextual information, and the attention mechanism is used to obtain more accurate vector representations. First, the word embedding of the words in the sentence is extracted from the BERT hidden layer. and As word pairs, and calculate the cosine similarity of the word pairs to obtain semantic similarity. The formula is as follows:

[0071]

[0072] The semantic weight between word pairs is measured by the following formula:

[0073]

[0074] in, Represents a node (V i ,V j) has a semantic relationship in the corpus, Min sem Indicates the minimum value of word pairs appearing in the corpus, Max sem Represents the maximum number of word pairs that occur in the corpus.

[0075] If the similarity value between two words exceeds the threshold ρ, it is determined that the two words have a semantic relationship in the current text, and the word pairs whose cosine similarity is greater than or equal to the threshold ρ are retained.

[0076] By constructing a syntactic masking matrix, this embodiment can prune the multi-head attention embedded in a specific syntactic structure, effectively mine non-adjacent but indirectly related high-order syntactic features in the text, enhance the ability of syntactic representation to obtain high-order information, and enable the model to comprehensively capture the syntactic structure information of the text, providing strong support for in-depth understanding and accurate judgment of text semantics.

[0077] (3) Sequence feature extraction

[0078] The sequence feature extraction submodule is used to extract the sequence feature vector. By integrating and processing the sequence weights between word nodes constructed based on point-by-point mutual information (PMI), comprehensively considering the local co-occurrence relationship between words, and converting the weight information corresponding to each word node into a vector representation, the sequence feature vector is obtained.

[0079] Specifically, a sliding window strategy is used to describe the sequence characteristics of text through the local co-occurrence relationship between words. A fixed-size sliding window window_size is defined on all documents in the corpus to collect co-occurrence statistics. The sequence weight between two word nodes is constructed based on the point-by-point mutual information PMI, that is:

[0080]

[0081] Among them, p(V i ,V j ) indicates a word pair (V i ,V j ) joint probability; p(V i ) and p(V j ) represent word V i and V j The marginal probability of N(V i ,V j ) is a word pair (V i ,V j ) appears in all windows; N window is the total number of sliding windows; It is the word V i Total number of occurrences in all windows.

[0082] (4) Text graph generation

[0083] Based on the extracted syntactic feature vector, the syntactic structure information contained in it is used to generate a syntactic graph to show the syntactic dependency of the text; based on the semantic feature vector, the semantic graph is generated with the semantic association information contained in it to present the semantic relationship of the text; through the sequence feature vector, the sequence graph is generated with the help of the word order relationship information represented by it. The syntactic graph, semantic graph and sequence graph together constitute the text graph.

[0084] Construct a text graph, map the words in the corpus to nodes, and determine the edges by the sequence weight, the syntactic weight obtained by the syntactic masking matrix, and the semantic weight, respectively, to represent the order relationship, syntactic dependency, and semantic association, and obtain a sequence graph, a syntactic graph, and a semantic graph. Define the text graph as G = (V, E), where V is a node set, representing a word or word in the text, and E is an edge set.

[0085] 2. Syntactic and semantic similarity separation module

[0086] In order to solve the problem that the distribution of syntactic and semantic features is intertwined and affects the context, a KL regularizer and a differential regularizer are constructed to enable the model to learn unique information from syntax and semantics. Figure 5 As shown in the figure, by separating syntactic and semantic similarities, the model's ability to distinguish differences in syntactic and semantic distributions during training is enhanced. Specifically, the following steps are included:

[0087] A KL regularizer is constructed to measure the difference between syntactic and semantic probability distributions in order to optimize the structures of the two. The syntactic graph represents the syntactic dependency between words, while the semantic graph represents their semantic associations. The adjacency matrices of the syntactic graph and the semantic graph are constructed respectively, and after normalizing them to probability distributions, the KL divergence of the syntactic adjacency matrix to the semantic adjacency matrix and the KL divergence of the semantic adjacency matrix to the syntactic adjacency matrix are calculated. The formula is:

[0088]

[0089] in, Represents the element of the i-th row of the semantic adjacency matrix; Represents the element of the i-th row of the syntactic adjacency matrix; represents the KL divergence of the syntactic adjacency matrix to the semantic adjacency matrix; represents the KL divergence of the semantic adjacency matrix to the syntactic adjacency matrix; f(·) represents the normalization of the adjacency matrix; R kl represents the KL regularization term.

[0090] To ensure syntactic and semantic learning, the adjacency matrix A sem and A synWithout excessive overlap, a differential regularizer is constructed to calculate the feature difference norm of syntax and semantics, and compare the differences of corresponding nodes or edges in the text graph, so as to achieve the separation of local syntax and semantic similarity. D The differential regularization term is:

[0091]

[0092] Where F represents the Frobenius norm of the matrix.

[0093] In order to effectively optimize the structure of syntax and semantics during model training, so that the adjacency matrix of syntax and semantic learning does not overlap excessively, and to achieve the separation of syntax and semantic similarity, it is necessary to introduce appropriate constraints. The KL regularizer can measure the difference in probability distribution between syntax and semantics, and the differential regularizer can calculate the feature difference norm of syntax and semantics. The two provide constraints from the perspectives of distribution difference and feature difference, respectively. Therefore, the KL regularization term and the differential regularization term are introduced into the loss function. By controlling the parameter update during the model training process, the model can better distinguish the differences in syntax and semantic distribution during the learning process, thereby achieving the separation of syntax and semantic similarity and obtaining the corrected syntax graph and the corrected semantic graph.

[0094] This embodiment effectively separates syntactic and semantic similarities by calculating the KL divergence and feature difference norm of the syntactic graph and the semantic graph, so that the model can more accurately distinguish syntactic and semantic features during feature extraction, reduce the difficulty of feature learning, and improve the accuracy and efficiency of feature extraction.

[0095] 3. Dual pooling feature enhancement module

[0096] Based on the modified syntactic graph, modified semantic graph and sequence graph, syntactic features, semantic features and sequence features are extracted, and a dual-pooling feature enhancement model is constructed. Smooth features and significant features are extracted through mean pooling and maximum pooling. A feedforward neural network is constructed to calculate the difference in importance between the pooled features and the original features, and a gating mechanism is introduced to dynamically control the update of features, so as to extract complex and diverse propaganda skills for news propaganda intention detection.

[0097] As a specific implementation method, Figure 6 As shown in Figure 1, feature pooling is divided into two parts: mean pooling and max pooling. seq , semantic feature A sem and syntactic features A syn Calculate the average pooling A separately mean and max pooling A max , taking semantic features as an example, the formula is as follows:

[0098]

[0099] Smooth features and significant features are introduced into the feedforward neural network, and finally these features are integrated in the form of weight fusion. FFN means that the feedforward neural network performs nonlinear transformation on the input features and calculates the importance score a of the features. c :

[0100]

[0101] in, Represents smooth features or salient features of semantics.

[0102] The importance scores of the features are weighted and summed, and the model fuses the pooled features to generate enhanced feature representations:

[0103]

[0104] in, and They respectively represent the importance scores corresponding to the mean pooling of semantic features and the importance scores corresponding to the maximum pooling of semantic features.

[0105] Finally, in order to dynamically control the fusion effect of features, the model introduces a gating mechanism to control the weight of original features and fusion features by updating the gate f, and updating the feature representation to obtain the semantic features enhanced by the double pooling features:

[0106] f=σ(W·[A sem ;Q]+b) (16)

[0107]

[0108] Among them, σ and b represent parameters.

[0109] Similarly, a similar method is used to perform double-pooling feature enhancement on sequence features, and the importance scores of the features are weighted and summed. The model fuses the pooled features to generate a sequence enhanced feature representation.

[0110] f=σ(W·[A seq ;Q]+b) (18)

[0111] Update the feature representation to obtain the sequence features after double pooling feature enhancement:

[0112]

[0113] Similarly, a similar method is used to perform double-pooling feature enhancement on syntactic features, and the importance scores of the features are weighted and summed. The model fuses the pooled features to generate syntactically enhanced feature representations.

[0114] f=σ(W·[A syn ;Q]+b) (20)

[0115] Update the feature representation to obtain the syntactic features enhanced by the double pooling feature:

[0116]

[0117] The enhanced semantic features, sequence features, and syntactic features are obtained and fused based on the pooling layer to obtain the final feature vector.

[0118] This embodiment effectively expands the node information update range by extracting the smooth features and significant features of the text as the feature enhancement layer, increases the effective information available to the model, enables the model to learn richer text features, and alleviates the problem of sparse text graph features.

[0119] 4. Classifier

[0120] Based on the final feature vector, the classifier makes the final classification judgment to determine whether the text is agitational or not, and presents the text result to the user in a visual way.

[0121] In the method for identifying news propaganda intentions proposed above in this embodiment, a graph convolutional neural network based on feature enhancement is constructed, in which feature extraction, syntactic and semantic similarity separation, and text feature enhancement are performed in sequence. By extracting the sequence features, semantic features, and syntactic features of the text, and constructing a text feature graph based on the extracted features to obtain a feature graph vector, the constructed graph is analyzed to enhance the representation of word dependency relationships, deeply explore the propaganda features hidden in the text, and improve the accuracy of propaganda intention detection.

[0122] Embodiment 2

[0123] This embodiment provides a news propaganda intention detection system, including:

[0124] A data acquisition unit is configured to acquire the text to be detected and pre-process the text to be detected;

[0125] The intention detection unit is configured to input the preprocessed text to be detected into the trained propaganda intention prediction model to obtain the propaganda intention prediction result;

[0126] Among them, the propaganda intention prediction model includes a text feature extraction module, a syntactic and semantic similarity separation module, a dual-pooling feature enhancement module and a classifier; the text feature extraction module is used to extract text features of the input text and construct a text graph, including a syntactic graph, a semantic graph and a sequence graph; the syntactic and semantic similarity separation module is used to perform similarity separation based on the syntactic graph and the semantic graph to obtain a modified syntactic graph and a modified semantic graph; the dual-pooling feature enhancement module is used to perform feature enhancement on the modified syntactic graph, the modified semantic graph and the sequence graph respectively, and fuse the enhanced features to obtain a final feature vector; the classifier is used to obtain the propaganda intention prediction result based on the final feature vector.

[0127] Embodiment 3

[0128] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in the method for detecting news propaganda intention as described in the first embodiment are implemented.

[0129] Embodiment 4

[0130] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the method for detecting news propaganda intention as described in the first embodiment are implemented.

[0131] The steps or modules involved in the above embodiments 2 to 4 correspond to those in embodiment 1. For the specific implementation, please refer to the relevant description of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0132] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for detecting news propaganda intentions, characterized in that: include: Obtain the text to be detected and preprocess the text to be detected; The pre-processed text to be detected is input into the trained propaganda intention prediction model to obtain the propaganda intention prediction result; Among them, the propaganda intention prediction model includes a text feature extraction module, a syntactic and semantic similarity separation module, a dual-pooling feature enhancement module and a classifier; the text feature extraction module is used to extract text features of the input text and construct a text graph, including a syntactic graph, a semantic graph and a sequence graph; the syntactic and semantic similarity separation module is used to perform similarity separation based on the syntactic graph and the semantic graph to obtain a modified syntactic graph and a modified semantic graph; the dual-pooling feature enhancement module is used to perform feature enhancement on the modified syntactic graph, the modified semantic graph and the sequence graph respectively, and fuse the enhanced features to obtain a final feature vector; the classifier is used to obtain the propaganda intention prediction result based on the final feature vector.

2. A method for detecting news propaganda intention as claimed in claim 1, characterized in that: The text feature extraction module is used to extract text features of the input text and construct a text graph, including a syntactic graph, a semantic graph and a sequence graph; specifically, it includes: The text feature extraction module includes a syntactic feature extraction submodule, a semantic feature extraction submodule and a sequence feature extraction submodule, which are respectively used to extract the syntactic feature vector, the semantic feature vector and the sequence feature vector of the text to be detected; Based on the syntactic feature vector, semantic feature vector and sequence feature vector, corresponding syntactic graph, semantic graph and sequence graph are generated respectively.

3. A method for detecting news propaganda intention as claimed in claim 2, characterized in that: The syntactic feature extraction submodule is used to extract syntactic feature vectors, specifically: Extract the syntactic dependencies between word pairs in the text to be detected and embed them into each head of the multi-head attention as additional features; Use a multi-head attention mechanism to construct an adjacency matrix to obtain node connection relationships; the number of the adjacency matrix is ​​the same as the number of attention heads; Construct a syntactic mask for each attention head according to the preset syntactic distance, and control the focus range of the attention head by reducing the connection weight of nodes beyond the syntactic distance; The features learned by each attention head are integrated to enhance the syntactic representation of nodes that are not directly adjacent but indirectly related, and a syntactic feature vector is obtained.

4. A method for detecting news propaganda intention as claimed in claim 2, characterized in that: The semantic feature extraction submodule is used to extract semantic feature vectors, specifically: Extract word embeddings of word pairs in the text sentence to be detected; Calculate the cosine similarity of word pairs based on their word embeddings, and retain the word embeddings of word pairs whose similarity exceeds a preset threshold as semantic feature vectors; The sequence feature extraction submodule is used to extract sequence feature vectors, specifically: A sliding window is set for the text to be detected. According to the co-occurrence times of each pair of words in the window and the point-by-point mutual information, the weights between word nodes are constructed to mine the sequential dependencies in the text and obtain the sequence feature vector.

5. A method for detecting news propaganda intention as claimed in claim 1, characterized in that: The syntactic and semantic similarity separation module is used to perform similarity separation based on the syntactic graph and the semantic graph to obtain a modified syntactic graph and a modified semantic graph; Construct the adjacency matrix of the syntactic graph and the semantic graph and normalize them into probability distribution; Input the KL regularizer to calculate the divergence of the syntactic adjacency matrix to the semantic adjacency matrix and the KL divergence of the semantic adjacency matrix to the syntactic adjacency matrix; Input differential regularizer, calculate the syntactic and semantic feature difference norms based on the syntactic adjacency matrix and the semantic adjacency matrix; By comparing the differences between corresponding nodes or edges in the adjacency matrix, the local syntactic and semantic similarity can be separated; The KL regularization term and the differential regularization term are introduced into the loss function. By controlling the parameter update during the model training process, the syntactic and semantic similarity separation is achieved, and the corrected syntactic graph and the corrected semantic graph are obtained.

6. A method for detecting news propaganda intention as claimed in claim 1, characterized in that: The dual-pooling feature enhancement module is used to perform feature enhancement on the modified syntactic graph, the modified semantic graph and the sequence graph respectively, and fuse the enhanced features to obtain the final feature vector; Specifically include: Extracting the revised text features of the revised syntactic graph, revised semantic graph and sequence graph, including syntactic features, semantic features and sequence features; The corrected text features are input into the dual-pooling feature enhancement module to extract smooth features and salient features through mean pooling and maximum pooling respectively; The double-pooled features are input into the feed-forward neural network and gating mechanism for enhancement; The enhanced features are fused through the pooling layer to obtain the final feature vector.

7. A method for detecting news propaganda intention as claimed in claim 6, characterized in that: The double-pooled features are input into the feedforward neural network and the gating mechanism for enhancement, specifically including: The double-pooled features are input into the feedforward neural network for nonlinear transformation and calculation of the importance scores of the features. Perform weighted summation on the importance scores of the features to generate preliminary enhanced features; The preliminary enhanced features are input into the gating mechanism, and the preliminary enhanced features are updated to obtain the enhanced features by controlling the weight of the revised text features and the preliminary enhanced features.

8. A news propaganda intention detection system, characterized in that: include: A data acquisition unit is configured to acquire the text to be detected and pre-process the text to be detected; The intention detection unit is configured to input the preprocessed text to be detected into the trained propaganda intention prediction model to obtain the propaganda intention prediction result; Among them, the propaganda intention prediction model includes a text feature extraction module, a syntactic and semantic similarity separation module, a dual-pooling feature enhancement module and a classifier; the text feature extraction module is used to extract text features of the input text and construct a text graph, including a syntactic graph, a semantic graph and a sequence graph; the syntactic and semantic similarity separation module is used to perform similarity separation based on the syntactic graph and the semantic graph to obtain a modified syntactic graph and a modified semantic graph; the dual-pooling feature enhancement module is used to perform feature enhancement on the modified syntactic graph, the modified semantic graph and the sequence graph respectively, and fuse the enhanced features to obtain a final feature vector; the classifier is used to obtain the propaganda intention prediction result based on the final feature vector.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of a method for detecting news propaganda intentions as described in any one of claims 1 to 7 are implemented.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method for detecting news propaganda intentions as described in any one of claims 1 to 7 are implemented.