Social Media Cognitive Threat Detection Method and System

By collecting and preprocessing text data on sensitive topics on social media platforms, using multi-level cognitive threat detection methods and knowledge graph technology, the problems of high concealment and difficulty in traceability on social media are solved, and rapid detection and effective blockade of the spread of cognitive threats are achieved.

CN116244446BActive Publication Date: 2025-06-20Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211732859.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2025-06-20
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

Cognitive threats on social media platforms are highly concealed, difficult to trace, and difficult to regulate cross-platform. It is difficult for existing technologies to effectively identify, trace and combat cognitive threats.

Method used

By collecting and preprocessing sensitive topic text data from social media platforms, multi-level cognitive threat detection methods, including sentiment analysis, named entity recognition and entity relationship extraction, we build a cognitive threat dissemination knowledge map to realize detection, traceability and confrontation of cognitive threats.

Benefits of technology

It realizes rapid detection and traceability of cognitive threats, improves detection efficiency and information volume, effectively blocks the spread of cognitive threats, and purifies cyberspace.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116244446B_ABST
    Figure CN116244446B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of network security technology, and particularly relates to a method and system for detecting cognitive threats in social media, which collects sensitive topic text data on a network platform and performs preprocessing operations on the data; for the preprocessed sensitive topic text data, multi-level cognitive threat detection is carried out to obtain cognitive threat topic text; the named entity recognition and entity relationship extraction of the cognitive threat topic text are used to construct a cognitive threat propagation knowledge graph; based on the cognitive threat propagation knowledge graph, user tracing, event tracing, and organization tracing are carried out for the propagation of the cognitive threat topic text. The present invention uses the underlying sentiment tendency to identify cognitive threats for topic text related to specific topics and sensitive events. Compared with traditional manual evidence collection, the identification cycle is greatly shortened, the detection information volume and efficiency are improved, it has good security and feasibility, a high accuracy rate for threat determination, good detection effects, and broad application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security, and in particular relates to a social media cognitive threat detection method and system. Background Art

[0002] Cognition refers to the process of people acquiring knowledge or applying knowledge, or the process of information processing, which includes sensation, perception, memory, thinking, imagination and language. Cognitive threats are based on the purposeful, inflammatory, concealed, directional and untrue information input to individuals. Through the continuous influence and solidification of the individual's cognitive process, the individual's distorted, unconventional, reverse and negative cognition is formed, or the individual's normal cognitive system is changed to deviate from the core value system of society. The emergence of self-media platforms and social platforms has provided a breeding ground for the breeding and spread of cognitive threats. Cyberspace has become the main battlefield for the confrontation of cognitive threats. Social networks have the characteristics of anonymous identity, "freedom" of speech, high real-time and fast dissemination. Most of the users are young people who are not sensitive to social issues and are easily penetrated by cognition. However, the problems of high concealment of cognitive threats, difficulty in tracing the source and difficulty in cross-platform supervision need to be solved urgently. It is urgent to identify, trace and confront cognitive threat information from a technical point of view. In view of the problems that cognitive threats are highly concealed, difficult to trace, and difficult to supervise across platforms, how to curb cognitive threats technically has become an urgent need to purify cyberspace. Summary of the invention

[0003] To this end, the present invention provides a social media cognitive threat detection method and system, which targets topic texts related to specific themes and sensitive events and uses the emotional tendencies behind them to identify cognitive threats. Compared with traditional manual evidence, it greatly shortens the identification cycle and improves the amount of detection information and efficiency.

[0004] According to the design scheme provided by the present invention, a social media cognitive threat detection method is provided, which includes the following contents:

[0005] Collect text data on sensitive topics on online platforms and pre-process the data;

[0006] For the pre-processed sensitive topic text data, cognitive threat topic text is obtained through multi-level cognitive threat detection, wherein the multi-level cognitive threat detection includes: primary detection of dividing the sensitive topic text data into cognitive threat topic text and initial suspected cognitive threat topic text, intermediate detection of classifying the initial suspected cognitive threat topic text into cognitive threat topic text, suspected cognitive threat topic text and non-cognitive threat topic text, and ultimate detection of obtaining cognitive threat topic text from suspected cognitive threat topic text through manual annotation;

[0007] Construct a cognitive threat propagation knowledge graph by performing named entity recognition and entity relationship extraction on cognitive threat topic texts;

[0008] Based on the cognitive threat propagation knowledge graph, trace the sources of users, events, and organizations for the propagation of cognitive threat topic texts.

[0009] As the social media cognitive threat detection method in the present invention, further, collect sensitive topic text data from network platforms and perform preprocessing operations on the data, including:

[0010] First, distribute and collect sensitive topic text information and related user data from network platforms according to the user authorization information library;

[0011] Then, for the collected text information, merge the title and the body, use the redundancy detection algorithm to remove redundant information, deduplicate the relevant comments, clean and transform the text noise data, and perform word segmentation on the text using a word segmentation system.

[0012] As the social media cognitive threat detection method in the present invention, further, in the primary detection, use the sentiment analysis method to divide the sensitive topic text data into cognitive threat topic texts and initial suspected cognitive threat topic texts. Among them, the process of dividing by the sentiment analysis method includes:

[0013] First, construct a basic sentiment dictionary based on the known sentiment dictionary and using the word frequency statistics method, and expand the sentiment dictionary by performing correlation statistics between the words in the text data and the words in the basic sentiment dictionary;

[0014] Next, take the text in the sensitive topic text data as the unit and the sentiment words as the delimiters, perform sentiment weight statistics on the sentences between each delimiter, and judge the sentiment polarity of the text according to the proportion of the negative sentiment weight in all sentiment word weights;

[0015] Then, divide the sensitive topic text data into cognitive threat topic texts and initial suspected cognitive threat topic texts according to the sentiment polarity of the text.

[0016] As the social media cognitive threat detection method in the present invention, further, construct a basic sentiment dictionary based on the known sentiment dictionary and using the word frequency statistics method, including:

[0017] First, select a series of sentiment words from the known sentiment dictionary, sort the sentiment words according to the search engine click-through rate in the series of sentiment words, and select several sentiment words according to the click-through rate popularity;

[0018] Next, select the sentiment words with the highest relevance to the theme based on word frequency statistics, and jointly form a basic sentiment dictionary using the selected several sentiment words and sentiment words;

[0019] Then, the basic sentiment dictionary is expanded using synonyms and candidate words with sentiment tendencies.

[0020] As the social media cognitive threat detection method in the present invention, further, sentiment weight statistics are performed on the sentences between each delimiter, and the sentiment polarity of the text is judged based on the proportion of the negative sentiment weight in all sentiment word weights, including:

[0021] First, for the sentences between delimiters, sentiment tendencies are statistically analyzed through sentiment word analysis, negation word analysis, adverb analysis, fixed collocation word analysis, transition word analysis, and exclamation sentence analysis respectively;

[0022] Then, the sum of the negative sentiment tendency values of all clauses included in the text and the sum of the absolute values of the overall sentiment weights are statistically analyzed, and the sentiment polarity of the text is judged using the proportion of the negative sentiment word weight in all sentiment word weights in the text.

[0023] As the social media cognitive threat detection method in the present invention, further, in the intermediate detection, the initial suspected cognitive threat topic text is classified into cognitive threat topic text, suspected cognitive threat topic text, and non-cognitive threat topic text using deep learning methods, and the classification process includes:

[0024] A deep learning model is constructed and pre-trained using a training data set with labeled tags. Among them, the deep learning model includes a BERT model for word vector representation of the input and a BiLSTM model for cognitive threat detection of the input word vectors;

[0025] The initial suspected cognitive threat topic text is input into the pre-trained deep learning model, and the deep learning model is used to obtain the cognitive threat probability value, and the cognitive threat topic text, suspected cognitive threat topic text, and non-cognitive threat topic text in the initial suspected cognitive threat topic text are determined through the cognitive threat probability value.

[0026] As the social media cognitive threat detection method in the present invention, further, for the preprocessed sensitive topic text data, cognitive threat topic text is obtained through multi-level cognitive threat detection, and it also includes: using the sentiment analysis method to evaluate the cognitive threat influence degree based on the overall sentiment tendency in the comment area of the cognitive threat topic text.

[0027] As the social media cognitive threat detection method in the present invention, further, a cognitive threat propagation knowledge graph is constructed through named entity recognition and entity relationship extraction of the cognitive threat topic text, including:

[0028] Construct a named entity extraction model and optimize the model using the adversarial training method. Among them, the named entity recognition model includes an encoder for mapping input characters to the real number space and mining potential semantics, a BiLSTM neural network layer for extracting context semantic information by capturing forward and backward bidirectional features in the vector transformed by the encoder, and a CRF conditional random field layer for taking the bidirectional features extracted by the BiLSTM neural network layer as input and generating character corresponding labels in combination with the Bioes annotation paradigm.

[0029] Use the cognitive threat topic text as the input of the optimized named entity extraction model, and use the named entity extraction model to identify the entity categories and relationships in the cognitive threat topic text.

[0030] As the social media cognitive threat detection method of the present invention, further, by constructing a cognitive threat propagation knowledge graph through named entity recognition and entity relationship extraction of the cognitive threat topic text, two named entity extraction models with pipeline connections are built. Among them, the first named entity extraction model uses a single-label multi-classification task method to identify entities in the cognitive threat topic text, and the second named entity extraction model uses a multi-label multi-classification task method to take the input of the first named entity extraction model as input to identify the relationships between entities.

[0031] Further, the present invention also provides a social media cognitive threat detection system, including: a data collection server, multiple cognitive threat identification servers, a knowledge graph server, and a web server. Among them,

[0032] The data collection server is used to collect sensitive topic text data on the network platform and perform preprocessing operations on the data;

[0033] Multiple cognitive threat identification servers are used to obtain cognitive threat topic text through multi-level cognitive threat detection for the preprocessed sensitive topic text data. Among them, the multiple cognitive threat identification servers specifically include: a primary identification server for dividing the sensitive topic text data into cognitive threat topic text and initial suspected cognitive threat topic text, a secondary identification server for classifying the initial suspected cognitive threat topic text into cognitive threat topic text, suspected cognitive threat topic text, and non-cognitive threat topic text, and an ultimate identification server for obtaining cognitive threat topic text from the suspected cognitive threat topic text through manual annotation;

[0034] The knowledge graph server is used to construct a cognitive threat propagation knowledge graph through named entity recognition and entity relationship extraction of the cognitive threat topic text;

[0035] The web server is used to perform user traceability, event traceability, and organization traceability on the spread of cognitive threat topic text based on the cognitive threat propagation knowledge graph and using the web interface.

[0036] Advantages of the present invention:

[0037] The present invention can rely on online social platforms such as Weibo, Zhihu, and WeChat official accounts to perform multi-dimensional sentiment analysis on the crawled sensitive topic texts and their comments to detect cognitive threat topic texts. By identifying cognitive threat-related named entities and extracting the relationships between entities, a cognitive threat propagation knowledge graph is constructed. The knowledge graph is used to realize the visualization display of the traceability of cognitive threat propagation users, the traceability of cognitive threat propagation events, and the traceability of cognitive threat propagation organizations. Through the discovery of implicit relationships, the prediction of cognitive threat propagation is realized. Key accounts, groups, organizations, and users are monitored in real time, providing cognitive countermeasure analysis, blocking the spread of cognitive threats, deepening the supervision of network cognitive threats, deterring network illegal acts related to cognitive threats, and effectively purifying the network space. Brief Description of the Drawings

[0038] Figure 1 It is a schematic diagram of the social media cognitive threat detection process in the embodiment;

[0039] Figure 2 It is a schematic diagram of the cognitive threat detection and measurement process in the embodiment;

[0040] Figure 3 It is a schematic diagram of the adversarial training process in the embodiment;

[0041] Figure 4 It is a schematic diagram of the hierarchical structure for visualizing the construction of the knowledge graph in the embodiment. Detailed Embodiment

[0042] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and technical solutions.

[0043] The development of the network environment makes the implementation of cognitive domain threats easier and more feasible, and can be implemented alone or jointly in multiple dimensions and at multiple levels, thus affecting the entire social value form. In the embodiments of this case, refer to Figure 1 as shown, a social media cognitive threat detection method is provided, including:

[0044] S101. Collect sensitive topic text data from the network platform and perform preprocessing operations on the data;

[0045] S102. For the pre - processed sensitive topic text data, obtain the cognitive threat topic text through multi - level cognitive threat detection, where the multi - level cognitive threat detection includes: primary detection of dividing the sensitive topic text data into cognitive threat topic text and initial suspected cognitive threat topic text, intermediate detection of classifying the initial suspected cognitive threat topic text into cognitive threat topic text, suspected cognitive threat topic text, and non - cognitive threat topic text, and ultimate detection of obtaining the cognitive threat topic text from the suspected cognitive threat topic text through manual annotation;

[0046] S103. Construct a cognitive threat propagation knowledge graph through named entity recognition and entity relationship extraction of the cognitive threat topic text;

[0047] S104. Based on the cognitive threat propagation knowledge graph, conduct user traceability, event traceability, and organization traceability for the propagation of the cognitive threat topic text.

[0048] Relying on online social platforms such as Weibo, Zhihu, and WeChat official accounts, conduct multi - dimensional sentiment analysis on the crawled sensitive topic text and its comments to achieve the detection of cognitive threat topic text, construct a cognitive threat propagation knowledge graph by identifying cognitive threat - related named entities and extracting relationships between entities, and use the knowledge graph to achieve user traceability of cognitive threat propagation. By constructing a cognitive threat knowledge graph, predict the cognitive threat propagation path and provide cognitive countermeasure analysis. It can not only achieve accurate identification in the context of short texts on social media platforms but also maintain a high accuracy rate when identifying long texts in the media, making news media responsible for their words and deeds and effectively deterring some unscrupulous media.

[0049] As a preferred embodiment, further, collect sensitive topic text data on the network platform and perform pre - processing operations on the data, including:

[0050] First, distributedly collect sensitive topic text information and related user data on the network platform according to the user authorization information library;

[0051] Then, for the collected text information, merge the title and the text body, use the redundancy detection algorithm to remove redundant information, perform duplicate removal on relevant comments, clean and transform text noise data, and perform word segmentation on the text using a word segmentation system.

[0052] A large amount of sensitive topic text data on social platforms such as Weibo and official accounts can be obtained through the API interface, which includes ten fields such as article titles, text bodies, and comments. For convenient further processing, perform data pre - processing operations of merging the title and the text body, removing redundant information through the redundancy detection algorithm, duplicate removal of comments, cleaning and transformation of noise data, and word segmentation through ICTCLAS. The redundancy detection algorithm for removing redundant information can be designed to include the following steps:

[0053] Step 1: Divide the text into sentences according to punctuation marks;

[0054] Step 2: Get the first 5 sentences of the article after sentence segmentation. If there are words such as "Follow us" or "Click ** font", delete the sentence and keep the rest.

[0055] Step 3: Get the first 10 sentences of the article after sentence segmentation. If there are any sentences containing "Edit", "Preliminary Review", "Click to Read", etc., delete the sentence and keep the rest.

[0056] Step4: Re-merge the retained sentences into text.

[0057] It should be noted that the data collected in this case can be sensitive topic text data from various online social platforms. It can be collected through the Weibo platform, and it will also be randomly collected through social platforms such as Zhihu and WeChat public accounts. The main process of Weibo data collection may include: user authorization, acquisition of newly released Weibo, Weibo information update, and user information acquisition. User authorization is completed through Oauth2, and the acquisition of newly released Weibo, Weibo information update, and user information acquisition are completed by automatically calling the official public API interface of Weibo.

[0058] In the primary detection, sentiment analysis methods can be used to divide sensitive topic text data into cognitive threat topic text and initial suspected cognitive threat topic text. The sentiment analysis method division process includes:

[0059] Firstly, a basic sentiment dictionary is constructed based on known sentiment dictionaries and word frequency statistics are used. Then, the sentiment dictionary is expanded by counting the correlation between words in text data and words in the basic sentiment dictionary.

[0060] Next, we use the text in the sensitive topic text data as the unit and the sentiment words as the delimiters to calculate the sentiment weights of the sentences between each delimiter, and judge the sentiment polarity of the text based on the proportion of negative sentiment weights in all sentiment word weights.

[0061] Then, the sensitive topic text data is divided into cognitive threat topic text and initial suspected cognitive threat topic text according to the sentiment polarity of the text.

[0062] Multi-dimensional sentiment analysis is a process of analyzing and processing the sentiment polarity, sentiment degree, and sentiment category of subjective texts with emotional colors by using natural language processing and text mining techniques. An important research direction in the field of NLP is sentiment analysis. Correct and effective sentiment analysis can quickly obtain the positive or negative emotions expressed by people from texts, help discover the underlying sentiment tendencies behind the texts, and further isolate the potential political threats and cognitive threats with cultural infiltration nature hidden in the vast amount of information. Sentiment analysis tasks can be classified into discourse-level, sentence-level, word or phrase-level according to the analysis granularity; they can be classified into text-based sentiment analysis and review-based sentiment analysis according to the types of texts processed, and can be divided into sub-problems such as sentiment classification, sentiment retrieval, and sentiment extraction according to the types of tasks studied. In the embodiments of this case, as Figure 2 shown, the basic process of cognitive threat recognition and measurement based on dynamic expansion of sentiment dictionaries and deep learning

[0063] For sentiment classification methods, they can generally be divided into sentiment dictionary-based classification methods and deep learning-based classification methods. Each type of method has its own characteristics and deficiencies. The sentiment dictionary-based method refers to using a sentiment dictionary marked with sentiment polarity to perform sentiment polarity quantification calculation on texts. This method uses a series of rules and sentiment dictionaries for classification. First, the words in the sentiment dictionary are matched with the words in the text to be analyzed, and then the sentiment value of the sentence is obtained through calculation. Finally, the obtained sentiment value is used as the judgment basis for the sentiment tendency classification of the sentence. Although the accuracy of this method is relatively high, the cost of constructing the sentiment dictionary is large, and the sentiment dictionary-based method does not consider the connection between words in the text and lacks semantic information; the deep learning-based method regards sentiment classification as a special text classification and uses artificial annotation and machine learning methods to perform sentiment classification on texts. The deep learning-based method uses the marked data and labels, which are all manually marked, and then uses deep learning methods to perform sentiment analysis on texts. Common machine learning methods include Naive Bayes (NB), decision tree, Support Vector Machine (SVM), etc. The quality of the effect of this method mainly depends on the quantity and quality of the manually annotated data, so it is greatly affected by people's subjective consciousness and requires a large amount of labor.

[0064] Aiming at the respective characteristics of the two methods, the two methods of sentiment dictionary-based and deep learning are combined and optimized, and a multi-dimensional sentiment analysis method combining dynamic expansion of sentiment dictionaries and deep learning is proposed, so as to overcome the respective shortcomings of the two methods and achieve a high accuracy rate.

[0065] Among them, a basic sentiment dictionary is constructed based on the known sentiment dictionary and using the word frequency statistics method, including:

[0066] First, select a series of sentiment words from the known sentiment dictionary, sort the sentiment words according to the click-through rate in the search engine in the series of sentiment words, and select several sentiment words according to the click-through rate popularity;

[0067] Next, select the sentiment words with the highest relevance to the theme based on word frequency statistics, and jointly form a basic sentiment dictionary using the selected several sentiment words and sentiment vocabulary;

[0068] Then, expand the basic sentiment dictionary using synonyms and candidate words with sentiment tendencies.

[0069] Furthermore, conduct sentiment weight statistics on the sentences between each delimiter, and judge the sentiment polarity of the text according to the proportion of the negative sentiment weight in the weights of all sentiment words, including:

[0070] First, for the sentences between delimiters, statistically analyze the sentiment tendencies respectively through sentiment word analysis, negative word analysis, adverb analysis, fixed collocation word analysis, transition word analysis, and exclamation sentence analysis;

[0071] Then, statistically analyze the sum of the negative sentiment tendency values of all clauses contained in the text and the sum of the absolute values of the overall sentiment weights, and judge the sentiment polarity of the text using the proportion of the negative sentiment word weights in the weights of all sentiment words in the text.

[0072] A series of sentiment words can be selected from HowNet of CNKI, input them into the search engine one by one, sort the sentiment words according to the size of the click-through rate (hits value) returned by the search engine, select several sentiment words with the highest click-through rate as basic sentiment words. In addition, adopt a method based on word frequency statistics to semi-automatically select basic sentiment vocabulary with higher relevance to the theme to jointly form a basic sentiment dictionary. Since most of the words with sentiment components in the text are adjectives, verbs, and some nouns, after preprocessing, only need to conduct word frequency statistics on the automatic text with sufficient entries, and then for several words with higher word frequencies, select the 20 positive sentiment words with the highest word frequencies and the 20 negative sentiment words with the highest word frequencies, and jointly form a basic sentiment dictionary with the general basic sentiment vocabulary.

[0073] Since the basic sentiment dictionary expresses relatively strong sentiment tendencies, assign a sentiment tendency value of -1 to the negative sentiment words in the basic sentiment dictionary. The vocabulary in the basic sentiment dictionary is small and it is impossible to contain all the words with sentiment tendencies that appear in the text set. Therefore, it is necessary to expand the basic sentiment dictionary to build a relatively complete sentiment dictionary. It can be expanded by adding synonyms and adding candidate words with sentiment tendencies.

[0074] Adding synonyms can help identify sentiment words more broadly and expand the basic sentiment dictionary using existing thesaurus. However, to improve the algorithm performance of sentiment tendency calculation, it is still necessary to manually screen out commonly used synonyms. After expansion, the number of words in the sentiment dictionary increases to 256, and the sentiment tendency value of the synonyms of negative sentiment words can be set to -1.

[0075] It is very difficult to construct a complete and omission-free sentiment dictionary. However, by analyzing the correlation between each word in the text set and the words in the sentiment dictionary and including highly correlated words in the dictionary, a sentiment dictionary with a wider coverage can be effectively constructed.

[0076] The Pointwise Mutual Information method can be used to calculate the correlation between sentiment words in the candidate word dictionary to determine whether to add them to the sentiment dictionary. The Pointwise Mutual Information method calculates the correlation between words based on the mutual information theory. Its basic idea is to statistically calculate the co-occurrence probability of two words wordi and wordj in the text. The greater the co-occurrence probability, the higher the correlation between the two words. The calculation formula is as follows:

[0077]

[0078] Where p(wordi^wordj) is the co-occurrence probability of wordi and wordj in the text, and the calculation method is as follows:

[0079]

[0080] Where n represents the total number of clauses in the text, numSentence(wordi, wordj) represents the number of clauses that contain both wordi and wordj. P(wordi) and P(wordj) respectively represent the proportion of the number of clauses containing wordi and wordj in the text in the total number of clauses. The calculation formula is as follows:

[0081]

[0082]

[0083] Where numSentence(wordi) represents the number of clauses containing wordi in the text. In the above formula, PMI(wordi, wordj) represents the amount of information of the other variable that can be obtained when one of the variables wordi and wordj appears, fully reflecting the statistical correlation between wordi and wordj: when PMI is greater than 0, it means that the two words are correlated, and the greater the PMI value, the stronger the correlation; when PMI is 0, it means that the two words are statistically independent; when PMI is less than 0, it means that the two words are mutually exclusive.

[0084] After segmenting the words using the ICTCLAS system, the part-of-speech property of the words can be obtained. Then, calculate the SO-PMI values of two candidate words restricted by word.propertyal ∈ {a, d, an, ag, al} and word.propertyal ∈ {vn, vd, vi, vg, vl}. Words with other parts of speech are directly regarded as neutral words. This method aims to solve the problem that when adding relevant words, some words without emotional tendency co-occur with positive or negative emotional words with a very high probability, resulting in being wrongly introduced into the emotional dictionary, causing unnecessary overhead in the performance of emotional classification and reducing the accuracy, and improving the efficiency of the extended dictionary algorithm. The specific calculation of the SO-PMI value of two candidate words word is as follows: Calculate the PMI value between the candidate word and the positive basic dictionary, calculate the PMI value between the candidate word and the negative words, and finally subtract the two to obtain the SO-PMI value of the candidate word. The calculation formula is as follows:

[0085] SO-PMI(word) =

[0086] ∑ posWord∈posWords PMI(word, posWord) - ∑ negWord∈negWords PMI(word, negWord)

[0087] The relationship between the SO-PMI value and the emotional tendency can be adjusted to

[0088]

[0089] In summary, the following is a summary of the emotional dictionary extension method:

[0090] posWords:

[0091] If word is a positive word in the basic emotional dictionary, then word is included in posWords;

[0092] If word is a synonym of a positive word in the basic emotional dictionary, then word is included in posWords;

[0093] If word meets the formula word.propertyal ∈ {a, d, an, ag, al} or word.propertyal ∈ {vn, vd, vi, vg, vl}, and 1.36 < SO-PMI(word) < 23, word is included in posWords.

[0094] Similarly, negWords:

[0095] If word is a negative word in the basic sentiment dictionary, then word is included in negWords;

[0096] If word is a synonym of a negative word in the basic sentiment dictionary, then word is included in negWords;

[0097] If word satisfies word.propertyal ∈ {a, d, an, ag, al} or word.propertyal ∈ {vn, vd, vi, vg, vl}, and -16 < SO-PMI(word) < -1, then word is included in negWords.

[0098] Based on the sentiment dictionary, taking each text sentence S as a unit and using each sentiment word WS in the sentence as a separator, calculate the sentiment weight of the sentence segment phrase (WSi-1, WSi) between two separators. The sentence segment phrase (WSi-1, WSi) contains the word WSI but does not contain the word WSi-1; This model consists of 5 modules, namely: analysis of sentiment words, analysis of negative words, analysis of adverbs, analysis of fixed collocations and sentences, analysis of transition words, and analysis of exclamatory sentences.

[0099] Analysis of sentiment words: For each word word in the text to be analyzed, scan the sentiment dictionary to determine whether word exists in the sentiment dictionary. If it exists, then regard word as a sentiment word and read the sentiment tendency value of this word from the negative sentiment dictionary and return it; If it does not exist, then regard word as a neutral word and return 0. Repeat this process until the words in the entire text set are judged. By calculating the sentiment tendency value of each word, we obtain accurate sentiment words (i.e., words with weights not equal to 0), and filter out sentiment words that do not play a sentiment role in a specific sentence (i.e., sentiment words with weights equal to 0).

[0100] Analysis of negative words: When there is a sentiment word Wsi in the sentence, calculate the number of negative words negNum(Wsi-1, Wsi) between Wsi and the previous separator Wsi-1 (i.e., in a sentence segment). If negNum is odd, then the sentiment value of this clause is the opposite of the sentiment tendency value of the sentiment word; Otherwise, keep the original sentiment tendency value.

[0101] Analysis of adverbs: Determine whether the word is in the adverb dictionary. If it is, obtain the adverb sentiment intensity from the adverb dictionary, and multiply the corresponding weight by the current sentiment tendency value of the clause as the sentiment weight of the clause.

[0102] Analysis of transition words: Start scanning backward from the current sentiment word Wsi to find the next sentiment word Wsi+1. During this process, if a transition word is scanned, reverse the weight(phrase(Wsi-1, Wsi)) so that the sentiment tendency of phrase(Wsi-1, Wsi) aligns with the sentiment tendency of the subsequent clause phrase(Wsi, Wsi+1) after the transition word.

[0103] Analysis of exclamatory sentences: For the analysis of exclamatory sentences, we use the exclamation mark "!" as the identifier of the exclamatory sentence and denote it as exc. The method for calculating its sentiment weight is as follows: When the exclamation mark is scanned, we search backward to find the nearest sentiment word Wsi-1 to the exclamation mark, and use the sentiment tendency value of Wsi-1 as the weight of exc.

[0104] Calculate the sum of the negative sentiment tendency values weight(S) of all clauses contained in a text S and the sum of the absolute values of the overall sentiment weights total(S). Calculate the proportion scale(S) of the negative sentiment word weight among all sentiment word weights in the text, scale(S) = weight(S) / total(S). Determine the sentiment polarity of the text S based on scale(S), and make a preliminary judgment on the nature of the cognitive threat of the text based on the sentiment polarity. Texts with scale(S) in the range of [0.68 - 1] are considered to have a high probability of being cognitive threats; texts with scale(S) in the range of [0 - 0.68) are suspected cognitive threat topic texts. Thus, the first-stage classification of the cognitive threat of the text is achieved.

[0105] As a preferred embodiment, further, in the intermediate detection, use the deep learning method to classify the initial suspected cognitive threat topic text into cognitive threat topic text, suspected cognitive threat topic text, and non-cognitive threat topic text. The classification process includes:

[0106] Construct a deep learning model and perform pre-training using a training data set with labeled tags. Among them, the deep learning model includes a BERT model for word vector representation of the input and a BiLSTM model for cognitive threat detection of the input word vectors;

[0107] Input the initial suspected cognitive threat topic text into the pre-trained deep learning model, use the deep learning model to obtain the cognitive threat probability value, and determine the cognitive threat topic text, suspected cognitive threat topic text, and non-cognitive threat topic text in the initial suspected cognitive threat topic text through the cognitive threat probability value.

[0108] In the embodiments of this case, considering the characteristics of the two methods respectively, a multi-dimensional sentiment analysis method combining the sentiment dictionary-based method and the deep learning-based method is proposed and optimized. By combining these two methods, the disadvantages of each method are overcome, and a relatively high accuracy is achieved.

[0109] Cognitive threat recognition based on sentiment analysis performs sentiment analysis on text in two stages. In the first stage, existing sentiment dictionaries such as HowNet and BosonNLP can be referred to, and a basic sentiment dictionary can be constructed using the word frequency statistics method. The sentiment tendency of candidate words is judged by calculating the statistical correlation between the candidate words and the words in the basic sentiment dictionary, realizing the dynamic expansion of the sentiment dictionary. Based on the sentiment dictionary, negation dictionary, and degree adverb dictionary, for each piece of text S, with each sentiment word WS in the sentence as a separator, the sum of negative sentiment weights weight(S) and the sum of the absolute values of sentiment word weights total(S) are calculated for the phrase (WSi-1, WSi) between two separators. scale(S) is defined as the proportion of the negative sentiment weight in all sentiment word weights. According to the size of scale(S), the sentiment polarity of text S is judged, and a preliminary judgment on the nature of the cognitive threat of this text is made based on the sentiment polarity, completing the preliminary recognition of cognitive threats. The statistical results of the analysis of a large number of experimental texts collected show that texts with a negative sentiment weight scale(S) in the range of [0.68 - 1] are very likely to be cognitive threats; texts with a negative sentiment weight scale(S) in the range of [0 - 0.68) are suspected cognitive threat topic texts; texts with a sentiment tendency value in the range of [0 - 0.68) are initially classified as suspected cognitive threats and are subject to the second-stage discrimination process. In the second stage, sentiment analysis can be carried out using the BERT+BiLSTM deep learning model as the core. The sentiment tendency of the text is further analyzed to complete the re-recognition of cognitive threats. The word vectors pre-trained by BERT (Bidirectional Encoder Representations from Transformers) are used to replace the word vectors trained in the traditional way, and the segmented text is transformed into multi-dimensional word vectors. The bidirectional long short-term memory network (BiLSTM) model, which can solve the problems of short-term and long-term dependencies, is used as the core of the sentiment tendency analysis of this section. Using the manually annotated cognitive threat topic text set and the known cognitive threat topic texts under the same theme as the training set, the BERT+BiLSTM model is trained. The trained model is used to further analyze the sentiment of the suspected cognitive threat topic texts obtained in the first stage, and the Softmax classifier is used to divide the text into a set of definite cognitive threat topic texts, a set of suspected cognitive threat topic texts, and a set of non-cognitive threat topic texts.

[0110] For the texts whose sentiment analysis results based on the dynamically extended sentiment dictionary in the first stage are suspected of cognitive threats, a second-stage discrimination process is carried out, and the BERT+BiLSTM deep learning model is used as the core for further identification of cognitive threats. First, model training is carried out, and the training process is as follows: First, the training data set can be manually annotated to mark whether it has the nature of cognitive threats. After word segmentation, the BERT model is used to represent its word vectors, and finally the converted vectors are passed into the BiLSTM neural network. According to the cognitive threat samples, a BiLSTM model covering the cognitive threat samples is trained. The texts whose processing results in the first stage are suspected of cognitive threats are vectorized by BERT words, and the converted vectors are respectively passed into the cognitive threat model. Through this model, a cognitive threat probability value will be obtained. Through a large number of text experiments, it is shown that the texts with a training result probability in the range of (0.68-1] can be determined to have the nature of cognitive threats, the probability in the range of (0.32-0.68] is suspected of cognitive threats and requires manual discrimination, and the probability in the range of [0-0.32] is non-cognitive threat.

[0111] The experimental results show that when the data set contains nearly 5,000 Weibo text data, the accuracy rates of cognitive threat recognition using the sentiment tendency analysis methods based solely on deep learning and solely on the sentiment dictionary are 67.9% and 83.27% respectively, and the accuracy rate of the comprehensive cognitive threat recognition based on multi-dimensional sentiment analysis in this case is 89.9%, which is relatively better.

[0112] Furthermore, in the embodiments of this case, for the preprocessed sensitive topic text data, multi-level cognitive threat detection is used to obtain cognitive threat topic texts, and it also includes: using the sentiment analysis method to evaluate the cognitive threat impact degree according to the overall sentiment tendency in the comment area of the cognitive threat topic text.

[0113] The cognitive threat impact degree is defined by the overall sentiment tendency in the comment area of the text that has been identified as a cognitive threat. After merging all the comment texts under a text that has been determined to be a cognitive threat, after data preprocessing and text word segmentation, the above-mentioned cognitive threat recognition method based on the sentiment dictionary can be used to judge the overall sentiment tendency of the comment area text. Taking the comment orientation triggered by the text with the nature of cognitive threat as the basis for judging the threat degree, and making an evaluation of the threat degree of the text according to the comment sentiment analysis results. The threat degree can be divided into the first, second, and third levels from high to low. The cognitive threat degree of the text with the proportion of the overall negative sentiment weight of the comment text in the overall sentiment word weight of the text in the range of [0.68-1] is defined as the first level; the threat degree of the text with the proportion of the overall negative sentiment weight of the comment text in the overall sentiment word weight of the text in the range of [0.32-0.68) is defined as the second level; the cognitive threat degree of the text with the proportion of the overall negative sentiment weight of the comment text in the overall sentiment word weight of the text in the range of [0-0.32) is defined as the third level. The analysis results can provide an important reference for coping with cognitive threats.

[0114] As a preferred embodiment, further, a cognitive threat propagation knowledge graph is constructed by named entity recognition and entity relationship extraction of cognitive threat topic texts, including:

[0115] A named entity extraction model is constructed and the model is optimized using an adversarial training method. Among them, the named entity recognition model includes an encoder for mapping input characters to a real number space and mining latent semantics, a BiLSTM neural network layer for extracting context semantic information by capturing forward and backward bidirectional features in the vector transformed by the encoder, and a CRF conditional random field layer for taking the bidirectional features extracted by the BiLSTM neural network layer as input and generating character corresponding labels in combination with the Bioes annotation paradigm.

[0116] The cognitive threat topic text is used as the input of the optimized named entity extraction model, and the named entity extraction model is used to identify the entity categories and relationships in the cognitive threat topic text.

[0117] Currently, it is relatively common to use methods based on statistical machine learning to implement named entity recognition tasks. The named entity extraction model in the embodiments of this case adopts the BERT-BiLSTM-CRF model, which is an end-to-end deep learning model that does not require manual feature induction and is developed based on the BiLSTM-CRF model, and can meet the current requirements of Chinese address parsing and address element annotation tasks. This model consists of an encoder (Transformer), a BiLSTM neural network layer, and a conditional random field (CRF) layer from bottom to top. The Transformer encoder is a Chinese BERT model based on character level, which maps the input Chinese address characters to a low-dimensional dense real number space and mines the latent semantics contained in various address elements in the Chinese address; the BiLSTM neural network layer takes the character vector transformed by the encoder as input and captures the forward (from left to right) and backward (from right to left) bidirectional features of the Chinese address sequence, and can fully obtain the semantic information of the context; the CRF conditional random field layer belongs to a probabilistic graphical model, takes the bidirectional features extracted by the upstream BiLSTM as input, and generates labels corresponding to each character in the address in combination with the Bioes annotation paradigm, so as to further parse the Chinese address into various address elements according to the labels, and the problem of the sequence is considered in the calculation process, which can greatly improve the recognition effect of named entities.

[0118] There is no established standard for entities in the cognitive threat field, and most of the existing network naming recognition tasks are only for network public opinion recognition. In the embodiments of this case, the crawled data is analyzed, and according to the cognitive threat recognition requirements, 6 types of entities in the cognitive threat field can be set, namely user, time, address, platform, organization, and hot event, as shown in Table 1:

[0119] Table 1 Cognitive Threat Entity Types

[0120]

[0121]

[0122] Entity annotation is the most important issue in the named entity recognition task and is also the basis for model training. The commonly used annotation methods are of two types: BIO and BIOES. Although the BIOES annotation method provides more information, it requires more tags to be predicted. Due to the limited number of datasets constructed in this case, the effect of using the BIOES annotation method may be affected. In the BIO annotation system, the "B" can be used to mark the start of an entity, the "I" to mark the inside of an entity, and the "O" to mark non-entities. The tags for each type of entity include "start" and "inside". Therefore, the named entity recognition dataset constructed in this case can be set to 13 tags.

[0123] There are complex relationships among the entities in the cognitive threat user knowledge graph. When extracting knowledge from cognitive threat topic texts, attention should also be paid to the dependency relationships between adjacent tags. However, since BiLSTM is good at processing long-distance text information and cannot handle the dependency relationships between adjacent tags, therefore, in the knowledge extraction of the cognitive threat user knowledge graph, CRF (Conditional Random Field) can be combined. Based on the preliminary predicted tags corresponding to each word output by BiLSTM, the output scores are corrected through the relationships of adjacent tags to obtain an optimal predicted sequence.

[0124] The CRF layer takes the output scores of the upper-layer BILSTM as input and outputs the most likely predicted annotation sequence that conforms to the annotation transition constraint conditions. For any sequence X = (x1, x2,..., xn); here it is assumed that P is the output score matrix of BiLSTM, and the size of P is n×k, where n is the number of words and k is the number of tags. Pij represents the score of the jth tag for the ith word. For the predicted sequence Y = (y1, y2,..., yn), its score function is:

[0125]

[0126] A represents the transition score matrix, Aij represents the score for the transition from tag i to tag j, and the size of A is k + 2. The probability of the predicted sequence Y is:

[0127]

[0128] Taking the logarithm at both ends gives the likelihood function of the predicted sequence:

[0129]

[0130] In the formula, \(\widetilde{Y}\) represents the true annotation sequence, and \(YX\) represents all possible annotation sequences. The output sequence with the maximum score is obtained after decoding:

[0131]

[0132] The CRF layer outputs the optimal label sequence of the cognitive threat topic text, focusing on the words corresponding to labels such as the forwarding user, forwarding time, forwarding location, forwarding platform, text summary information, text theme, etc. These are the basis for building the knowledge graph of cognitive threat users and conducting traceability processes such as tracing the forwarding user, tracing the forwarding process, tracing the forwarding time, tracing the forwarding platform, etc., and for inferring the relationship of cognitive threat forwarding users.

[0133] When using BERT and its variants, since they have been pre-trained and the parameters have reached a good level, in order to maintain the training effect, a lower learning rate should be adopted; on the contrary, since its downstream tasks have not been pre-trained, if a lower learning rate is set, not only will the training process be slow, but it will also be difficult to synchronize with BERT training. Therefore, in the embodiments of this case, a strategy of setting the learning rate in layers can be adopted: for the upstream BERT pre-training layer, a smaller learning rate is set, and for the downstream layer, a larger learning rate is set.

[0134] During the model training process, when the loss value decreases gradually and flattens out, if a relatively large learning rate is still adopted, it will cause the model to oscillate back and forth near the global optimum point when converging to the global optimum point. To ensure that the loss function finally always remains within a very close range to the optimal value and gradually approaches the optimal value, a learning rate decay strategy needs to be adopted, that is, to reduce the step size of parameter update. In this case, a learning decay strategy can be set: when the model effect does not improve during the training process, reduce the learning rate, which can effectively improve the model accuracy.

[0135] BERT-BiLSTM-CRF is used as a named entity recognition model. However, due to the local instability of neural networks, even a tiny perturbation may cause a large error in the model. Therefore, in the embodiments of this case, an adversarial training method is adopted to optimize the model. Adversarial training can improve the robustness of the model by inputting tiny perturbations into the model, and can achieve the effects of alleviating the defects of the local instability of neural networks and improving the robustness of the model. See Figure 3 as shown. During the training process, first BERT will generate an initial vector for the input text, and then add some perturbations to the initial vector to generate adversarial samples. These adversarial samples, as variants of the original samples, are likely to mislead the model. The initial vector and the adversarial samples will be input into BiLSTM for training together, and the neural network will learn more robust parameters during the training process to resist the attack of adversarial samples.

[0136] As a preferred embodiment, further, in constructing the cognitive threat propagation knowledge graph by performing named entity recognition and entity relationship extraction on cognitive threat topic texts, two named entity extraction models with pipeline connections can be established. Among them, the first named entity extraction model uses a single-label multi-classification task method to identify entities in cognitive threat topic texts, and the second named entity extraction model uses a multi-label multi-classification task method to take the input of the first named entity extraction model as input to identify the relationships between entities.

[0137] Knowledge fusion is an important task in constructing a domain knowledge graph. By aligning, associating, and merging multiple related entities into a whole, its main work is divided into two parts: entity unification and entity disambiguation. Due to the political and aggressive characteristics of cognitive threat topic texts, there are problems with the non-unification of entities identified by named entity recognition. Therefore, entity unification and entity disambiguation are required. Among them, entity unification refers to different entity examples with the same meaning, which need to be unified.

[0138] The ambiguity of a named entity refers to the situation where an entity reference can correspond to multiple real-world entities. Due to the richness and complexity of Chinese semantics, the meaning represented by the same word may be different in different contexts. Therefore, entity disambiguation is required. In this case, a link-based entity disambiguation method can be adopted to link the entity reference to the corresponding entity in the knowledge base. After entity unification, the final valid entities can be obtained.

[0139] Graph databases are good at processing large amounts of complex, interconnected, and low-structured data. These data change rapidly and require frequent queries - in relational databases, these queries would result in a large number of table joins, thus causing performance problems. And due to the performance degradation problems that occur during querying in traditional ones such as RDBMS, in the embodiments of this case, a persistent engine Neo4j that supports full transactions can be adopted. It provides large-scale scalability and can process billions of node-relationship-attribute graphs on a single machine, and can be extended to multiple machines for parallel operation. At the same time, it focuses on solving the performance degradation problem. By modeling data around the graph, Neo4j traverses nodes and edges at the same speed, and its traversal speed has nothing to do with the amount of data that makes up the graph.

[0140] Such as Figure 4As shown in the figure, when implementing the visualization and display of the knowledge graph based on Neo4j, a set of visualization graph elements can be defined in the Neo4j program. The schema of cognitive threat propagation can be mainly expressed by type and property. Entities such as users, events, addresses, platforms, hot events, sentiment tags, and threat intents are defined. In terms of the definition of relationships, the following relationships can be defined: text - hot event, user - text, forwarding, user - organization, text - sentiment tag, text - threat intent, etc. It can be represented by triples as: <text, text - hot event, hot event>, <user, user - cognitive threat topic text, text>, <user, forwarding, user>, <user, user - organization, user>, <text, text - threat intent, threat intent>, etc.

[0141] By constructing a cognitive threat propagation knowledge graph, user tracing, event tracing, and organization tracing of cognitive threat propagation are realized. Through the mining of implicit relationships, real - time monitoring of key accounts, groups, organizations, and users can provide a basis for the accurate positioning and directional blocking of cognitive threats.

[0142] Furthermore, based on the above - mentioned method, an embodiment of the present invention also provides a social media cognitive threat detection system, including: a data collection server, multiple cognitive threat discrimination servers, a knowledge graph server, and a web server, where,

[0143] The data collection server is used to collect sensitive topic text data of the network platform and perform pre - processing operations on the data;

[0144] Multiple cognitive threat discrimination servers are used to obtain cognitive threat topic text through multi - level cognitive threat detection for the pre - processed sensitive topic text data. Among them, the multiple cognitive threat discrimination servers specifically include: a primary discrimination server that divides the sensitive topic text data into cognitive threat topic text and initial suspected cognitive threat topic text, a secondary discrimination server that classifies the initial suspected cognitive threat topic text into cognitive threat topic text, suspected cognitive threat topic text, and non - cognitive threat topic text, and an ultimate discrimination server that obtains cognitive threat topic text from the suspected cognitive threat topic text through manual annotation;

[0145] The knowledge graph server is used to construct a cognitive threat propagation knowledge graph through named entity recognition and entity relationship extraction of cognitive threat topic text;

[0146] The web server is used to perform user tracing, event tracing, and organization tracing of the propagation of cognitive threat topic text based on the cognitive threat propagation knowledge graph and using a web interaction interface.

[0147] The front end can be designed based on JavaScript and use Echart to achieve data visualization. Multiple cognitive threat identification servers can be set to have a decentralized distributed feature and jointly be responsible for the detection and measurement of cognitive threat information. Only the text that is recognized as a cognitive threat by multiple servers will be determined as a cognitive threat topic text. After the text is determined as a cognitive threat topic text, the cognitive threat identification server uploads the text information and the text cognitive threat attribute measurement information to the distributed network. The distributed structure not only improves the accuracy of cognitive threat detection but also enhances the anti-risk ability of the system. The damage of one server will not affect the operation of the entire system.

[0148] The knowledge graph server corresponds to the cognitive threat knowledge extraction module and the cognitive threat propagation knowledge graph construction module. The knowledge graph server automatically accesses the distributed network through a smart contract, extracts the text information of the detected cognitive threat topic text, performs named entity recognition and relationship extraction on cognitive threat entities such as users, time, address, organization, forwarding platform, and hot events related to the cognitive threat topic text, and constructs a cognitive threat propagation knowledge graph through Neo4j.

[0149] The human-computer interaction interface can use the Web to realize the bridge for users to interact with data, and use the Echart visualization tool to convert abstract data and relationships into intuitive charts.

[0150] In the embodiments of this case, the system has good security and feasibility, the modular operation complexity is relatively low, and it is easy to maintain. And through experimental data verification, the accuracy rate of threat determination for topic texts within a hundred words reaches 93%, which is relatively high and the detection effect is good. In addition, the solution of this case has a very broad application scenario and can be used in aspects such as news media supervision, network public opinion supervision, and combating illegal acts.

[0151] Unless otherwise specifically stated, the relative steps, numerical expressions, and numerical values of the components and steps described in these embodiments do not limit the scope of the present invention.

[0152] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0153] The units and method steps of each example described in combination with the embodiments disclosed in this specification can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.

[0154] Those of ordinary skill in the art can understand that all or part of the steps in the above methods can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software functional module. The present invention is not limited to any specific form of the combination of hardware and software.

[0155] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for detecting cognitive threats in social media, characterized in that, It includes the following content: Collect sensitive topic text data on the network platform and perform preprocessing operations on the data; For the preprocessed sensitive topic text data, obtain cognitive threat topic text through multi-level cognitive threat detection. Among them, multi-level cognitive threat detection includes: primary detection of dividing sensitive topic text data into cognitive threat topic text and initial suspected cognitive threat topic text, intermediate detection of classifying initial suspected cognitive threat topic text into cognitive threat topic text, suspected cognitive threat topic text, and non-cognitive threat topic text, and ultimate detection of obtaining cognitive threat topic text from suspected cognitive threat topic text through manual annotation; and in the primary detection, use sentiment analysis methods to divide sensitive topic text data into cognitive threat topic text and initial suspected cognitive threat topic text. The process of sentiment analysis method division includes: first, construct a basic sentiment dictionary based on the known sentiment dictionary and using word frequency statistics method, and expand the sentiment dictionary by performing correlation statistics between the words in the text data and the words in the basic sentiment dictionary; then, take the text in the sensitive topic text data as the unit and the sentiment words as the delimiters, perform sentiment weight statistics on the sentences between each delimiter, and judge the sentiment polarity of the text according to the proportion of the negative sentiment weight in all sentiment word weights; then, divide the sensitive topic text data into cognitive threat topic text and initial suspected cognitive threat topic text according to the sentiment polarity of the text; in the intermediate detection, use deep learning methods to classify the initial suspected cognitive threat topic text into cognitive threat topic text, suspected cognitive threat topic text, and non-cognitive threat topic text. The classification process includes: constructing a deep learning model and performing pre-training using a training data set with labeled tags. The deep learning model includes a BERT model for word vector representation of the input and a BiLSTM model for cognitive threat detection of the input word vectors; input the initial suspected cognitive threat topic text into the pre-trained deep learning model, use the deep learning model to obtain the cognitive threat probability value, and determine the cognitive threat topic text, suspected cognitive threat topic text, and non-cognitive threat topic text in the initial suspected cognitive threat topic text through the cognitive threat probability value; Construct a cognitive threat propagation knowledge graph through named entity recognition and entity relationship extraction of cognitive threat topic text; Based on the cognitive threat propagation knowledge graph, conduct user traceability, event traceability, and organization traceability for the propagation of cognitive threat topic text.

2. The method for detecting cognitive threats in social media according to claim 1, characterized in that, Collect sensitive topic text data on the network platform and perform preprocessing operations on the data, including: First, distributedly collect sensitive topic text information and relevant user data on the network platform according to the user authorization information library; Then, for the collected text information, merge the title and the text body, use the redundancy detection algorithm to remove redundant information, perform deduplication processing on relevant comments, clean and transform text noise data, and perform word segmentation processing on the text using a word segmentation system.

3. The method for detecting cognitive threats in social media according to claim 1, characterized in that, Construct a basic sentiment dictionary based on the known sentiment dictionary and using word frequency statistics method, including: First, select a series of sentiment words from the known sentiment dictionary, sort the sentiment words according to the search engine click-through rate in the series of sentiment words, and select several sentiment words according to the click-through rate popularity; Next, select the sentiment words with the highest relevance to the theme based on word frequency statistics, and use the selected several sentiment words and sentiment words to jointly form a basic sentiment dictionary; Then, expand the basic sentiment dictionary using synonyms and candidate words with sentiment tendencies.

4. The method for detecting cognitive threats in social media according to claim 1, characterized in that, Perform sentiment weight statistics on the sentences between each delimiter, and judge the sentiment polarity of the text according to the proportion of the negative sentiment weight in the weights of all sentiment words, including: First, for the sentences between delimiters, statistically analyze the sentiment tendencies respectively through sentiment word analysis, negative word analysis, adverb analysis, fixed collocation word analysis, transition word analysis, and exclamation sentence analysis; Then, statistically analyze the sum of the negative sentiment tendency values of all clauses included in the text and the sum of the absolute values of the overall sentiment weights, and use the proportion of the negative sentiment word weights in the weights of all sentiment words in the text to judge the sentiment polarity of the text.

5. The method for detecting cognitive threats in social media according to claim 1, characterized in that, For the preprocessed sensitive topic text data, obtain the cognitive threat topic text through multi-level cognitive threat detection, and also include: using the sentiment analysis method to evaluate the cognitive threat impact degree according to the overall sentiment tendency in the comment area of the cognitive threat topic text.

6. The method for detecting cognitive threats in social media according to claim 1, characterized in that, Construct a cognitive threat propagation knowledge graph through named entity recognition and entity relationship extraction of the cognitive threat topic text, including: Construct a named entity extraction model, and use the adversarial training method to optimize the model. Among them, the named entity recognition model includes an encoder for mapping input characters to the real number space and mining potential semantics, a BiLSTM neural network layer for extracting context semantic information by capturing the forward and backward bidirectional features in the vector transformed by the encoder, and a CRF conditional random field layer for taking the bidirectional features extracted by the BiLSTM neural network layer as input and generating character corresponding labels in combination with the Bioes annotation paradigm; Use the optimized named entity extraction model as the input of the cognitive threat topic text, and use the named entity extraction model to identify the entity categories and relationships in the cognitive threat topic text.

7. The method for detecting cognitive threats in social media according to claim 6, characterized in that, In constructing the cognitive threat propagation knowledge graph through named entity recognition and entity relationship extraction of the cognitive threat topic text, build two named entity extraction models connected in a pipeline. Among them, the first named entity extraction model uses the single-label multi-classification task method to identify the entities in the cognitive threat topic text, and the second named entity extraction model uses the multi-label multi-classification task method to take the input of the first named entity extraction model as the input to identify the relationships between entities.

8. A system for detecting cognitive threats in social media, characterized in that, Implemented based on the method described in claim 1, including: a data collection server, multiple cognitive threat identification servers, a knowledge graph server, and a web server, where, The data collection server is used to collect sensitive topic text data on the network platform and perform preprocessing operations on the data; Multiple cognitive threat identification servers are used to obtain cognitive threat topic texts through multi-level cognitive threat detection for preprocessed sensitive topic text data. Specifically, the multiple cognitive threat identification servers include: a primary identification server that divides the sensitive topic text data into cognitive threat topic texts and initial suspected cognitive threat topic texts, a secondary identification server that classifies the initial suspected cognitive threat topic texts into cognitive threat topic texts, suspected cognitive threat topic texts, and non-cognitive threat topic texts, and an ultimate identification server that obtains cognitive threat topic texts from the suspected cognitive threat topic texts through manual annotation; A knowledge graph server is used to construct a cognitive threat propagation knowledge graph through named entity recognition and entity relationship extraction of cognitive threat topic texts; A web server is used to perform user traceability, event traceability, and organization traceability of the spread of cognitive threat topic texts based on the cognitive threat propagation knowledge graph and using a web interface.

Citation Information

Patent Citations

  • Text data-oriented threat intelligence knowledge graph construction method

    CN110717049A