A method for constructing a knowledge graph
By automatically extracting knowledge on unlabeled political theory corpus text and combining expert annotation and pre-trained language models, the problem that the existing technology cannot effectively extract political theory knowledge and associations is solved, and a political theory knowledge graph construction with high accuracy and coverage is achieved.
Patent Information
- Application Number
- CN202210729506.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-06-24
AI Technical Summary
The existing automatic construction methods of knowledge graphs cannot effectively extract longer conceptual knowledge in political theory, and cannot automatically extract the correlation between political theory, and there are shortcomings in accuracy and coverage.
By automatically extracting political theory knowledge on unlabeled political theory corpus text, combining expert labeling and pre-training language models, the political knowledge extraction model is trained, and the correlation between knowledge points is calculated based on co-occurrence and semantic similarity.
It realizes automatic extraction of political theory knowledge under no labeling conditions, constructs a political theory knowledge graph, and improves the accuracy and coverage of the knowledge graph, effectively integrates expert knowledge and improves the credibility of the knowledge base.
Smart Images

Figure CN115221335B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of natural language processing and Internet applications, and particularly relates to a method for constructing a knowledge graph. This method can automatically extract political theory knowledge and construct a political theory knowledge graph without knowledge annotation; in the case of text with knowledge annotation, an extraction model for political theory knowledge can be trained using the annotation data based on a pre-trained language model. Based on co-occurrence and semantic similarity, the relationships between political theories are calculated. In addition, this system can incorporate experts' knowledge into the knowledge graph and can also use experts' knowledge to iteratively optimize the knowledge extraction ability of the model. Background Art
[0002] There are a very large number of theories in the political field, and people have a great demand for learning and querying. However, for political theories, there is currently only a knowledge graph manually sorted by experts. Such a knowledge graph, due to being manually sorted by people, will not be very large in scale and cannot cover a wide range. Currently, there is still no good system that can structurally manage the theoretical knowledge in the large-scale political field. Structural management means that for each political theory, it is classified into one or more knowledge points, and there are clear relationships between the knowledge points.
[0003] A knowledge graph is a graph structure that shows the knowledge structure relationship and can be used to store and manage knowledge and its connections. With the development of machine learning and artificial intelligence, machines have achieved excellent performance in the task of automatically constructing knowledge graphs. Especially in general knowledge fields, such as the knowledge in Wikipedia, a large amount of knowledge has been automatically constructed by machines, helping people save a lot of time.
[0004] However, the existing methods for automatically constructing knowledge graphs are not applicable to the political theory field. There are the following three reasons:
[0005] 1. The knowledge extracted by the existing methods for automatically constructing knowledge graphs is not political knowledge. The existing methods for automatically constructing knowledge graphs regard entities or attributes as knowledge, usually a single word, while the knowledge in political theories is mostly proper nouns, usually a concept. The existing methods for extracting knowledge cannot extract these longer concepts.
[0006] 2. The knowledge graph of political theories has a relatively high requirement for the accuracy of knowledge due to its sensitivity. Limited by the existing technology, the knowledge graph completely automatically constructed by machines cannot meet the requirement of 100% accuracy. Therefore, this system is required to incorporate experts' knowledge and replace the knowledge automatically extracted by machines with lower confidence with experts' knowledge with higher confidence.
[0007] 3. There are logical connections between political theories, including knowledge points with a hierarchical relationship, knowledge points belonging to the same field, and similar knowledge points. Existing knowledge graph construction methods cannot automatically extract the associations between political theories. Summary of the Invention
[0008] One of the objectives of the present invention is to provide a method for automatically constructing a knowledge graph, which can achieve automatic extraction of political theory knowledge, integration of expert-annotated knowledge, and automatic discovery of the associations between political theory knowledge. The present invention uses statistical methods to extract candidate knowledge, and trains a knowledge extraction model after expert annotation, so as to better extract theoretical knowledge and solve the problem that the existing knowledge extraction methods cannot extract these long concepts. Based on the co-occurrence score and semantic similarity score between knowledge points, the present invention calculates the association between theoretical knowledge by weighted calculation, and solves the problem that existing knowledge graph construction methods cannot automatically extract the associations between political theories.
[0009] The technical solution of the present invention is as follows:
[0010] A method for constructing a knowledge graph, the steps of which include:
[0011] 1) Automatically extract political theory knowledge from unannotated political theory corpus texts;
[0012] 2) Screen the political theory knowledge extracted in step 1), and annotate the screened political theory knowledge as the training text for training a political knowledge extraction model;
[0013] 3) Use the training text to train the political knowledge extraction model;
[0014] 4) Use the trained political knowledge extraction model to extract knowledge from the corpus to obtain political theory knowledge;
[0015] 5) For any two political theory knowledge in the political theory knowledge obtained in step 4), calculate the co-occurrence degree and semantic similarity of the two political theory knowledge in the corpus. If the co-occurrence degree or semantic similarity is not zero, connect an edge between the two political theory knowledge, and calculate the association score between the two political theory knowledge based on the co-occurrence degree and semantic similarity as the weight of the edge, so as to obtain the knowledge graph corresponding to the corpus;
[0016] 6) Align the knowledge system with a hierarchical structure annotated by experts with the knowledge graph generated in step 5),
[0017] Integrate the hierarchical relationship between the subject words annotated by experts in the knowledge system into the knowledge graph.
[0018] Further, the method for automatically extracting political theory knowledge from unannotated political theory corpus texts includes:
[0019] 11) Tokenize each sentence S in a political theory corpus text A in the corpus to obtain a token list w = {w1, w2,..., w n} and the corresponding part-of-speech list t = {t1, t2,..., t n}; w n is the nth token in sentence S, and t n is the part of speech of w n .
[0020] 12) Combine adjacent k tokens in the token list w to obtain multiple candidate phrases k-gram; calculate the tf-idf scores of each k-gram in the political theory corpus text A when k takes different values.
[0021] 13) Add up the tf-idf scores of each candidate phrase k-gram in each political theory corpus text in the corpus to obtain the final tf-idf score of the candidate phrase k-gram, and select several candidate phrases with the largest final tf-idf scores as the automatically extracted political theory knowledge.
[0022] Further, the association score includes the co-occurrence degree score of two political theory knowledge, the semantic similarity score between two political theory knowledge, and the expert annotation score; among them, for two political theory knowledge i, j, if they co-occur in n1 sentences, n2 paragraphs, and n3 texts in each political theory corpus text in the corpus, then their co-occurrence degree score is C ij =(a * n1 + b * n2 + c * n3) p , where a, b, c are the weights corresponding to sentence co-occurrence, paragraph co-occurrence, and text co-occurrence, and p is the total number of texts in the corpus; the semantic similarity score S ij between two political theory knowledge i, j is calculated through a semantic similarity model; if two political theory knowledge i, j are co-occurrence annotated l times by experts in the same sentence, then their expert annotation score is Z ij =z * l; the final association score of two political theory knowledge i, j is: R ij =c * C ij +s * S ij +z * Z ij .
[0023] Further, the optimization function used when training the political knowledge extraction model based on a large-scale language model is the maximum likelihood optimization function.
[0024] Further, the political knowledge extraction model includes a large-scale pre-trained language model BERT and a conditional random field model; the political theory corpus text is input into the large-scale pre-trained language model BERT to obtain the word encoding of each word, which is used as the input of the conditional random field model. The conditional random field model outputs the probability that the sentence sequence is labeled with different tags; the tag sequence with the highest probability is selected for decoding to obtain the political theory knowledge contained in the corresponding sentence.
[0025] Further, the method for integrating the hyponymy relationships between the subject words annotated by experts in the knowledge system into the knowledge graph includes:
[0026] 61) Knowledge alignment: If the subject words annotated by experts are character-identical to the extracted subject words, then these two words are considered to be the same political theory knowledge; otherwise, the subject words annotated by experts are considered to be new political theory knowledge and merged into the extracted knowledge base.
[0027] 62) Aggregation of associations between knowledge: Aggregate the associations in the knowledge graph obtained in step 5) and the associations annotated by experts.
[0028] 63) Association score between subject words and knowledge: According to the subject word w theme annotated by experts, the association score between w robot and the extracted political theory knowledge w theme is weighted and obtained by summing the association relationship scores between the corresponding related words and w robot ; where experts annotate a set of related keywords for each subject word; finally, the fused knowledge graph contains edges with association scores, edges with hyponymy relationships, and edges with expert association scores.
[0029] Further, the knowledge system is a three-level knowledge system, including first-level knowledge subject words, second-level knowledge subject words, and third-level knowledge subject words.
[0030] A server, characterized in that it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in the above method.
[0031] A computer-readable storage medium, on which a computer program is stored, characterized in that the steps of the above method are implemented when the computer program is executed by a processor.
[0032] The embodiments of the present invention provide a method and system for constructing a knowledge graph of political theory, including the following steps:
[0033] · Automatically extract political theory knowledge from unannotated political theory corpus text; the extracted political theory knowledge is the concept mentioned above.
[0034] · Experts judge the quality of the extracted political theory knowledge, screen out the extracted political theory knowledge and label it, which is used as the training text for training the political knowledge extraction model;
[0035] · Use the training text to train a political knowledge extraction model based on a large-scale language model;
[0036] · Use the trained political knowledge extraction model to extract knowledge from the political theory corpus to obtain political theory knowledge;
[0037] · Calculate the correlation relationship between knowledge points based on the co-occurrence degree and semantic similarity of political theory knowledge in the corpus;
[0038] · Integrate the knowledge graph by combining the topic words and knowledge structure system annotated by experts.
[0039] Furthermore, automatically extract political theory knowledge from unannotated texts, including: segmenting political articles, calculating the tf-idf scores of all 2-gram (text composed of two adjacent words) to k-gram (text composed of k adjacent words), and screening the candidates based on the strategy of political knowledge to obtain the n-gram with the largest tf-idf score as the result of automatically extracted political theory knowledge.
[0040] Furthermore, experts need to screen the automatically extracted political theory knowledge, including: removing low-quality political theory knowledge and adding key political theory knowledge that has not been extracted.
[0041] Furthermore, use the training text to train a political knowledge extraction model based on a large-scale language model, where the political knowledge extraction model is a trained sequence annotation model (including the large-scale pre-trained language model BERT and the conditional random field model); first, combine the large-scale pre-trained language model BERT and the conditional random field model to obtain an initial sequence annotation model; use the word encoding obtained by passing the text through the pre-trained model as the input of the conditional random field model, and the conditional random field model outputs the probability that the sentence sequence is labeled with different labels. The parameters of the large-scale pre-trained language model BERT are initialized with the pre-trained parameters (the relevant parameters can be downloaded from the Internet), and the parameters of the conditional random field model are randomly initialized.
[0042] Further, use a political knowledge extraction model to extract knowledge from the political theory corpus, obtaining political theory knowledge, including: adding [CLS] and [SEP] before and after each sentence in the political theory article and inputting it into the sequence annotation model to obtain the vector representation of each word. Input the vector representations of all words in the same sentence into the fully connected layer to obtain the label probability of each word. Feed the label probability of each word into the conditional random field to obtain the probability scores of different label sequences for the entire sentence. Define the transition probability from label A to label B as the probability q(A,B) that label A is followed by B. Then, the probability P(T) that sentence S is marked as the label sequence T=t1,t2,...,t n is:
[0043]
[0044] Further, based on the co-occurrence degree and semantic similarity of political theory knowledge in the corpus, calculate the association relationship between knowledge, including: the association score between two political knowledge consists of three parts, namely the co-occurrence degree score of knowledge in the theory article, the semantic similarity score between knowledge, and the expert annotation score between knowledge generated according to expert annotation.
[0045] Further, combine the subject words and knowledge structure system annotated by experts to integrate the knowledge graph, including: align the knowledge system with upper and lower level structures annotated by experts with the knowledge in the knowledge graph automatically generated by the machine, and incorporate the upper and lower level relationships between the subject words annotated by experts in the knowledge system into the knowledge graph.
[0046] The advantages of the present invention are as follows:
[0047] 1. Automatically extract knowledge points from unannotated text, saving effort for expert annotation;
[0048] 2. It is relatively well-matched in the scenario of political theory knowledge, and can extract political theory knowledge from the text and construct the association between theoretical knowledge;
[0049] 3. It can integrate the knowledge of experts, making the content of the entire knowledge base more accurate and more credible. Brief Description of the Drawings
[0050] Figure 1 It is a schematic diagram of input and output results in an embodiment of the present invention. Detailed Embodiment
[0051] The technical solutions in the embodiments of the present invention will be clearly and completely described below. It can be understood that the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.
[0052] The example of the present invention obtains a data set based on political theory texts and a knowledge system designed by experts. Those skilled in the art should clearly understand that other candidate information sets can also be adopted in the specific implementation process.
[0053] Specifically, this example comes from the content of an optional political work and the transcript of the News Broadcast.
[0054] Step 1: Automatically extract political theory knowledge from unannotated texts:
[0055] Segment the sentences of the political article. The jieba software package existing in the art is used for segmentation. The sentence S is segmented into a list of words w = {w1, w2,..., w n}, and at the same time, the part-of-speech of each word is predicted to obtain a list of part-of-speech of each word t = {t1, t2,..., t n}.
[0056] Combine the words in the word list w to obtain all the results of k-grams. Specifically, the results of k-grams in the sentence S are {w1w2...w k , w2w3...w k+1 ,..., w i w i+1 ...w i+k-1 ,..., w n-k+1 w n-k+2 ...w n}. In practice, we take all k-grams where 1 <= k <= 3 as alternative results.
[0057] Among all the k-grams, a further screening is also required to obtain the final segmentation result. The screening is carried out according to the part-of-speech combination rules in the phrase, and the screening rules are as follows:
[0058] · For a phrase composed of a single word with a length less than two characters, delete it from the result;
[0059] · For a phrase containing punctuation, delete it from the result;
[0060] · For a phrase containing conjunctions, adverbs, interjections, onomatopoeias, prepositions, pronouns, auxiliary words, modal particles, and time words, delete it from the result;
[0061] · For a phrase containing empty content words such as ['change', 'hope', 'propose', 'include', 'promote', 'improve', 'build','strengthen', 'continue'], delete it from the result;
[0062] · For all phrases in the form of noun + verb, delete them from the result;
[0063] · For phrases whose last word is a number or an adjective, delete them from the results.
[0064] · For all two-word phrases, if they are in the structure of verb + noun or noun + verb, delete them from the results.
[0065] · For phrases that start or end with a locative word, delete them from the results.
[0066] · For phrases that do not contain a noun, delete them from the results.
[0067] · For phrases whose first word is a verb, delete them from the results.
[0068] After the above steps, for each article, the results of all phrases in the article can be calculated as candidates for extracting political theory knowledge.
[0069] Next, calculate the tf-idf of each k-gram phrase in the corpus. The term frequency tf refers to the frequency of a word in an article, and its expression is:
[0070]
[0071] where md is the number of times the k-gram phrase appears in the article, and nd is the total number of phrases in this article.
[0072] The expression of the inverse document frequency idf is:
[0073]
[0074] where p is the total number of documents in the corpus, and q is the number of documents in the corpus that contain this phrase.
[0075] The expression for calculating the tf-idf score is:
[0076] tf-idf = tf × idf
[0077] Among the tf-idf score results of the phrases calculated in an article, for two phrases a and b with the same score, if b is completely included in a, then lower the score of b to retain the longer phrase as political theory knowledge.
[0078] Add up the tf-idf scores of a phrase in each article to obtain the score of this phrase in the entire knowledge base as a piece of political theory knowledge. Sort the scores of all phrases and select the top k with the largest scores to obtain the set of political knowledge automatically extracted by the machine.
[0079] Step 2: Experts judge the quality of the extracted political theory knowledge and screen high-quality results:
[0080] Experts conduct manual evaluation on the extracted political theory knowledge and delete unreasonable results from the automatically extracted political theory knowledge. At the same time, if there are key political theories that are missed, they will be specially supplemented. For the supplemented political theory results, the algorithm will specifically improve the strategy to avoid subsequent omissions.
[0081] Step 3: Use the results of machine automatic extraction of political knowledge screened by experts to train a political knowledge extraction model based on a large-scale language model:
[0082] Based on the political knowledge screened by experts, supervised training data is obtained through processing: Using BIO tags, the first character of the political theory knowledge is labeled as B, the remaining characters in the political theory knowledge are labeled as I, and the remaining characters are labeled as O. The article is divided into sentences, and each sentence is used as a sample to input into the pre-trained model. The sentence sequence is input into the pre-trained model to obtain the representation of each character. The representation of the sentence is passed through a fully connected layer, and the normalized raw probability is obtained through softmax, that is, the probability of each character belonging to the three tags. The raw probability sequence is input into the conditional random field to obtain the label sequence corresponding to the maximum sequence probability after multiplying by the transition probability. During training, maximum likelihood optimization is adopted, and the optimization function is:
[0083]
[0084] where t i represents the true label of the i-th character, and p i (t i ) represents the probability that the i-th character is the label t i . During prediction, the label sequence with the highest probability is selected. Through decoding, all the political theory knowledge contained in the sentence is obtained.
[0085] Step 4: Calculate the association relationship between knowledge based on the co-occurrence and semantic similarity of political theory knowledge in the corpus:
[0086] The association score between two political knowledge consists of three parts, namely the co-occurrence score of the knowledge in theoretical articles, the semantic similarity score between the knowledge, and the knowledge connection score generated by expert annotation.
[0087] The co-occurrence score of the knowledge in theoretical articles is obtained by weighting the scores of the two knowledge collinear at different granularities. For two pieces of knowledge, if they co-occur in n1 sentences, n2 paragraphs, and n3 articles, then their co-occurrence score is C ij =(a*n1 + b*n2 + c*n3) p . Where a, b, and c are the weights of sentence co-occurrence, paragraph co-occurrence, and article co-occurrence respectively.
[0088] After obtaining the co-occurrence scores for each pair of knowledge in the knowledge base, we will find the maximum value C of the co-occurrence scores max , and divide all co-occurrence scores by C max , so that the range of co-occurrence scores is normalized between 0 and 1.
[0089] The semantic similarity scores between knowledge are calculated through a semantic similarity model. By encoding the knowledge i, j model, the vectorized representation E of the knowledge is obtained i , E j , and then calculate the cosine similarity between the two vectors to obtain the semantic similarity score:
[0090]
[0091] The knowledge connection scores generated by expert annotation are naturally generated through expert annotation. If knowledge i, j are co-occurrence annotated l times by an expert in the same sentence, then the expert annotation score between knowledge i, j is:
[0092] Z ij = z * l
[0093] Similar to the normalization of co-occurrence scores, the expert annotation scores will also be normalized to the interval between 0 and 1.
[0094] The final association score between knowledge is:
[0095] R ij = c * C ij + s * S ij + z * Z ij
[0096] Step Five: Integrate the knowledge graph by combining the subject words and knowledge structure system annotated by experts:
[0097] Experts have annotated a three-level knowledge system for each field, namely the first-level knowledge subject words, the second-level knowledge subject words, and the third-level knowledge subject words. This step requires integrating the knowledge system annotated by experts with the knowledge base extracted by the machine. The integration of the expert knowledge system and the machine-extracted knowledge base can be divided into three steps: alignment of knowledge, aggregation of associations between knowledge, and calculation of associations between subject words and knowledge.
[0098] The alignment of knowledge is obtained through character matching. If the subject words annotated by experts are exactly the same as the subject words extracted by the machine, or there are only differences in punctuation or spaces, then these two words are considered to be the same knowledge. If the subject words annotated by experts do not correspond to the knowledge extracted by the machine, then this annotated subject word is considered to be new knowledge. Finally, the complete set of knowledge is obtained and merged into the knowledge base extracted by the machine.
[0099] The aggregation of the associations between knowledge is to aggregate the associations extracted by the machine (there is an association between two pieces of knowledge with a connection relationship in the knowledge graph) and the associations annotated by experts. The experts' annotations will mark the hyponymy relationships between knowledge. For example, if a first-level topic term contains k second-level topic terms, then k knowledge pairs with hyponymy relationships are naturally induced. We take the hyponymy relationships annotated by experts as a new type of relationship and merge them into the knowledge base extracted by the machine.
[0100] The association between a topic term and knowledge is to specifically calculate the expert association score between the topic term annotated by experts and the knowledge extracted by the machine in a new dimension. The topic term w annotated by experts theme and the knowledge point w extracted by the machine robot The association score between them is weighted by the sum of all association relationship scores between w theme and w robot The association relationship score specifically refers to the association between a topic term and knowledge and is obtained through expert annotation. The expert annotates a set of relevant keyword sets {w r1 , w r2 ,..., w rn} for each topic term. Then the expression for calculating the expert association score is:
[0101]
[0102] In the finally obtained knowledge base, it contains edges with association scores, edges with hyponymy relationships, and edges with expert association scores, so as to obtain a political theory knowledge graph based on the connection relationships between knowledge.
[0103] Although specific embodiments of the present invention are disclosed for illustrative purposes, the purpose is to help understand the content of the present invention and implement it accordingly. Those skilled in the art can understand that: without departing from the spirit and scope of the present invention and the appended claims, various substitutions, changes, and modifications are possible. Therefore, the present invention should not be limited to the content disclosed in the best embodiments, and the scope of protection claimed by the present invention is subject to the scope defined by the claims.
Claims
1. A method for constructing a knowledge graph, the steps of which include: 1) Automatically extract political theory knowledge from unannotated political theory corpus texts; 2) Screen the political theory knowledge extracted in step 1), and annotate the screened political theory knowledge, as the training text for training the political knowledge extraction model; 3) Use the training text to train the political knowledge extraction model; 4) Use the trained political knowledge extraction model to extract knowledge from the corpus to obtain political theory knowledge; 5) For any two political theory knowledge in the political theory knowledge obtained in step 4), calculate the co-occurrence degree and semantic similarity of the two political theory knowledge in the corpus. If the co-occurrence degree or semantic similarity is not zero, connect an edge between the two political theory knowledge, and calculate the association score between the two political theory knowledge based on the co-occurrence degree and semantic similarity as the weight of the edge, so as to obtain the knowledge graph corresponding to the corpus; 6) Align the knowledge system with the upper and lower structure annotated by experts with the knowledge graph generated in step 5), and integrate the upper and lower relationships between the subject words annotated by experts in the knowledge system into the knowledge graph; among them, the method of integrating the upper and lower relationships between the subject words annotated by experts in the knowledge system into the knowledge graph includes: 61) Knowledge alignment: If the subject words annotated by experts are the same as the extracted subject words in terms of characters, it is considered that the two words are the same political theory knowledge; otherwise, it is considered that the subject words annotated by experts are new political theory knowledge and merge them into the extracted knowledge base; 62) Aggregation of associations between knowledge: Aggregate the associations in the knowledge graph obtained in step 5) and the associations annotated by experts; 63) Association score between the subject term and knowledge: Based on the relevant terms corresponding to the subject term w marked by experts theme and the extracted political theory knowledge w robot The association score is obtained by weighting the sum of the association relationship scores theme and robot ; among them, experts have marked a set of relevant keywords for each subject term; the finally fused knowledge graph contains edges with association scores, edges with hyponymy relationships, and edges with expert association scores.
2. The method according to claim 1, characterized in that, The method for automatically extracting political theory knowledge from unannotated political theory corpus texts includes: 11) Tokenize each sentence S in a political theory corpus text A within the corpus to obtain a token list w = {w1, w2,..., w n} and the corresponding part-of-speech list t = {t1, t2,..., t n}; w n is the nth token in sentence S, and t n is the part of speech of w n ; 12) Combine adjacent k word segments in the word segment list w to obtain multiple alternative phrases k-gram; calculate the tf-idf scores of each k-gram in the political theory corpus text A when k takes different values; 13) Add up the tf-idf scores of each alternative phrase k-gram in each political theory corpus text in the corpus, to obtain the final tf-idf score of the alternative phrase k-gram, and select several alternative phrases with the largest final tf-idf scores as the automatically extracted political theory knowledge.
3. The method according to claim 1, wherein The associated score includes the co-occurrence degree score of two political theory knowledge, the semantic similarity score between two political theory knowledge, and the expert annotation score; among them, for two political theory knowledge i and j, if they co-occur in n1 sentences, n2 paragraphs, and n3 texts of each political theory corpus text in the corpus, then their co-occurrence degree score is C ij =(a * n1 + b * n2 + c * n3) p , where a, b, and c are the weights corresponding to sentence co-occurrence, paragraph co-occurrence, and text co-occurrence, and p is the total number of texts in the corpus; the semantic similarity score S between two political theory knowledge i and j is calculated through a semantic similarity model ij ; if two political theory knowledge i and j are co-occurrence annotated l times by experts in the same sentence, then their expert annotation score is Z ij =z * l; the final associated score of two political theory knowledge i and j is: R ij =x * C ij +s * S ij +z * Z ij .
4. The method according to claim 1 or 2 or 3, characterized in that, The optimization function used when training the political knowledge extraction model based on a large-scale language model is the maximum likelihood optimization function.
5. The method according to claim 1 or 2 or 3, characterized in that The political knowledge extraction model includes a large-scale pre-trained language model BERT and a conditional random field model; input the political theory corpus text into the large-scale pre-trained language model BERT to obtain the word encoding of each character and use it as the input of the conditional random field model. The conditional random field model outputs the probability that the sentence sequence is labeled with different labels; select the label sequence with the largest probability to decode and obtain the political theory knowledge contained in the corresponding sentence.
6. The method according to claim 1, wherein The knowledge system is a three-level knowledge system, including first-level knowledge subject words, second-level knowledge subject words, and third-level knowledge subject words.
7. A server, characterized in that, It includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in any one of the methods recited in claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of any one of the methods recited in claims 1 to 6 are implemented.
Citation Information
Patent Citations
Cross-language multi-source vertical domain knowledge graph construction method
CN112199511A
Domain Expertise Determination
US20100088331A1