A method for constructing knowledge graph of ethnic cultural information resources
By using the Chinese word segmentation system and user-defined lexicon library, we extract word segmentation and attributes of ethnic cultural information resources, and build a knowledge graph of ethnic cultural information resources, solving the problem that the existing technology cannot effectively build a Chinese knowledge graph, and achieving effective management and utilization of ethnic cultural information.
Patent Information
- Application Number
- CN201910042744.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-01-17
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2039-01-17
AI Technical Summary
The existing technology cannot effectively build a knowledge map of ethnic cultural information resources, especially in the Chinese environment, and lacks a method for construction of ethnic cultural information resources.
The Chinese word segmentation system and user-defined vocabulary database are used to perform word segmentation, part-of-speech annotation, manual word segmentation and attribute extraction of ethnic minority word segmentation data to construct a domain knowledge graph, and through repetitive detection and storage, a final knowledge graph is formed.
By constructing a knowledge graph of national cultural information resources, the national cultural information on the Internet is expressed in a form closer to the human cognitive world, which is easy to manage and utilize, and the problem of building a Chinese knowledge graph is solved.
Smart Images

Figure BDA0001948115290000021 
Figure BDA0001948115290000041 
Figure FDF0000014157690000011
Abstract
Description
Technical Field
[0001] The present invention relates to a method for constructing a knowledge graph of ethnic cultural information resources, and belongs to the technical field of knowledge graphs. Background Art
[0002] National culture is the spiritual wealth of a nation. Preserving national culture records can not only allow future generations to inherit excellent culture, but also make the national culture leave a strong mark in history. Through digital technology, traditional national culture is transformed into digital coding form, and then stored, transmitted, copied, reproduced and even created, making it appear as a "living culture". This will become an inevitable trend in protecting national cultural heritage and an innovative idea for the digital development of ethnic minority culture.
[0003] Knowledge graph technology has attracted widespread attention from scholars in recent years. Through knowledge graphs, information on the Internet can be expressed in a form that is closer to the human cognitive world, and it provides a better way to organize, manage and use information.
[0004] In order to promote the gradual development of national cultural information resources, it is necessary to construct a knowledge graph of national cultural information resources. However, the current foreign methods for constructing English knowledge graphs cannot be fully applicable to the construction of Chinese knowledge graphs. There are also few knowledge graphs for national cultural information resources. There is an urgent need for a method to construct a knowledge graph for national cultural information resources. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a method for constructing a knowledge graph of ethnic cultural information resources to solve the above-mentioned problems.
[0006] The technical solution of the present invention is: a method for constructing a knowledge graph of ethnic cultural information resources, firstly using a Chinese word segmentation system and a user-defined vocabulary to segment and tag the entry data in the collected ethnic minority dictionary data, then detecting the segmented and part-of-speech tagged entry data, if the number of consecutive segmented words is not less than a set threshold, performing manual segmentation, and adding the manual segmentation results to the user-defined vocabulary of the Chinese word segmentation system until there are no new words, then extracting attributes from the correctly segmented entry data to construct a domain knowledge graph, again performing a repeatability check on the domain knowledge graph, deleting duplicate data, linking the stored domain knowledge graph with resources, and finally storing it.
[0007] The specific steps are:
[0008] Step 1: Collect minority vocabulary data, build a minority vocabulary database, use the Chinese word segmentation system and user-defined vocabulary to segment and tag the vocabulary data in the collected minority vocabulary database, and remove punctuation marks;
[0009] Step 2: Then, the data after word segmentation and part-of-speech tagging is tested. If the number of consecutive word segmentations is not less than the set threshold, manual word segmentation is performed, and the manual word segmentation results are added to the user-defined word library of the Chinese word segmentation system. Step 1 is repeated until there are no new words.
[0010] Step 3: Extract attributes from the correctly segmented data to build a domain knowledge graph;
[0011] Step 4: Check the domain knowledge graph for duplication, delete duplicate data, and store it;
[0012] Step 5: Link the stored domain knowledge graph with resources.
[0013] The word segmentation system in step 1 and step 2 is the NLPIR Chinese word segmentation system
[0014] The specific method for detecting the text data after word segmentation and part-of-speech tagging in step 2 is:
[0015] ①Define the word segmentation result set S(S1,S2,……,S m );
[0016] ②For each word segmentation result S in the set S i Count the words and get the result of the word count of the set C (C1, C2, ..., C m ), where C i =len(S i ), and 1≤i≤m;
[0017] ③Set the threshold k to satisfy 2≤k≤m;
[0018] ④ Extract a subset P from S, P satisfies equations (1) and (2)
[0019]
[0020] j-i+1≤k<m (2)
[0021] Description in S i To S j There are k consecutive words with the number of characters being 1 at the position of , and by setting the k value, it is considered that the consecutive words with the number of characters being 1 are a new word x, x = {S i ,S i+1 …S i+k},S i ∈S;
[0022] ④ Define the new word set W as W = (x1, x2…x n) and manually review W lines. If they are new words, they are added to the user-defined word library.
[0023] The threshold k is set from large to small. When it is set for the first time, k=m, and then it decreases gradually until k=1. After each threshold setting, step 2 is repeated until all new words are added to the user-defined vocabulary.
[0024] The purpose of the manual word segmentation operation in step 2 is to determine whether a combination of k consecutive words, all of which are single characters, is a new word.
[0025] In step 3, the attribute extraction is performed one by one according to the word segmentation results and part-of-speech tagging, and all the contents are subjected to attribute extraction, and the attribute names are marked to form a triple of "topic-attribute name-attribute value", i.e., the knowledge graph.
[0026] The repeatability tests in step 4 are divided into the following types:
[0027] Type 1: The same attribute of the same entity has multiple attribute values. If an attribute value contains other attribute values, this eliminates the contained attribute value;
[0028] Type 2: The same attribute of the same entity has multiple attribute values. If the attribute values are mutually exclusive, the number of attribute values is used to determine the value. The value with more attribute values is retained and submitted for manual review.
[0029] Type 3: The same attribute of the same entity has multiple attribute values. If the attribute values are mutually exclusive and the number of attribute values is the same, they will be submitted for manual review.
[0030] The storage of the domain knowledge graph in step 4 simulates the way in which a graph database (such as Neo4j) stores knowledge graphs using a relational database.
[0031] The relational database structure is designed as follows:
[0032] Node table( serial number , node name, node label)
[0033] Entity Table( serial number , node number, entity name)
[0034] Attribute Name Table ( serial number , node number, entity number, attribute name)
[0035] Attribute Value Table( serial number , node number, entity number, attribute name number, attribute value)
[0036] Relationship table( serial number , starting node number, target node number, relationship)
[0037] The beneficial effect of the present invention is that by constructing a knowledge graph of ethnic cultural information resources, the ethnic cultural information on the Internet is expressed in a form that is closer to the human cognitive world, making it easier to manage and utilize the ethnic cultural information. DETAILED DESCRIPTION
[0038] The present invention will be further described below in conjunction with specific implementation modes.
[0039] A method for constructing a knowledge graph of ethnic cultural information resources, comprising:
[0040] Step 1: Collect minority vocabulary data, build a minority vocabulary database, use the Chinese word segmentation system and user-defined vocabulary to segment and tag the vocabulary data in the collected minority vocabulary database, and remove punctuation marks;
[0041] Step 2: Then, the data after word segmentation and part-of-speech tagging is tested. If the number of consecutive word segmentations is not less than the set threshold, manual word segmentation is performed, and the manual word segmentation results are added to the user-defined word library of the Chinese word segmentation system. Step 1 is repeated until there are no new words.
[0042] Step 3: Extract attributes from the correctly segmented data to build a domain knowledge graph;
[0043] Step 4: Check the domain knowledge graph for duplication, delete duplicate data, and store it;
[0044] Step 5: Link the stored domain knowledge graph with resources.
[0045] The word segmentation system in step 1 and step 2 is the NLPIR Chinese word segmentation system
[0046] The specific method for detecting the text data after word segmentation and part-of-speech tagging in step 2 is:
[0047] ①Define the word segmentation result set S(S1,S2,……,S m );
[0048] ②For each word segmentation result S in the set S i Count the words and get the result of the word count of the set C (C1, C2, ..., C m ), where C i =len(S i ), and 1≤i≤m;
[0049] ③Set the threshold k to satisfy 2≤k≤m;
[0050] ④ Extract a subset P from S, P satisfies equations (1) and (2)
[0051]
[0052] j-i+1≤k<m (2)
[0053] Description in S i To S j There are k consecutive words with the number of characters being 1 at the position of , and by setting the k value, it is considered that the consecutive words with the number of characters being 1 are a new word x, x = {S i ,S i+1 …S i+k},S i ∈S;
[0054] ④ Define the new word set W as W = (x1, x2…x n ) and manually review W lines. If they are new words, they are added to the user-defined word library.
[0055] The threshold k is set from large to small. When it is set for the first time, k=m, and then it decreases gradually until k=1. After each threshold setting, step 2 is repeated until all new words are added to the user-defined vocabulary.
[0056] In step 3, the attribute extraction is performed one by one according to the word segmentation results and part-of-speech tagging, and all the contents are subjected to attribute extraction, and the attribute names are marked to form a triple of "topic-attribute name-attribute value", i.e., the knowledge graph.
[0057] The repeatability tests in step 4 are divided into the following types:
[0058] Type 1: The same attribute of the same entity has multiple attribute values. If an attribute value contains other attribute values, this eliminates the contained attribute value;
[0059] Type 2: The same attribute of the same entity has multiple attribute values. If the attribute values are mutually exclusive, the number of attribute values is used to determine the value. The value with more attribute values is retained and submitted for manual review.
[0060] Type 3: The same attribute of the same entity has multiple attribute values. If the attribute values are mutually exclusive and the number of attribute values is the same, they will be submitted for manual review.
[0061] The storage of the domain knowledge graph in step 4 simulates the way of storing knowledge graphs in graph databases using relational databases.
[0062] Example 1: Entry content: [Mengpeng River] A tributary of the right bank of the Nanpeng River. Located in Zhenkang County, Lincang City, it borders the Nujiang River in the north and Myanmar in the west and south.
[0063] 1. Word Segmentation: Use a Chinese word segmentation system to segment the entry into: "
/ Meng / Peng / River /
[Punctuation Mark] / Meng[Noun] / Peng[Verb] / River[Noun] /
[0064] 2. Detection: Define the word segmentation result set as S(Meng, Peng, River, South, Peng, River, Right, Shore, Tributary, Located, Lincang City, Zhenkang County, North, and, Nu River, Adjacent, West, West, and, South, Myanmar, Border). Count the number of characters for each word segmentation result in set S to obtain the set of character counts C(1, 1, 1, 1, 1, 1, 1, 1, 2, 2, 3, 3, 2, 1, 2, 2, 2, 2, 1, 2, 2, 2). From this, we know that m = 22. Set the k value to 22 and perform operations according to step 2 until the k value decreases to 3. When both 3 - 1 + 1 ≤ 3 < 22 are satisfied, consider the consecutive words with a character count of 1 as a new word x = {Meng, Peng, River}, that is, find that "Meng / Peng / River" is a set where all consecutive single characters have a count of 3. Define all the found words as set W = (Mengpeng River, Nanpeng River). After manual review, confirm that these are proper nouns and add them to the user-defined vocabulary.
[0065] 3. The result after re-word segmentation is: "Mengpeng River / Southpeng River / right / shore / tributary / Located in Lincang City / Zhenkang County / north / and / Nu River / adjacent / west / west / and / south / Myanmar / border" and there is no longer a set of consecutive single characters with a count of 2.
[0066] 4. Use the segmentation results to perform attribute annotation. For example, "Mengpeng River / Nanpeng River / right / bank / tributary / located / Lincang City / Zhenkang County / north / adjacent to / Nujiang River / west / west / and / south / Myanmar / border" forms a series of triples after segmentation and annotation:
[0067] ① Mengpeng River - affiliated river - tributary of the right bank of Nanpeng River;
[0068] ② Mengpeng River - Address - Zhenkang County, Lincang City, bordering Nujiang River in the north and Myanmar in the west and south;
[0069] ③ Mengpeng River - adjacent river - adjacent to Nujiang River in the north;
[0070] ④ Mengpeng River - border area - west, west and southern Myanmar border;
[0071] 5. Perform duplicate attribute detection.
[0072] Type 1: The same attribute of the same entity has multiple attribute values. If an attribute value contains other attribute values, this eliminates the contained attribute value. For example:
[0073] ① Mengpeng River - Address - Zhenkang County, Lincang City, adjacent to Nujiang River in the north and bordering Myanmar in the west and south;
[0074] ② Mengpeng River——Address——Zhenkang County, Lincang City;
[0075] Then, eliminate ② and keep ①;
[0076] Type 2: The same attribute of the same entity has multiple attribute values. If the attribute values are mutually exclusive, the number of attribute values will be used for judgment. The one with more attribute values will be retained and submitted for manual review. For example:
[0077] ① Mengpeng River - Address - Zhenkang County, Lincang City, adjacent to Nujiang River in the north and bordering Myanmar in the west and south;
[0078] ② Mengpeng River——Address——Zhenkang County, Lincang City;
[0079] ③ Mengpeng River - Address - Cangyuan County, Lincang City;
[0080] Then, eliminate ③, retain ①②, and submit for manual review, and use other information for supplementary verification;
[0081] Type 3: The same attribute of the same entity has multiple attribute values. If the attribute values are mutually exclusive and the number of attribute values is the same, then all are submitted for manual review. For example:
[0082] ① Mengpeng River——Address——Zhenkang County, Lincang City;
[0083] ②Mengpeng River - Address - Cangyuan County, Lincang City;
[0084] Then it is fully submitted to manual review and supplemented and verified using other materials.
[0085] 6. Link the knowledge graph with resources. Since all knowledge graphs are constructed by extracting from resources, a unique resource address for each resource is formed. Hyperlinks are added to the resources for each attribute of the knowledge graph for attribute verification and resource viewing.
[0086] 7. Store the knowledge graph using a relational database. For example:
[0087] Node table: N001, River, river; N002, Land, land
[0088] Entity table: E001, N001, Mengpeng River
[0089] Attribute name table: P001, N001, E001, Address
[0090] Attribute value table: V001, N001, E001, P001, Zhenkang County, Lincang City
[0091] Relationship table: R001, N001, N002, Irrigation
[0092] Example 2: Entry content: [Kaxie], a Lahu name, which is the "Lu Chang" in the Wa name and was an official position name in the Wa region during the period of the Three Buddhas.
[0093] 1. Word segmentation: Use a Chinese word segmentation system to segment the entry into: "[
/ Ka / Xie /
[Punctuation mark] / Ka[Noun] / Xie[Classifier] /
[0094] 2. Detection: Define the set of word segmentation results as S (Ka, Xie, La, Hu, Ming, Ji, Wa, Ming, Zhong, De, Lu, Chang, Wei, San, Buddha, Period, Wa Ethnic Group, Region, Official Position, Name). Count the number of characters for each word segmentation result in set S to obtain the set word count result C (1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 1, 2, 1). It can be known from this that m = 21. Set the k value to 21 and perform operations according to step 2 until the k value is reduced to 2, and 2 - 1 + 1 ≤ 2 < 21 are both satisfied. Consider the word segmentations with a continuous character count of 1 as a new word x = {Ka, Xie}, that is, find that "Ka / Xie" is a set where the continuous single characters are both 2. Define all the found words as set W = (Ka Xie, La Hu, Lu Chang). After manual review, confirm that these are proper nouns and add them to the user-defined word library.
[0095] 3. The result after re-word segmentation is: "Ka Xie / La Hu / Ming / / Ji / Wa / Ming / Zhong / De / Lu Chang / Wei / San / Buddha / Period / Wa Ethnic Group / Region / De / Official Position / Name" and there is no longer a set of continuous single characters with a count of 2.
[0096] 4. Use the word segmentation results for attribute annotation. For example, after word segmentation and annotation of "Ka Xie / La Hu / Ming / Ji / Wa / Ming / Zhong / De / Lu Chang / Wei / San / Buddha / Period / Wa Ethnic Group / Region / De / Official Position / Name", a series of triples are formed:
[0097] ① Ka Xie - Source Ethnic Group - La Hu Ming;
[0098] ② Ka Xie - Wa Name - Lu Chang;
[0099] ③ Ka Xie - Period - Three Buddha Period;
[0100] ④ Ka Xie - Region - Wa Ethnic Group Region
[0101] ⑤ Ka Xie - Explanation - Official Position Name in the Three Buddha Period of the Wa Ethnic Group Region;
[0102] 5. Perform duplicate attribute detection.
[0103] Type 1: If there are multiple attribute values for the same attribute of the same entity and one of the attribute values contains other attribute values, eliminate the included attribute value. For example:
[0104] ① Ka Xie - Explanation - Official Position Name in the Three Buddha Period of the Wa Ethnic Group Region;
[0105] ② Ka Xie - Explanation - Official Position Name;
[0106] Then, eliminate ② and retain ①;
[0107] Type 2: The same attribute of the same entity has multiple attribute values. If the attribute values are mutually exclusive, the number of attribute values will be used for judgment. The one with more attribute values will be retained and submitted for manual review. For example:
[0108] ① Kaxie——Explanation——The name of an official position in the Wa area during the Three Buddhas period;
[0109] ②Ka Xie——Explanation——The name of an official position in the Wa area;
[0110] ③Ka Xie——Explanation——Official title in Cangyuan area;
[0111] Then, eliminate ③, retain ①②, and submit for manual review, and use other information for supplementary verification;
[0112] Type 3: The same attribute of the same entity has multiple attribute values. If the attribute values are mutually exclusive and the number of attribute values is the same, then all cases will be submitted for manual review. For example:
[0113] ① Kaxie - area - Wa area;
[0114] ②Kaxie-area-Cangyuan area;
[0115] It will be submitted entirely to manual review, and other information will be used for additional verification.
[0116] 6. Link the knowledge graph with resources. Since all knowledge graphs are constructed by extracting resources, a unique resource address is formed for each resource. For each attribute of the knowledge graph, a hyperlink is added to the resource to facilitate attribute verification and resource viewing.
[0117] 7. Use relational database to store knowledge graph. For example:
[0118] Node table: N001, official title, river; N002, location
[0119] Entity table: E001, N001, K001
[0120] Attribute name table: P001, N001, E001, period
[0121] Attribute value table: V001, N001, E001, P001, Three Buddha Period
[0122] Relationship table: R001, N001, N002, subordinate
[0123] The specific embodiments of the present invention have been described in detail above, but the present invention is not limited to the above embodiments, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.
Claims
1. A method for constructing a knowledge graph of ethnic cultural information resources, characterized by: Step 1: Collect minority vocabulary data, build a minority vocabulary database, use the Chinese word segmentation system and user-defined vocabulary to segment and tag the vocabulary data in the collected minority vocabulary database, and remove punctuation marks; Step 2: Then, the data after word segmentation and part-of-speech tagging is tested. If the number of consecutive word segmentations is not less than the set threshold, manual word segmentation is performed, and the manual word segmentation results are added to the user-defined word library of the Chinese word segmentation system. Step 1 is repeated until there are no new words. Step 3: Extract attributes from the correctly segmented data to build a domain knowledge graph; Step 4: Check the domain knowledge graph for duplication, delete duplicate data, and store it; Step 5: Link the stored domain knowledge graph with resources; The word segmentation system in step 1 and step 2 is the NLPIR Chinese word segmentation system; The specific method for detecting the text data after word segmentation and part-of-speech tagging in step 2 is: ①Define the word segmentation result set S(S1,S2,……,S m ); ②For each word segmentation result S in the set S i Count the words and get the result of the word count of the set C (C1, C2, ..., C m ), where C i =len(S i ), and 1≤i≤m; ③Set the threshold k to satisfy 2≤k≤m; ④ Extract a subset P from S, P satisfies equations (1) and (2) j-i+1≤k<m (2) Description in S i To S j There are k consecutive words with the number of characters being 1 at the position of , and by setting the k value, it is considered that the consecutive words with the number of characters being 1 are a new word x, x = {S i ,S i+1 …S i+k },S i ∈S; ④ Define the new word set W as W = (x1, x2…x n ), and manually review W lines. If they are new words, they are added to the user-defined word library; The threshold k is set from large to small. When it is set for the first time, k=m, and then it decreases gradually until k=1. After each threshold is set, step 2 is repeated until all new words are added to the user-defined vocabulary. The repeatability tests in step 4 are divided into the following types: Type 1: The same attribute of the same entity has multiple attribute values. If an attribute value contains other attribute values, this eliminates the contained attribute value; Type 2: The same attribute of the same entity has multiple attribute values. If the attribute values are mutually exclusive, the number of attribute values is used to determine the value. The value with more attribute values is retained and submitted for manual review. Type 3: The same attribute of the same entity has multiple attribute values. If the attribute values are mutually exclusive and the number of attribute values is the same, they will be submitted for manual review.
2. The method for constructing a knowledge graph of ethnic cultural information resources according to claim 1, characterized in that: In step 3, the attribute extraction is classified one by one according to the word segmentation results and part-of-speech tagging, and all the contents are subjected to attribute extraction, and the attribute names are marked to form a triple of "topic-attribute name-attribute value", that is, a knowledge graph.
3. The method for constructing a knowledge graph of ethnic cultural information resources according to claim 1, characterized in that: The storage of the domain knowledge graph in step 4 simulates the way of storing knowledge graphs in graph databases using relational databases.
Citation Information
Patent Citations
Tibetan language entity knowledge information extraction method
CN104133848A
Natural-language-processing method of ancient-tablature and ancient-culture knowledge graph
CN108509420A
Determination of product attributes and values using a product entity graph
US9607098B2
Method and device for creating knowledge graph
CN107665252A
Knowledge graph construction method and system
CN108694177A