Teaching subject knowledge graph construction method based on language model
Through the Bert language model, the knowledge graph construction method of teaching subjects is constructed using attention matrix and cosine distance to determine the relationship words, which solves the problem of high efficiency and low cost of building knowledge graphs in the field of education, and realizes the low-cost and efficient construction of knowledge graphs in teaching subjects.
Patent Information
- Application Number
- CN202510493500.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When building knowledge graphs in the field of education, there are problems of high cost and inefficiency, especially when facing complex and diverse knowledge relationships in different disciplines and grades, traditional methods are difficult to adapt to open relationship extraction.
The pre-trained Bert language model is adopted to extract seed entities in the teaching subject field, build a collection of seed entities, and use the attention matrix and cosine distance to determine the relational words, combine search engines to verify the rationality of the relational words, and build a knowledge graph for teaching subjects.
It has achieved low-cost and efficient construction of teaching subject knowledge graphs, improved the accuracy and practicality of knowledge relationships, and is suitable for the mining of complex and diverse knowledge relationships in the field of education and teaching.
Smart Images

Figure CN120373446A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for constructing a knowledge graph of teaching subjects based on a language model, belonging to the technical fields of computer and natural language processing. Background Art
[0002] Currently, the application of knowledge graphs is becoming more and more extensive. However, as a highly structured data, the construction process of knowledge graphs is time-consuming and laborious. Especially in the field of education, for the knowledge graphs constructed for teaching subjects, due to the possible differences in the relationship types of different disciplines, different grades, and even different chapters, in traditional solutions, usually only fixed types of relationship extraction are performed on unstructured documents to construct knowledge graphs, and for teaching data, these solutions are obviously unrealistic;
[0003] With the rise of large language models, many solutions will try to use large models to assist in generating knowledge graphs; in such solutions, there is a practice of only using prompts to let the large model generate entity relationships, such as giving (s, r) and letting the large model output o. However, due to the lack of constraints from real corpora, the large model is very prone to imperceptible hallucination problems. And if corpora are provided and the large model is used to extract entities and relationships, then under large-scale data, the time-consuming and cost problems will be difficult to control; a more common solution is to collect a large amount of unstructured corpora and complete the construction of the graph through natural language processing technologies such as entity recognition and relationship extraction; however, such solutions often require training a set of entity recognition and relationship extraction models by themselves to achieve better results. On the one hand, this brings a costly annotation process. On the other hand, such models are usually for fixed types of relationships, but if faced with open relationship extraction (that is, the types of relationships are not fixed), this solution is also difficult to handle. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, device, and system for constructing a knowledge graph of teaching subjects based on a language model, which can quickly and low-costly construct a set of subject knowledge graphs by using a pre-trained Bert language model.
[0005] To achieve the above object / to solve the above technical problems, the present invention is implemented by the following technical solutions.
[0006] On the one hand, the present invention provides a method for constructing a knowledge graph of teaching subjects based on a language model, including:
[0007] Extracting corresponding knowledge points in the field of teaching subjects as seed entities and constructing a seed entity set;
[0008] Merging synonyms in the seed entity set;
[0009] Count the seed entities that co-occur in the same paragraph in the merged seed entity set; establish a binary tuple (s, o) for all seed entities with co-occurrences greater than once, and record all statements containing the corresponding binary tuple (s, o). Perform clustering on all statements, and retain several sentences closest to the cluster center in each cluster of statements, where s is the head entity and o is the tail entity;
[0010] Input the binary tuple (s, o) of the seed entities after clustering and the retained statements containing this binary tuple (s, o) into a pre-trained Bert model, and output the relationship word r between the seed entities of the binary tuple (s, o);
[0011] According to the relationship word r and the corresponding binary tuple (s, o), construct a triple (s, r, o) of the corresponding seed entities:
[0012] Among them: The processing steps of the pre-trained Bert model are specifically as follows:
[0013] Extract the attention matrix of the original statement through the attention structure of Transformer, and find the k candidate relationship words most relevant to the binary tuple (s, o) in the original statement through the attention matrix;
[0014] Replace the candidate relationship words in the original statement in turn, calculate the cosine distance between the vector of the replaced statement and the vector of the original statement, and select the candidate relationship word with the largest cosine distance after replacement as the final relationship word;
[0015] Among them: The replacement method is to randomly select one from the existing candidate relationship words and replace the candidate relationship word in the statement.
[0016] Furthermore, extracting the corresponding knowledge points in the field of teaching subjects as seed entities and constructing a seed entity set specifically includes:
[0017] Use a CSV / JSON parser to take the field values that conform to the predetermined format specification from the structured text as the first seed entity;
[0018] Match semi-structured content through regular expressions as the second seed entity;
[0019] Use the jieba word segmentation tool to segment the unstructured text, and use the TF-IDF keyword extraction algorithm to screen keywords as the third seed entity;
[0020] Construct a seed entity set through the first seed entity, the second seed entity, and the third seed entity.
[0021] Furthermore, it is necessary to perform text coreference resolution on the unstructured text, specifically including:
[0022] Use Bert to extract features from sentences with keywords and pronouns in unstructured text, and obtain the word vectors corresponding to the pronouns;
[0023] Calculate the cosine distance between the word vector of the pronoun and all keywords in turn. When the cosine distance is less than the specified threshold, use this keyword to replace the pronoun;
[0024] Among them: The cosine distance calculation formula is: ,
[0025] Among them: is the pronoun word vector, is the keyword word vector;
[0026] represents the dot product of the pronoun word vector and the keyword word vector;
[0027] represents the norm of the pronoun word vector;
[0028] represents the norm of the keyword word vector.
[0029] Furthermore, the merging of synonyms in the seed entity set specifically includes:
[0030] Extract features from all sentences where the seed entities are located, and obtain the word vectors corresponding to the seed entities;
[0031] Perform hierarchical clustering on the obtained word vectors of the seed entities;
[0032] After clustering, select the seed entities in the clusters with the number of entities less than 5 for merging.
[0033] Furthermore, the attention matrix of the sentence is extracted through the attention structure of Transformer, and the k candidate relationship words most relevant to the sentence and the binary group are found, specifically including:
[0034] ;
[0035] < ;
[0036] , i = 1, 2,..., z - 1;
[0037] , j = 1, 2,..., t - 1;
[0038] Among them: is the entity word The relevance to the relational word r is the head entity s or the tail entity ;
[0039] is the th word in the sentence, is the th word in the sentence,
[0040] is the first word in the head entity word t and is the i-th word in the sentence, is the first word in the relational word r and is the j-th word in the sentence;
[0041] z and t are the total number of words of the entity words and the total number of words of the relational word r respectively;
[0042] represents the word and the word in the attention matrix;
[0043] The calculation formula is as follows:
[0044] ;
[0045] Where:
[0046] n is the total number of words in the sentence, and m is the number of iterations;
[0047] comes from the output of the attention structure in the Bert network, that is, the attention matrix , is a symmetric matrix, that is = ;
[0048] Select the k candidate relational words with the highest relevance to the entity word .
[0049] Furthermore, before replacing the candidate relational words, filter the candidate relational words, specifically including:
[0050] By maintaining a stop word list, remove the relational words that appear in it;
[0051] Require that the relational word appears more than 5 times in all corpora, otherwise filter it;
[0052] Using the search engine interface, search with the head entity and the relationship as keywords, and determine whether the tail entity appears within the first 150 characters of the search results. If it exists, keep it; if not, filter it out.
[0053] Further, after obtaining the final relationship words, perform unification processing, specifically including:
[0054] Establish a dictionary mapping or co-occurrence mapping for different expressions of the final relationship words;
[0055] Among them: The dictionary mapping is specifically: maintain a thesaurus, and map relationship words that are synonyms of each other to standard relationship words;
[0056] Co-occurrence mapping: For two triples (s, r1, o) and (s, r2, o), if the two relationships r1 and r2 frequently co-occur under the same seed entities s and o, set the relationships r1 and r2 as synonymous relationship words; and map the final relationship words to the relationship word that appears most frequently.
[0057] In a second aspect, the present invention provides a teaching subject knowledge graph construction device based on a language model, including:
[0058] A construction module for extracting corresponding knowledge points in the teaching subject field as seed entities and constructing a seed entity set;
[0059] A merging module for merging synonyms in the seed entity set;
[0060] A clustering module for counting the seed entities that co-occur in the same paragraph in the merged seed entity set; establish a binary tuple (s, o) for all seed entities with co-occurrence greater than once, and record all statements containing the corresponding binary tuple (s, o), perform clustering processing on all statements, and keep several sentences closest to the cluster center in each cluster of statements, where s is the head entity and o is the tail entity;
[0061] A processing module for inputting the binary tuple (s, o) of the clustered seed entities and the retained statements containing the binary tuple (s, o) into a pre-trained Bert model, and outputting the relationship word r between the binary tuple (s, o) seed entities;
[0062] According to the relationship word r and the corresponding binary tuple (s, o), construct a triple (s, r, o) of the corresponding seed entity:
[0063] Among them: The processing steps of the pre-trained Bert model are specifically:
[0064] Extract the attention matrix of the original statement through the attention structure of the Transformer, and find the k candidate relationship words in the original statement that are most relevant to the binary tuple (s, o) through the attention matrix;
[0065] Replace the candidate relationship words in the original statement in turn, calculate the cosine distance between the vector of the replaced statement and the vector of the original statement, and select the candidate relationship word with the largest cosine distance after replacement as the final relationship word;
[0066] Among them: the replacement method is to randomly select one from the existing candidate relationship words and replace the candidate relationship word in the statement.
[0067] In a third aspect, the present invention provides a teaching subject knowledge graph construction system based on a language model, including:
[0068] A memory for storing computer programs / instructions;
[0069] A processor for executing the computer programs / instructions to implement the steps of the above-mentioned teaching subject knowledge graph construction method based on a language model.
[0070] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer programs / instructions are stored, and when the computer programs / instructions are executed by a processor, the steps of the above-mentioned teaching subject knowledge graph construction method based on a language model are implemented.
[0071] Compared with the prior art, the beneficial effects achieved by the present invention: The present invention innovatively uses the attention matrix generated by Bert to determine candidate relationship words, applies the supervised learning method to the construction of the education and teaching knowledge graph, overcomes the problem of the traditional method requiring a large amount of labeled data, greatly improves the construction efficiency, reduces the cost, and is especially suitable for mining complex and diverse and constantly changing knowledge relationships in the field of education and teaching;
[0072] The present invention determines the final relationship word through carefully designed relationship word replacement and semantic deviation degree evaluation, fully considering the accuracy and rigor requirements of education and teaching knowledge expression, effectively avoiding the wrong selection of relationship words, and ensuring the reliability of knowledge association in the knowledge graph;
[0073] The present invention cleverly uses the search engine interface to verify the rationality of the relationship word, combines the vast information resources of the Internet, further enhances the accuracy and practicality of the knowledge graph relationship, and provides a more accurate and comprehensive knowledge association for education and teaching. Description of the Drawings
[0074] Figure 1 is the flowchart of knowledge graph generation in an embodiment of the present invention;
[0075] Figure 2 Schematic diagram of the attention matrix in an embodiment of the present invention. DETAILED DESCRIPTION
[0076] It should be noted that:
[0077] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. The embodiments of the present invention and the technical features in the embodiments may be combined with each other unless there is a conflict.
[0078] The term "and / or" is only a description of the association relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " generally indicates that the related objects are in an "or" relationship.
[0079] Example 1
[0080] like Figure 1 An embodiment shown in the figure provides a method for constructing a teaching subject knowledge graph based on a language model, including:
[0081] Step 1: Extract the corresponding knowledge points in the teaching subject area as seed entities and construct a seed entity set;
[0082] The extracting of corresponding knowledge points in the teaching subject field as seed entities and constructing a seed entity set specifically includes:
[0083] Step 1.1: Collect structured text through standard API interfaces (such as Wikipedia API); use crawler technology to crawl semi-structured text from encyclopedia websites on the web; use PDFMiner or OCR technology to extract unstructured text from textbook PDF files. For example:
[0084] Parse the structured fields from the CSV term list that comes with the mathematics textbook, and extract fields such as "metaphor", "personification", "subject", and "predicate" as the first seed entities.
[0085] Regular expressions are used to match semi-structured content in encyclopedia pages. For example, matching "{rhetorical device} refers to {description}" can obtain "metaphor", "using similar things to describe another thing" and so on as the second seed entity.
[0086] Step 1.2: Unstructured texts need to be processed for reference disambiguation, including:
[0087] Use Bert to extract features from sentences in unstructured text that contain keywords and pronouns, and obtain the word vectors corresponding to the pronouns;
[0088] Calculate the cosine distance between the word vector of the pronoun and all keywords in turn. When the cosine distance is less than the specified threshold, use this keyword to replace the pronoun;
[0089] Among them: The cosine distance calculation formula is: ,
[0090] Among them: is the pronoun word vector, is the keyword word vector;
[0091] represents the dot product of the pronoun word vector and the keyword word vector;
[0092] represents the modulus of the pronoun word vector;
[0093] represents the modulus of the keyword word vector.
[0094] In the sentence "It connects the noumenon and the figure of speech through words such as 'like' and'seemingly', and this technique can make the language more vivid", "it" refers to "metaphor"; Use the word vectors extracted by Bert to calculate the cosine distance between "it" and the word vectors of other words in the sentence in turn, and find the word with the smallest distance, and finally get its pronoun "metaphor", and replace "it" with "metaphor".
[0095] Step 1.3: Use the CSV / JSON parser to take the field values that conform to the predetermined format specification from the structured text as the first seed entity; match the semi-structured content through regular expressions as the second seed entity; use the jieba word segmentation tool to segment the unstructured text, and use the TF-IDF keyword extraction algorithm to screen keywords as the third seed entity; construct a seed entity set through the first seed entity, the second seed entity and the third seed entity;
[0096] Step 2: Merge the synonyms in the seed entity set, specifically including:
[0097] Extract features from the sentences where all seed entities are located, and obtain the word vectors corresponding to the seed entities;
[0098] Perform hierarchical clustering on the obtained word vectors of the seed entities;
[0099] After clustering, select the seed entities in the clusters with less than 5 entities in the cluster for merging; In the result of hierarchical clustering, it is found that "metaphor" and "analogy" are divided into one cluster and the number of entities in the cluster is less than 5, then "analogy" is merged into the standard entity "metaphor".
[0100] Step 3: Count the seed entities that co-occur in the same paragraph in the merged seed entity set, establish a binary tuple (s, o) for all seed entities with co-occurrences greater than once, and record all the statements containing the corresponding binary tuple (s, o). Perform clustering on all the statements, and retain several sentences closest to the cluster center in each cluster of statements, where s is the head entity and o is the tail entity; when the co-occurrence count of the co-occurring entity pair (metaphor, noumenon) is 20 times, then record all the sentences in which these two entities co-occur, and perform clustering on these sentences. For example, "The three essential elements that a metaphor must possess: noumenon, vehicle, and metaphorical word", "A metaphorical sentence generally consists of three parts: noumenon, vehicle, and metaphorical word.", "A metaphor consists of three parts: noumenon, metaphorical word, and vehicle", etc. Sentences with similar expressions like these will be grouped into one category, and only one or two representative sentences will be selected from the sentences grouped into one category (the purpose is to avoid repeatedly extracting the same relational words).
[0101] Step 4: Input the binary tuple (s, o) of the clustered seed entities and the retained statements containing the binary tuple (s, o) into a pre-trained Bert model, and output the relational word r between the binary tuple (s, o) seed entities;
[0102] Step 4.1: The specific operation of inputting the binary tuple (s, o) of the clustered seed entities and the retained statements containing the binary tuple (s, o) into a pre-trained Bert model is as follows:
[0103] Step 4.11: As Figure 2 shown, extract the attention matrix of the original statement through the attention structure of Transformer, and find the 3 candidate relational words most relevant to the original statement and the binary tuple (s, o), specifically including:
[0104] The specific expression is:
[0105] ;
[0106] < ;
[0107] , i = 1, 2,..., z - 1;
[0108] , j = 1, 2,..., t - 1;
[0109] Where: is the relevance between the entity word and the relational word r, is the head entity s or the tail entity ;
[0110] is the th word in the sentence, is the th word in the sentence,
[0111] is the first word in the head entity word t and is the i-th word in the sentence, is the first word in the relation word r and is the j-th word in the sentence;
[0112] z and t are the total number of words of the entity word and the total number of words of the relation word r respectively;
[0113] represents the word and the word in the attention matrix;
[0114] The calculation formula is as follows:
[0115] ;
[0116] Where:
[0117] n is the total number of words in the sentence, and m is the number of iterations;
[0118] comes from the output of the attention structure in the Bert network, that is, the attention matrix , is a symmetric matrix, that is = ;
[0119] Select the 3 candidate relation words with the highest correlation with the entity word .
[0120] Step 4.12: Replace the candidate relation words in the original sentence in turn, calculate the cosine distance between the sentence vector after replacement and the original sentence vector, and select the candidate relation word with the largest cosine distance after replacement as the final relation word, specifically including:
[0121] Replace the candidate relation words with a word randomly selected from the existing relation words in turn, and calculate the cosine distance between the sentence vector after replacement and the original sentence vector using Bert (i.e., the offset distance), and its calculation formula is:
[0122] ;
[0123] Where: and represent the norms of the sentence vector after replacement and the original sentence vector respectively, represents the dot product of the sentence vector after replacement and the original sentence vector; when is very large, it indicates that the semantic difference between the replaced sentence and the original sentence is relatively large. At this time, the replaced word is determined as the relational word.
[0124] When the input binary tuple (metaphor, noumenon) and the sentence "A metaphor consists of three parts: noumenon, metaphorical word, and vehicle"; locate the key relational words through the attention matrix of Bert, such as "metaphorical word", "vehicle", "consist of". Replacement verification: For example, replace the words "metaphorical word", "vehicle", and "consist of" in the original sentence with "classical poetry" respectively. By calculating the semantic shift of the sentence, it is found that replacing "consist of" has the greatest impact on the semantics of the original sentence. Therefore, "consist of" is used as the final candidate relational word.
[0125] Step 4.13: Before replacing the candidate relational word, filter the candidate relational word, specifically including:
[0126] First, by maintaining a stop word list, remove the candidate relational words that appear in it;
[0127] Secondly, require that the candidate relational word appears more than 5 times in all corpora, otherwise filter it;
[0128] Finally, use the search engine interface to search with the head entity and the relational word as keywords, and judge whether the tail entity appears within the first 150 characters of the search results. If it exists, keep it; if it does not exist, filter it.
[0129] Step 4.14: After obtaining the final relational word, perform a unification process, specifically including:
[0130] Establish a dictionary mapping or co-occurrence mapping for different expressions of the final relational word;
[0131] Among them: The dictionary mapping is specifically: maintain a synonym table, and map the relational words that are synonyms of each other to the standard relational word;
[0132] Co-occurrence mapping: For two triples (s, r1, o) and (s, r2, o), if the two relations r1 and r2 frequently co-occur under the same seed entities s and o, set the relations r1 and r2 as synonymous relational words; and map the final relational word to the most frequently occurring relational word.
[0133] Step 4.2: According to the relational word r and the corresponding binary tuple (s, o), construct the corresponding triple (s, r, o) of the seed entity.
[0134] Example 2
[0135] This embodiment provides a teaching subject knowledge graph construction device based on a language model, including:
[0136] A construction module, configured to extract corresponding knowledge points in the teaching subject field as seed entities and construct a seed entity set;
[0137] A merging module, configured to merge synonyms in the seed entity set;
[0138] A clustering module, configured to count the seed entities that co-occur in the same paragraph in the merged seed entity set; establish a binary tuple (s, o) for all seed entities with co-occurrence greater than once, and record all statements containing the corresponding binary tuple (s, o), perform clustering processing on all statements, and retain several sentences closest to the cluster center in each cluster of statements, where s is the head entity and o is the tail entity;
[0139] A processing module, configured to input the binary tuple (s, o) of the clustered seed entities and the retained statements containing the binary tuple (s, o) into a pre-trained Bert model, and output the relationship word r between the binary tuple (s, o) seed entities;
[0140] According to the relationship word r and the corresponding binary tuple (s, o), construct a triple (s, r, o) of the corresponding seed entity:
[0141] Wherein: The processing steps of the pre-trained Bert model are specifically:
[0142] Extract the attention matrix of the original statement through the attention structure of Transformer, and find the k candidate relationship words most relevant to the binary tuple (s, o) from the attention matrix;
[0143] Successively replace the candidate relationship words in the original statement, calculate the cosine distance between the vector of the replaced statement and the vector of the original statement, and select the candidate relationship word with the largest cosine distance after replacement as the final relationship word;
[0144] Wherein: The replacement method is to randomly extract one from the existing candidate relationship words and replace the candidate relationship word in the statement.
[0145] Embodiment 3
[0146] This embodiment provides a teaching subject knowledge graph construction system based on a language model, including:
[0147] A memory, configured to store computer programs / instructions;
[0148] A processor, configured to execute the computer programs / instructions to implement the steps of the above-mentioned teaching subject knowledge graph construction method based on a language model.
[0149] Example 4
[0150] This embodiment provides a computer-readable storage medium, on which computer programs / instructions are stored. When the computer programs / instructions are executed by a processor, the steps of the above-mentioned method for constructing a knowledge graph of teaching subjects based on a language model are implemented.
[0151] The present invention can automatically complete the entities and relationships of the knowledge graph by utilizing the capabilities of a large language model, make full use of the information mined by the pre-trained large language model in a large amount of text, quickly establish a set of knowledge graphs, and this process is completely unsupervised without fine-tuning the language model, which is suitable for quickly establishing a set of knowledge graphs in the field of teaching subjects.
[0152] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0153] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0154] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing the instructions executed on the computer or other programmable apparatus for implementing the steps of the function specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps of the function specified in one block or a plurality of blocks.
[0156] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.
Claims
1. A method for constructing a knowledge graph of teaching subjects based on a language model, characterized in that Including: Extract the corresponding knowledge points in the teaching subject field as seed entities, and construct a set of seed entities; Merge the synonyms in the set of seed entities; Count the seed entities that co-occur in the same paragraph in the merged set of seed entities; establish a binary tuple (s, o) for all seed entities with co-occurrences greater than once, and record all statements containing the corresponding binary tuple (s, o), and perform clustering processing on all statements, retaining several sentences closest to the cluster center in each cluster of statements, where: s is the head entity and o is the tail entity; Input the binary tuple (s, o) of the seed entities after clustering and the retained statements containing the binary tuple (s, o) into a pre-trained Bert model, and output the relationship word r between the binary tuple (s, o) seed entities; Construct a triple (s, r, o) of the corresponding seed entities according to the relationship word r and the corresponding binary tuple (s, o): Where: The processing steps of the pre-trained Bert model are specifically: Extract the attention matrix of the original statement through the attention structure of Transformer, and find the k candidate relationship words most relevant to the binary tuple (s, o) from the attention matrix; Replace the candidate relationship words in the original statement in turn, calculate the cosine distance between the vector of the replaced statement and the vector of the original statement, and select the candidate relationship word with the largest cosine distance after replacement as the final relationship word; Where: The replacement method is to randomly select one from the existing candidate relationship words and replace the candidate relationship word in the statement.
2. The method for constructing a teaching subject knowledge graph based on a language model according to claim 1, wherein Specifically including: The extraction of the corresponding knowledge points in the teaching subject field as seed entities and the construction of a set of seed entities specifically include: Use a CSV / JSON parser to take the field values that conform to the predetermined format specification from the structured text as the first seed entities; Match semi-structured content through regular expressions as the second seed entities; Use the jieba word segmentation tool to segment the unstructured text, and use the TF-IDF keyword extraction algorithm to screen keywords as the third seed entities; Construct a set of seed entities through the first seed entities, the second seed entities, and the third seed entities.
3. The method for constructing a teaching subject knowledge graph based on a language model according to claim 2, wherein The unstructured text needs to be processed for text reference disambiguation, specifically including: Use Bert to extract features from the sentences in the unstructured text that contain keywords and pronouns, and obtain the word vectors corresponding to the pronouns; Calculate the cosine distance between the word vectors of the pronouns and all keywords in turn. When the cosine distance is less than the specified threshold, use the keyword to replace the pronoun; Among them: The cosine distance calculation formula is: , Wherein: is the pronoun word vector, is the keyword word vector; Indicates the dot product of the pronoun word vector and the keyword word vector; Denote the norm of the pronoun vector; Represents the norm of the keyword word vector.
4. The method for constructing a teaching subject knowledge graph based on a language model according to claim 1, wherein The merging of synonyms in the set of seed entities specifically includes: Extract features from the sentences where all seed entities are located, and obtain the word vectors corresponding to the seed entities; Perform hierarchical clustering processing on the obtained word vectors of the seed entities; After clustering, select the seed entities in the clusters with the number of entities less than 5 in the cluster for merging.
5. The method for constructing a teaching subject knowledge graph based on a language model according to claim 1, wherein The attention matrix of the statement is extracted through the attention structure of the Transformer, and k candidate relation words r that are most relevant to the statement and the binary tuple (s, o) are found through the attention matrix. The specific expression is as follows: ; < ; , where i = 1, 2, …… z - 1; , j = 1, 2, …… t - 1; Wherein: is an entity word the relevance to the relational word r is the head entity s or the tail entity ; is the th character in the statement, is the th character in the statement, is the first character in the head entity word t and is the i-th character in the sentence, is the first character in the relation word r and is the j-th character in the sentence; z and t are the total number of entity words and the total number of relational words r, respectively; Representing character And character The distance value in the attention matrix; The calculation formula is as follows: ; Where: n is the total number of words in the sentence, and m is the number of iterations; Output from the attention structure in the Bert network, i.e., the attention matrix , is a symmetric matrix, i.e., = ; Select the k candidate relationship words with the highest relevance to the entity word .
6. The method for constructing a teaching subject knowledge graph based on a language model according to claim 5, wherein Before replacing the candidate relation words, filtering processing is performed on the candidate relation words, specifically including: By maintaining a stop word list, removing the candidate relation words that appear in it; Requiring that the candidate relation words appear more than 5 times in all corpora, otherwise filtering; Using the search engine interface, searching with the head entity and the relation word as keywords, and judging whether the tail entity appears within the first 150 characters of the search results. If it exists, it is retained; if not, it is filtered.
7. The method for constructing a teaching subject knowledge graph based on a language model according to claim 1, wherein After obtaining the final relation words, uniform processing is performed, specifically including: Establishing a dictionary mapping or co-occurrence mapping for different expressions of the final relation words; Among them: The dictionary mapping is specifically: maintaining a synonym table, and mapping relation words that are synonyms of each other to standard relation words; Co-occurrence mapping: For two triples (s, r1, o) and (s, r2, o), if the two relations r1 and r2 co-occur frequently under the same seed entities s and o, set the relations r1 and r2 as synonymous relation words; and map the final relation words to the relation word that appears most frequently.
8. An apparatus for constructing a knowledge graph of teaching subjects based on a language model, characterized in that Including: A construction module for extracting corresponding knowledge points in the teaching subject field as seed entities and constructing a seed entity set; A merging module for merging synonyms in the seed entity set; A clustering module for counting the seed entities that co-occur in the same paragraph in the merged seed entity set; establishing a binary tuple (s, o) for all seed entities with co-occurrence greater than once, and recording all statements containing the corresponding binary tuple (s, o), performing clustering processing on all statements, and retaining several sentences closest to the cluster center in each cluster of statements, where s is the head entity and o is the tail entity; A processing module for inputting the binary tuple (s, o) of the clustered seed entities and the retained statements containing the binary tuple (s, o) into a pre-trained Bert model, and outputting the relation word r between the binary tuple (s, o) seed entities; According to the relation word r and the corresponding binary tuple (s, o), constructing a triple (s, r, o) of the corresponding seed entity: Among them: The processing steps of the pre-trained Bert model are specifically: Extracting the attention matrix of the original statement through the attention structure of the Transformer, and finding k candidate relation words that are most relevant to the original statement and the binary tuple (s, o) through the attention matrix; Successively replacing the candidate relation words in the original statement, calculating the cosine distance between the vector of the replaced statement and the vector of the original statement, and selecting the candidate relation word with the largest cosine distance after replacement as the final relation word; Among them: The replacement method is to randomly select one from the existing candidate relation words and replace the candidate relation word in the statement.
9. A teaching subject knowledge graph construction system based on a language model, characterized in that Including: A memory for storing computer programs / instructions; A processor for executing the computer program / instructions to implement the steps of the method for constructing a knowledge graph of teaching subjects based on a language model according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method for constructing a knowledge graph of teaching subjects based on a language model according to any one of claims 1-7 are implemented.
Citation Information
Cited By
Key reference resolution method, device, medium and product in speech-to-text
CN120748408B