Knowledge graph construction method, device, equipment and product
By extracting the catalog from the teaching materials and matching it with the text, the teaching knowledge graph is constructed, and the accuracy and inefficiency caused by manual annotation is solved, and a more efficient and accurate knowledge graph construction is achieved.
Patent Information
- Application Number
- CN202510847893.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-24
AI Technical Summary
In the prior art, the construction of educational knowledge graphs relies on manual annotation, and there are problems of insufficient accuracy and inefficiency.
By extracting the catalog from teaching materials, generating a knowledge point hierarchical catalog, matching it with the content of the text, extracting entities and relationships, and building a teaching knowledge graph.
It improves the accuracy and efficiency of building a knowledge graph, and can more accurately reflect the logical levels and correlation between knowledge points, which is suitable for the needs of teaching scenarios.
Smart Images

Figure CN120353940A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular, to a method, apparatus, device, and product for constructing a knowledge graph. Background Art
[0002] A knowledge graph can integrate massive and scattered information into a structured and interconnected network, thereby improving the efficiency and accuracy of information retrieval, data analysis, and intelligent decision-making. Therefore, constructing a high-quality knowledge graph is of great significance in multiple fields.
[0003] Currently, the construction of knowledge graphs mainly relies on manual annotation and collation by domain experts. Taking the education field as an example, experts usually identify and define key knowledge points and concepts based on teaching syllabuses, teaching materials, and their own teaching experience, and sort out the relationships between them, thereby constructing the structure of the knowledge graph. However, this manual-dependent method often has problems of insufficient accuracy and low efficiency. Summary of the Invention
[0004] Based on the above technical status quo, this application provides a method, apparatus, device, and product for constructing a knowledge graph, which can improve the accuracy and efficiency of knowledge graph construction.
[0005] To achieve the above technical objectives, the present application specifically proposes the following technical solutions: According to the first aspect of the embodiments of the present application, a method for constructing a knowledge graph is provided, including: extracting the table of contents in teaching materials to obtain multiple sets of hierarchical knowledge point directories, where each set of hierarchical knowledge point directories includes at least one-level directory; matching the last-level directory under each set of hierarchical knowledge point directories with the text in the teaching materials to obtain the text content corresponding to the last-level directory under each set of hierarchical knowledge point directories; and constructing a teaching knowledge graph based on the extraction of entities and entity relationships from the text content corresponding to the last-level directory under each set of hierarchical knowledge point directories.
[0006] In some implementation manners, the extracting the table of contents in teaching materials to obtain multiple sets of hierarchical knowledge point directories includes: identifying a first keyword and a second keyword in the teaching materials, where the first keyword represents the starting position of the table of contents, and the second keyword represents the ending position of the table of contents; extracting each title located between the first keyword and the second keyword; and dividing each title into the corresponding-level directory according to the format of each title and the preset format corresponding to each level of directory in each set of hierarchical knowledge point directories to obtain the multiple sets of hierarchical knowledge point directories.
[0007] In some implementations, the main text includes various main text segments. The step of matching the last-level directory under each group of knowledge point hierarchical directories with the main text in the teaching materials to obtain the main text content corresponding to the last-level directory under each group of knowledge point hierarchical directories includes: performing the following steps for each group of knowledge point hierarchical directories respectively: determining the similarity between the last-level directory under this group of knowledge point hierarchical directories and each of the main text segments; determining the main text segment with the highest similarity to the last-level directory under this group of knowledge point hierarchical directories as the main text content corresponding to the last-level directory under this group of knowledge point hierarchical directories; or merging at least one main text segment with a similarity to the last-level directory under this group of knowledge point hierarchical directories greater than or equal to a first similarity threshold to obtain the main text content corresponding to the last-level directory under this group of knowledge point hierarchical directories.
[0008] In some implementations, after determining the main text segment with the highest similarity to the last-level directory under this group of knowledge point hierarchical directories as the content corresponding to the last-level directory under this group of knowledge point hierarchical directories, the method further includes: if the similarity between the first main text segment and the last-level directory under another group of knowledge point hierarchical directories is greater than the similarity between the first main text segment and the last-level directory under this group of knowledge point hierarchical directories, generating an association relationship between the other group of knowledge point hierarchical directories and this group of knowledge point hierarchical directories; wherein, the similarity between the first main text segment and the last-level directory under this group of knowledge point hierarchical directories is greater than or equal to a second similarity threshold.
[0009] In some implementations, after merging at least one main text segment with a similarity to the last-level directory under this group of knowledge point hierarchical directories greater than or equal to a first similarity threshold to obtain the main text content corresponding to the last-level directory under this group of knowledge point hierarchical directories, the method further includes: for the main text segment corresponding to the similarity that is less than the first similarity threshold and greater than or equal to a third similarity threshold among all the similarities, if the similarity between the second main text segment and the last-level directory under another group of knowledge point hierarchical directories is greater than the similarity between the second main text segment and the last-level directory under this group of knowledge point hierarchical directories, generating an association relationship between the other group of knowledge point hierarchical directories and this group of knowledge point hierarchical directories; wherein, the similarity between the second main text segment and the last-level directory under this group of knowledge point hierarchical directories is less than the first similarity threshold and greater than or equal to the third similarity threshold.
[0010] In some implementations, each text segment contains key elements, where the key elements include at least one of a title, a table, a picture, a paragraph, and a label; among them, determining the similarity between the last-level directory in the hierarchical directory of this set of knowledge points and each text segment includes: determining the similarity between the last-level directory in the hierarchical directory of this set of knowledge points and each key element in each text segment; according to the weight coefficients corresponding to the respective key elements, performing weighted summation on the similarity between the last-level directory in the hierarchical directory of this set of knowledge points and each key element in each text segment to obtain the similarity between the last-level directory in the hierarchical directory of this set of knowledge points and each text segment.
[0011] In some implementations, the teaching materials include multiple teaching materials, and the teaching knowledge graph includes a teaching knowledge graph corresponding to each teaching material. The method further includes: identifying entities representing the same object and relationships representing the same semantics in the multiple teaching knowledge graphs; associating the entities representing the same object in the multiple teaching knowledge graphs and associating the relationships representing the same semantics to obtain a comprehensive teaching knowledge graph.
[0012] According to the second aspect of the embodiments of the present application, there is provided a knowledge graph construction device, including: an extraction unit, configured to extract the table of contents in the teaching materials to obtain multiple sets of hierarchical directories of knowledge points, where each set of hierarchical directories of knowledge points includes at least one level of directory; a matching unit, configured to match the last-level directory in each set of hierarchical directories of knowledge points with the text in the teaching materials to obtain the text content corresponding to the last-level directory in each set of hierarchical directories of knowledge points; a construction unit, configured to construct a teaching knowledge graph based on the extraction of entities and entity relationships from the text content corresponding to the last-level directory in each set of hierarchical directories of knowledge points.
[0013] According to the third aspect of the embodiments of the present application, there is provided an electronic device, including a memory and a processor; the memory is connected to the processor and is configured to store a program; the processor is configured to implement the knowledge graph construction method as described in the first aspect by running the program in the memory.
[0014] According to the fourth aspect of the embodiments of the present application, there is provided a computer program product, including computer program instructions, where when the computer program instructions are run by a processor, the processor is caused to execute: the knowledge graph construction method as described in the first aspect.
[0015] A method, apparatus, device, and product for constructing a knowledge graph provided by an embodiment of the present application extract a table of contents from teaching materials to obtain a hierarchical table of knowledge points including at least one level of table of contents, and for the last level of table of contents under each group of hierarchical tables of knowledge points, match the corresponding text content in the teaching materials. Subsequently, entity and relationship extraction are performed on the text content corresponding to the last level of table of contents under each group of hierarchical tables of knowledge points, and finally a structured teaching knowledge graph is constructed. This technical solution is based on the existing table of contents structure in teaching materials as the framework of the knowledge hierarchy, realizes accurate positioning and content matching of the text content, thereby providing a structured basis for subsequent knowledge extraction and graph construction. Compared with the traditional method of relying on manual collation of knowledge structures, it can improve the accuracy and context relevance of knowledge extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0017] Figure 1 It is a flowchart of a method for constructing a knowledge graph provided by an embodiment of the present application.
[0018] Figure 2 It is a flowchart of a process for extracting a table of contents provided by an embodiment of the present application.
[0019] Figure 3 It is a flowchart of a process for matching a table of contents with text provided by an embodiment of the present application.
[0020] Figure 4 It is a flowchart of another process for matching a table of contents with text provided by an embodiment of the present application.
[0021] Figure 5 It is a flowchart of determining similarity based on structured text provided by an embodiment of the present application.
[0022] Figure 6 It is a schematic diagram of structured processing provided by an embodiment of the present application.
[0023] Figure 7 It is a schematic structural diagram of a knowledge graph construction apparatus provided by an embodiment of the present application.
[0024] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The technical solution proposed in the embodiments of this application is applicable to any application scenarios in the field of education, such as auxiliary teaching scenarios, personalized learning scenarios, intelligent Q&A scenarios, knowledge management and discovery scenarios, and education research scenarios. Specifically: In the auxiliary teaching scenario, the knowledge graph can help teachers efficiently organize and manage teaching resources, and assist in formulating teaching plans and progress arrangements. For example, by combining the teaching syllabus with the current knowledge mastery of students, relevant knowledge points, teaching cases, and exercise questions can be automatically recommended to improve teaching quality and efficiency.
[0026] In the personalized learning scenario, based on the individual learning situation and interest preferences of students, the knowledge graph can recommend personalized learning paths and resources to help students study and review efficiently. For example, appropriate course content, exercise training, and extended reading materials can be recommended according to the student's level.
[0027] In the intelligent Q&A scenario, by combining natural language processing technology, the knowledge graph can provide intelligent Q&A services to quickly respond to education-related questions raised by students or teachers. For example, when a student asks "What is Newton's second law?", detailed explanations and practical application examples can be provided through the entities and relationships in the knowledge graph.
[0028] In the knowledge management and discovery scenario, the knowledge graph helps to systematically organize and store educational knowledge, facilitating knowledge retrieval, sharing, and dissemination. At the same time, by analyzing the structure and relationships in the knowledge graph, new knowledge associations and laws can be discovered to promote the innovation of educational content.
[0029] In the education research scenario, it can provide rich data support and analysis tools for education researchers, assisting in educational theory research, teaching method research, subject knowledge structure research, etc. For example, researchers can use the knowledge graph to analyze the impact of different teaching strategies on students' knowledge mastery, or explore the internal connections between various subject knowledge, thereby promoting the development of educational practice and theory.
[0030] The technical solution provided in the embodiments of this application can be exemplarily applied to hardware devices such as processors, electronic devices, and servers (including cloud servers), or packaged as a software program to be run. When the hardware device executes the processing process of the technical solution in the embodiments of this application, or when the above software program is run, the automatic splitting of the target task and the automatic invocation of the application program interfaces required for the task can be achieved to complete the purpose of the target task. The embodiments of this application only provide an exemplary introduction to the specific processing process of the technical solution of this application, and do not limit the specific implementation form of the technical solution of this application. Any technical implementation form that can execute the processing process of the technical solution of this application can be adopted by the embodiments of this application.
[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0032] Before introducing the solution of the present application, the related technologies will be introduced first: A knowledge graph is a structured semantic knowledge base that represents entities (such as people, places, objects, concepts, etc.) in the real world and their relationships with each other in the form of a graph. A knowledge graph usually contains nodes and edges. Nodes represent entities, and edges represent various semantic relationships between entities. Through this highly organized way, it can not only capture and express complex knowledge structures, but also help machines better understand natural language, achieve complex reasoning and learning tasks, and thus play an important role in multiple fields such as search engine optimization, recommendation systems, natural language processing, and intelligent question answering.
[0033] Specifically, applying the knowledge graph to a search engine can provide more accurate and relevant search results by understanding the context meaning of a user's query; applying the knowledge graph to a recommendation system can provide personalized content recommendations for users based on user preferences and behavior patterns; applying the knowledge graph to natural language processing can support machines to understand and analyze the deep semantics of language and improve text processing capabilities; applying the knowledge graph to an intelligent question answering system can use the structured knowledge in the graph to achieve fast and accurate question answering; applying the knowledge graph to data integration can uniformly integrate data from different sources into a structured knowledge framework to improve the consistency and usability of the data.
[0034] An educational knowledge graph is a structured knowledge representation method constructed for the education field. Its core lies in organizing and storing knowledge points, concepts, resources, etc. in the education field in the form of a graph, making it have good interpretability and operability, facilitating machine recognition and processing, and thus providing a basis for promoting education informatization and intelligence.
[0035] However, the current construction of educational knowledge graphs mainly relies on manual annotation and collation by domain experts, which has the problems of insufficient accuracy and low efficiency in the constructed knowledge graphs.
[0036] In view of this, embodiments of the present application are dedicated to providing a method, apparatus, device, and product for constructing a knowledge graph. Specifically, the method extracts the table of contents from teaching materials to obtain multiple sets of hierarchical knowledge point directories. Subsequently, the last-level directories under each set of hierarchical knowledge point directories are matched with the main text in the teaching materials, thereby realizing the association between each set of hierarchical knowledge point directories and the main text part. Finally, based on this association information, entities and the relationships between them are further extracted to construct an accurate teaching knowledge graph, which can not only improve the accuracy of the constructed knowledge graph, but also greatly improve the construction efficiency. Detailed descriptions will be given one by one in the following embodiments.
[0037] Exemplary Method Figure 1 is a flowchart of a method for constructing a knowledge graph provided by an embodiment of the present application. As Figure 1 shown, the method for constructing a knowledge graph provided in this embodiment includes steps S101 - S103: S101. Extract the table of contents from the teaching materials to obtain multiple sets of hierarchical knowledge point directories, where each set of hierarchical knowledge point directories includes at least one level of directory.
[0038] In this embodiment, the teaching materials include textbooks, curriculum outlines, academic papers, and various teaching-related materials, etc. These teaching materials contain knowledge points, concepts, terms in various subject fields and the relationships between them, constituting rich knowledge resources.
[0039] The table of contents refers to an overview of each part of the content listed in the teaching materials, which can help quickly understand the entire content structure and conveniently locate the interested part. The table of contents is usually organized by chapters or topics, indicating the title of each part and the corresponding page number or location information.
[0040] When constructing a knowledge graph, based on the table of contents in teaching materials (such as textbooks, curriculum outlines, etc.), a hierarchical knowledge point structure can be constructed. This hierarchical knowledge point structure can not only summarize the content framework of the teaching materials, but also display the logical associations between knowledge points through hierarchical relationships, helping to further analyze and process these knowledge points and the relationships between them.
[0041] Among them, each set of hierarchical knowledge point directories can include at least one level of directory. Taking any set of hierarchical knowledge point directories as an example, it can only include a first-level directory, or can include a first-level directory and its subordinate multi-level subdirectories, such as a second-level directory, or further include a second-level directory and a third-level directory.
[0042] For the specific implementation method of extracting the table of contents from the above teaching materials, reference can be made to the detailed introduction of the embodiments shown below Figure 2 as follows: Figure 2This is a flowchart of a directory extraction process provided by an embodiment of the present application. As Figure 2 shown, when step S101 extracts the directory in the teaching materials to obtain multiple groups of hierarchical knowledge point directories, it specifically includes the following sub-steps S201 - S203: S201. Identify the first keyword and the second keyword in the teaching materials. The first keyword represents the starting position of the directory, and the second keyword represents the ending position of the directory.
[0043] In specific implementation, OCR (Optical Character Recognition) technology can be used to recognize the text content in the teaching materials, so as to locate the starting position identifier and the ending position identifier of the directory.
[0044] Among them, the first keyword can be words such as "Directory"; the second keyword can be words such as "Reference Documents", "Appendix", "Answers", or "Acknowledgements".
[0045] When keywords such as "Directory" are recognized, it is determined as the starting point of the directory; when keywords such as "Reference Documents", "Appendix", "Answers", and "Acknowledgements" are recognized, it is determined as the ending point of the directory.
[0046] S202. Extract each title located between the first keyword and the second keyword.
[0047] In this step, OCR technology will continue to be used to recognize and extract each title between the starting position and the ending position of the directory, providing data support for the subsequent generation of a structured directory.
[0048] S203. According to the format of each title and the preset format corresponding to each level of the directory in each group of hierarchical knowledge point directories, divide each title into the corresponding level of the directory to obtain multiple groups of hierarchical knowledge point directories.
[0049] Since directories at different levels usually have different format features (such as font, font size, numbering, indentation method, etc.), the format of the extracted title can be compared with the preset format corresponding to each level of the directory in each group of hierarchical knowledge point directories, so as to determine the directory level to which each title belongs.
[0050] For example, the preset format corresponding to each level of the directory in the hierarchical knowledge point directory is specifically as follows: The first-level title uses Arial font, font size 16, bold, and has no numbering; the second-level title uses Arial font, font size 14, bold, and is numbered with Roman numerals (I, II, III...); the third-level title uses Times New Roman font, font size 12, not bold, and is numbered with Arabic numerals (1, 2, 3...) and parentheses.
[0051] Assume that title A uses Arial font, size 16, bold, and has no numbering; title B uses Arial font, size 14, bold, and is numbered with Roman numeral "I"; title C uses Times New Roman font, size 12, not bold, and is numbered with "(1)". Then title A conforms to the preset format of a first-level title, title B conforms to the preset format of a second-level title, and title C conforms to the preset format of a third-level title. Therefore, title A, title B, and title C are the first-level title, the second-level title, and the third-level title respectively.
[0052] In the above way, each title can be classified into the corresponding directory level to obtain multiple sets of knowledge point classification directories.
[0053] The above steps S201 - S203 introduce the specific implementation method of generating a directory from teaching materials. However, in some embodiments, the directory in the teaching materials may also exist in the form of an independent directory file. For this case, the present application also provides the following processing method: For example, before executing step S201, first determine whether there is an independent directory file in the teaching materials; if there is an independent directory file, directly read the directory file to obtain multiple sets of knowledge point classification directories; if there is no independent directory file, execute the extraction process of steps S201 - S203 to obtain multiple sets of knowledge point classification directories.
[0054] After obtaining multiple sets of knowledge point classification directories from the teaching materials, the multiple sets of knowledge point classification directories can be associated with the text content in the teaching materials, so as to facilitate subsequent knowledge extraction. The specific association process can continue to refer to Figure 1 for the detailed introduction of step S102 as follows: S102. Match the last-level directory under each set of knowledge point classification directories with the text in the teaching materials to obtain the text content corresponding to the last-level directory under each set of knowledge point classification directories.
[0055] By dividing the text content in the teaching materials through the knowledge point classification directories, it can be divided into parts corresponding to each set of knowledge point classification directories, so that in the subsequent knowledge extraction process, it is possible to focus on the text fragments related to specific knowledge points, avoid the interference of irrelevant information, and thus improve the accuracy and relevance of knowledge extraction.
[0056] In some embodiments, vector matching technology can be adopted to achieve semantic-level matching between the hierarchical directory of knowledge points and the text fragments. For example, by vectorizing the directory and the text respectively and using semantic similarity calculation methods, the text content most relevant to the semantics of each directory can be accurately identified. This method can not only capture the surface matching relationship between texts, but also deeply explore their potential semantic associations, so as to more accurately locate the text content corresponding to the directory.
[0057] Figure 3 FIG. is a flowchart of a matching process between a directory and text provided by an embodiment of the present application. As Figure 3 shown, when performing step S102, it specifically includes: respectively performing the following sub-steps S301-S302 for each group of hierarchical directories of knowledge points: S301. Determine the similarity between the last-level directory under this group of hierarchical directories of knowledge points and each text fragment.
[0058] In some embodiments, step S301 includes: performing vectorization processing on the last-level directory under this group of hierarchical directories of knowledge points to obtain the corresponding directory vector; performing vectorization processing on each text fragment in the teaching material respectively to obtain the corresponding text fragment vector for each text fragment; calculating the similarity between the directory vector and each text fragment vector to obtain the similarity between the last-level directory under this group of hierarchical directories of knowledge points and each text fragment.
[0059] Among them, the similarity between the directory vector and each text fragment vector can be measured by similarity calculation methods such as cosine similarity, Euclidean distance, Manhattan distance, Pearson correlation coefficient, and Jaccard similarity.
[0060] Alternatively, a vector matching model based on deep learning (such as a pre-trained language model like BERT) can also be used to generate semantic vectors for the last-level directory and each text fragment respectively and perform similarity calculation.
[0061] S302. Determine the text fragment with the highest similarity to the last-level directory under this group of hierarchical directories of knowledge points as the text content corresponding to the last-level directory under this group of hierarchical directories of knowledge points.
[0062] After obtaining the similarity between the last-level directory under the hierarchical directory of this group of knowledge points and each text segment, the text segments with a similarity greater than or equal to the second similarity threshold can be sorted in descending order of similarity, and the text segment corresponding to the highest similarity is determined as the text content corresponding to the last-level directory under the hierarchical directory of this group of knowledge points.
[0063] In some other embodiments, steps S301 and S302 can also be implemented through a large language model, that is, to calculate the similarity between the last-level directory under the hierarchical directory of this group of knowledge points and each text segment through the large language model, and the text segment with the highest similarity to the last-level directory under the hierarchical directory of this group of knowledge points is determined as the text content corresponding to the last-level directory under the hierarchical directory of this group of knowledge points, as the final output result of the model.
[0064] To further improve the accuracy of knowledge extraction, after associating multiple groups of hierarchical directories of knowledge points with text content through the above steps S301 and S302, the construction of the association relationship between the hierarchical directories of knowledge points across groups can also be provided. Continue to refer to Figure 3 , after step S302, the following step S303 can also be included: S303. If the similarity between the first text segment and the last-level directory under the hierarchical directory of other groups of knowledge points is greater than the similarity between the first text segment and the last-level directory under the hierarchical directory of this group of knowledge points, an association relationship between the hierarchical directories of other groups of knowledge points and the hierarchical directory of this group of knowledge points is generated.
[0065] Among them, the similarity between the first text segment and the last-level directory under the hierarchical directory of this group of knowledge points is greater than or equal to the second similarity threshold.
[0066] In this step, the first text segment can be understood as at least one text segment that is greater than or equal to the second similarity threshold except for the highest similarity in the similarity ranking result of the last-level directory under the hierarchical directory of this group of knowledge points.
[0067] If the similarity between the first text segment and the last-level directory under the hierarchical directory of this group of knowledge points is greater than or equal to the second similarity threshold, and the similarity between the first text segment and the last-level directory of a certain other group of knowledge points is higher than its similarity in the original group directory, it is considered that this text segment is more suitable for the knowledge system structure of this other group, and its relevance to the original group directory is relatively high.
[0068] For the sake of easy understanding, the following further illustrates step S303 with specific examples: In some examples, it is assumed that knowledge extraction is performed on text segment A to obtain entity A. The similarity between entity A and section 1 is 0.62, and the similarity between entity A and section 2 is 0.78. If the second similarity threshold is set to 0.5, since the similarity between entity A and section 2 is higher than the threshold, a triple of "section 2 contains entity A" is created; at the same time, since its similarity with section 1 also exceeds the threshold, entity A is associated with section 1 across directories to reflect its potential semantic connection between multiple knowledge points.
[0069] In other examples, suppose that knowledge extraction is performed on text segment A to obtain entity B and entity C. Among them, the similarity between entity B and chapter 1 is 0.63, and the similarity between entity C and chapter 2 is 0.71, and 0.5 is still used as the second similarity threshold. At this time, the triples of "Chapter 1 contains entity B" and "Chapter 2 contains entity C" are generated respectively, and the cross-segment association relationship between entity B and entity C is further established, thereby reflecting the possible cross-knowledge point semantic connection between different entities.
[0070] Through the above method, not only can the structured mapping between knowledge points and entities be achieved, but also the cross-level and cross-segment associations between entities can be effectively explored, thereby improving the integrity and semantic expression ability of the constructed teaching knowledge graph, and facilitating subsequent cross-topic knowledge recommendation, auxiliary learning path planning and other application scenarios.
[0071] Figure 4 Another flowchart of the process of matching a directory with a text provided in the embodiment of the present application. Figure 4 As shown, when executing step S102, it specifically includes: executing the following sub-steps S401-S402 for each group of knowledge point classification catalogs: S401: Determine the similarity between the last level directory under the group of knowledge point hierarchical directories and each text segment.
[0072] The specific implementation method of step S401 is the same as that of the aforementioned step S301. For details, please refer to the detailed introduction of the specific implementation method of step S301, which will not be repeated here.
[0073] S402: merge at least one text segment whose similarity with the last level directory under the group of knowledge point hierarchical directories is greater than or equal to a first similarity threshold to obtain text content corresponding to the last level directory under the group of knowledge point hierarchical directories.
[0074] After obtaining the similarity between the last-level directory under the hierarchical directory of this group of knowledge points and each text segment, the text segments with a similarity greater than or equal to the second similarity threshold can be sorted in descending order of similarity, and at least one text segment with a similarity greater than or equal to the first similarity threshold is taken, and at least one text segment after determination and merging is determined as the text content corresponding to the last-level directory under the hierarchical directory of this group of knowledge points.
[0075] Among them, the first similarity threshold is greater than or equal to the second similarity threshold.
[0076] After associating multiple groups of hierarchical directories of knowledge points with the text content through the above steps S401 and S402, in order to further improve the accuracy of knowledge extraction, this embodiment can also provide an implementation method for cross-directory association. Referring to Figure 4 , after step S402, the following step S403 can also be included: S403. If among the text segments corresponding to the similarities that are less than the first similarity threshold and greater than or equal to the third similarity threshold in each similarity, there is a second text segment whose similarity with the last-level directory under another group of hierarchical directories of knowledge points is greater than the similarity of the second text segment with the last-level directory under this group of hierarchical directories of knowledge points, then an association relationship between the other group of hierarchical directories of knowledge points and this group of hierarchical directories of knowledge points is generated.
[0077] Among them, the first similarity threshold is greater than the third similarity threshold. The similarity between the second text segment and the last-level directory under this group of hierarchical directories of knowledge points is less than the first similarity threshold and greater than or equal to the third similarity threshold.
[0078] In this step, the second text segment can be understood as at least one text segment in the similarity sorting result of the last-level directory under this group of hierarchical directories of knowledge points, whose similarity with the last-level directory under this group of hierarchical directories of knowledge points is less than the first similarity threshold and greater than or equal to the third similarity threshold, and whose similarity with the last-level directory under another group of hierarchical directories of knowledge points is greater than the similarity of the second text segment with the last-level directory under this group of hierarchical directories of knowledge points.
[0079] That is to say, if the similarity between the second text segment and the last-level directory under this group of hierarchical directories of knowledge points is less than the first similarity threshold and greater than or equal to the third similarity threshold, and the similarity between the second text segment and the last-level directory of a certain other group of hierarchical directories of knowledge points is higher than its similarity in the original group directory, then it is considered that this text segment is more suitable for the knowledge system structure of the other group, and its relevance to the original group directory is relatively high.
[0080] Therefore, an association relationship record across groups can be generated, which includes the directory identifiers of the source group and the target group, the text segment IDs, and the similarity values between the two. These association relationships can be used in subsequent application scenarios such as knowledge graph construction, cross-topic knowledge recommendation, and auxiliary learning path planning.
[0081] Continue to refer to Figure 1 , after step S102, step S103 may further be included.
[0082] S103. Construct a teaching knowledge graph based on the extraction of entities and entity relationships from the text content corresponding to the last-level directory under the hierarchical directory of each group of knowledge points.
[0083] In some embodiments, the specific implementation manner of step S103 includes: preprocessing the text content corresponding to the last-level directory under the hierarchical directory of each group of knowledge points to obtain preprocessed text content; performing entity recognition and relationship extraction based on the preprocessed text content to obtain a plurality of triples; and constructing a teaching knowledge graph based on the plurality of triples.
[0084] Among them, the preprocessing includes: splitting the text content corresponding to the last-level directory under the hierarchical directory of each group of knowledge points into words or lexical units, removing high-frequency words that are not very helpful for understanding the text content, such as "de" (of), "shi" (is), etc., and assigning grammatical roles to each word, such as noun or verb.
[0085] When performing entity recognition and relationship extraction, rule-based methods, supervised learning methods, or deep learning methods can be used. Among them, the rule-based method refers to matching entities and their relationships in the text through predefined grammar rules and patterns. For example, by writing regular expressions or using dependency syntactic analysis to analyze specific relationship patterns.
[0086] The supervised learning method requires a labeled training data set, which contains labeled entities and the relationships between them. Then, machine learning algorithms (such as support vector machine SVM, decision tree, etc.) are used to train the model to automatically identify the entity relationships in the text content.
[0087] The deep learning method also requires a labeled training data set, which contains labeled entities and the relationships between them. Then, deep learning algorithms, such as convolutional neural network (CNN), recurrent neural network (RNN) and its variants (such as LSTM, GRU), and Transformer architecture (such as BERT), are used to train the neural network model to automatically identify the entity relationships in the text content.
[0088] After obtaining entities and their relationships, the entities can be used as nodes in the knowledge graph, and the relationships can be used as the connection relationships between the nodes. That is, if there is a relationship between two entities, a connection line is used to connect the corresponding nodes of the two entities; otherwise, no connection is made. This is used to indicate whether there is a relationship between two entities.
[0089] This embodiment adopts a content matching mechanism based on directory guidance, that is, using the existing directory system in the teaching materials as the framework of the knowledge hierarchy to locate and match the corresponding text content, and based on this, subsequent knowledge extraction and graph construction are carried out. Compared with the traditional method that relies on manual collation of the knowledge structure, it can improve the accuracy and efficiency of knowledge graph construction. In addition, by extracting the entities and their relationships of the text content corresponding to the last-level directory under each group of knowledge point hierarchical directories, the constructed teaching knowledge graph not only has rich semantic information, but also can accurately reflect the logical hierarchy and correlation relationships between knowledge points, thus better meeting the knowledge organization and retrieval needs in the teaching scenario.
[0090] In order to more accurately measure the semantic relevance between the knowledge point hierarchical directory and the text content, this embodiment can also adopt a structured processing mechanism for teaching materials. That is, by identifying and extracting key elements in the text content and performing weighted calculations in combination with their importance in knowledge expression, the accuracy and relevance of similarity judgment are improved. Specifically, when determining the similarity between the last-level directory under the group of knowledge point hierarchical directories and each text segment based on the structured text, the following embodiments can be provided: Figure 5 This is a flowchart for determining similarity based on structured text provided by an embodiment of the present application. As Figure 5 shown, when step S301 determines the similarity between the last-level directory under the group of knowledge point hierarchical directories and each text segment, it may further include the following steps S501 and step S502: S501. Determine the similarity between the last-level directory under the group of knowledge point hierarchical directories and each key element in each text segment.
[0091] Figure 6 This is a schematic diagram of the structured processing provided by an embodiment of the present application. As Figure 6 shown, through structured analysis of the text content in the teaching materials, various key elements are identified and extracted, including at least one of titles, tables, pictures, paragraphs, and annotations. Among them, annotations include notes, footnotes, emphasis marks, etc.
[0092] After that, vectorize the last-level directory under the hierarchical directory of this group of knowledge points and various key elements in each text segment respectively to obtain the directory vector and the key element vectors corresponding to each key element in each text segment. Among them, traditional word vector models such as TF-IDF, Word2Vec, GloVe, etc., or pre-trained language models based on deep learning (such as BERT, RoBERTa, ERNIE, etc.) can be used to vectorize various key elements. For non-text key elements (such as pictures), their semantic descriptions can be extracted through image recognition technology or vector encoding can be performed using multi-modal models.
[0093] Furthermore, for each text segment, calculate the similarity between the directory vector and the key element vectors corresponding to each key element in this text segment to obtain the similarity between the last-level directory under the hierarchical directory of this group of knowledge points and each key element under each text segment.
[0094] Among them, the similarity between the directory vector and each key element can be measured by similarity calculation methods such as Cosine Similarity, Euclidean Distance, Manhattan Distance, Pearson Correlation Coefficient, and Jaccard Similarity.
[0095] Or, an end-to-end semantic matching model (such as BERT-based Sentence Pair Classification) can also be used to directly output the similarity between the two.
[0096] S502. According to the weight coefficients corresponding to each key element, perform weighted summation on the similarity between the last-level directory under the hierarchical directory of this group of knowledge points and each key element in each text segment to obtain the similarity between the last-level directory under the hierarchical directory of this group of knowledge points and each text segment.
[0097] After obtaining the similarity between each key element in each text segment and the knowledge point directory, the weight coefficients corresponding to each key element can also be used to reflect the relative importance of different key elements in knowledge expression. And perform weighted summation on the key element similarities to obtain the final similarity. The formula is as follows: Sim(final)=w1*Sim(title)+w2*Sim(paragraph)+w3*Sim(table)+w4*Sim(image) + w5*Sim(annotation); Among them, w1 to w5 are the weight coefficients corresponding to each key element, and they satisfy: w1 + w2 + w3 + w4 + w5 = 1.
[0098] In practical applications, the weight coefficients can be dynamically adjusted according to the types of teaching materials (such as textbooks, lecture notes, PPTs, etc.), subject fields (for example, liberal arts may focus on paragraphs, while science may focus on charts), and user feedback. The weight allocation can also be automatically optimized through machine learning methods.
[0099] In some examples, the title is usually highly general and a highly condensed core content of the text, with high indicativeness. Its value range can be set to 0.25 - 0.35; the text paragraphs carry the main knowledge information and are the core source of knowledge extraction. Its value range can be set to 0.20 - 0.30; tables are used to structurally display data or comparison information, with moderate importance. Its value range can be set to 0.10 - 0.20. Pictures are usually used to assist understanding, and their semantic importance needs to be judged in combination with captions or context. Its value range can be set to 0.05 - 0.15. Annotations such as highlights and comments provide supplementary explanations, but the information density is low. Its value range can be set to 0.05 - 0.10.
[0100] It should be noted that the weight allocation method introduced in the above examples is only for illustrative purposes and is not used to limit the specific allocation method of the weight coefficients of this application. In the actual application process, the corresponding weight parameters can be flexibly set according to specific business requirements, application scenarios, or data characteristics.
[0101] In this embodiment, by structuring the text content into multiple key elements and assigning different semantic weights to them, the semantic association relationship between the knowledge point catalog and the text fragments can be more accurately represented, thereby improving the accuracy of knowledge extraction.
[0102] After processing various teaching materials through the above embodiments to obtain a series of independent teaching knowledge graphs, in order to achieve unified representation and interoperability of knowledge and further improve the integration and utilization efficiency of educational resources, the knowledge graphs generated based on different teaching materials can also be merged and complemented to generate a comprehensive knowledge graph.
[0103] Among them, the merging of the knowledge graphs includes: identifying the entities representing the same object and the relationships representing the same semantics in multiple teaching knowledge graphs; associating the entities representing the same object in multiple teaching knowledge graphs and associating the relationships representing the same semantics to obtain a comprehensive teaching knowledge graph.
[0104] In this embodiment, it is first necessary to identify entities representing the same object in multiple teaching knowledge graphs. Specifically, it includes: comparing and analyzing the names, description information, and association relationships of entities in different graphs. For example, if two graphs respectively contain knowledge points about "Newton's second law", it is necessary to determine whether these two knowledge points represent the same concept and perform standardization processing on them (such as unifying naming conventions) to ensure the consistency of the concept. Among them, natural language processing techniques, such as text similarity calculation or semantic analysis, can be used to determine whether these two knowledge points point to the same concept.
[0105] For the relationships between entities, identification and matching also need to be carried out. Taking "Newton's second law" as an example, its mathematical expression relationships with "force", "mass", and "acceleration" may have different expression forms or levels of detail in different knowledge graphs. By comparing and analyzing the specific content of these relationships, ensure that they can be correctly understood and associated in the new comprehensive graph, so as to ensure the consistency and coherence of the knowledge system.
[0106] After determining the entities representing the same object and their relationships, it is also necessary to extract relevant entities and relationships from the original graphs and integrate them into the comprehensive teaching knowledge graph. Specifically, it includes: creating new nodes and edges, or adjusting the existing structure to reflect a more comprehensive knowledge system.
[0107] To further improve the integrity of the knowledge graph, after completing the merging operation of the knowledge graph, a knowledge graph completion operation can also be performed.
[0108] Specifically, if it is identified that there is a lack of direct connection between some entities in the merged graph, but it is determined based on existing knowledge that such a connection should exist, then the "missing links" can be automatically filled through inference algorithms to enhance the coherence and depth of the knowledge graph.
[0109] In addition, authoritative data sources in other related fields can also be introduced as supplements, such as academic databases, encyclopedias, etc., to add more dimensional information to the comprehensive knowledge graph.
[0110] In addition, a user feedback channel can also be set up. By collecting the actual usage experiences and suggestions provided by users, it helps to discover and correct potential problems, making the comprehensive knowledge graph more in line with educational needs and improving its applicability and user satisfaction.
[0111] Exemplary device Corresponding to the above knowledge graph construction method, an embodiment of the present application also provides a knowledge graph construction device. Figure 7 It is a schematic structural diagram of a knowledge graph construction device provided by an embodiment of the present application. As Figure 7As shown in the figure, the knowledge graph construction device provided by the embodiment of the present application includes: an extraction unit 701, a matching unit 702, and a construction unit 703; wherein, the extraction unit 701 is configured to extract the table of contents in the teaching materials to obtain multiple sets of hierarchical knowledge point directories, and each set of hierarchical knowledge point directories includes at least one-level directory; the matching unit 702 is configured to match the last-level directory under each set of hierarchical knowledge point directories with the main text in the teaching materials to obtain the main text content corresponding to the last-level directory under each set of hierarchical knowledge point directories; the construction unit 703 is configured to construct a teaching knowledge graph based on the extraction of entities and entity relationships from the main text content corresponding to the last-level directory under each set of hierarchical knowledge point directories.
[0112] In some embodiments, when the extraction unit 701 extracts the table of contents in the teaching materials to obtain multiple sets of hierarchical knowledge point directories, it specifically includes: identifying a first keyword and a second keyword in the teaching materials, where the first keyword represents the starting position of the table of contents, and the second keyword represents the ending position of the table of contents; extracting each title located between the first keyword and the second keyword; and dividing each title into the corresponding-level directory according to the format of each title and the preset format corresponding to each level of directory in each set of hierarchical knowledge point directories, so as to obtain the multiple sets of hierarchical knowledge point directories.
[0113] In some embodiments, the main text includes each main text segment. When the matching unit 702 matches the last-level directory under each set of hierarchical knowledge point directories with the main text in the teaching materials to obtain the main text content corresponding to the last-level directory under each set of hierarchical knowledge point directories, it specifically includes: performing the following steps for each set of hierarchical knowledge point directories respectively: determining the similarity between the last-level directory under the set of hierarchical knowledge point directories and each main text segment; determining the main text segment with the highest similarity to the last-level directory under the set of hierarchical knowledge point directories as the main text content corresponding to the last-level directory under the set of hierarchical knowledge point directories; or merging at least one main text segment with a similarity greater than or equal to a first similarity threshold to the last-level directory under the set of hierarchical knowledge point directories to obtain the main text content corresponding to the last-level directory under the set of hierarchical knowledge point directories.
[0114] In some embodiments, after determining the body segment with the highest similarity to the last-level directory in the hierarchical directory of this group of knowledge points as the content corresponding to the last-level directory in the hierarchical directory of this group of knowledge points, the matching unit 702 is further configured to perform the following steps: If the similarity between the first body segment and the last-level directory in the hierarchical directory of other groups of knowledge points is greater than the similarity between the first body segment and the last-level directory in the hierarchical directory of this group of knowledge points, an association relationship between the hierarchical directory of other groups of knowledge points and the hierarchical directory of this group of knowledge points is generated; wherein, the similarity between the first body segment and the last-level directory in the hierarchical directory of this group of knowledge points is greater than or equal to the second similarity threshold.
[0115] The matching unit 702, after merging at least one body segment whose similarity to the last-level directory in the hierarchical directory of this group of knowledge points is greater than or equal to the first similarity threshold to obtain the body content corresponding to the last-level directory in the hierarchical directory of this group of knowledge points, the matching unit 702 is further configured to perform the following steps: For the body segment corresponding to the similarity that is less than the first similarity threshold and greater than or equal to the third similarity threshold among the various similarities, if the similarity between the second body segment and the last-level directory in the hierarchical directory of other groups of knowledge points is greater than the similarity between the second body segment and the last-level directory in the hierarchical directory of this group of knowledge points, an association relationship between the hierarchical directory of other groups of knowledge points and the hierarchical directory of this group of knowledge points is generated; wherein, the similarity between the second body segment and the last-level directory in the hierarchical directory of this group of knowledge points is less than the first similarity threshold and greater than or equal to the third similarity threshold.
[0116] The matching unit 702, each body segment contains key elements, and the key elements include at least one of a title, a table, a picture, a paragraph, and a label; wherein, when the matching unit 702 determines the similarity between the last-level directory in the hierarchical directory of this group of knowledge points and each body segment, it specifically includes: determining the similarity between the last-level directory in the hierarchical directory of this group of knowledge points and each key element in each body segment; according to the weight coefficients corresponding to the various key elements, performing weighted summation on the similarity between the last-level directory in the hierarchical directory of this group of knowledge points and each key element in each body segment to obtain the similarity between the last-level directory in the hierarchical directory of this group of knowledge points and each body segment.
[0117] In some embodiments, the teaching materials include a plurality of teaching materials, and the teaching knowledge graph includes a corresponding teaching knowledge graph for each teaching material. The construction unit 703 is further configured to perform the following steps: identifying entities representing the same object and relationships representing the same semantics in the plurality of teaching knowledge graphs; associating the entities representing the same object in the plurality of teaching knowledge graphs and associating the relationships representing the same semantics to obtain a comprehensive teaching knowledge graph.
[0118] The knowledge graph construction device provided in this embodiment belongs to the same inventive concept as the knowledge graph construction method provided in the above embodiments of the present application. It can execute the knowledge graph construction method provided in any of the above embodiments of the present application, and has corresponding functional modules and beneficial effects for executing the knowledge graph construction method. For technical details not described in detail in this embodiment, reference may be made to the specific processing content of the knowledge graph construction method provided in the above embodiments of the present application, which will not be elaborated here.
[0119] The functions implemented by the above extraction unit 701, matching unit 702, and construction unit 703 can be implemented by the same or different processors respectively, which is not limited in the embodiments of the present application.
[0120] It should be understood that the units in the above device can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, instructions are stored in the memory, and the processor calls the instructions stored in the memory to implement any of the above methods or the functions of each unit of the device. The processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory inside the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of a hardware circuit, and the functions of some or all of the units can be realized through the design of the hardware circuit. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are realized through the design of the logical relationship between the components in the circuit; again, in another implementation, the hardware circuit can be realized by a PLD. Taking an FPGA as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured through a configuration file to realize the functions of some or all of the above units. All units of the above device can be all realized in the form of a processor calling software, or all realized in the form of a hardware circuit, or some realized in the form of a processor calling software, and the remaining part realized in the form of a hardware circuit.
[0121] In the embodiments of the present application, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and running capabilities, such as a CPU, microprocessor, GPU, or DSP, etc.; in another implementation, the processor can implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of this hardware circuit is fixed or can be reconfigured. For example, the processor is a hardware circuit implemented by an ASIC or PLD, such as an FPGA, etc. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a type of ASIC, such as an NPU, TPU, DPU, etc.
[0122] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0123] In addition, each unit in the above device can be integrated in whole or in part, or can be independently implemented. In one implementation, these units are integrated together and implemented in the form of an SOC. The SOC can include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The types of the at least one processor can be different, such as including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0124] Exemplary electronic device Embodiments of the present application propose an electronic device. Refer to Figure 8 As shown, the electronic device includes: A memory 200 and a processor 210; Wherein, the memory 200 is connected to the processor 210 and is used to store programs; The processor 210 is used to implement the knowledge graph construction method disclosed in any of the above embodiments by running the programs stored in the memory 200.
[0125] Specifically, the above electronic device may further include: a bus, a communication interface 220, an input device 230, and an output device 240.
[0126] The processor 210, the memory 200, the communication interface 220, the input device 230, and the output device 240 are interconnected through the bus. Among them: The bus may include a path for transmitting information between various components of the computer system.
[0127] The processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the solution of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0128] The processor 210 may include a main processor, and may also include a baseband chip, a modem, etc.
[0129] The memory 200 stores a program for implementing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include program code, and the program code includes computer operation instructions. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, etc.
[0130] The input device 230 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.
[0131] The output device 240 may include a device for allowing information to be output to a user, such as a display screen, a printer, a speaker, etc.
[0132] The communication interface 220 may include a device of any transceiver type for communicating with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0133] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement each step of any one of the knowledge graph construction methods provided in the above embodiments of the present application.
[0134] An embodiment of the present application also proposes a chip, which includes a processor and a data interface. The processor reads and runs a program stored on a memory through the data interface to execute the knowledge graph construction method introduced in any of the above embodiments. For the specific processing process and its beneficial effects, reference can be made to the embodiments of the knowledge graph construction method described above.
[0135] Exemplary Computer Program Product and Storage Medium In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps in the knowledge graph construction method according to various embodiments of the present application described in any of the above embodiments of this specification.
[0136] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0137] Furthermore, an embodiment of the present application may also be a storage medium on which a computer program is stored. The computer program is executed by a processor to perform the steps in the knowledge graph construction method according to various embodiments of the present application described in any of the above embodiments of this specification, and specifically may implement the steps of the above method embodiments.
[0138] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps may be in other sequences or performed simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0139] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments may be referred to each other. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts may refer to the partial description of the method embodiments.
[0140] The steps in the methods of the embodiments of the present application may be adjusted, combined, and deleted according to actual needs. The technical features recorded in each embodiment may be replaced or combined.
[0141] The modules and sub-modules in the devices and terminals in the embodiments of the present application may be combined, divided, and deleted according to actual needs.
[0142] In several embodiments provided by the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices or modules, and can be in electrical, mechanical or other forms.
[0143] The modules or sub-modules described as separate components may or may not be physically separated. The components as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place, or can be distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0144] In addition, each functional module or sub-module in various embodiments of the present application can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above-mentioned integrated modules or sub-modules can be implemented in the form of hardware, or can be implemented in the form of software functional modules or sub-modules.
[0145] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0146] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software units executed by a processor, or a combination of the two. The software units can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0147] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0148] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for constructing a knowledge graph, characterized in that, Including: Extract the table of contents from the teaching materials to obtain multiple sets of hierarchical knowledge point catalogs, where each set of hierarchical knowledge point catalogs includes at least a first-level catalog; Match the last-level catalog under each set of hierarchical knowledge point catalogs with the main text in the teaching materials to obtain the main text content corresponding to the last-level catalog under each set of hierarchical knowledge point catalogs; Based on the extraction of entities and entity relationships from the main text content corresponding to the last-level catalog under each set of hierarchical knowledge point catalogs, construct a teaching knowledge graph.
2. The method according to claim 1, wherein The extracting the table of contents from the teaching materials to obtain multiple sets of hierarchical knowledge point catalogs includes: Identify the first keyword and the second keyword in the teaching materials, where the first keyword represents the starting position of the table of contents, and the second keyword represents the ending position of the table of contents; Extract each title located between the first keyword and the second keyword; According to the format of each title and the preset format corresponding to each level of catalog in each set of hierarchical knowledge point catalogs, divide each title into the catalogs at the corresponding levels to obtain the multiple sets of hierarchical knowledge point catalogs.
3. The method according to claim 1, wherein The main text includes each main text segment. The matching the last-level catalog under each set of hierarchical knowledge point catalogs with the main text in the teaching materials to obtain the main text content corresponding to the last-level catalog under each set of hierarchical knowledge point catalogs includes: For each set of hierarchical knowledge point catalogs, perform the following steps respectively: Determine the similarity between the last-level catalog under this set of hierarchical knowledge point catalogs and each main text segment; Determine the main text segment with the highest similarity to the last-level catalog under this set of hierarchical knowledge point catalogs as the main text content corresponding to the last-level catalog under this set of hierarchical knowledge point catalogs; or, merge at least one main text segment with a similarity greater than or equal to the first similarity threshold to the last-level catalog under this set of hierarchical knowledge point catalogs to obtain the main text content corresponding to the last-level catalog under this set of hierarchical knowledge point catalogs.
4. The method according to claim 3, wherein After determining the main text segment with the highest similarity to the last-level catalog under this set of hierarchical knowledge point catalogs as the content corresponding to the last-level catalog under this set of hierarchical knowledge point catalogs, the method further includes: If the similarity between the first main text segment and the last-level catalog under other sets of hierarchical knowledge point catalogs is greater than the similarity between the first main text segment and the last-level catalog under this set of hierarchical knowledge point catalogs, generate an association relationship between the other set of hierarchical knowledge point catalogs and this set of hierarchical knowledge point catalogs; where the similarity between the first main text segment and the last-level catalog under this set of hierarchical knowledge point catalogs is greater than or equal to the second similarity threshold.
5. The method according to claim 3, wherein After merging at least one main text segment with a similarity greater than or equal to the first similarity threshold to the last-level catalog under this set of hierarchical knowledge point catalogs to obtain the main text content corresponding to the last-level catalog under this set of hierarchical knowledge point catalogs, the method further includes: For the text segments corresponding to the similarities that are less than the first similarity threshold and greater than or equal to the third similarity threshold among all the similarities, if the similarity between the second text segment and the last-level directory under the knowledge point classification directory of other groups is greater than the similarity between the second text segment and the last-level directory under the knowledge point classification directory of this group, an association relationship between the knowledge point classification directory of other groups and the knowledge point classification directory of this group is generated; Among them, the similarity between the second text segment and the last-level directory under the knowledge point classification directory of this group is less than the first similarity threshold and greater than or equal to the third similarity threshold.
6. The method according to claim 3 or 4 or 5, characterized in that, Each text segment contains key elements, and the key elements include at least one of title, table, picture, paragraph, and annotation; Among them, the determination of the similarity between the last-level directory under the knowledge point classification directory of this group and each text segment includes: Determining the similarity between the last-level directory under the knowledge point classification directory of this group and each key element in each text segment; According to the weight coefficients corresponding to the respective key elements, the similarities between the last-level directory under the knowledge point classification directory of this group and each key element in each text segment are weighted and summed to obtain the similarity between the last-level directory under the knowledge point classification directory of this group and each text segment.
7. The method according to any one of claims 1-5, characterized in that, The teaching materials include multiple teaching materials, and the teaching knowledge graph includes the teaching knowledge graph corresponding to each teaching material respectively. The method further includes: Identifying entities representing the same object and relationships representing the same semantics in multiple teaching knowledge graphs; Associating the entities representing the same object in the multiple teaching knowledge graphs and associating the relationships representing the same semantics to obtain a comprehensive teaching knowledge graph.
8. A knowledge graph construction device, characterized in that, Including: An extraction unit for extracting the table of contents from the teaching materials to obtain multiple groups of knowledge point classification directories, where each group of knowledge point classification directories includes at least one level of directory; A matching unit for matching the last-level directory under each group of knowledge point classification directories with the text in the teaching materials to obtain the text content corresponding to the last-level directory under each group of knowledge point classification directories; A construction unit for constructing a teaching knowledge graph based on the extraction of entities and entity relationships from the text content corresponding to the last-level directory under each group of knowledge point classification directories.
9. An electronic device, characterized in that, Including a memory and a processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the method according to any one of claims 1 to 7 by running the programs in the memory.
10. A computer program product, characterized in that, Including computer program instructions, and when the computer program instructions are run by the processor, the processor implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method for establishing mapping knowledge domain based on book catalogue
CN103729402A
Knowledge graph construction method and device and computer readable storage medium
CN114238654A
Cognitive platform for knowledge extraction from heterogenous data sources and the method thereof
US20230186111A1
Service assurance in 5g networks using key performance indicator navigation tool
US20240114363A1
Feedback data graph generation method and refrigerator
WO2023246849A1