Method and device for extracting knowledge spectrogram for textbook and storage medium
Through a large language model, a global knowledge graph is constructed, which solves the problems of inefficiency and difficulty in dealing with multi-source heterogeneous data in the existing technology, and realizes efficient and automated knowledge graph construction.
Patent Information
- Application Number
- CN202510034082.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-13
AI Technical Summary
The existing technology relies on manual annotation, is inefficient, and is difficult to process multi-source heterogeneous textbook information, and cannot effectively build a knowledge graph.
By obtaining multi-source heterogeneous data, separating text and non-text textbook data, using a large language model to identify directory structure, entities, association relationships and attribute information, constructing a sub-graph of chapter knowledge points, and fusing them into a global knowledge graph to associate knowledge points in non-text data.
It significantly reduces manual participation, improves the automation level and accuracy of knowledge graph construction, realizes hierarchical construction from chapter subgraphs to global graphs, and maintains semantic consistency during the fusion process.
Smart Images

Figure CN119990289A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of large model technology, and specifically relates to a method, device and storage medium for extracting knowledge spectra from teaching materials. Background Art
[0002] With the development of educational informatization, the data processing of textbook content has become increasingly important. As an effective way to organize and represent knowledge, knowledge graphs are increasingly widely used in the field of education. However, existing knowledge graph construction methods often rely on manual annotation, which is inefficient and difficult to process multi-source heterogeneous textbook information.
[0003] The existing technical solutions are as follows:
[0004] (1) Rule-based method: This method identifies entities and relations in text by pre-defining a series of rules. For example, regular expressions are used to identify entities such as dates, places, and knowledge points in a specific format. This method is simple and direct, but has poor adaptability, requires a large number of manually defined rules, and has difficulty processing complex and ambiguous text.
[0005] (2) Statistical-based methods: This method identifies entities and relationships by counting the co-occurrence frequency of entities in the text, context information, etc. For example, the term frequency-inverse document frequency (TF-IDF) is used to identify keywords. This method has a certain processing capability for large-scale data, but its accuracy is limited by the generalization ability of the statistical model.
[0006] (3) Machine learning-based methods: This method uses machine learning models to identify entities and relationships. For example, classifiers such as support vector machines (SVM) and naive Bayes are used for entity recognition. This method has higher accuracy than the first two methods, but requires a large amount of labeled data to train the model.
[0007] Therefore, how to effectively process multi-source heterogeneous teaching material information and reduce manual intervention while improving processing efficiency is a topic worthy of study for technical personnel in this field. Summary of the invention
[0008] In view of the above analysis, the embodiments of the present invention aim to provide a method, device and storage medium for extracting knowledge spectra from textbooks, so as to solve the problem that the existing technology relies on manual annotation, is inefficient and has difficulty in processing multi-source heterogeneous textbook information.
[0009] In a first aspect of the present application, a method for extracting a knowledge spectrum graph for a teaching material is provided, comprising:
[0010] Acquire multi-source heterogeneous data; the multi-source heterogeneous data at least includes text, picture, and video multimodal data;
[0011] Separating the multi-source heterogeneous data to obtain text teaching material data and non-text teaching material data;
[0012] Inputting the text teaching material data into a large language model, identifying the directory structure of the text teaching material data; and obtaining the positioning content information of each chapter according to the directory structure;
[0013] A large language model is used to process the positioning content information of each chapter to identify the entities, associations and attribute information in the chapter; based on the entities, associations and attribute information, a knowledge point subgraph of each chapter is constructed;
[0014] Traverse to obtain the knowledge point subgraphs of all chapters, and merge the knowledge point subgraphs of all chapters into a global knowledge graph;
[0015] The non-text teaching material data is identified by using a large language model to determine the knowledge points mentioned in the non-text teaching material data; the corresponding knowledge point nodes are matched from the global knowledge graph, and the non-text teaching material data is associated with the corresponding knowledge point nodes in the global knowledge graph.
[0016] Optionally, the adopting of a large language model to identify the non-text teaching material data to determine the knowledge points mentioned in the non-text teaching material data includes:
[0017] The non-text teaching material data includes video teaching material data; obtaining video subtitle text of the video teaching material data; summarizing the video subtitle text using a large language model, and using the text description of the main content as the knowledge points mentioned in the video teaching material data;
[0018] The non-text teaching material data includes picture teaching material data; obtaining picture title text of the picture teaching material data; using a large language model to summarize the picture title text, and using the text description of the main content as the knowledge point mentioned in the picture teaching material data.
[0019] Optionally, matching corresponding knowledge point nodes from the global knowledge graph and associating the non-text teaching material data with the knowledge point nodes corresponding to the global knowledge graph includes:
[0020] Perform semantic similarity matching between the knowledge points mentioned in the non-text teaching material data and the knowledge point nodes of the global knowledge graph, and when the semantic similarity exceeds a preset threshold, use the corresponding knowledge point node as the knowledge node matched by the non-text teaching material data;
[0021] Add a relationship to the knowledge point node corresponding to the global knowledge graph to mark the content of the non-text teaching material data.
[0022] Optionally, matching corresponding knowledge point nodes from the global knowledge graph includes:
[0023] Obtain the source chapter of the non-text textbook data;
[0024] Corresponding knowledge point nodes are matched from the global knowledge graph, and a higher weight is given to the knowledge point nodes matched in the source chapter than to the knowledge point nodes matched in other chapters.
[0025] Optionally, the entities, associations, and attribute information in the identified chapters include:
[0026] Use a large language model to perform semantic analysis on the positioning content information of each chapter and output semantic entities; use language tools to identify proper names and terms; exclude non-knowledge points and retain entities related to the main body of the chapter;
[0027] Based on the extracted entities and the corresponding context text, a large language model is used to obtain the association relationship between entities;
[0028] The context of each entity is summarized to obtain entity attribute information; and the relationship attribute information of the association relationship is generated using a large language model.
[0029] Optionally, the step of fusing the knowledge point subgraphs of all chapters into a global knowledge graph includes:
[0030] Use semantic analysis technology to calculate the similarity of knowledge points in different chapters and select entity pairs with high semantic similarity;
[0031] Judge the screened similar entity pairs;
[0032] If the entity names and entity descriptions of similar entity pairs are similar, the entity pairs are directly merged;
[0033] If only one of the entity name and entity description of the similar entity pair is similar, a large language model is used to summarize and generate a description; relevant personnel confirm whether to merge based on the description;
[0034] After merging the entities, the relations with high semantic similarity under the same entity are merged.
[0035] Optionally, after fusing the knowledge point subgraphs of all chapters into a global knowledge graph, the method further includes:
[0036] The global knowledge graphs of different books are integrated according to the logic of the same subject to form a course knowledge graph that includes multiple books in the same field or a professional knowledge graph that includes multiple different fields.
[0037] Optionally, the fusion of the global knowledge graphs of different books according to the logic of the same subject includes:
[0038] Unify the semantic expressions of the differences in expressions and terminology in different books; merge repeated or similar nodes;
[0039] Identify knowledge points with similar semantics in different book graphs and align them;
[0040] Identify and unify the relationships between knowledge points in different books and perform relationship alignment;
[0041] The aligned knowledge points and relationships are merged into a unified knowledge graph to obtain a fused global knowledge graph; where nodes with similar semantics are merged into one node, and a comprehensive attribute set is created for the merged node; and relationships between the same nodes are merged;
[0042] Identify the association between knowledge points in different books, obtain the logical relationship between subjects, and supplement the fused global knowledge graph based on the logical relationship.
[0043] According to a second aspect of the present application, there is provided a device for extracting a knowledge spectrum graph for a teaching material, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method for extracting a knowledge spectrum graph for a teaching material according to any one of the above-described methods is implemented.
[0044] According to a third aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for extracting a knowledge spectrum graph for teaching materials according to any one of the above-mentioned methods is implemented.
[0045] The method for extracting knowledge graphs from textbooks provided in this application introduces heterogeneous data classification and targeted processing mechanisms. It is no longer limited to text content, but integrates pictures and video information into the knowledge graph construction framework, which is more comprehensive. From catalog recognition to entity and relationship extraction, and then to multimodal knowledge point matching, the large language model plays a core role in each step, significantly reducing manual participation and improving the automation level and accuracy of the overall knowledge graph construction process. The hierarchical construction from chapter subgraphs to global graphs is realized, and semantic consistency is maintained during the fusion process. In addition, the knowledge points in non-text data are identified using a large language model and associated with the nodes of the text graph, providing more dimensional information support and enhancing the coverage and practicality of the knowledge graph. In addition, the present application also provides a device and storage medium for extracting knowledge graphs from textbooks with the above-mentioned technical effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0047] Figure 1 A flowchart of a specific implementation method of the method for extracting knowledge spectrograms from teaching materials provided in this application;
[0048] Figure 2 A schematic diagram of the process of fusing all the knowledge point subgraphs of each chapter into a global knowledge graph;
[0049] Figure 3 A schematic diagram of the process of integrating the global knowledge graphs of different books according to the logic of the same subject;
[0050] Figure 4 A flowchart of another specific implementation of the method for extracting knowledge spectrograms from teaching materials provided in this application;
[0051] Figure 5 This is a structural block diagram of a device for extracting knowledge spectrograms from teaching materials. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. It should be noted that, in the absence of conflict, the embodiments in the present disclosure and the features in the embodiments can be combined, separated, interchanged and / or rearranged with each other. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0053] The terms used here are for the purpose of describing specific embodiments, and are not intended to be restrictive. As used here, unless the context clearly indicates otherwise, the singular forms "one (kind, person)" and "said (the)" are also intended to include plural forms. In addition, when the terms "comprise" and / or "include" and their variations are used in this specification, it is explained that there are stated features, integral bodies, steps, operations, parts, assemblies and / or their groups, but it is not excluded that there are or add one or more other features, integral bodies, steps, operations, parts, assemblies and / or their groups. It should also be noted that, as used here, the terms "substantially", "approximately" and other similar terms are used as approximate terms and not as degree terms, so that they are used to explain the inherent deviations of the measured values, calculated values and / or the values provided that will be recognized by those of ordinary skill in the art.
[0054] A flowchart of a specific implementation method of the method for extracting knowledge spectrograms from textbooks provided in this application is as follows Figure 1 As shown, the method specifically includes:
[0055] S101: Acquire multi-source heterogeneous data; the multi-source heterogeneous data at least includes text, picture, and video multimodal data.
[0056] Multi-source heterogeneous data includes text, image, and video data. Text data can include: the main text content, title, directory, index, etc. of the textbook. Image data can include: formula diagrams, flow charts, structure diagrams, etc. Video data can include: teaching videos, experimental demonstrations, etc.
[0057] S102: Separate the multi-source heterogeneous data to obtain text teaching material data and non-text teaching material data.
[0058] Automatically identify data formats and classify them. For text-based textbook data, extract chapters and text. For non-text-based textbook data, extract content as textual information through technologies such as OCR and video-to-text conversion.
[0059] S103: Input the text teaching material data into a large language model, identify the directory structure of the text teaching material data; and obtain the positioning content information of each chapter according to the directory structure.
[0060] The large language model can use GLM-4, or other large models, such as GPT-4, Wenxin Yiyan, Tongyi Qianwen, etc.
[0061] Use a large language model to parse the directory structure and automatically generate chapter location information. Locate the start and end positions of each chapter according to the directory and associate them to obtain the location content information.
[0062] S104: Using a large language model to process the positioning content information of each chapter, identifying entities, associations, and attribute information in the chapter; and constructing a knowledge point subgraph for each chapter based on the entities, associations, and attribute information.
[0063] A large language model is used to semantically parse the positioning content information of each chapter and output semantically important entities; language tools are used to identify proper names and terms; non-knowledge points are excluded and entities related to the main body of the chapter are retained.
[0064] Based on the extracted entities and the corresponding context text, a large language model is used to obtain the association relationships between entities, such as causal, hierarchical, and associated relationships.
[0065] The context of each entity is summarized to obtain entity attribute information; and the relationship attribute information of the association relationship is generated using a large language model.
[0066] Based on the extracted knowledge points, relations and attributes, a chapter-level knowledge point subgraph is generated.
[0067] S105: Traverse to obtain the knowledge point subgraphs of all chapters, and merge the knowledge point subgraphs of all chapters into a global knowledge graph.
[0068] Traverse all chapter subgraphs, and integrate the knowledge point subgraphs of all chapters into a global knowledge graph based on entity alignment and relationship merging methods.
[0069] S106: Use a large language model to identify the non-text teaching material data and determine the knowledge points mentioned in the non-text teaching material data; match corresponding knowledge point nodes from the global knowledge graph and associate the non-text teaching material data with the corresponding knowledge point nodes in the global knowledge graph.
[0070] Among them, the non-text teaching material data includes video teaching material data; the video subtitle text of the video teaching material data is obtained; the video subtitle text is summarized by using a large language model, and the text description of the main content is used as the knowledge point mentioned in the video teaching material data.
[0071] Among them, the non-text teaching material data includes picture teaching material data; the picture title text of the picture teaching material data is obtained; the picture title text is summarized by using a large language model, and the text description of the main content is used as the knowledge point mentioned in the picture teaching material data.
[0072] Matching corresponding knowledge point nodes from the global knowledge graph and associating the non-text teaching material data to the corresponding knowledge point nodes of the global knowledge graph includes: performing semantic similarity matching on the knowledge points mentioned in the non-text teaching material data and the knowledge point nodes of the global knowledge graph, and when the semantic similarity exceeds a preset threshold, using the corresponding knowledge point nodes as the knowledge nodes matched by the non-text teaching material data; adding a relationship to the corresponding knowledge point nodes of the global knowledge graph to mark the content of the non-text teaching material data.
[0073] The similarity matching scheme used in the embodiment of the present application is vector similarity. Other non-vector similarity schemes, such as jaccard similarity, can be used to achieve the same effect as the present application.
[0074] Matching corresponding knowledge point nodes from the global knowledge graph includes: obtaining the source chapter of the non-text teaching material data; matching corresponding knowledge point nodes from the global knowledge graph, and assigning a higher weight to the knowledge point nodes matched in the source chapter than to the knowledge point nodes matched in other chapters, so that the knowledge point nodes in the source chapter can be matched preferentially.
[0075] The method for extracting knowledge graphs from textbooks provided in this application introduces heterogeneous data classification and targeted processing mechanisms. It is no longer limited to text content, but integrates pictures and video information into the knowledge graph construction framework, which is more comprehensive. From catalog recognition to entity and relationship extraction, and then to multimodal knowledge point matching, the large language model plays a core role in each step, significantly reducing manual participation and improving the automation level and accuracy of the overall knowledge graph construction process. The hierarchical construction from chapter subgraphs to global graphs is realized, and semantic consistency is maintained during the fusion process. In addition, the use of large language models to identify knowledge points in non-text data and associate them with nodes of the text graph provides more dimensional information support and enhances the coverage and practicality of the knowledge graph.
[0076] In the above embodiment, if Figure 2 As shown in Figure 1, the process of fusing all the knowledge point subgraphs of each chapter into a global knowledge graph specifically includes:
[0077] S201: Use semantic analysis technology to calculate the similarity of knowledge points in different chapters and select entity pairs with high semantic similarity.
[0078] Semantic analysis technology can calculate similarity by using word vectors, sentence vectors, semantic embedding generated by large models, etc. Entity pairs with high semantic similarity are selected, such as two knowledge points with similar names, or descriptions with overlapping content.
[0079] S202: judging the screened similar entity pairs.
[0080] S203: If the entity names and entity descriptions of similar entity pairs are similar, the entity pairs are directly merged.
[0081] For example, "Newton's Laws of Motion" and "Newton's Laws of Motion" have highly consistent entity names and entity descriptions, so they can be merged into one node.
[0082] S204: If only one of the entity name and the entity description of the similar entity pair is similar, a large language model is used to summarize and generate a description; and relevant personnel confirm whether to merge based on the description.
[0083] The large model is used to summarize the content, extract similarities, and generate descriptions. Domain experts will then make manual judgments to confirm whether to merge.
[0084] S205: After the entities are merged, the relations with high semantic similarity under the same entity are merged.
[0085] If two relationship edges (such as "definition" and "concept") under the same entity (such as "Newton's laws of motion") have high semantic similarity, they can be merged into one relationship edge.
[0086] On the basis of any of the above embodiments, after fusing the knowledge point sub-graphs of all chapters into a global knowledge graph, it also includes: fusing the global knowledge graphs of different books according to the logic of the same subject to form a course knowledge graph including multiple books in the same field or a professional knowledge graph including multiple different fields.
[0087] like Figure 3 As shown in Figure 1, the process of integrating the global knowledge graphs of different books according to the logic of the same subject specifically includes:
[0088] S301: Unify the semantic expressions of the differences in expressions and terminology in different books; merge repeated or similar nodes.
[0089] Unify semantic expressions to eliminate differences in expressions and terminology in different books. Merge duplicate or similar nodes to reduce graph redundancy.
[0090] S302: Identify semantically similar knowledge points in different book graphs and align the knowledge points.
[0091] S303: Identify and unify the relationships between knowledge points in different books and perform relationship alignment.
[0092] S304: Merge the aligned knowledge points and relationships into a unified knowledge graph to obtain a fused global knowledge graph; wherein, nodes with similar semantics are merged into one node, and a comprehensive attribute set is created for the merged node; and relationships between the same nodes are merged.
[0093] S305: Identify the association between knowledge points in different books, obtain the logical relationship between subjects, and supplement the fused global knowledge graph based on the logical relationship.
[0094] From chapter map to professional map, it is a multi-level process of knowledge abstraction and integration:
[0095] From chapter map to book map: form a book knowledge system by integrating chapter knowledge points to ensure rigorous logic and complete coverage.
[0096] From book map to course map: involving multiple books in the same field, the process involves unified terminology, standardized relationships, and elimination of redundancy.
[0097] From course map to professional map: further expand to interdisciplinary fields, such as integrating relevant concepts in "physics" and "engineering mechanics" to form a comprehensive professional knowledge map.
[0098] The flowchart of another specific implementation method of the method for extracting knowledge spectrogram from textbooks provided in this application is as follows Figure 4 As shown, the process specifically includes:
[0099] S401: Acquire multi-source heterogeneous data.
[0100] S402: Preprocess multi-source heterogeneous data.
[0101] First, the multi-source heterogeneous data is separated. The input data in this process is: original textbook data, involving multiple formats such as DOC, PDF, PPT, etc. The separation process specifically includes: separating the text data and image data in the original textbook data; converting non-TXT format text to TXT format files; storing the image data with its corresponding title name. The output is: TXT format file of the original textbook data, and the image data storage path.
[0102] Furthermore, the text data is preprocessed. This process includes: inputting the separated text teaching material data, converting the non-TXT format texts such as DOC, PDF, PPT into TXT format files, cleaning out invalid texts such as prefaces, references, and garbled characters, etc., and outputting the preprocessed text teaching material data.
[0103] Preprocessing is performed on the video data. The process includes: inputting the video subtitle text. Using the large language model to summarize the video subtitle text, the text description of the main content is output as the knowledge point mentioned in the picture teaching material data. It is understandable that the output can be in json format.
[0104] Preprocess the image data. This process includes: inputting the image title text. Summarizing the image title text using a large language model, and outputting the text description of the main content as the knowledge points mentioned in the image teaching material data. It is understandable that the output can be in json format.
[0105] S403: Construct a knowledge graph.
[0106] First, the textbook directory is identified. The preprocessed text textbook (in file units) is input into the large language model, and the large language model is used to identify the directory structure of the textbook and extract the chapter titles and subtitles. According to the directory structure, the content range of each chapter is located to obtain the location content information of each chapter.
[0107] According to the positioning content information of each chapter, the entities, association relationships and attribute information in the chapter are identified; based on the entities, association relationships and attribute information, a knowledge point subgraph of each chapter is constructed.
[0108] Entity recognition: Use a large language model to identify the core knowledge points in the chapter as entities.
[0109] Relation extraction: Use a large language model to identify the associations between knowledge points.
[0110] Attribute extraction: Use a large language model to extract descriptive attributes for knowledge point entities, such as definitions, examples, etc.
[0111] Construct subgraphs: Construct knowledge point subgraphs based on entities, relationships, and attributes to obtain knowledge point subgraphs for each chapter (in the form of mind maps).
[0112] S404: Perform knowledge point sub-graph fusion.
[0113] The knowledge point subgraphs of all chapters are processed as follows:
[0114] Entity alignment: Based on entity attributes, a large language model and similarity matching scheme are used to identify and merge different nodes describing the same entity.
[0115] Relation merging: Combine the large language model and similarity matching scheme to merge different edges describing the same relation.
[0116] Conflict Resolution: Handling conflicts between entities and relationships.
[0117] Build a global graph: merge all subgraphs into a global knowledge graph.
[0118] Format conversion: organize the graph data into the "head entity-relationship / attribute-tail entity" format.
[0119] A global knowledge graph is obtained, which can be in the format of a csv file.
[0120] S405: Supplement heterogeneous information into the knowledge graph.
[0121] For video textbook data, compare the preprocessed video to the textbook (json data), and recall the candidate knowledge point names from the global graph based on the knowledge points mentioned in json. Use the large language model to determine the best match among the candidate knowledge points. Mount the video to the knowledge point node in the form of a link and output it as a csv file.
[0122] For the image textbook data, compare the preprocessed images to the textbook (json data), and recall the candidate knowledge point names from the global graph based on the knowledge points mentioned in json. Use the large language model to determine the best match among the candidate knowledge points. Mount the image on the knowledge point node in the form of a link and output it as a csv file.
[0123] This application is centered on core knowledge points, and the number and content of nodes are screened to ensure that the description of key knowledge points is complete without causing information fragmentation due to too many details. Improve generalization in knowledge structure design and support interdisciplinary applications. The introduction of large language models automates the processes of semantic extraction, node induction, and relationship generation, greatly reducing the workload and complexity of manual participation, improving processing efficiency, and improving accuracy. This application is designed to be compatible with multi-source heterogeneous teaching material information, and can effectively process teaching material content of different formats and structures to achieve information unification and integration. Through the in-depth understanding and analysis of the large language model, this application can more accurately identify and extract knowledge points in the textbooks, ensuring the accuracy and completeness of the knowledge graph.
[0124] In addition, the present application provides a device for extracting knowledge spectrograms from teaching materials, such as Figure 5 As shown in the structural block diagram of the device for extracting knowledge spectra from textbooks, the device specifically includes a memory 51 and a processor 52. The memory 51 stores a computer program. When the computer program is executed by the processor 52, it implements the method for extracting knowledge spectra from textbooks according to any of the above-mentioned methods.
[0125] In addition, the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method for extracting a knowledge spectrum graph for teaching materials according to any one of the above-described methods is implemented.
[0126] Computer readable storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0127] The professionals should also be further aware that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0128] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0129] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the specific implementation method of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A method for extracting a knowledge spectrum graph for teaching materials, characterized in that: include: Acquire multi-source heterogeneous data; the multi-source heterogeneous data at least includes text, picture, and video multimodal data; Separating the multi-source heterogeneous data to obtain text teaching material data and non-text teaching material data; Inputting the text teaching material data into a large language model, identifying the directory structure of the text teaching material data; and obtaining the positioning content information of each chapter according to the directory structure; A large language model is used to process the positioning content information of each chapter to identify the entities, associations, and attribute information in the chapter; Based on the entities, association relationships and attribute information, construct a knowledge point subgraph for each chapter; Traverse to obtain the knowledge point subgraphs of all chapters, and merge the knowledge point subgraphs of all chapters into a global knowledge graph; The non-text teaching material data is identified by using a large language model to determine the knowledge points mentioned in the non-text teaching material data; the corresponding knowledge point nodes are matched from the global knowledge graph, and the non-text teaching material data is associated with the corresponding knowledge point nodes in the global knowledge graph.
2. The method for extracting knowledge spectrogram from teaching materials according to claim 1, characterized in that: The adopting of a large language model to identify the non-text teaching material data and determining the knowledge points mentioned in the non-text teaching material data includes: The non-text teaching material data includes video teaching material data; obtaining video subtitle text of the video teaching material data; summarizing the video subtitle text using a large language model, and using the text description of the main content as the knowledge points mentioned in the video teaching material data; The non-text teaching material data includes picture teaching material data; obtaining picture title text of the picture teaching material data; using a large language model to summarize the picture title text, and using the text description of the main content as the knowledge point mentioned in the picture teaching material data.
3. The method for extracting knowledge spectrogram from teaching materials according to claim 2, characterized in that: The matching of corresponding knowledge point nodes from the global knowledge graph and associating the non-text teaching material data with the knowledge point nodes corresponding to the global knowledge graph comprises: Perform semantic similarity matching between the knowledge points mentioned in the non-text teaching material data and the knowledge point nodes of the global knowledge graph, and when the semantic similarity exceeds a preset threshold, use the corresponding knowledge point node as the knowledge node matched by the non-text teaching material data; Add a relationship to the knowledge point node corresponding to the global knowledge graph to mark the content of the non-text teaching material data.
4. The method for extracting knowledge spectrogram from teaching materials according to claim 3 is characterized in that: The matching of corresponding knowledge point nodes from the global knowledge graph includes: Obtain the source chapter of the non-text textbook data; Corresponding knowledge point nodes are matched from the global knowledge graph, and a higher weight is given to the knowledge point nodes matched in the source chapter than to the knowledge point nodes matched in other chapters.
5. The method for extracting knowledge spectrogram from teaching materials according to claim 3, characterized in that: The entities, relationships and attribute information in the identification section include: Use a large language model to perform semantic analysis on the positioning content information of each chapter and output semantic entities; use language tools to identify proper names and terms; exclude non-knowledge points and retain entities related to the main body of the chapter; Based on the extracted entities and the corresponding context text, a large language model is used to obtain the association relationship between entities; The context of each entity is summarized to obtain entity attribute information; and the relationship attribute information of the association relationship is generated using a large language model.
6. The method for extracting knowledge spectrogram from teaching materials according to claim 1, characterized in that: The fusion of all the knowledge point subgraphs of the chapters into a global knowledge graph includes: Use semantic analysis technology to calculate the similarity of knowledge points in different chapters and select entity pairs with high semantic similarity; Make judgments on the screened similar entity pairs; If the entity names and entity descriptions of similar entity pairs are similar, the entity pairs are directly merged; If only one of the entity name and entity description of the similar entity pair is similar, a large language model is used to summarize and generate a description; relevant personnel confirm whether to merge based on the description; After merging the entities, the relations with high semantic similarity under the same entity are merged.
7. The method for extracting knowledge spectrogram from teaching materials according to any one of claims 1 to 6, characterized in that: After the knowledge point subgraphs of all chapters are merged into a global knowledge graph, the following steps are also included: The global knowledge graphs of different books are integrated according to the logic of the same subject to form a course knowledge graph that includes multiple books in the same field or a professional knowledge graph that includes multiple different fields.
8. The method for extracting knowledge spectrogram from teaching materials according to claim 7, characterized in that: The fusion of the global knowledge graphs of different books according to the logic of the same subject includes: Unify the semantic expressions of the differences in expressions and terminology in different books; merge repeated or similar nodes; Identify knowledge points with similar semantics in different book graphs and align them; Identify and unify the relationships between knowledge points in different books and perform relationship alignment; The aligned knowledge points and relationships are merged into a unified knowledge graph to obtain a fused global knowledge graph; where nodes with similar semantics are merged into one node, and a comprehensive attribute set is created for the merged node; and relationships between the same nodes are merged; Identify the association between knowledge points in different books, obtain the logical relationship between subjects, and supplement the fused global knowledge graph based on the logical relationship.
9. A device for extracting knowledge spectrograms from teaching materials, characterized in that: It includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, it implements the method for extracting a knowledge spectrum graph for teaching materials according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 8. ZSP240687CN Method for extracting knowledge spectrum from textbooks.
Citation Information
Cited By
Digital course generation method and system based on large language model
CN120471029A
Data analysis method, device and equipment
CN121351977A
Educational resource big data visual analysis method and system based on knowledge graph
CN121658669A
Knowledge graph-based education resource big data visual analysis method and system
CN121658669B