Education knowledge graph construction method, device and equipment

By automatically extracting subject concepts and generating multimodal educational knowledge graphs using large language models (LLMs), the problems of complex training and insufficient multimodal support in existing technologies are solved. This enables high-precision, multi-dimensional educational knowledge expression and resource integration, thereby improving learning efficiency and teaching quality.

CN121809631APending Publication Date: 2026-04-07BEIJING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing methods for constructing educational knowledge graphs rely on manual annotation and deep learning models, which are complex to train, costly, and difficult to adapt to data changes. Furthermore, they are mainly based on a single text modality and lack multimodal information support, thus failing to meet the needs of multimodal educational applications.

Method used

Large Language Models (LLMs) are used to extract subject concepts, and multimodal information is combined to generate an educational knowledge graph. Subject concepts are automatically extracted through semantic understanding and reasoning capabilities, and multimodal data alignment and fusion are performed to generate multidimensional knowledge representations.

Benefits of technology

It achieves high-precision and complete extraction of subject concepts, breaks through the limitations of text modality, supports the deep integration of multimodal educational resources, and improves learning efficiency and teaching quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809631A_ABST
    Figure CN121809631A_ABST
Patent Text Reader

Abstract

The invention provides an educational knowledge graph construction method, device and equipment. The method comprises the following steps: acquiring multi-source heterogeneous educational data; based on a large language model, subject concepts are extracted from the education data; and according to the subject concept and the multi-modal information corresponding to the subject concept, generating a multi-modal education knowledge graph. According to the method provided by the embodiment of the invention, on one hand, automatic and high-precision extraction of subject concepts is realized by relying on semantic understanding and reasoning capabilities of the large language model LLMs, and compared with a traditional method depending on manual data annotation and special deep learning model training, the method does not need a complex model training process and high annotation cost, and is high in efficiency and high in efficiency. And meanwhile, the precision and integrity of subject concept extraction can be effectively improved. And on the other hand, the method breaks through the limitation of a traditional text modal knowledge graph, the deep fusion of multi-modal education resources is realized by generating the multi-modal education knowledge graph and expanding single text knowledge into multi-dimensional knowledge expression, so that the traditional education mode is promoted to be upgraded to intelligentization and multi-modality, and the intellectualization and the multi-modality are improved. And powerful support is provided for improving learning efficiency and teaching quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of educational knowledge graph technology, and in particular to a method, apparatus and device for constructing an educational knowledge graph. Background Technology

[0002] In related technologies, traditional educational knowledge bases extract subject concepts and knowledge points from textbooks. Although these knowledge bases can effectively store subject knowledge, they typically save this knowledge directly as static files in JSON format. This limits the expressive power and application scope of the knowledge base, making it difficult to realize its full value in knowledge sharing, personalized teaching and tutoring, and intelligent education systems. Summary of the Invention

[0003] This invention provides a method, apparatus, and device for constructing an educational knowledge graph. On one hand, relying on the semantic understanding and reasoning capabilities of Large Language Models (LLMs), it achieves automated and high-precision extraction of subject concepts. Compared to traditional methods that rely on manually labeled data and training dedicated deep learning models, this method eliminates the need for complex model training processes and high labeling costs, while effectively improving the accuracy and completeness of subject concept extraction. On the other hand, this method overcomes the limitations of traditional text-based modal knowledge graphs. By generating a multimodal educational knowledge graph, it expands single textual knowledge into multi-dimensional knowledge expressions, achieving deep integration of multimodal educational resources. This, in turn, promotes the upgrading of traditional education models towards intelligence and multimodality, providing strong support for improving learning efficiency and teaching quality.

[0004] This invention provides a method for constructing an educational knowledge graph, comprising the following steps.

[0005] Acquire multi-source heterogeneous educational data; Based on a large language model, subject concepts are extracted from the educational data. A multimodal educational knowledge graph is generated based on the subject concepts and the corresponding multimodal information.

[0006] According to the present invention, an educational knowledge graph construction method is provided, wherein the extraction of subject concepts from the educational data based on a large language model includes: The educational data is input into a large language model, which outputs candidate subject concepts. Large language models trained with different corpora output the determination results and reasons for whether the candidate subject concepts belong to subject concepts; based on the determination results and reasons, the scoring results corresponding to each candidate subject concept are determined. Based on the scoring results of the candidate subject concepts, a subject concept is determined from the candidate subject concepts using a voting mechanism.

[0007] According to the educational knowledge graph construction method provided by the present invention, after extracting subject concepts from the educational data based on a large language model, the method further includes: The subject concepts and concept networks are matched to obtain the concepts that are semantically associated with the subject concepts; The subject concept is expanded based on the concept that is semantically associated with the subject concept.

[0008] According to a method for constructing an educational knowledge graph provided by the present invention, the step of generating a multimodal educational knowledge graph based on the subject concept and the corresponding multimodal information of the subject concept includes: Align the subject concepts and their corresponding multimodal information to generate a multimodal educational knowledge graph.

[0009] According to a method for constructing an educational knowledge graph provided by the present invention, the step of aligning the subject concept and the corresponding multimodal information to generate a multimodal educational knowledge graph includes: Input the subject concepts and Wikipedia content into the large language model to obtain textual explanation information aligned with the subject concepts; The subject concept and the educational data are input into a large language model to obtain image information aligned with the subject concept; Based on the text explanation information aligned with the subject concept, audio information aligned with the subject concept is obtained; Based on the time information of the occurrence of the subject concept in the educational resources, video information aligned with the subject concept is obtained; Based on the text explanation information aligned with the subject concept, the image information aligned with the subject concept, the audio information aligned with the subject concept, and the video information aligned with the subject concept, the multimodal information aligned with the subject concept is obtained; A multimodal educational knowledge graph is generated based on the multimodal information aligned with the subject concepts.

[0010] According to a method for constructing an educational knowledge graph provided by the present invention, the step of generating a multimodal educational knowledge graph based on multimodal information aligned with the subject concepts includes: Based on the multi-source heterogeneous educational data, obtain subject information and multiple subject knowledge points corresponding to the subject information; Based on the subject information, the relationships between the various subject knowledge points, the relationships between the subject knowledge points and the subject concepts, and the multimodal information aligned with the subject concepts, a hierarchical multimodal educational knowledge graph is generated.

[0011] The present invention also provides an educational knowledge graph construction device, comprising the following modules: The acquisition module is used to acquire multi-source heterogeneous educational data; An extraction module is used to extract subject concepts from the educational data based on a large language model. The generation module is used to generate a multimodal educational knowledge graph based on the subject concept and the corresponding multimodal information.

[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the educational knowledge graph construction method as described above.

[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the educational knowledge graph construction method as described above.

[0014] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the educational knowledge graph construction method as described above.

[0015] The educational knowledge graph construction method, apparatus, and device provided by this invention, on the one hand, rely on the semantic understanding and reasoning capabilities of large language models (LLMs) to achieve automated and high-precision extraction of subject concepts. Compared with traditional methods that rely on manually labeled data and training dedicated deep learning models, this method eliminates the need for complex model training processes and high labeling costs, while effectively improving the accuracy and completeness of subject concept extraction. On the other hand, this method breaks through the limitations of traditional text-based modal knowledge graphs. By generating multimodal educational knowledge graphs, it expands single textual knowledge into multi-dimensional knowledge expressions, achieving deep integration of multimodal educational resources. This, in turn, promotes the upgrading of traditional education models towards intelligence and multimodality, providing strong support for improving learning efficiency and teaching quality. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 This is one of the flowcharts illustrating the educational knowledge graph construction method provided by this invention.

[0018] Figure 2This is the second flowchart of the educational knowledge graph construction method provided by the present invention.

[0019] Figure 3 This is a schematic diagram of the educational knowledge graph construction device provided by the present invention.

[0020] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0022] The following is combined Figures 1-4 The present invention describes a method, apparatus, and device for constructing an educational knowledge graph.

[0023] To facilitate a clearer understanding of the technical solutions of the various embodiments of this application, some technical content related to the various embodiments of this application will be introduced first.

[0024] Educational knowledge graphs are a method of storing subject-specific knowledge. They record relationships between concepts, between concepts and knowledge points, and the attributes of concepts using triples. They aim to improve the efficiency of organizing and retrieving educational resources, supporting personalized learning and intelligent education. Simultaneously, educational knowledge graphs can provide precise data analysis and insights for academic research and educational decision-making.

[0025] Educational knowledge graph construction is the process of organizing and presenting knowledge points, concepts, and their relationships in the field of education in a structured and visual form. It utilizes technologies such as natural language processing and machine learning to extract and integrate knowledge content from textbooks, courses, academic literature, and teaching resources. The resulting graphs can help teachers and students achieve precise teaching, personalized learning, and the exploration of knowledge connections.

[0026] Table 1 Education Dataset Educational datasets time Data source Organizational methods automation Multimodal automation LectureBank 2019 MOOCs knowledge base MOOCcube 2020 MOOCs knowledge base MOOCubeX 2021 MOOCs knowledge base MoocRadar 2023 MOOCs knowledge base KnowEdu 2018 Curriculum Standard MOOC knowledge graph MEduKG 2022 Teaching books knowledge graph GPTEKG 2024 MOOCs knowledge graph ✓ LLM4EduKG 2024 Teaching books knowledge graph ✓ EduMKG 2025 MOOCs + Curriculum Standards + Concept Networks + Wikipedia knowledge graph ✓ ✓ ✓ As shown in Table 1, traditional educational knowledge bases such as LectureBank, MOOCcube, MOOCubeX, and MoocRadar extract subject concepts and knowledge points from textbooks and MOOCs. While these knowledge bases can effectively store subject knowledge, they typically save this knowledge directly as static JSON files, rather than constructing a dynamically queryable and expandable knowledge graph based on relationships and semantic information. For example, they cannot reflect the logical connections, semantic relationships, and evolutionary processes between subject knowledge. This limitation restricts the expressive power and application scope of the knowledge base, making it difficult to realize its full value in knowledge sharing, personalized tutoring, and intelligent education systems.

[0027] To address these issues, KnowEdu and MEduKG, two methods for constructing educational knowledge graphs, extract subject knowledge from primary and secondary school curriculum standards, books, and MOOCs by manually annotating datasets and manually training BERT-based subject concept extraction models. However, due to the dynamic and constantly updated nature of knowledge in the educational field, the trained models exhibit insufficient adaptability when faced with new data.

[0028] In recent years, constructing educational knowledge graphs using large language models has become a research hotspot. For example, GPTEKG uses ChatGPT to extract subject concepts from MOOCs content and construct a graph, while LLM4EduKG extracts subject knowledge points from textbooks based on large language models, while simultaneously mining the relationships between knowledge points and concepts to construct a single-text modality educational knowledge graph. These methods have improved the efficiency and coverage of knowledge extraction to some extent. However, they also have several significant problems in their implementation. First, these methods do not fully consider the potential "illusionary output" problem when using large language models. Large language models sometimes generate content that is inconsistent with reality or does not exist in the input information, which directly affects the accuracy and consistency of the subject concepts extracted from the model, thus weakening the scientific validity and reliability of the knowledge graph. Second, the knowledge graphs generated by these construction methods are mainly for a single text modality, lacking support for other modal information and making it difficult to cover multimodal data in modern education, such as images, videos, and audio. This limitation prevents the application of knowledge graphs from meeting the needs of educational scenarios such as multimodal exercise generation and multimodal question answering, which are crucial areas for the development of intelligent education systems. Furthermore, the lack of multimodal information also restricts the application potential of knowledge graphs in terms of interpretability of the learning process and diversity of learning content.

[0029] Current methods for constructing educational knowledge graphs use deep learning models to extract knowledge points. However, deep learning models require tedious and expensive manual annotation of training data, and their adaptability to new data is poor. Furthermore, these methods extract knowledge points and subject concepts from textual modal data such as textbooks, curriculum standards, and learners' learning data, ignoring the multimodal nature of subject concepts—that is, each subject corresponds to related images, audio, video, and other information. This means the knowledge graph cannot support multimodal teaching applications. Simultaneously, single-text modal data is not conducive to learners' understanding and mastery of knowledge. Moreover, while these methods extract educational knowledge from classroom resources such as textbooks and curriculum standards, the scope and depth of this knowledge are relatively limited, making it difficult to fully cover the rich and diverse situations and needs in educational practice.

[0030] In summary, current methods for constructing educational knowledge graphs primarily rely on manually training deep learning models to extract subject knowledge, mainly in text mode, from multi-source heterogeneous data. However, these methods struggle to address the issue of insufficient model generalization ability caused by data variations. Furthermore, the lack of multimodal educational resources such as images, videos, and audio significantly limits the application potential of subject knowledge in multimodal representation. To address these problems, we propose using a large language model approach. Large language models demonstrate powerful contextual understanding and semantic reasoning capabilities in natural language processing tasks, enabling more efficient extraction and deep integration of subject concepts from multimodal data, thus compensating for the shortcomings of existing methods in manual data annotation, knowledge reasoning, understanding, and expression. Although existing deep learning methods have achieved some success in multimodal knowledge alignment, their limited versatility and cross-domain transferability make them difficult to adapt to different educational data scenarios.

[0031] Figure 1 This is one of the flowcharts illustrating the educational knowledge graph construction method provided by the present invention. The method includes the following: Step 101: Obtain multi-source heterogeneous educational data.

[0032] Specifically, in this embodiment, multi-source heterogeneous educational data is first acquired. Optionally, a web crawler script can be designed to automatically crawl teaching videos and supporting courseware from open-source MOOC platforms. Optionally, multi-source data such as curriculum standard documents, related concept data in ConceptNet, and educational knowledge from Wikipedia can also be acquired to provide rich and diverse data source support for the knowledge graph.

[0033] Step 102: Extract subject concepts from educational data based on the large language model.

[0034] Specifically, after acquiring multi-source heterogeneous educational data, this embodiment of the application relies on the semantic understanding and reasoning capabilities of large language models (LLMs) to extract subject concepts from the educational data, thereby achieving automated and high-precision extraction of subject concepts.

[0035] It should be noted that, compared to traditional methods that rely on manual annotation and training of dedicated deep learning models, this application eliminates the need for complex model training and high annotation costs, effectively improving the accuracy and completeness of concept extraction. Furthermore, the large language model exhibits strong adaptability to multi-source heterogeneous educational data, enabling rapid extraction of subject concepts in new scenarios without retraining the model. This effectively addresses the problem of poor data adaptability in traditional models and meets the needs of dynamic updates in educational knowledge.

[0036] Step 103: Generate a multimodal educational knowledge graph based on the subject concepts and the corresponding multimodal information.

[0037] Specifically, after extracting subject concepts from educational data based on a large language model, this application can accurately associate multiple modal information such as text, video, audio, and images for subject concepts, ultimately generating a multimodal educational knowledge graph. This enables deep integration of multimodal educational resources, breaks through the limitations of traditional text-based knowledge graphs, meets the needs of intelligent education scenarios such as multimodal exercise generation and multimodal question answering, promotes the upgrading of traditional education models to intelligence and multimodality, and helps improve learning efficiency and teaching quality.

[0038] The method described above, on the one hand, relies on the semantic understanding and reasoning capabilities of Large Language Models (LLMs) to achieve automated and high-precision extraction of subject concepts. Compared to traditional methods that rely on manually labeled data and training dedicated deep learning models, this method eliminates the need for complex model training processes and high labeling costs, while effectively improving the accuracy and completeness of subject concept extraction. On the other hand, this method breaks through the limitations of traditional text-based modal knowledge graphs. By generating multimodal educational knowledge graphs, it expands single textual knowledge into multidimensional knowledge expressions, achieving deep integration of multimodal educational resources. This, in turn, promotes the upgrading of traditional education models towards intelligence and multimodality, providing strong support for improving learning efficiency and teaching quality.

[0039] In some embodiments, based on a large language model, subject concepts are extracted from educational data, including: Input educational data into a large language model and output candidate subject concepts; Large language models trained with different corpora output the judgment results and reasons for whether candidate subject concepts belong to subject concepts; based on the judgment results and reasons, the scoring results corresponding to each candidate subject concept are determined. Based on the scoring results of the candidate subject concepts, and using a voting mechanism, the subject concepts are determined from the candidate subject concepts.

[0040] Specifically, in this embodiment, educational data is first input into a large language model to obtain candidate subject concepts. Optionally, for subtitles extracted from MOOC videos, LLMs can be used to merge paragraphs with the same semantics, and then prompt words can be used to guide LLMs to extract a sufficient number of candidate concepts according to a specified format, which greatly reduces human input and improves concept extraction efficiency.

[0041] Optionally, after extracting candidate subject concepts, in this embodiment, large language models trained on different corpora output the judgment results and reasons for whether the candidate subject concepts belong to subject concepts; based on the judgment results and reasons, the scoring results corresponding to each candidate subject concept are determined. Optionally, multiple large language models trained on different corpora can be selected, such as an LLM trained on education-specific corpora and an LLM trained on general corpora. Each model judges whether the candidate concept belongs to a valid subject concept and generates judgment reasons. Then, the large language model, such as GPT-4o, is used as the scoring model. The multi-model judgment results and judgment reasons corresponding to each candidate concept are input into the scoring model. The scoring model uses the subject suitability of the judgment results, the logical rigor of the judgment reasons, and the consistency of the multi-model results as the core evaluation dimensions to comprehensively quantify and score the candidate concepts, outputting a scoring result in the range of 0-1. That is, this application effectively compensates for the judgment bias caused by the limitation of the training corpus of a single model by cross-judging large models trained on different corpora, and effectively improves the accuracy of subject concepts.

[0042] Optionally, after determining the scoring results corresponding to each candidate subject concept, in this embodiment of the application, subject concepts are determined from the candidate subject concepts based on the scoring results and a voting mechanism, ultimately obtaining an accurate set of subject concepts. Optionally, a self-consistency voting mechanism can be used, and the confidence level of the final concept selection can be set to 0.7 to extract valid subject concepts.

[0043] For example, in this embodiment of the application, the extraction of subject concepts is divided into three stages: A1: Concept Extraction Stage. For subtitles extracted from MOOC videos, LLMs are first used to merge paragraphs with the same semantics, and then prompt words are used to guide LLMs to extract a sufficient number of candidate concepts according to a specified format.

[0044] A2: Concept Validation Stage. Addressing the issue that single LLMs may exhibit biases between their training corpora and correct answers due to varying training focuses, this application proposes a self-validating multi-LLMs concept validation mechanism. First, LLMs trained on different corpora determine whether each candidate concept belongs to a specific subject area and generate corresponding justifications. Next, the LLMs score the concepts based on their judgments and justifications.

[0045] A3: Concept Integration Stage. Utilizing a self-consistency voting mechanism, the voting results of multiple LLMs on candidate concepts are integrated, and the confidence level for the final concept selection is set to 0.7 to extract effective subject concepts.

[0046] The method described in the above embodiments, employing a three-stage automated concept extraction approach of "extraction-verification-integration," fully leverages the deep semantic understanding and powerful logical reasoning capabilities of large language models to achieve accurate, comprehensive, and efficient automated extraction of subject concepts. This method not only mines explicit information from text but also reasons about implicit logical connections, ensuring the completeness and accuracy of subject concepts. Simultaneously, it possesses a fully automatic concept correction mechanism that filters out erroneous concepts, ultimately yielding accurate subject concepts. From a technical advantage perspective, this method avoids the problems of complex training processes, high data annotation costs, and the illusionary output of large language models inherent in traditional deep learning concept extraction methods, eliminating the need for complex deep learning model training or costly manual annotation. Furthermore, it effectively addresses the pain points of traditional deep learning models' poor adaptability to new data and the need for retraining, meeting the demands of dynamic updates in educational knowledge. While achieving efficient and automated extraction of subject concepts, it also effectively guarantees the accuracy of those concepts.

[0047] In some embodiments, after extracting subject concepts from educational data based on a large language model, the method further includes: Match subject concepts with concept networks to obtain concepts semantically related to subject concepts; Expand the concept of a discipline based on the concepts that are semantically related to the discipline concept.

[0048] Specifically, after extracting subject concepts from educational data based on a large language model, this application can mine and extract related extended concepts from the ConceptNet based on the extracted subject concepts, thereby enriching the diversity and depth of the concept set. Optionally, to ensure the high accuracy and relevance of the extracted extended concepts, a large language model is introduced as a concept validator. Its powerful language understanding and reasoning capabilities are used to screen and verify related concepts, improving credibility and ensuring that the extended subject concepts not only conform to the logic of subject knowledge but also have a strong correlation with existing subject concepts, avoiding knowledge graph redundancy or knowledge bias caused by ineffective expansion.

[0049] The method described above extracts subject concepts from educational data based on a large language model, and then mines related extended concepts from the ConceptNet, thereby effectively overcoming the knowledge limitations of educational data sources, significantly broadening the knowledge coverage dimension of subject concepts, and making up for the problems of isolated concepts and single knowledge dimensions in traditional methods. Moreover, this application also uses a large language model to screen and verify the extended subject concepts, ensuring the strong correlation between the extended subject concepts and existing subject concepts and the consistency of subject knowledge logic, effectively avoiding knowledge graph redundancy or knowledge deviation caused by invalid extensions.

[0050] In some embodiments, a multimodal educational knowledge graph is generated based on subject concepts and the corresponding multimodal information, including: Align subject concepts with their corresponding multimodal information to generate a multimodal educational knowledge graph.

[0051] Specifically, in this embodiment, by aligning subject concepts with their corresponding multimodal information, a precise association between subject concepts and multimodal data is achieved. This ensures that each modal of data forms a semantically consistent, logically related, and error-free correspondence with the subject concept, ultimately expanding subject concepts from a single textual form to a multi-dimensional knowledge representation encompassing text, images, audio, and video. Through this precise multi-dimensional alignment and the generation of a multimodal educational knowledge graph, users can simultaneously access strongly related text, audio explanations, and experimental videos when querying subject concepts. This multi-sensory collaborative understanding deepens knowledge retention and effectively solves the problems of abstract and difficult-to-understand traditional single-text knowledge and vague memory points.

[0052] In some embodiments, aligning subject concepts and their corresponding multimodal information to generate a multimodal educational knowledge graph includes: By inputting subject concepts and Wikipedia content into a large language model, textual explanations aligned with the subject concepts are obtained. By inputting subject concepts and educational data into a large language model, image information aligned with subject concepts is obtained. Based on the textual explanation information aligned with the subject concept, audio information aligned with the subject concept is obtained; Based on the timing information of the appearance of subject concepts in educational resources, video information aligned with the subject concepts is obtained; Based on the textual explanation information aligned with the subject concept, the image information aligned with the subject concept, the audio information aligned with the subject concept, and the video information aligned with the subject concept, the multimodal information aligned with the subject concept is obtained; A multimodal educational knowledge graph is generated based on the multimodal information aligned with subject concepts.

[0053] Specifically, in this embodiment, an innovative multimodal alignment method that integrates semantic and structural features is adopted, aiming to achieve accurate and efficient cross-modal data alignment based on the characteristics of different modal attributes. This method promotes the deep fusion and collaborative processing of multimodal information by combining the semantic information and structural features of the data.

[0054] For example, the multimodal concept alignment method based on semantics and structure in this application embodiment is as follows: Text alignment: LLMs are used to generate a concise, accurate, and authoritative text explanation for each extracted concept, and this explanation is compared and improved with Wikipedia content to ensure the rigor and reliability of the generated text explanation, effectively solving semantic ambiguity and the illusion problem that may occur in the generated content by LLMs.

[0055] Image alignment: Leveraging the powerful visual recognition and semantic reasoning capabilities of MLLMs, the GPT-4o model is used to align the extracted images and concepts, and the Gemini-1.5Pro model is used as a validator to ensure the accuracy of the alignment. Optionally, a multimodal large language model can be used to generate semantic interpretations for each image in the educational resources, and then, based on these semantic interpretations, the alignment of concepts and images can be achieved from a semantic dimension.

[0056] Audio alignment: Based on the text explanation of each concept, a TTS model is used to generate audio explanations of the concepts, thereby aligning the concepts with the audio.

[0057] Video alignment: Based on the timestamps extracted from MOOCs, each concept is aligned with its corresponding video timestamp, achieving alignment between concepts and videos. In other words, it starts from the structural dimension, using timestamps to determine the specific segment of the video in which the concept resides, thereby achieving a precise association with the video modal data.

[0058] The method described above analyzes, aligns, and integrates multimodal data (text, images, videos, and audio) related to subject concepts from two key dimensions: semantics and structure. This method not only focuses on the representation accuracy of single-modal knowledge but also comprehensively considers the intrinsic correlation and semantic consistency between different modalities. This allows for the construction of more accurate and comprehensive associations among multimodal knowledge, improving the expressive power of cross-modal data and significantly enhancing the correlation and consistency between multimodal information.

[0059] In some embodiments, a multimodal educational knowledge graph is generated based on multimodal information aligned with subject concepts, including: Based on multi-source heterogeneous educational data, acquire subject information and the corresponding multiple subject knowledge points; Based on subject information, the relationships between various subject knowledge points, the relationships between subject knowledge points and subject concepts, and the multimodal information aligned with subject concepts, a hierarchical multimodal educational knowledge graph is generated.

[0060] Specifically, in this embodiment, a large language model is used to semantically classify multi-source data, identify the subject areas corresponding to the educational data, and for each subject's educational data, the large language model is used to extract and verify the knowledge points of each subject. Furthermore, based on all knowledge points, concepts, and their multimodal data, relationships are established between subjects, knowledge points, concepts, and multimodal data. Subjects, subject concepts, and knowledge points are scientifically layered and stored hierarchically, achieving a four-layer hierarchical association of subject-knowledge point-concept-multimodal information. This generates a hierarchical multimodal educational knowledge graph, accurately and efficiently restoring the internal logic of educational knowledge and effectively solving the problems of fragmented knowledge and chaotic logical associations in traditional knowledge graphs.

[0061] Optionally, to ensure the connectivity and practical application value of the graph, this application uses a knowledge symbolization organization method to assign unique identifiers to knowledge points, exercises, concepts, and their four modal information. Furthermore, a bidirectional mutual indexing file system is established between knowledge and concepts, and between concepts and modal data, for all knowledge points, concepts, and their modal data to improve access efficiency and support flexible multimodal retrieval operations.

[0062] The method described in the above embodiments scientifically hierarchizes disciplines, discipline concepts, and knowledge points. This approach not only constructs a semantic hierarchy of concepts but also systematically classifies and integrates multimodal data. Furthermore, through a cross-modal mutual indexing mechanism, it innovatively achieves efficient cross-retrieval between different modalities. This mechanism significantly improves the efficiency and depth of knowledge access, providing strong technical support for the utilization and application development of educational resources.

[0063] For example, such as Figure 2 As shown in the figure, this application provides a method for constructing an educational knowledge graph, the specific process of which is as follows: First, web crawlers are used to automatically acquire videos and courseware from MOOC platforms. Simultaneously, open-source image extraction programs are used to efficiently extract image resources from the courseware. Furthermore, OCR technology is incorporated to accurately extract knowledge points from the course standards, further optimizing the efficiency of information integration and knowledge mining.

[0064] Then, a three-stage "extraction-verification-integration" method is used to achieve efficient and automated extraction of subject concepts. First, a candidate concept set C is extracted from the MOOC subtitles using this method. Then, starting from candidate set C, it is linked to the ConceptNet network to explore subject concepts associated with the concepts in set C, and these associated concepts are added to candidate set C, thereby further expanding and enriching the number of concepts. Finally, after the above steps of processing and optimization, the target concept set D is obtained.

[0065] Next, for the four modal attributes of text, image, video and audio of all concepts in the concept set D, a semantic-constructive multimodal alignment method combining concept semantics and data structure is proposed to link concepts with their corresponding four modal information.

[0066] Finally, to optimize the organization and storage of the knowledge graph, knowledge points are first extracted from MOOC resources and curriculum standards, and then processed into fine-grained hierarchical layers. Subsequently, these knowledge points are associated with corresponding concepts, and each concept is further linked to its four modal characteristics. Simultaneously, knowledge points are also associated with their corresponding exercises, collectively constructing a hierarchical and structured knowledge graph. In other words, a large language model is used to automatically extract subject knowledge points from curriculum standards and MOOCs, and semantically associate them with the extracted concepts, thereby forming a hierarchical knowledge graph that supports learners in better mastering and understanding the hierarchical relationships of subject knowledge.

[0067] For example, symbolic data identification methods can be used to uniquely identify and hierarchically manage knowledge points, concepts, and their multimodal attributes, while simultaneously connecting concepts with related knowledge points, thereby solving the challenges of organizing and accessing multimodal educational knowledge. Based on the constructed index table, it can efficiently support the retrieval of multimodal educational knowledge and the generation of cross-modal questions, realizing a hierarchical organization method and mutual indexing storage of educational knowledge graphs. This significantly improves the accessibility of the graphs, supports complex applications such as multimodal retrieval, and provides strong technical support for educational resource integration, intelligent teaching, and personalized learning, facilitating the integration and efficient utilization of intelligent educational resources.

[0068] The method described above integrates multi-source data from open-source education platforms, fully leveraging the powerful semantic parsing, knowledge extraction, and reasoning capabilities of large language models in natural language processing. Combined with multimodal data processing technologies, it achieves the integration of multi-source data, the consolidation of multimodal information, and the processing of multidisciplinary knowledge. Through the complementarity and alignment of multimodal data, a comprehensive, automated, intelligent, and hierarchical multimodal educational knowledge graph is constructed. This not only improves the quality of knowledge extraction and expression and enriches the breadth and depth of educational resources, but also effectively integrates multimodal information such as text, images, videos, and audio, enhancing the relevance of educational resources and the accuracy of knowledge expression, making the management, retrieval, and application of educational knowledge more efficient. This method not only provides new ideas and tools for research and practice in the field of education and offers strong technical support for intelligent applications in multimodal education, but also contributes to personalized learning recommendations, the improvement of educational assessment systems, and the intelligentization of content creation, significantly improving learning efficiency and teaching quality in educational scenarios. In educational applications, it can promote innovation in traditional education models, foster educational equity, and provide technical support for personalized learning, possessing broad social value and market potential.

[0069] The educational knowledge graph construction apparatus provided by this invention is described below. The educational knowledge graph construction apparatus described below can be referred to in correspondence with the educational knowledge graph construction method described above. The educational knowledge graph construction apparatus of this application embodiment is as follows: Figure 3 As shown, it includes: Module 310 is used to acquire multi-source heterogeneous educational data; Extraction module 320 is used to extract subject concepts from educational data based on a large language model; The generation module 330 is used to generate a multimodal educational knowledge graph based on subject concepts and the corresponding multimodal information.

[0070] Figure 4 A schematic diagram of the physical structure of an electronic device is provided. This electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440. The processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can invoke logical instructions stored in the memory 430 to execute an educational knowledge graph construction method. This method includes: acquiring multi-source heterogeneous educational data; extracting subject concepts from the educational data based on a large language model; and generating a multimodal educational knowledge graph based on the subject concepts and their corresponding multimodal information.

[0071] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0072] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the educational knowledge graph construction method provided by the above methods. The method includes: acquiring multi-source heterogeneous educational data; extracting subject concepts from the educational data based on a large language model; and generating a multimodal educational knowledge graph based on the subject concepts and the multimodal information corresponding to the subject concepts.

[0073] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the educational knowledge graph construction method provided by the above methods, the method comprising: acquiring multi-source heterogeneous educational data; extracting subject concepts from the educational data based on a large language model; and generating a multimodal educational knowledge graph based on the subject concepts and the multimodal information corresponding to the subject concepts.

[0074] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0075] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing an educational knowledge graph, characterized in that, include: Acquire multi-source heterogeneous educational data; Based on a large language model, subject concepts are extracted from the educational data. A multimodal educational knowledge graph is generated based on the subject concepts and the corresponding multimodal information.

2. The method for constructing an educational knowledge graph according to claim 1, characterized in that, The extraction of subject concepts from the educational data based on the large language model includes: The educational data is input into a large language model, which outputs candidate subject concepts. Large language models trained with different corpora output the determination results and reasons for whether the candidate subject concepts belong to subject concepts; based on the determination results and reasons, the scoring results corresponding to each candidate subject concept are determined. Based on the scoring results of the candidate subject concepts, a subject concept is determined from the candidate subject concepts using a voting mechanism.

3. The method for constructing an educational knowledge graph according to claim 1, characterized in that, After extracting subject concepts from the educational data based on the large language model, the method further includes: The subject concepts and concept networks are matched to obtain the concepts that are semantically associated with the subject concepts; The subject concept is expanded based on the concept that is semantically associated with the subject concept.

4. The method for constructing an educational knowledge graph according to any one of claims 1-3, characterized in that, The step of generating a multimodal educational knowledge graph based on the subject concept and the corresponding multimodal information includes: Align the subject concepts and their corresponding multimodal information to generate a multimodal educational knowledge graph.

5. The method for constructing an educational knowledge graph according to claim 4, characterized in that, The step of aligning the subject concept and its corresponding multimodal information to generate a multimodal educational knowledge graph includes: Input the subject concepts and Wikipedia content into the large language model to obtain textual explanation information aligned with the subject concepts; The subject concept and the educational data are input into a large language model to obtain image information aligned with the subject concept; Based on the text explanation information aligned with the subject concept, audio information aligned with the subject concept is obtained; Based on the time information of the occurrence of the subject concept in the educational resources, video information aligned with the subject concept is obtained; Based on the text explanation information aligned with the subject concept, the image information aligned with the subject concept, the audio information aligned with the subject concept, and the video information aligned with the subject concept, the multimodal information aligned with the subject concept is obtained; A multimodal educational knowledge graph is generated based on the multimodal information aligned with the subject concepts.

6. The method for constructing an educational knowledge graph according to claim 5, characterized in that, The process of generating a multimodal educational knowledge graph based on the multimodal information aligned with the subject concepts includes: Based on the multi-source heterogeneous educational data, obtain subject information and multiple subject knowledge points corresponding to the subject information; Based on the subject information, the relationships between the various subject knowledge points, the relationships between the subject knowledge points and the subject concepts, and the multimodal information aligned with the subject concepts, a hierarchical multimodal educational knowledge graph is generated.

7. An educational knowledge graph construction device, characterized in that, include: The acquisition module is used to acquire multi-source heterogeneous educational data; An extraction module is used to extract subject concepts from the educational data based on a large language model. The generation module is used to generate a multimodal educational knowledge graph based on the subject concept and the corresponding multimodal information.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the educational knowledge graph construction method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the educational knowledge graph construction method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the educational knowledge graph construction method as described in any one of claims 1 to 6.