Teaching video knowledge graph generation method and device, and related product

By extracting multimodal data from teaching videos and aligning them with curriculum standards, the teaching video knowledge graph is constructed, and the problem of insufficient correlation between multimodal data and curriculum standards in the existing technology is solved, and the comprehensiveness and teaching standardization of knowledge point extraction are achieved.

CN120338069APending Publication Date: 2025-07-18SOUTH CHINA NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510416630.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing teaching video knowledge graph is constructed to ignore the relationship between multimodal data and curriculum standards, resulting in the knowledge graph deviating from the teaching theme, the knowledge system is loose and disorderly, and it is difficult to provide accurate knowledge support for teaching.

Method used

Extract multimodal data from teaching videos, align it with curriculum standard data through the search enhancement generation (RAG) framework, generate video knowledge points entities and relationships, and build teaching video knowledge graphs.

Benefits of technology

It improves the comprehensiveness of knowledge point extraction and teaching normativeness, enhances the breadth of the knowledge graph and the complex relationship modeling ability, and provides more comprehensive knowledge support for teaching and learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338069A_ABST
    Figure CN120338069A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method and a device for generating a teaching video knowledge graph, and a related product. The method comprises the following steps: acquiring teaching video data; searching for curriculum standard data corresponding to the teaching video data, and constructing an internal subject knowledge database according to the curriculum standard data; extracting first multi-modal data in the teaching video data, preprocessing the first multi-modal data, and performing modal semantic alignment enhancement on the preprocessed data through a retrieval enhancement generation (RAG) framework to generate second multi-modal data; wherein a retrieval module of a retrieval enhancement generation (RAG) framework retrieves in an internal subject knowledge database; extracting and generating video knowledge point entity data according to the second multi-modal data; extracting a relationship between knowledge points according to the video knowledge point entity data, and generating video entity relationship set data; and constructing a teaching video knowledge graph according to the video knowledge point entity data and the video entity relationship set data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing technologies, and in particular, to a method and apparatus for generating a knowledge graph of teaching videos, and related products. Background Art

[0002] As a key carrier for knowledge dissemination, the rich knowledge contained in teaching videos urgently needs to be effectively organized and deeply mined. The knowledge graph provides the possibility for the structured presentation and intelligent application of teaching video knowledge.

[0003] Currently, the construction of the knowledge graph usually relies on text single-modal data in data extraction, only extracting knowledge elements from text information such as the subtitles and speech (transcribed text) of teaching videos, ignoring the massive situational information and implicit knowledge clues contained in multi-modal data such as video images, courseware demonstrations, and blackboard writing. In the knowledge graph construction process, existing construction practices seriously lack effective association and deep integration with curriculum standards. As the core guiding principle of subject teaching, curriculum standards clearly define the knowledge framework, ability requirements, and literacy goals, and are the fundamental basis for teaching activity design and evaluation. However, in the process of knowledge graph construction, curriculum standards are often marginalized and not deeply embedded in the construction process. The construction mainly focuses on the extraction of surface knowledge from video content, does not accurately anchor the core points of curriculum standards, does not sort out the knowledge levels and association contexts according to the standard logical framework, resulting in the generated knowledge graph deviating from the teaching main idea, the knowledge system being loose and disordered, key knowledge missing, and the depth and shallowness being unbalanced, and it is difficult to provide a solid, accurate, and effective knowledge support framework for teaching planning, personalized learning path customization, and teaching quality evaluation, seriously hindering the improvement of teaching accuracy and efficiency in the process of educational intelligence, and restricting the in-depth development and optimal utilization of teaching resources. Summary of the Invention

[0004] The purpose of the present invention is to address the defects existing in the prior art, and provide a method and apparatus for generating a knowledge graph of teaching videos, and related products, which extract multi-modal data from teaching video data, and accurately align the video knowledge point data extracted from the multi-modal data based on curriculum standard data, thereby improving the comprehensiveness of knowledge point extraction and teaching standardization.

[0005] To achieve the above object, a first aspect of an embodiment of the present invention provides a method for generating a knowledge graph of teaching videos, the method comprising:

[0006] Obtain teaching video data;

[0007] Search for curriculum standard data corresponding to the teaching video data, and construct an internal subject knowledge database according to the curriculum standard data;

[0008] Extract the first multimodal data from the teaching video data, preprocess the first multimodal data, and enhance the modal semantic alignment of the preprocessed data through a Retrieval-Augmented Generation (RAG) framework to generate second multimodal data; wherein, the retrieval module of the Retrieval-Augmented Generation (RAG) framework retrieves in the internal subject knowledge database;

[0009] Extract video knowledge point entity data according to the second multimodal data;

[0010] Extract the relationships between knowledge points according to the video knowledge point entity data to generate video entity relationship set data;

[0011] Construct a teaching video knowledge graph according to the video knowledge point entity data and the video entity relationship set data.

[0012] Further, finding the curriculum standard data corresponding to the teaching video data and constructing an internal subject knowledge database specifically includes:

[0013] Extract the first standard core knowledge point data and knowledge unit data of the curriculum standard data;

[0014] Optimize the first standard core knowledge point data through a Retrieval-Augmented Generation (RAG) framework to generate second standard core knowledge point data and description data of the second standard core knowledge point;

[0015] Extract the relationships between knowledge points according to the second standard core knowledge point data to generate standard core knowledge point relationship set data;

[0016] Structurally store the second standard core knowledge point data, the description data of the second standard core knowledge point, and the standard core knowledge point relationship set data in a graph database to generate an internal subject knowledge database.

[0017] Further, the first multimodal data includes subtitle data, voice data, image data, and blackboard writing data.

[0018] Further, the preprocessing of the first multimodal data to generate second multimodal data specifically includes:

[0019] Perform data cleaning and time alignment on the first multimodal data in sequence to generate first multimodal preprocessed data;

[0020] Convert the first multimodal preprocessed data into second multimodal preprocessed data in a unified format;

[0021] Enhance the modal semantic alignment of the second multi-modal preprocessed data through the Retrieval-Augmented Generation (RAG) framework to generate the second multi-modal data.

[0022] Furthermore, extracting and generating video knowledge point entity data based on the second multi-modal data specifically includes:

[0023] Extract potential video knowledge point data from the second multi-modal data through a named entity recognition model;

[0024] Enhance the knowledge points of the potential video knowledge point data through the Retrieval-Augmented Generation (RAG) framework to obtain video knowledge point entity data; the retrieval module of the Retrieval-Augmented Generation (RAG) framework retrieves in the subject knowledge database;

[0025] Verify and complete the video knowledge point data in the video knowledge point entity data through the subject knowledge database; the subject knowledge database includes an internal subject knowledge database and an external subject knowledge database.

[0026] Furthermore, extracting the relationships between knowledge points based on the video knowledge point entity data to generate video entity relationship set data specifically includes:

[0027] Extract semantic relationship data, time association relationship data, and implicit relationship data based on the video knowledge point entity data;

[0028] Perform cross-verification on the semantic relationship data, time association relationship data, and implicit relationship data to generate video entity relationship set data.

[0029] Furthermore, constructing a teaching video knowledge graph based on the video knowledge point entity data and the video entity relationship set data specifically includes:

[0030] Construct a preliminary teaching video knowledge graph based on the video knowledge point entity data and the video entity relationship set data;

[0031] Optimize the node representations in the preliminary teaching video knowledge graph through an embedding optimization objective function to generate the final teaching video knowledge graph.

[0032] The second aspect of the embodiments of the present invention provides a device for generating a teaching video knowledge graph, and the device includes:

[0033] A data acquisition module, configured to acquire teaching video data;

[0034] A first data processing module, configured to find the curriculum standard data corresponding to the teaching video data and construct an internal subject knowledge database according to the curriculum standard data;

[0035] A second data processing module extracts first multi-modal data from the teaching video data, preprocesses the first multi-modal data, and enhances the modal semantic alignment of the preprocessed data through a Retrieval-Augmented Generation (RAG) framework to generate second multi-modal data; wherein, the retrieval module of the Retrieval-Augmented Generation (RAG) framework retrieves in the internal subject knowledge database;

[0036] A third data processing module is configured to extract and generate video knowledge point entity data according to the second multi-modal data;

[0037] A fourth data processing module is configured to extract the relationships between knowledge points according to the video knowledge point entity data to generate video entity relationship set data;

[0038] A knowledge graph construction module is configured to construct a teaching video knowledge graph according to the video knowledge point entity data and the video entity relationship set data.

[0039] A third aspect of the embodiments of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;

[0040] The processor is used to be coupled with the memory, read and execute instructions in the memory to implement the method described in the first aspect above;

[0041] The transceiver is coupled with the processor, and the processor controls the transceiver to send and receive messages.

[0042] A fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, which stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the method described in the first aspect above.

[0043] The embodiments of the present invention provide a method and apparatus for generating a teaching video knowledge graph, related products, which extract multi-modal data from teaching video data, and accurately align the video knowledge point data extracted from the multi-modal data based on curriculum standard data, thereby improving the comprehensiveness of knowledge point extraction and teaching standardization. Description of the Drawings

[0044] Figure 1 It is a flowchart of a method for generating a teaching video knowledge graph provided in Embodiment 1 of the present invention;

[0045] Figure 2 It is a schematic structural diagram one of an apparatus for generating a teaching video knowledge graph provided in Embodiment 2 of the present invention;

[0046] Figure 3 It is a schematic structural diagram two of an apparatus for generating a teaching video knowledge graph provided in Embodiment 2 of the present invention;

[0047] Figure 4 This is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. Detailed implementation manners

[0048] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0049] A method and device for generating a knowledge graph of teaching videos, and related products provided in the embodiments of the present invention extract multimodal data from teaching video data, improving the comprehensiveness of knowledge point extraction; accurately aligning the video knowledge point data extracted from the multimodal data based on curriculum standard data, improving teaching standardization; through the Retrieval-Augmented Generation (RAG) framework, retrieving external knowledge bases to expand knowledge point descriptions and supplement implicit knowledge points and complex associations, improving the breadth of the knowledge graph and enhancing its ability to model complex knowledge relationships, providing more comprehensive knowledge support for teaching and learning; through visual display, the application of teaching resources and knowledge systems is more intuitive and convenient, helping teachers, students and education administrators achieve more effective teaching activities and learning support.

[0050] Embodiment 1

[0051] Figure 1 This is a flowchart of a method for generating a knowledge graph of teaching videos provided in Embodiment 1 of the present invention. The technical solution of the present invention will be described below with reference to Figure 1 , and the technical solution of the present invention will be illustrated with specific embodiments.

[0052] Step 110: Obtain teaching video data.

[0053] Specifically, the teaching video data is obtained by the path and name of the teaching video, or by calling from the resource directory of the teaching resource management system. The obtained target teaching video data includes the teaching video content and the corresponding theme name of the teaching video. During the obtaining process, the system performs format verification on the called information to ensure that the video format conforms to the processable range of the system, such as common formats like MP4 and AVI, and the theme name has no special characters or garbled characters, ensuring the availability of the data.

[0054] Step 120: Search for curriculum standard data corresponding to the teaching video data, and construct an internal subject knowledge database according to the curriculum standard data.

[0055] In a possible implementation, step 120 specifically includes steps A1 - A4;

[0056] Step A1: Extract the first standard core knowledge point data and knowledge unit data of the curriculum standard data.

[0057] Specifically, perform text parsing on the curriculum standard data, including unifying the text format, removing redundant symbols, etc.; use natural language processing technology to split the parsed text into sentences, and extract the first standard core knowledge point data, which is a set of multiple first standard core knowledge points. Among them, the knowledge unit data is the knowledge unit corresponding to each extracted first standard core knowledge point. In a specific example, from the curriculum standard data of junior high school mathematics, core knowledge points such as "the solution method of linear equations with one unknown" and "the theorem of the sum of interior angles of a triangle" can be extracted as elements in the set of the first standard core knowledge point data.

[0058] Step A2: Optimize the first standard core knowledge point data through the Retrieval-Augmented Generation (RAG) framework to generate the second standard core knowledge point data and the description data of the second standard core knowledge point.

[0059] Retrieval-Augmented Generation (RAG) is a natural language processing framework that combines retrieval and generation technologies. The process of this framework is as follows: the user asks a question, relevant context is retrieved, and an answer is generated. When retrieving relevant context, the system sends the user's query to the retrieval module, and the retrieval module finds the context information related to the query from the knowledge base; when generating an answer, the user's query is combined with the retrieved context information and input into the generation module. The generation module further processes the data through a generation model based on this information to ensure that the generated content is coherent, accurate, and rich in information. The generation model used in the generation module can be an autoregressive model, a Transform-based generation model, a diffusion model, an energy model, a probabilistic graphical model, etc., or a combination of them. Due to the self-attention mechanism, the Transform-based generation model can capture long-range dependencies well and performs excellently in the field of natural language processing. The RAG framework can update the knowledge reserve of the model in a timely manner by retrieving the external knowledge base, improve the accuracy and relevance of the output answer, and at the same time reduce the possibility of the large language model generating false information.

[0060] Specifically, the first standard core knowledge point data is input into the retrieval module of the RAG framework. The retrieval module uses an embedding model, such as the sentence embedding model Sentence-BERT. The first standard core knowledge point data is converted into computer-processable embedding vectors to represent semantic features. The Sentence-BERT model can map text to a low-dimensional vector space, making texts with similar semantics closer in the vector space. The embedding vectors are used to query in the external subject knowledge database, calculate the similarity between vectors, such as cosine similarity, and retrieve documents related to the first standard core knowledge point data. Among them, the external subject knowledge database can be the subject database on the National Smart Education Platform and other authoritative educational resource databases, etc. The generation module is based on the attention mechanism of the transformer, Monte Carlo tree search, and neuro-symbolic hybrid reasoning method to further deeply mine and refine the retrieved documents related to the first standard core knowledge point data and the first standard core knowledge point data, and generate the second standard core knowledge point data and the description data of the second standard core knowledge point.

[0061] Step A3, extract the relationships between knowledge points according to the second standard core knowledge point data to generate the standard core knowledge point relationship set data.

[0062] Step A4, structurally store the second standard core knowledge point data, the description data of the second standard core knowledge point, and the standard core knowledge point relationship set data using a graph database to generate an internal subject knowledge database.

[0063] Specifically, the second standard core knowledge point data and the description data of the second standard core knowledge point are used as basic storage units and structurally stored using a graph database. The second standard core knowledge point is represented as a node in the graph, and the relationships between knowledge points are represented as edges in the graph. Attributes are added to each node and edge according to the description data of the second standard core knowledge point, such as the type of knowledge point (geometry, algebra, etc.), importance level (core examination point, general knowledge point), strength of the relationship (close relationship, weak association, etc.), etc. The internal subject knowledge database generated by the structured graph storage is convenient for subsequent query and analysis by the retrieval module. The optional graph database is Neo4j.

[0064] Step 130, extract the first multimodal data from the teaching video data, preprocess the first multimodal data, and perform modal semantic alignment enhancement on the preprocessed data through the Retrieval-Augmented Generation (RAG) framework to generate the second multimodal data.

[0065] The retrieval module of the retrieval enhancement generation (RAG) framework searches the internal subject knowledge database. The first multimodal data includes subtitle data, voice data, image data, and blackboard data. The first multimodal data is not limited here, and can also include physical display data, etc.

[0066] In a possible implementation manner, extracting the first multimodal data from the teaching video data in step 130 specifically includes steps B1-B4:

[0067] Step B1, extracting subtitle data. If the teaching video data contains a subtitle track, the text content of the subtitle is directly extracted through a video parsing tool; if there is no subtitle track but there is an external subtitle file, the file is automatically identified and loaded for text extraction; or, optical character recognition (OCR) tools are used to identify and extract frame by frame.

[0068] Step B2, language data extraction: Use video editing software or special tools to directly extract from the audio track of the teaching video data, or use a microphone to directly collect when recording the video.

[0069] Step B3, image data extraction. Among them, it can be images in various forms, such as courseware images produced by software such as PowerPoint and Keynote, and courseware images can contain various elements such as text, pictures, charts and animations. Through video image extraction technology, images in teaching videos are captured frame by frame in chronological order. For courseware images, accurate screening and extraction are performed based on the characteristics of the courseware in the video, such as fixed layout, color, etc. Image classification algorithms are used to distinguish different types of images, such as pure text images, icon images, etc., so as to facilitate subsequent processing. If it is an independent courseware, the original file is directly obtained.

[0070] Step B4, extracting blackboard writing data. Using OCR technology to identify the blackboard writing content appearing in the teaching video data. By performing text recognition on the image of the blackboard or whiteboard, extract all the contents of the blackboard writing, and proofread the extracted blackboard writing content.

[0071] In a possible implementation, the first multimodal data is preprocessed in step 130, and the preprocessed data is subjected to modality semantic alignment enhancement through a retrieval augmentation generation (RAG) framework to generate the second multimodal data, specifically including steps C1-C3:

[0072] Step C1, performing data cleaning and time alignment on the first multimodal data in sequence to generate first multimodal preprocessed data.

[0073] Here, the first multimodal data is first cleaned to remove irrelevant information and noise from each modal data, such as advertising logos in subtitles, environmental noise in voice, watermarks in images, etc. At the same time, errors generated during voice recognition and image OCR recognition are corrected, such as typos, character recognition errors, etc.; then, based on the timestamp information of the teaching activity progress in the video, the subtitles, voice, courseware, blackboard writing and other modal data are synchronously processed, and a time index table is established to correspond different modal data one by one according to the actual situation.

[0074] In a specific example, the teaching video is a physics experiment teaching video. When the teacher explains the experimental steps in the period of 30 minutes to 40 minutes, the corresponding letters, voices, experimental operation images and blackboard contents are associated in the time index table to ensure that the data of each modality are consistent in the time dimension in subsequent analysis.

[0075] For videos with inaccurate or missing timestamps, time series analysis algorithms are used for estimation and calibration. For example, if the timestamp is lost due to editing, the time series analysis algorithm can be used to infer the exact time of each modality data by analyzing the logical sequence and duration of the teaching activities in the video.

[0076] Step C2: converting the first multimodal preprocessed data into second multimodal preprocessed data in a unified format.

[0077] Specifically, for voice data, a speech recognition model, such as Wav2Vec2.0, is applied to transcribe it into text data. During the transcription process, post-processing technologies, such as language model error correction and context matching, are used to improve the accuracy of the transcribed text data. For example, during the transcription of voice, the system can identify and correct accents, homophones, and other problems to ensure that the transcribed text data is consistent with the actual content of the explanation. For image data, OCR recognition is performed to extract text information from the image. For example, in a physics class, the image may contain charts or formulas. These text information are extracted through OCR recognition and converted into structured data. At the same time, target detection algorithms based on deep learning, such as YOLO and Faster R-CNN, perform object detection, identify key visual elements in the image, such as people, objects, icons, etc., and classify them. For another example, in a history class, historical figures and event images appearing in the teaching video are identified and marked.

[0078] Step C3, performing modal semantic alignment enhancement on the second multimodal preprocessed data through a retrieval augmentation generation (RAG) framework to generate second multimodal data.

[0079] Specifically, the second multi-modal preprocessed data is input into the embedding model of the search module, which is converted into an embedding vector. Through this embedding vector, queries are made in the internal disciplinary knowledge database to obtain semantic association data between multiple modalities. The generation model generates the second multi-modal data based on the semantic association data between multiple modalities and the second multi-modal preprocessed data. Among them, the generation model fuses and adjusts different modal data to further ensure the consistency of multiple modules in terms of semantics and time.

[0080] In a specific example, the teaching video includes speech data, caption data, and blackboard writing data. The generation model combines the three modal data to ensure that the three modal data are presented synchronously on the time axis and are highly consistent semantically. For example, when the instructor explains a certain mathematical theorem, the content of the caption and speech will be precisely aligned with the corresponding blackboard writing image to generate high-quality multi-modal data to support the subsequent construction of the knowledge graph.

[0081] Step 140, extract video knowledge point entity data according to the second multi-modal data.

[0082] In a possible implementation manner, step 140 specifically includes steps D1 - D2:

[0083] Step D1, extract potential video knowledge point data from the second multi-modal data through a named entity recognition model.

[0084] Specifically, various texts in the second multi-modal data are analyzed through a named entity recognition (NER) model, and entities related to knowledge points are extracted. In a specific example, the video discusses "Newton's Three Laws of Motion". The NER model will identify "Newton" as a person's name and "laws of motion" as a disciplinary term and mark them as potential video knowledge points.

[0085] Step D2, enhance the knowledge points of the potential video knowledge point data through a retrieval-augmented generation (RAG) framework; the retrieval module of the retrieval-augmented generation (RAG) framework retrieves in the disciplinary knowledge database.

[0086] Specifically, the retrieval module of the RAG framework converts potential video knowledge point data into embedding vectors and retrieves them in the subject knowledge base to query supplementary information related to these potential video knowledge points and generate knowledge point supplementary data. The subject knowledge base includes an internal subject knowledge database and an external subject knowledge database. In a specific example, if the potential video knowledge point is "Newton's three laws", the definition, formula, and application examples of Newton's laws are queried in the subject knowledge database through the embedding vector. Based on the knowledge point supplementary data, the generation model of the generation module completes and verifies the potential video knowledge point data to generate video knowledge point entity data, where the video knowledge point entity data includes video knowledge point data and video knowledge point description data.

[0087] In an alternative solution, after step D2, it further includes:

[0088] Step D3, verifying and completing the video knowledge point data in the video knowledge point entity data through the subject knowledge database; the subject knowledge database includes an internal subject knowledge database and an external subject knowledge database.

[0089] Specifically, the video knowledge point data is matched with the subject knowledge database to verify whether the extracted knowledge points conform to the subject's definition and standards, and modify and supplement the unmatched content.

[0090] Step 150, extracting the relationships between knowledge points based on the video knowledge point entity data to generate video entity relationship set data.

[0091] Specifically, semantic relationship data, time correlation relationship data, and implicit relationship data are extracted from the video knowledge point entity data; the semantic relationship data, time correlation relationship data, and implicit relationship data are cross-verified to generate video entity relationship set data.

[0092] Among them, for the extraction of semantic relationship data, specifically, a relationship classification model is used to extract relationships from the video knowledge point entity data to generate semantic relationship data in text modality, which characterizes the associations between knowledge points at the semantic level, such as causal relationships, inclusion relationships, and parallel relationships. In a specific example, the teaching video explains "Newton's second law", and the relationship classification model can identify the causal relationship between "Newton's second law" and "the relationship between force and acceleration".

[0093] Among them, for the extraction of time correlation relationship data, specifically, a time series model is used to capture the associations of the second multi-modal data in the time dimension. For example, after receiving the second multi-modal data that has been semantically and temporally aligned as input, the long short-term memory network time series model analyzes the changes and associations of the data on the time axis to extract the time correlation relationship.

[0094] Among them, the extraction of implicit relationship data specifically involves inputting the text information, key visual elements, and image data in the image data of the first multimodal data into a model or algorithm for mining implicit relationships in the image modality. Through data analysis and processing by this model or algorithm, the hidden relationships in the image are mined. Among them, the text information is the text information in the image extracted by OCR, such as formulas, titles, annotations, etc. in the image; the key visual elements are the key visual elements after classification in the image recognized by the object detection algorithm based on deep learning, such as people, objects, icons, etc.; the image data is the image data in the first multimodal preprocessed data, such as courseware pictures, blackboard writing pictures, etc. in the teaching video. In a specific example, the teaching video shows a formula and the corresponding physical graph, and the image relationship model can extract the corresponding relationship between the formula and the graph, for example, matching the formula in Newton's law with the demonstrated mechanical graph.

[0095] Among them, cross-validation is performed on the semantic relationship data, time association relationship data, and implicit relationship data to generate video entity relationship set data. Specifically, the semantic relationship data, time association relationship data, and implicit relationship data are input into a cross-validation function or model for cross-validation to check the consistency and rationality among the three relationships, remove possible incorrect or conflicting relationships, and generate video entity relationship set data.

[0096] Step 160, construct a teaching video knowledge graph based on the video knowledge point entity data and the video entity relationship set data.

[0097] In a possible implementation manner, step 160 specifically includes steps E1 - E2:

[0098] Step E1, construct a preliminary teaching video knowledge graph based on the video knowledge point entity data and the video entity relationship set data.

[0099] Among them, the video knowledge point entity data is the nodes in the graph, and the video entity relationship set data is the edges in the graph. For example, in the teaching video, "Newton's second law" and "acceleration" in physics are explained. In the preliminary teaching video knowledge graph, "Newton's second law" and "acceleration" are used as nodes, and the causal relationship between the two is used as an edge to connect the two nodes.

[0100] Step E2, perform optimization processing on the node representations in the preliminary teaching video knowledge graph through an embedding optimization objective function to generate the final teaching video knowledge graph.

[0101] The embedding optimization objective function is shown in formula (1):

[0102]

[0103] Among them, E is the set of edges in the graph, Zi and Z j are the embedding representations of nodes i and j respectively, V is the set of nodes in the graph, and λ is the regularization coefficient used to control the weight of the regularization term. The first part of formula (1) aims to ensure that connected nodes in the graph are also as close as possible in the embedding space, thus preserving the topological structure of the graph. The second part prevents overfitting of the embedding representation through the regularization term, making the embedding representation smoother and more stable.

[0104] In an alternative solution, after step E2, it further includes:

[0105] Step E3, visualizing the teaching video knowledge graph.

[0106] In a possible implementation manner, when representing the nodes of the graph, each node represents a knowledge point, and the size, color, and shape of the node can be used to represent different attributes of the knowledge point. For example, the size can represent the importance degree of the knowledge point, the color can represent the type of the knowledge point, and the shape can represent other attributes of the knowledge point; when representing the edges of the graph, each edge represents the relationship between knowledge points, and the thickness, color, and label of the edge can be used to represent different attributes of the relationship. For example, the thickness represents the strength of the relationship, the color represents the type of the relationship, and a label can be added to the edge to directly display the specific type of the relationship; clicking on a node can display the description data of the knowledge point, and clicking on an edge can display the detailed information of the relationship; a search function is provided, and the user can enter keywords to quickly locate relevant knowledge points or relationships; a filtering function is provided, and the user can filter according to the type of knowledge point, importance degree, type of relationship, etc.; intelligent answers are given according to the user's questions about the knowledge points; according to the user's learning progress and interests, a learning path is automatically recommended.

[0107] Embodiment 2

[0108] Embodiment 2 of the present invention provides a generating device for a teaching video knowledge graph. Figure 2 FIG. 20 is one of the structural schematic diagrams of a generating device for a teaching video knowledge graph provided in Embodiment 2 of the present invention. The generating device 200 includes a data acquisition module 201, a first data processing module 202, a second data processing module 203, a third data processing module 204, a fourth data processing module 205, and a knowledge graph construction module 206.

[0109] The data acquisition module 201 is used to acquire teaching video data.

[0110] The first data processing module 202 is used to find the curriculum standard data corresponding to the teaching video data and construct an internal subject knowledge database according to the curriculum standard data.

[0111] Specifically, extract the first standard core knowledge point data and knowledge unit data of the curriculum standard data; optimize the knowledge points of the first standard core knowledge point data through the Retrieval-Augmented Generation (RAG) framework to generate the second standard core knowledge point data and the description data of the second standard core knowledge point; extract the relationships between the knowledge points according to the second standard core knowledge point data to generate the standard core knowledge point relationship set data; store the second standard core knowledge point data, the description data of the second standard core knowledge point, and the standard core knowledge point relationship set data in a graph database in a structured manner to generate an internal subject knowledge database.

[0112] The second data processing module 203 is used to extract the first multimodal data from the teaching video data, preprocess the first multimodal data, and enhance the modal semantic alignment of the preprocessed data through the Retrieval-Augmented Generation (RAG) framework to generate the second multimodal data; among them, the retrieval module of the Retrieval-Augmented Generation (RAG) framework retrieves in the internal subject knowledge database. Among them, the first multimodal data includes subtitle data, speech data, image data, and blackboard writing data.

[0113] Specifically, perform data cleaning and time alignment on the first multimodal data in sequence to generate the first multimodal preprocessed data; convert the first multimodal preprocessed data into the second multimodal preprocessed data in a unified format; enhance the modal semantic alignment of the second multimodal preprocessed data through the Retrieval-Augmented Generation (RAG) framework to generate the second multimodal data.

[0114] The third data processing module 204 is used to extract and generate video knowledge point entity data according to the second multimodal data.

[0115] Specifically, extract potential video knowledge point data from the second multimodal data through a named entity recognition model; enhance the knowledge points of the potential video knowledge point data through the Retrieval-Augmented Generation (RAG) framework to obtain video knowledge point entity data; the retrieval module of the Retrieval-Augmented Generation (RAG) framework retrieves in the subject knowledge database; verify and complement the video knowledge point data in the video knowledge point entity data through the subject knowledge database; the subject knowledge database includes an internal subject knowledge database and an external subject knowledge database.

[0116] The fourth data processing module 205 is used to extract the relationships between the knowledge points according to the video knowledge point entity data to generate the video entity relationship set data.

[0117] Specifically, extract semantic relationship data, time association relationship data, and implicit relationship data according to the video knowledge point entity data; perform cross-verification on the semantic relationship data, time association relationship data, and implicit relationship data to generate the video entity relationship set data.

[0118] The knowledge graph construction module 206 is used to construct a teaching video knowledge graph based on the video knowledge point entity data and the video entity relationship set data.

[0119] Specifically, a preliminary teaching video knowledge graph is constructed based on the video knowledge point entity data and the video entity relationship set data; the node representations in the preliminary teaching video knowledge graph are optimized through an embedding optimization objective function to generate the final teaching video knowledge graph.

[0120] Optionally, Figure 3 This is the second structural schematic diagram of a generating device for a teaching video knowledge graph provided in the second embodiment of the present invention. As Figure 3 shown, the generating device 200 further includes a visualization processing module 207.

[0121] The visualization processing module 207 is used to perform visualization processing on the teaching video knowledge graph.

[0122] A generating device for a teaching video knowledge graph provided in the second embodiment of the present invention can execute the method steps in the first embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0123] It should be noted that it should be understood that the division of each module of the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; they can also be partially implemented in the form of software called by a processing element and partially implemented in the form of hardware. For example, the first data processing module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the above determined modules. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together or independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the hardware of the processor element or the instruction in the form of software.

[0124] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. For another example, when a certain above module is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. For another example, these modules may be integrated together and implemented in the form of a System-on-a-chip (SOC).

[0125] In the above embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The above computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.). The above computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The above available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0126] Embodiment III

[0127] Figure 4Schematic diagram of a structure of an electronic device provided in Embodiment 3 of the present invention. The electronic device may be a terminal device or a server for implementing the method in Embodiment 1 above, or may be a terminal device or a server for implementing the method in Embodiment 1 above that is connected to the foregoing terminal device or server. As Figure 4 shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver operations of the transceiver 303. Various instructions may be stored in the memory 302 for completing various processing functions and implementing the processing steps described in the method of Embodiment 1 above. Preferably, the electronic device involved in the embodiment of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The foregoing communication port 306 is used for the electronic device to connect and communicate with other peripherals.

[0128] In Figure 4 The system bus 305 mentioned may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.

[0129] The foregoing processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0130] Embodiment 4

[0131] The embodiment of the present invention provides a computer-readable storage medium, in which instructions are stored, and when they run on a computer, the computer is caused to execute the method and processing procedure provided in Embodiment 1 above.

[0132] An embodiment of the present invention further provides a chip for executing instructions, and the chip is used to execute the processing steps described in the first method embodiment above.

[0133] Those skilled in the art should further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0134] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein can be implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0135] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for generating a knowledge graph of teaching videos, characterized in that, The method includes: Obtaining teaching video data; Searching for curriculum standard data corresponding to the teaching video data, and constructing an internal subject knowledge database according to the curriculum standard data; Extracting first multi-modal data from the teaching video data, preprocessing the first multi-modal data, and enhancing the modal semantic alignment of the preprocessed data through a Retrieval-Augmented Generation (RAG) framework to generate second multi-modal data; wherein, the retrieval module of the Retrieval-Augmented Generation (RAG) framework retrieves in the internal subject knowledge database; Extracting and generating video knowledge point entity data according to the second multi-modal data; Extracting the relationships between knowledge points according to the video knowledge point entity data to generate video entity relationship set data; Constructing a teaching video knowledge graph according to the video knowledge point entity data and the video entity relationship set data.

2. The method for generating a knowledge graph of teaching videos according to claim 1, wherein, The searching for curriculum standard data corresponding to the teaching video data and constructing an internal subject knowledge database according to the curriculum standard data specifically includes: Extracting first standard core knowledge point data and knowledge unit data of the curriculum standard data; Optimizing the knowledge points of the first standard core knowledge point data through a Retrieval-Augmented Generation (RAG) framework to generate second standard core knowledge point data and description data of the second standard core knowledge point; Extracting the relationships between knowledge points according to the second standard core knowledge point data to generate standard core knowledge point relationship set data; Structurally storing the second standard core knowledge point data, the description data of the second standard core knowledge point, and the standard core knowledge point relationship set data in a graph database to generate an internal subject knowledge database.

3. The method for generating a knowledge graph of teaching videos according to claim 1, wherein The first multi-modal data includes subtitle data, voice data, image data, and blackboard writing data.

4. The method for generating a knowledge graph of teaching videos according to claim 1, wherein The preprocessing of the first multi-modal data and enhancing the modal semantic alignment of the preprocessed data through a Retrieval-Augmented Generation (RAG) framework to generate second multi-modal data specifically includes: Successively performing data cleaning and time alignment on the first multi-modal data to generate first multi-modal preprocessed data; Converting the first multi-modal preprocessed data into second multi-modal preprocessed data in a unified format; Enhancing the modal semantic alignment of the second multi-modal preprocessed data through a Retrieval-Augmented Generation (RAG) framework to generate second multi-modal data.

5. The method for generating a knowledge graph of teaching videos according to claim 1, wherein The extracting and generating video knowledge point entity data according to the second multi-modal data specifically includes: Extracting potential video knowledge point data from the second multi-modal data through a named entity recognition model; Enhancing the knowledge points of the potential video knowledge point data through a Retrieval-Augmented Generation (RAG) framework to obtain video knowledge point entity data; the retrieval module of the Retrieval-Augmented Generation (RAG) framework retrieves in the subject knowledge database; Verifying and complementing the video knowledge point data in the video knowledge point entity data through the subject knowledge database; the subject knowledge database includes an internal subject knowledge database and an external subject knowledge database.

6. The method for generating a knowledge graph of teaching videos according to claim 1, wherein Extracting the relationships between knowledge points according to the video knowledge point entity data to generate video entity relationship set data specifically includes: Extracting semantic relationship data, time correlation relationship data, and implicit relationship data according to the video knowledge point entity data; Performing cross-validation on the semantic relationship data, time correlation relationship data, and implicit relationship data to generate video entity relationship set data.

7. The method for generating a knowledge graph of teaching videos according to claim 1, wherein Constructing a teaching video knowledge graph according to the video knowledge point entity data and the video entity relationship set data specifically includes: Constructing a preliminary teaching video knowledge graph according to the video knowledge point entity data and the video entity relationship set data; Optimizing the node representations in the preliminary teaching video knowledge graph through an embedding optimization objective function to generate a final teaching video knowledge graph.

8. A generating device for a teaching video knowledge graph, characterized in that, The device includes: A data acquisition module for acquiring teaching video data; A first data processing module for finding the curriculum standard data corresponding to the teaching video data and constructing an internal subject knowledge database according to the curriculum standard data; A second data processing module for extracting the first multimodal data from the teaching video data, preprocessing the first multimodal data, and performing modality semantic alignment enhancement on the preprocessed data through a Retrieval-Augmented Generation (RAG) framework to generate second multimodal data; wherein, the retrieval module of the Retrieval-Augmented Generation (RAG) framework retrieves in the internal subject knowledge database; A third data processing module for extracting and generating video knowledge point entity data according to the second multimodal data; A fourth data processing module for extracting the relationships between knowledge points according to the video knowledge point entity data to generate video entity relationship set data; A knowledge graph construction module for constructing a teaching video knowledge graph according to the video knowledge point entity data and the video entity relationship set data.

9. An electronic device, characterized in that, Including: A memory, a processor, and a transceiver; The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method according to any one of claims 1-7; The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1-7.

Citation Information

Cited By

  • Data processing method, electronic equipment, medium and product

    CN121635766A