Multi-modal knowledge extraction and association method and system based on recorded and broadcast class content

By constructing a multi-layered knowledge structure and multi-dimensional content archive, the problem of insufficient structured processing of recorded course content has been solved, enabling accurate extraction and related recommendations of knowledge, thereby improving learning efficiency and experience.

CN121502292APending Publication Date: 2026-02-10BEIJING JINGYEDA TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511610383.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

The structuring of recorded course content is rather crude, failing to construct a multi-layered knowledge system that combines macro, meso, and micro perspectives. This results in unreasonable knowledge granularity and insufficient utilization of multimodal information fusion. Existing methods neglect the collaborative analysis of speech features, visual events, and text content, leading to low accuracy and timeliness in knowledge extraction. They also fail to effectively identify dependencies and semantic similarities between knowledge points, impacting learning efficiency and experience.

Method used

By acquiring the audio stream of recorded courses and performing speech recognition, a multimodal text stream is generated, and a multi-level knowledge structure is constructed, including macro chapters, meso knowledge modules, and micro knowledge atoms. A multi-dimensional content archive is established, and a global intelligent search engine is built to achieve accurate identification and related recommendations of knowledge points.

Benefits of technology

It enables refined segmentation and structured organization of course content, improves the accuracy and timeliness of knowledge extraction, significantly enhances learning efficiency and personalized learning experience, and constructs a knowledge graph through deep fusion and collaborative analysis of multimodal information, supporting precise push of related content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121502292A_ABST
    Figure CN121502292A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal knowledge extraction and association method and system based on recorded and broadcast class content, and relates to the technical field of artificial intelligence, and the method comprises the following steps: obtaining an audio stream of a recorded and broadcast class, carrying out the voice recognition processing, obtaining a multi-modal text stream, constructing a multi-layer knowledge structure based on the multi-modal text stream, and carrying out the multi-layer knowledge extraction and association. The method comprises the following steps: establishing a multi-dimensional content file of microscopic knowledge atoms, constructing a global intelligent search engine based on a multi-level knowledge structure and the multi-dimensional content file, responding to a query instruction of a learner through the global intelligent search engine, and carrying out associated content recommendation. According to the method, a multi-layer knowledge system combining macroscopic chapters, mesoscopic knowledge modules and microscopic knowledge atoms is constructed, multi-modal information is deeply fused, a multi-dimensional content file is established for the microscopic knowledge atoms, and the knowledge graph is constructed based on the multi-dimensional content file, so that the dependency relationship and semantic association between knowledge points are effectively revealed, and the knowledge graph is established. And finally, the learning efficiency and the personalized learning experience are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for multimodal knowledge extraction and association based on recorded course content. Background Technology

[0002] In recent years, with the booming development of online education, recorded courses have become an important learning resource. However, traditional recorded courses are presented in a linear video format and lack a structured way of organizing knowledge, making it difficult for learners to quickly locate content and effectively understand the internal connections between knowledge points. This unstructured characteristic increases the cognitive load of learners, making them feel like they are searching for information in a book without an index, which seriously affects learning efficiency and experience and limits the full realization of the educational value of recorded courses.

[0003] In existing technologies, the structuring of recorded course content is rather crude, typically only achieving simple chapter divisions and failing to construct a multi-layered knowledge system combining macro, meso, and micro levels. This results in unreasonable knowledge granularity, making it difficult for learners to systematically grasp the course structure. Furthermore, the integration and utilization of multimodal information is insufficient; existing methods often neglect the collaborative analysis of speech features, visual events, and text content, leading to low accuracy and timeliness in knowledge extraction. The lack of multi-dimensional content archives and deep semantic associations for knowledge atoms makes it difficult to effectively identify dependencies and semantic similarities between knowledge points, limiting the recommendation capabilities of search engines and hindering the provision of intelligent, knowledge graph-based content recommendations. Ultimately, this impacts learners' learning efficiency and experience. Summary of the Invention

[0004] The technical problem addressed by this invention is that the structuring of recorded course content is rather crude, typically only achieving simple chapter divisions and failing to construct a multi-layered knowledge system combining macro, meso, and micro levels. This results in unreasonable knowledge granularity, making it difficult for learners to systematically grasp the course structure. Furthermore, the integration and utilization of multimodal information is insufficient; existing methods often neglect the collaborative analysis of speech features, visual events, and text content, leading to low accuracy and timeliness in knowledge extraction. The lack of multi-dimensional content archives and deep semantic associations for knowledge atoms makes it difficult to effectively identify dependencies and semantic similarities between knowledge points, limiting the recommendation capabilities of search engines and hindering the provision of intelligent, knowledge graph-based content recommendations. Ultimately, this negatively impacts learners' learning efficiency and experience.

[0005] To address the aforementioned technical problems, this invention provides the following technical solution: a method for multimodal knowledge extraction and association based on recorded course content, comprising the following steps: Step S1: Obtain the audio stream of the recorded course and perform speech recognition processing to obtain a multimodal text stream. Based on the multimodal text stream, construct a multi-level knowledge structure, which includes macro chapters, meso knowledge modules, and micro knowledge atoms. Step S2: Establish a multi-dimensional content archive for the microscopic knowledge atom; Step S3: Based on the multi-level knowledge structure and multi-dimensional content archives, construct a global intelligent search engine; Step S4: The global intelligent search engine responds to the learner's query and recommends related content.

[0006] As a preferred embodiment of the multimodal knowledge extraction and association method based on recorded course content described in this invention, step S1 specifically includes: The audio stream of the recorded course is acquired, and speech recognition processing is performed on the audio stream to generate preliminary text subtitles with timestamps. The preliminary text subtitles are then aligned with the PPT text accompanying the course using a dynamic time warping algorithm to obtain a multimodal text stream. The multimodal text stream is analyzed and the course content is segmented into layers to construct a multi-layered knowledge structure; The multi-level knowledge structure includes macro chapters, meso knowledge modules, and micro knowledge atoms.

[0007] As a preferred embodiment of the multimodal knowledge extraction and association method based on recorded course content described in this invention, the macro-level section specifically includes: Acquire teacher's audio stream and teacher's video stream, analyze the multimodal text stream, and determine chapter boundary points based on multimodal features, then segment macro chapters based on the chapter boundary points; The multimodal features include speech features and visual events; The method of determining chapter boundary points by combining multimodal features includes: Speech features are extracted from the teacher's speech stream, including prolonged silences, changes in speech rate, and chapter-leading words; Visual events are detected from the teacher's video stream, including slides or whiteboard areas with chapter titles being completely cleared and then rewritten. By integrating the speech features and the visual events, when the speech features and the visual events occur simultaneously within a preset time window, the current time point is determined as the chapter boundary point.

[0008] As a preferred embodiment of the multimodal knowledge extraction and association method based on recorded course content described in this invention, the meso-level knowledge module specifically includes: Based on the multimodal text stream, semantically coherent text paragraphs are identified through a text topic modeling algorithm, while visual cues in the teacher's video stream are detected. When the text paragraphs and visual cues reach a preset degree of consistency on the timeline, they are segmented into meso-level knowledge modules. The visual cues include switching PPT pages, the appearance of new knowledge blocks in the whiteboard area, and the teacher's gestures pointing to new content areas.

[0009] As a preferred embodiment of the multimodal knowledge extraction and association method based on recorded course content described in this invention, the micro-knowledge atoms specifically include: Key entities are extracted from the text of the meso-level knowledge module using named entity recognition, key sentence patterns are identified from the text of the meso-level knowledge module using syntactic analysis, and interactive segments of teacher-student questions and answers are located and extracted from the teacher's audio stream using speech activity detection and speaker separation technology. The independent knowledge points in the key entities, key sentence patterns and interactive segments are used as micro-level knowledge atoms. The key entities include core concepts, formulas, and theorems; The key sentence structures include definition sentences, example sentences, and question sentences.

[0010] As a preferred embodiment of the multimodal knowledge extraction and association method based on recorded course content described in this invention, step S2 specifically includes: The multi-dimensional content archive includes text vectors, semantic tags, and visual indexes; The text content of the microscopic knowledge atoms is semantically encoded to generate text vectors; Based on the text vectors, semantic tags are obtained through clustering analysis and association analysis; Keyframes of the video segments corresponding to the microscopic knowledge atoms are extracted, and image features are extracted from the keyframes to generate visual feature vectors, which serve as the visual index.

[0011] As a preferred embodiment of the multimodal knowledge extraction and association method based on recorded course content described in this invention, the step of obtaining semantic tags through cluster analysis and association analysis specifically includes: The semantic tags include topic tags, difficulty tags, and knowledge dependency tags; The knowledge dependency tags include pre-knowledge dependency tags and post-knowledge dependency tags; Based on the text vectors, a clustering algorithm is used to aggregate semantically similar micro-knowledge atoms into different clusters, and each cluster is assigned a topic tag that includes the core content. By analyzing the lexical complexity, sentence length, and terminology density of the text content of knowledge atoms, a comprehensive difficulty score is obtained. Based on the numerical range of the comprehensive difficulty score, the comprehensive difficulty score is mapped to a preset difficulty level to generate a difficulty label. By calculating the semantic similarity between the text vectors, when the semantic similarity is greater than a preset similarity threshold, knowledge atoms that appear earlier in the course are labeled as subsequent knowledge dependency tags, and knowledge atoms that appear later in the course are labeled as preceding knowledge dependency tags, thus obtaining knowledge dependency tags. The process of obtaining the overall difficulty score includes: The ratio of uncommon words in the text to the total number of words in the text is used as an indicator of word complexity. The ratio of the number of long and complex sentences in the text to the total number of sentences in the text is used as an indicator of sentence complexity. The ratio of the frequency of occurrence of specialized terms in a specific field to the total number of words in the text is used as an indicator of terminology density. The comprehensive difficulty score is obtained by weighting the vocabulary complexity index, sentence complexity index, and terminology density index.

[0012] As a preferred embodiment of the multimodal knowledge extraction and association method based on recorded course content described in this invention, step S3 specifically includes: Each micro-knowledge atom is treated as an independent node, and a multi-dimensional content file of the corresponding micro-knowledge atom is stored in each node. Connections are established between nodes to obtain different types of edges. The nodes and edges are integrated to form a knowledge graph, and the knowledge graph is used as the core index structure of the global intelligent search engine. The acquisition of different types of edges includes temporal edges, dependency edges, and semantic edges: Obtaining temporal edges involves: establishing directed edges between nodes with earlier and later times based on the appearance time of microscopic knowledge atoms in the course, representing the teaching order of the course; Obtaining dependency edges involves: establishing directed edges between nodes with preceding knowledge dependency labels and nodes with subsequent knowledge dependency labels based on the knowledge dependency labels between micro-knowledge atoms, representing the order of learning requirements; Obtaining semantic edges includes: establishing connections between node pairs whose semantic similarity is greater than a preset similarity threshold based on the semantic similarity between text vectors of micro-knowledge atoms, thereby representing the relevance of text content.

[0013] As a preferred embodiment of the multimodal knowledge extraction and association method based on recorded course content described in this invention, step S4 specifically includes: The system receives a query instruction from a learner, converts the query instruction into a query vector through semantic encoding, calculates the similarity between the query vector and the text vectors of each node in the knowledge graph, and takes the nodes with similarity greater than a preset similarity value as the initial search results. Based on the nodes corresponding to the initial search results, the system traverses the edge relationships of the nodes corresponding to the initial search results in the knowledge graph through the global intelligent search engine and recommends related content.

[0014] A multimodal knowledge extraction and association system based on recorded course content is applied to a multimodal knowledge extraction and association method based on recorded course content, including a processing module, an establishment module, a construction module, and a response module; The processing module is used to acquire the audio stream of the recorded course and perform speech recognition processing to obtain a multimodal text stream. Based on the multimodal text stream, a multi-level knowledge structure is constructed, which includes macro chapters, meso knowledge modules and micro knowledge atoms. The establishment module is used to establish a multi-dimensional content archive of the microscopic knowledge atom; The construction module is used to build a global intelligent search engine based on the multi-level knowledge structure and multi-dimensional content archive; The response module is used to respond to learners' query commands through the global intelligent search engine and recommend related content.

[0015] The beneficial effects of this invention are as follows: By constructing a multi-layered knowledge system that combines macro-level chapters, meso-level knowledge modules, and micro-level knowledge atoms, this invention achieves refined segmentation and structured organization of course content, making the knowledge granularity division more reasonable. This helps learners clearly grasp the course structure, deeply integrates multi-modal information such as voice, vision, and text, and accurately identifies knowledge boundaries through collaborative analysis, significantly improving the accuracy and timeliness of knowledge extraction. It establishes multi-dimensional content archives for micro-level knowledge atoms, including text vectors, semantic tags, and visual indexes, and constructs a knowledge graph based on this, effectively revealing the dependencies and semantic associations between knowledge points. This empowers a global intelligent search engine, enabling precise recommendations of related content based on learner queries, ultimately significantly improving learning efficiency and personalized learning experience. Attached Figure Description

[0016] Figure 1 This is a basic flowchart illustrating a method for multimodal knowledge extraction and association based on recorded course content, provided in one embodiment of the present invention. Figure 2 This is a schematic diagram of the basic process of a multimodal knowledge extraction and association system based on recorded course content, provided as an embodiment of the present invention. Detailed Implementation

[0017] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0018] Example 1, referring to Figure 1 As an embodiment of the present invention, a method for multimodal knowledge extraction and association based on recorded course content is provided, comprising the following steps: Step S1: Obtain the audio stream of the recorded course and perform speech recognition processing to obtain a multimodal text stream. Based on the multimodal text stream, construct a multi-level knowledge structure, which includes macro chapters, meso knowledge modules, and micro knowledge atoms.

[0019] Step S2: Establish a multi-dimensional content archive of microscopic knowledge atoms.

[0020] Step S3: Based on a multi-layered knowledge structure and multi-dimensional content archives, construct a global intelligent search engine.

[0021] Step S4: The global intelligent search engine responds to the learner's query and recommends related content.

[0022] Unstructured recorded course content is transformed into a structured knowledge system encompassing macro, meso, and micro levels through automated processing. This makes the originally linear and continuous video content indexable, comprehensible, and interconnected. By building a global intelligent search engine, learners no longer blindly drag and drop through long videos; they can accurately locate specific knowledge points and receive systematic recommendations of related content. This significantly shortens information retrieval time, deepens knowledge understanding, clarifies the technical path from raw audio and video streams to final intelligent applications, and provides a solid data and model foundation for more complex educational applications in the future.

[0023] Step S1 specifically includes: The audio stream of the recorded course is acquired, and speech recognition processing is performed on the audio stream to generate preliminary text subtitles with timestamps. The preliminary text subtitles are then aligned with the PPT text accompanying the course using a dynamic time warping algorithm to obtain a multimodal text stream.

[0024] Analyze multimodal text streams and segment course content into layers to construct a multi-layered knowledge structure.

[0025] The multi-layered knowledge structure includes macro chapters, meso knowledge modules, and micro knowledge atoms.

[0026] The initial text captions are timestamped. The timestamps convert continuous audio into a locationable text sequence, which serves as the time reference for all subsequent alignment, segmentation, and retrieval operations.

[0027] By generating timestamped subtitles through speech recognition and aligning them with PPT text using a dynamic time warping algorithm, the system effectively overcomes the errors that may occur with pure speech recognition and ensures precise temporal synchronization between text content and visual materials. This provides a high-quality data foundation for subsequent multimodal analysis, clarifies the specific operations for constructing multi-level knowledge structures, and can automatically identify the logical levels of courses. It breaks down lengthy recorded courses into easily manageable and learnable knowledge units, replacing traditional manual labeling and segmentation work.

[0028] The macro section specifically includes: Acquire teacher audio and video streams, analyze multimodal text streams, and determine chapter boundary points based on multimodal features. Then, segment macro chapters based on these chapter boundary points.

[0029] Multimodal features include speech features and visual events.

[0030] Determining chapter boundary points by combining multimodal features includes: Speech features were extracted from the teacher's speech stream, including prolonged silences, changes in speech rate, and chapter-leading words.

[0031] Visual events are detected in the teacher's video stream, including slides or whiteboard areas with chapter titles that are completely cleared and then rewritten.

[0032] By integrating voice features and visual events, when voice features and visual events occur simultaneously within a preset time window, the current time point is determined as the chapter boundary point.

[0033] Setting a preset time window is a crucial step in balancing the sensitivity and accuracy of the detection. Based on statistical analysis, by collecting a large amount of real recorded course data, the average time delay and standard deviation between when the teacher issues the chapter introduction and when switching to the title slide are statistically analyzed. The window is set to the average plus one or two standard deviations to cover the vast majority of normal situations.

[0034] By fusing speech features and visual events, a preset time window is initiated for verification when either a speech feature or a visual event is detected. Within this window, if a corresponding feature of another modality can be captured, a chapter boundary point is confirmed. If no match is found by the end of the window, the event is ignored as noise. This cross-validation logic ensures that the final segmentation decision is made only when auditory and visual cues occur concurrently in time.

[0035] This solution improves the robustness and accuracy of chapter division. Single-modal recognition may lead to misjudgments. By integrating speech features (silence, speech rate, and guiding words) and visual events (chapter title slides, clearing of whiteboard writing), the boundary is determined only when multiple modal features are triggered together. This greatly improves the accuracy and reliability of chapter division and simulates the teaching logic of human teachers. When teachers switch chapters, they usually provide prompts through both language and vision. This solution intelligently simulates this natural teaching behavior, making the division results more consistent with human cognitive habits and the real logical structure of the course.

[0036] The meso-level knowledge module specifically includes: Based on multimodal text streams, semantically coherent text paragraphs are identified through text topic modeling algorithms, while visual cues in the teacher's video stream are detected. When the text paragraphs and visual cues reach a preset degree of consistency on the timeline, they are segmented into meso-level knowledge modules.

[0037] Visual cues include PPT slide transitions, the appearance of new knowledge blocks in the whiteboard area, and teacher gestures pointing to new content areas.

[0038] The core function of text topic modeling algorithms is to automatically discover and aggregate semantically highly related sentences from a continuous text stream to form a meaningful semantic block. It does not rely on preset keywords, but identifies potential discussion topics by analyzing the co-occurrence patterns of words. This unsupervised learning method can objectively divide according to the inherent logic of the content itself, which is a key technology to ensure that each knowledge module has high cohesion and integrity in terms of topic.

[0039] The preset matching degree setting quantifies the time overlap or boundary proximity between text paragraphs and visual events, and then determines a judgment threshold through empirical or data-driven methods. The selection of this threshold finds a balance between missed cutting and incorrect cutting.

[0040] The system achieves semantic coherence recognition of knowledge modules. Through text topic modeling algorithms, it can identify paragraphs that are semantically closely related, ensuring that the segmented knowledge modules are complete and self-consistent in content. This avoids artificially splitting a complete concept and enhances the multimodal verification of the segmentation results. The system verifies the consistency between the text semantic analysis results and visual cues (PPT switching, new blackboard writing, teacher gestures) on the timeline, ensuring that the identified semantic modules match the teacher's actual teaching behavior and pace in the classroom. This makes the division of knowledge modules both consistent with text logic and teaching context.

[0041] Microscopic knowledge atoms specifically include: Key entities are extracted from the text of the meso-level knowledge module using named entity recognition, while key sentence patterns are identified from the text of the meso-level knowledge module using syntactic analysis. Interactive segments of teacher-student questions and answers are located and extracted from the teacher's audio stream using speech activity detection and speaker separation technology. The key entities, key sentence patterns, and independent knowledge points in the interactive segments are used as micro-level knowledge atoms.

[0042] Key entities include core concepts, formulas, and theorems.

[0043] Key sentence structures include definition sentences, example sentences, and question sentences.

[0044] Named entity recognition can accurately extract structured knowledge units such as core concepts, formulas, and theorems from text, transforming abstract knowledge points into atomic data that can be independently retrieved and understood, which is the foundation for building a knowledge system.

[0045] By analyzing sentence structure, key sentence patterns such as definitions, examples, and questions that carry specific teaching intentions are identified, thereby assigning semantic labels such as definitions, applications, or questions to knowledge atoms and understanding their logical functions.

[0046] The combination of speech activity detection and speaker separation technology can automatically locate and separate complete interactive segments of teacher-student questions and answers. These dynamic and contextualized questions and answers are themselves highly valuable micro-knowledge atoms, directly reflecting the doubts and difficulties in teaching.

[0047] It achieves refined and atomic extraction of knowledge points. Through named entity recognition, syntactic analysis, and interaction fragment detection, it can further extract the most core and independent knowledge units from knowledge modules. This atomic processing is a prerequisite for achieving accurate retrieval and related recommendations. It captures both the static and dynamic dimensions of knowledge, not only extracting static knowledge content but also paying special attention to dynamic teacher-student interaction fragments. These interaction fragments often contain doubts, difficulties, and error-prone points, which are extremely valuable learning materials. Using them as knowledge atoms greatly enriches the connotation of the knowledge base.

[0048] Step S2 specifically includes: The multi-dimensional content archive includes text vectors, semantic tags, and visual indexes.

[0049] Semantic encoding is performed on the text content of microscopic knowledge atoms to generate text vectors.

[0050] Semantic labels are obtained based on text vectors through clustering and association analysis.

[0051] Keyframes of video segments corresponding to microscopic knowledge atoms are extracted, and image features are extracted from the keyframes to generate visual feature vectors, which serve as visual indexes.

[0052] Text vectors transform unstructured text content into a mathematical form that computers can understand. Through semantic encoding, the generated vectors can capture the deep semantic information of knowledge atoms, rather than simply matching keywords. This makes it possible to calculate the similarity between knowledge atoms, which is the foundation for intelligent retrieval and association analysis.

[0053] Semantic tags are key to classifying and defining the functions of knowledge atoms. Clustering based on text vectors can automatically group atoms with similar content into one category. Through association analysis, the logical relationships between atoms can be discovered. These tags construct a structured network of knowledge, which goes beyond the atoms themselves and reveals the intrinsic connections between knowledge points.

[0054] Visual indexing establishes a direct bridge between each textualized knowledge atom and its original video footage. By extracting image features from keyframes, the generated visual vectors enable users not only to search for the textual definition of photosynthesis, but also to directly locate the specific illustrations drawn on the blackboard by the teacher when explaining the concept, achieving precise cross-modal positioning.

[0055] A comprehensive digital profile has been created for each knowledge point. Through text vectors, semantic tags, and visual indexes, an archive has been built for each micro-knowledge atom from three dimensions: semantic, attribute, and visual. This enables computers not only to read knowledge, but also to understand its meaning, classification, and appearance, supporting cross-modal intelligent retrieval. Text vectors support semantic search, and visual indexes support image-based video search. This multi-dimensional content archive is the key technical support for realizing the WYSIWYG intelligent search experience.

[0056] Obtaining semantic tags through cluster analysis and association analysis specifically includes: Semantic tags include topic tags, difficulty tags, and knowledge dependency tags.

[0057] Knowledge dependency tags include pre-knowledge dependency tags and post-knowledge dependency tags.

[0058] Based on text vectors, a clustering algorithm is used to aggregate semantically similar micro-knowledge atoms into different clusters, and each cluster is assigned a topic tag that includes the core content.

[0059] By analyzing the lexical complexity, sentence length, and terminology density of the text content of knowledge atoms, a comprehensive difficulty score is obtained. Based on the numerical range of the comprehensive difficulty score, the comprehensive difficulty score is mapped to a preset difficulty level to generate a difficulty label.

[0060] By calculating the semantic similarity between text vectors, when the semantic similarity is greater than a preset similarity threshold, knowledge atoms that appear earlier in the course are labeled as subsequent knowledge dependency tags, and knowledge atoms that appear later in the course are labeled as preceding knowledge dependency tags, thus obtaining knowledge dependency tags.

[0061] The overall difficulty score includes: The ratio of uncommon words in a text to the total number of words in the text is used as an indicator of lexical complexity.

[0062] The ratio of the number of long and complex sentences in a text to the total number of sentences in the text is used as an indicator of sentence complexity.

[0063] The ratio of the frequency of occurrence of specialized terms in a specific field to the total number of words in the text is used as an indicator of terminology density.

[0064] The overall difficulty score is obtained by weighting the vocabulary complexity index, sentence complexity index, and professional terminology density index.

[0065] The identification of long and complex sentences combines length and syntactic complexity. Long sentences are filtered out by setting a threshold for the number of words, and their structural complexity is detected by syntactic analysis techniques. Core indicators include whether there are multiple clauses, whether there are nested clauses, whether complex non-finite verb phrases, parenthetical phrases, or inverted structures are used. These structural features are the main reasons for the difficulty in understanding sentences.

[0066] The determination of uncommon words relies on comparison with a general high-frequency word list. A standard high-frequency word database is built in. When analyzing text, any word that does not appear in this database will be initially marked as an uncommon word.

[0067] The identification of specialized terminology in a particular field relies on a pre-built terminology dictionary or knowledge base specific to that discipline.

[0068] The core of the comprehensive difficulty score design lies in making the abstract concept of difficulty concrete and operational. It transforms the three core factors affecting comprehension—vocabulary complexity, sentence complexity, and density of professional terminology—into quantifiable indicators, and through weighted fusion, it ultimately yields a comprehensive value that can fully and objectively reflect the cognitive difficulty of knowledge atoms.

[0069] The calculation logic of the comprehensive difficulty score lies in the fact that it is not a simple addition of the three indicators, but a weighted fusion to simulate the differences in human perception of different difficulty factors in the cognitive process. Each of the three indicators, vocabulary complexity, sentence complexity, and terminology density, is assigned a weight coefficient, which reflects the contribution of this dimension to the overall comprehension difficulty. The actual calculated value of each indicator is multiplied by its corresponding weight to obtain the weighted score. The sum of these three weighted scores yields a quantitative comprehensive difficulty score. This score can more scientifically and comprehensively reflect the comprehensive cognitive load that a micro-level knowledge atom brings to the learner.

[0070] It achieves deep semantic annotation of knowledge, using tags across three dimensions: topic, difficulty, and dependency relationships. This transforms knowledge from isolated entities into entities imbued with context, hierarchy, and logical relationships, providing core data for constructing intelligent learning paths. It also establishes the foundation for quantifying learning difficulty and adaptive learning by using a set of calculable metrics (vocabulary, sentence structure, terminology density) to quantify the difficulty of knowledge points. This allows for an objective assessment of the challenge of different content, laying the groundwork for recommending appropriately challenging content based on learner levels and achieving personalized adaptive learning. Furthermore, it constructs a knowledge dependency network for the course, automatically inferring pre- and post-knowledge dependencies through temporal and semantic similarity. This connects scattered knowledge points into a directed acyclic graph, clearly revealing the logical order to follow when learning a course and effectively avoiding knowledge gaps caused by skipping steps in learning.

[0071] Step S3 specifically includes: Each micro-knowledge atom is treated as an independent node, and a multi-dimensional content file of the corresponding micro-knowledge atom is stored in each node. Connections are established between nodes to obtain different types of edges. The nodes and edges are integrated to form a knowledge graph, which is then used as the core index structure of the global intelligent search engine.

[0072] Retrieve different types of edges, including temporal edges, dependency edges, and semantic edges: Obtaining temporal edges involves establishing directed edges between nodes that appear earlier and nodes that appear later in the course, based on the appearance time of microscopic knowledge atoms in the course, to represent the teaching order of the course.

[0073] Obtaining dependency edges involves: establishing directed edges between nodes with preceding knowledge dependency labels and nodes with subsequent knowledge dependency labels based on the knowledge dependency labels between micro-knowledge atoms, representing the order of learning requirements.

[0074] Obtaining semantic edges includes: establishing connections between node pairs with semantic similarity greater than a preset similarity threshold based on the semantic similarity between text vectors of micro-knowledge atoms, representing the relevance of text content.

[0075] Knowledge graphs, as the core index, go beyond keyword matching, connecting isolated knowledge points into a semantic network through various relationships. This enables search engines to understand the logic and connections between knowledge points, supporting intelligent queries on what knowledge is needed before learning a particular concept.

[0076] Temporal edges are the temporal skeleton of a knowledge graph, connecting knowledge points in the order of instruction, providing learners with the most intuitive learning path, and are the foundation for ensuring the continuity of knowledge.

[0077] Dependency edges form the logical framework of a knowledge graph, revealing the order in which knowledge is learned. They clarify the necessary conditions for learning and are the core of achieving intelligent path planning and avoiding knowledge gaps.

[0078] Semantic edges are associative networks in knowledge graphs that break the limitations of chapters and time sequences, connecting related knowledge points. They promote the integration of knowledge and are key to achieving intelligent recommendation that allows for generalization.

[0079] A highly interconnected knowledge network was constructed, with a knowledge graph as the core index structure. Knowledge points were used as nodes, and temporal, dependency, and semantic relationships were used as edges, forming a network-like knowledge system. Compared with traditional linear lists or tree structures, this more accurately reflects the complex relationships between knowledge, endowing the search engine with reasoning and associative capabilities. The knowledge graph-based search engine can not only perform keyword matching, but also traverse and reason along the edges. In this embodiment, when searching for a concept, not only can the concept itself be found, but its preceding knowledge can also be found through dependency edges, and related concepts can be found through semantic edges, realizing a truly meaningful recommendation of related content.

[0080] Step S4 specifically includes: The system receives the learner's query command, converts it into a query vector through semantic encoding, calculates the similarity between the query vector and the text vector of each node in the knowledge graph, and takes the nodes with a similarity greater than a preset similarity value as the initial search results. Based on the nodes corresponding to the initial search results, the system traverses the edge relationships of the nodes corresponding to the initial search results in the knowledge graph through a global intelligent search engine and recommends related content.

[0081] The preset similarity value is set by first analyzing the data distribution of the similarity of all nodes to find a theoretical inflection point as a reference. Small-scale iterative tests are then conducted around this value, and the results are evaluated using manually labeled query sets. Finally, the value that performs best on the F1-Score metric is selected as the final threshold to achieve the best balance between precision and recall.

[0082] This approach achieves a leap from keyword search to semantic search, converting user queries into query vectors and calculating their similarity with node text vectors. It understands the true intent of the query, rather than simply matching literal words. In this example, searching for "how to calculate the area of ​​a circle" understands the query's intent and returns knowledge points containing the formula. It provides a systematic and in-depth learning path recommendation. After finding initial results, it further traverses the edges in the knowledge graph to recommend relevant content to learners. This recommendation is not random but based on the inherent logic of knowledge (temporal sequence, dependency, semantics), helping learners build a complete knowledge system and achieve an improvement from learning a single point to mastering a broader understanding.

[0083] Example 2, refer to Figure 2 This invention provides a multimodal knowledge extraction and association system based on recorded course content, including a processing module, an establishment module, a construction module, and a response module; The processing module is used to acquire the audio stream of the recorded course and perform speech recognition processing to obtain a multimodal text stream. Based on the multimodal text stream, a multi-level knowledge structure is constructed, which includes macro chapters, meso knowledge modules, and micro knowledge atoms.

[0084] The module is used to create multi-dimensional content archives of microscopic knowledge atoms.

[0085] The building module is used to construct a global intelligent search engine based on a multi-level knowledge structure and multi-dimensional content archive.

[0086] The response module is used to respond to learners' query commands through a global intelligent search engine and recommend related content.

[0087] This invention constructs a multi-layered knowledge system combining macro-level chapters, meso-level knowledge modules, and micro-level knowledge atoms. This system enables refined segmentation and structured organization of course content, resulting in a more rational division of knowledge granularity. It helps learners clearly grasp the course structure, deeply integrates multi-modal information such as speech, vision, and text, and accurately identifies knowledge boundaries through collaborative analysis. This significantly improves the accuracy and timeliness of knowledge extraction. It establishes multi-dimensional content archives for micro-level knowledge atoms, including text vectors, semantic tags, and visual indexes. Based on this, a knowledge graph is constructed, effectively revealing the dependencies and semantic connections between knowledge points. This empowers a global intelligent search engine, enabling precise recommendations of related content based on learner queries, ultimately significantly improving learning efficiency and personalized learning experience.

[0088] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0089] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for multimodal knowledge extraction and association based on recorded course content, characterized in that, Includes the following steps: Step S1: Obtain the audio stream of the recorded course and perform speech recognition processing to obtain a multimodal text stream. Based on the multimodal text stream, construct a multi-level knowledge structure, which includes macro chapters, meso knowledge modules, and micro knowledge atoms. Step S2: Establish a multi-dimensional content archive for the microscopic knowledge atom; Step S3: Based on the multi-level knowledge structure and multi-dimensional content archives, construct a global intelligent search engine; Step S4: The global intelligent search engine responds to the learner's query and recommends related content.

2. The multimodal knowledge extraction and association method based on recorded course content as described in claim 1, characterized in that, Step S1 specifically includes: The audio stream of the recorded course is acquired, and speech recognition processing is performed on the audio stream to generate preliminary text subtitles with timestamps. The preliminary text subtitles are then aligned with the PPT text accompanying the course using a dynamic time warping algorithm to obtain a multimodal text stream. The multimodal text stream is analyzed and the course content is segmented into layers to construct a multi-layered knowledge structure; The multi-level knowledge structure includes macro chapters, meso knowledge modules, and micro knowledge atoms.

3. The multimodal knowledge extraction and association method based on recorded course content as described in claim 2, characterized in that, The macro-level section specifically includes: Acquire teacher's audio stream and teacher's video stream, analyze the multimodal text stream, and determine chapter boundary points based on multimodal features, then segment macro chapters based on the chapter boundary points; The multimodal features include speech features and visual events; The method of determining chapter boundary points by combining multimodal features includes: Speech features are extracted from the teacher's speech stream, including prolonged silences, changes in speech rate, and chapter-leading words; Visual events are detected from the teacher's video stream, including slides or whiteboard areas with chapter titles being completely cleared and then rewritten. By integrating the speech features and the visual events, when the speech features and the visual events occur simultaneously within a preset time window, the current time point is determined as the chapter boundary point.

4. The multimodal knowledge extraction and association method based on recorded course content as described in claim 3, characterized in that, The meso-level knowledge module specifically includes: Based on the multimodal text stream, semantically coherent text paragraphs are identified through a text topic modeling algorithm, while visual cues in the teacher's video stream are detected. When the text paragraphs and visual cues reach a preset degree of consistency on the timeline, they are segmented into meso-level knowledge modules. The visual cues include switching PPT pages, the appearance of new knowledge blocks in the whiteboard area, and the teacher's gestures pointing to new content areas.

5. The multimodal knowledge extraction and association method based on recorded course content as described in claim 4, characterized in that, The microscopic knowledge atoms specifically include: Key entities are extracted from the text of the meso-level knowledge module using named entity recognition, key sentence patterns are identified from the text of the meso-level knowledge module using syntactic analysis, and interactive segments of teacher-student questions and answers are located and extracted from the teacher's audio stream using speech activity detection and speaker separation technology. The independent knowledge points in the key entities, key sentence patterns and interactive segments are used as micro-level knowledge atoms. The key entities include core concepts, formulas, and theorems; The key sentence structures include definition sentences, example sentences, and question sentences.

6. The multimodal knowledge extraction and association method based on recorded course content as described in claim 5, characterized in that, Step S2 specifically includes: The multi-dimensional content archive includes text vectors, semantic tags, and visual indexes; The text content of the microscopic knowledge atoms is semantically encoded to generate text vectors; Based on the text vectors, semantic tags are obtained through clustering analysis and association analysis; Keyframes of the video segments corresponding to the microscopic knowledge atoms are extracted, and image features are extracted from the keyframes to generate visual feature vectors, which serve as the visual index.

7. The multimodal knowledge extraction and association method based on recorded course content as described in claim 6, characterized in that, The acquisition of semantic tags through cluster analysis and association analysis specifically includes: The semantic tags include topic tags, difficulty tags, and knowledge dependency tags; The knowledge dependency tags include pre-knowledge dependency tags and post-knowledge dependency tags; Based on the text vectors, a clustering algorithm is used to aggregate semantically similar micro-knowledge atoms into different clusters, and each cluster is assigned a topic tag that includes the core content. By analyzing the lexical complexity, sentence length, and terminology density of the text content of knowledge atoms, a comprehensive difficulty score is obtained. Based on the numerical range of the comprehensive difficulty score, the comprehensive difficulty score is mapped to a preset difficulty level to generate a difficulty label. By calculating the semantic similarity between the text vectors, when the semantic similarity is greater than a preset similarity threshold, knowledge atoms that appear earlier in the course are labeled as subsequent knowledge dependency tags, and knowledge atoms that appear later in the course are labeled as preceding knowledge dependency tags, thus obtaining knowledge dependency tags. The process of obtaining the overall difficulty score includes: The ratio of uncommon words in the text to the total number of words in the text is used as an indicator of word complexity. The ratio of the number of long and complex sentences in the text to the total number of sentences in the text is used as an indicator of sentence complexity. The ratio of the frequency of occurrence of specialized terms in a specific field to the total number of words in the text is used as an indicator of terminology density. The comprehensive difficulty score is obtained by weighting the vocabulary complexity index, sentence complexity index, and terminology density index.

8. The multimodal knowledge extraction and association method based on recorded course content as described in claim 7, characterized in that, Step S3 specifically includes: Each micro-knowledge atom is treated as an independent node, and a multi-dimensional content file of the corresponding micro-knowledge atom is stored in each node. Connections are established between nodes to obtain different types of edges. The nodes and edges are integrated to form a knowledge graph, and the knowledge graph is used as the core index structure of the global intelligent search engine. The acquisition of different types of edges includes temporal edges, dependency edges, and semantic edges: Obtaining temporal edges involves: establishing directed edges between nodes with earlier and later times based on the appearance time of microscopic knowledge atoms in the course, representing the teaching order of the course; Obtaining dependency edges involves: establishing directed edges between nodes with preceding knowledge dependency labels and nodes with subsequent knowledge dependency labels based on the knowledge dependency labels between micro-knowledge atoms, representing the order of learning requirements; Obtaining semantic edges includes: establishing connections between node pairs whose semantic similarity is greater than a preset similarity threshold based on the semantic similarity between text vectors of micro-knowledge atoms, thereby representing the relevance of text content.

9. The multimodal knowledge extraction and association method based on recorded course content as described in claim 8, characterized in that, Step S4 specifically includes: The system receives a query instruction from a learner, converts the query instruction into a query vector through semantic encoding, calculates the similarity between the query vector and the text vectors of each node in the knowledge graph, and takes the nodes with similarity greater than a preset similarity value as the initial search results. Based on the nodes corresponding to the initial search results, the system traverses the edge relationships of the nodes corresponding to the initial search results in the knowledge graph through the global intelligent search engine and recommends related content.

10. A multimodal knowledge extraction and association system based on recorded course content, which is applied in the multimodal knowledge extraction and association method based on recorded course content as described in any one of claims 1-9, characterized in that, It includes a processing module, a creation module, a construction module, and a response module; The processing module is used to acquire the audio stream of the recorded course and perform speech recognition processing to obtain a multimodal text stream. Based on the multimodal text stream, a multi-level knowledge structure is constructed, which includes macro chapters, meso knowledge modules and micro knowledge atoms. The establishment module is used to establish a multi-dimensional content archive of the microscopic knowledge atom; The construction module is used to build a global intelligent search engine based on the multi-level knowledge structure and multi-dimensional content archive; The response module is used to respond to learners' query commands through the global intelligent search engine and recommend related content.

Citation Information

Patent Citations

  • Teaching outline generation method and device, storage medium and electronic equipment

    CN112232066A

  • Education knowledge-based resource management platform and management method

    CN118608344A

  • Personnel configuration CMS content feature extraction method and system and medium

    CN118886414A

  • Teaching video reconstruction method and system based on intelligent agent

    CN119271844A

  • Digital intelligent teaching recording and broadcasting system and method

    CN119379502A