Teaching video summary mind map generation system based on multi-modal retrieval enhancement generation
The generated teaching video summary mind map system is solved through multimodal retrieval, which is not very transferable, inaccurate node extraction and wrong hierarchical relationships in the existing technology, and has achieved improvements in multilingual adaptability and accuracy, and the generated mind map is more reliable.
Patent Information
- Application Number
- CN202510276213.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art has problems such as poor language transferability, inaccurate node extraction and hierarchical relationship errors when generating mind maps, especially when processing video modal information, which fails to effectively utilize visual information.
The teaching video summary mind map system is adopted to generate multimodal retrieval enhancement, and the audio and visual information are separated through the video acquisition module, the Whisper-NER model is used for text transcription and entity extraction, combined with the GLM4 model for visual feature extraction and missing frame search, and integrated multimodal information to generate mind maps.
It improves the language transferability, node extraction accuracy and hierarchical relationship of mind maps, enhances the reliability of generated mind maps, is suitable for teaching videos in multiple languages, and ensures the reliability of generated results through improved evaluation indicators.
Smart Images

Figure CN120256680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electrical testing, and specifically relates to a teaching video summary mind map generation system based on multi-modal retrieval enhanced generation. Background Art
[0002] To solve the problem of automatic mind map generation, Abdeen M et al. first proposed the M2Gen method, which automatically converts the input text into a mind map based on components such as a lexical analyzer and a syntactic analyzer; Purwarianti A et al. migrated the M2Gen method to Indonesian and added a tag acquisition step in the generation stage to improve the accuracy of node generation, expanding the language domain of the existing method; Vimalaksha A et al. first focused on the mind map generation task of video lectures and realized a multi-modal generation method from the video (audio) modality to the text modality by means of video segmentation, audio transcription, data cleaning, etc. The nodes, semantic relationships, etc. obtained in the previous steps are input into a visualization tool to achieve an automatic video summary mind map.
[0003] However, the M2Gen method proposed by Abdeen M et al. is only applicable to English texts and has weak portability; although the method of Purwarianti A et al. has migrated the language of the M2Gen method, it is still limited to the text modality and cannot handle the generation task from the video modality to the mind map. In addition, this method uses traditional natural language processing techniques, focusing on the analysis of syntactic and semantic dependency relationships in the text, lacking external prior knowledge of the mutual relationships between nodes in the objective world, so there is a problem of incorrect node hierarchical relationships in the generated mind map; although the video lecture mind map generation method of Vimalaksha A et al. has extended the method to the input modality of the video, it actually only uses the audio information in the video and does not analyze and effectively utilize the visual information in the video. Therefore, there is also a problem of incorrect hierarchical relationships between nodes in the generated mind map. In summary, the existing methods mainly have problems such as weak language portability, inaccurate node extraction, and incorrect hierarchical relationships. Summary of the Invention
[0004] The object of the present invention is to provide a teaching video summary mind map generation system based on multi-modal retrieval enhanced generation, including: a video acquisition module, a signal separation module, a text transcription module, a frame sampling module, an entity extraction module, a mind map node extraction module, a visual feature extraction module, a visual feature missing judgment module, a missing frame search module, and a mind map generation module.
[0005] The video acquisition module is used to acquire the original video signal.
[0006] The signal separation module is used to separate audio information and visual information from the original video signal.
[0007] The text transcription module is used to transcribe the audio information to obtain text information.
[0008] The frame sampling module is used to uniformly sample the visual information to construct a set of video frames.
[0009] The entity extraction module is used to extract entities and the relationships between entities from the text information.
[0010] The mind map node extraction module is used to extract mind map nodes from the text information.
[0011] The visual feature extraction module is used to extract visual features from the set of video frames to obtain visual features.
[0012] The visual feature missing judgment module is used to judge whether the visual features extracted by the visual feature extraction module are complete. If so, the extracted visual features are input into the mind map generation module. If not, the extracted visual features are input into the missing frame search module.
[0013] The missing frame search module is used to select missing frames from the visual information, add the missing frames to the set of video frames, and input the set of video frames into the visual feature extraction module.
[0014] The mind map generation module is used to integrate entities and the relationships between entities, mind map nodes, and visual features to generate a mind map.
[0015] Furthermore, the text information is transcribed using the Whisper-NER model.
[0016] Furthermore, the text information is as follows:
[0017] R speech =(S, σ) (1)
[0018] In the formula, S represents the audio information, σ represents the text transcription method, and R speech represents the text information.
[0019] Furthermore, the entities and the relationships between entities are extracted using the Whisper-NER model.
[0020] The output of the Whisper-NER model is as follows:
[0021] y s =Decoder(y 1:s-1 , h, t) (2)
[0022] Wherein, s is the output sequence index, and t is the entity label. y s is the output sequence. Decoder is the decoder. y 1:s-1 is the set of the previous s - 1 output sequences.
[0023] Among them, the hidden state h is as follows:
[0024] h = Encoder(x) (3)
[0025] Wherein, x is the audio information. Encoder is the audio encoder.
[0026] The loss function of the Whisper - NER model is as follows:
[0027]
[0028] Wherein, is the loss function. n is the total number of output sequences. y* is the true sequence. y is the predicted output sequence to be minimized. P(y s = y s *|y 1:s-1 , h, t) is the standard cross - entropy between the output sequence y s and the true sequence y s *.
[0029] Furthermore, the relationships between entities include the syntactic and semantic relationships of the text.
[0030] Furthermore, the mind map nodes are extracted by the GLM4 model.
[0031] The GLM4 model extracts appropriate granularity nodes according to the input natural language, and adjusts the GLM4 model by combining the prompt words containing the original text and the granularity.
[0032] The mind map nodes are as follows:
[0033] N = (R speech , λ) (5)
[0034] Wherein, N is the mind map node. λ is the node extraction method. R speech represents the text information.
[0035] Furthermore, the visual features include the title in the video presentation, the content of the text, the font and position information of the text, and the hierarchical relationship of the mind map nodes.
[0036] Furthermore, whether the extracted visual features are complete is judged by the confidence score output by the GLM4 model.
[0037] If the confidence score output by the GLM4 model is greater than the preset threshold, it is determined that the extracted visual features are incomplete. If the confidence score output by the GLM4 model is less than or equal to the preset threshold, it is determined that the extracted visual features are complete.
[0038] The GLM4 model combines text information and the timeline for self-questioning. If the visual features are incomplete, it performs self-retrieval based on the visual features, text information, and timeline to select the missing frames.
[0039] Furthermore, the visual features are extracted by an OCR model.
[0040] The visual features are as follows:
[0041] R img =(I, ρ), I = φ(video) (6)
[0042] In the formula, φ represents the frame selection method, I represents the set of video frames, ρ represents the visual extraction method, and the visual extraction method includes the optical character recognition method. R img represents the extracted visual features.
[0043] Furthermore, the entities and the relationships between entities, the mind map nodes, and the visual features are integrated by the GLM4 model.
[0044] The mind map is as follows:
[0045] M=(G, R img =(I, ρ), R speech =(S, σ), N=(R speech ,(λ)) (7)
[0046] In the formula, M is the generated mind map. G is the mind map generation method. φ represents the frame selection method, I represents the set of video frames, ρ represents the visual extraction method, and R img represents the extracted visual features. N is the mind map node. λ is the node extraction method. R speech represents the text information. S represents the audio information, σ represents the text transcription method, and R speech represents the text information.
[0047] The technical effects of the present invention are beyond doubt. Based on the mind map technology for summarizing teaching videos enhanced by multi-modal retrieval, the present invention addresses the problems existing in the existing methods, such as weak language transferability, the inability of single-modal (text-modal) methods to meet the requirements of video summarization tasks, and the ineffective utilization of visual-modal information in multi-modal methods. Improvements are made in aspects such as language transferability, node extraction accuracy, and hierarchical relationship organization. By using a large language model as the data integration center, relying on its rich multi-lingual prior knowledge to solve the language transfer problem; constructing a "video - mind map" generation architecture enhanced by multi-modal retrieval, accurately extracting mind map nodes according to human intentions, and efficiently retrieving visual information in teaching videos to enhance the generation ability of the large model, and then performing feature fusion of audio and video modalities, ensuring the accuracy of node extraction and the correctness of hierarchical relationships.
[0048] In terms of language transferability, the present invention uses the multi-lingual pre-trained large language model GLM4 for multi-modal fusion to guide the generation of mind maps. Benefiting from its large amount of training data, the large language model can translate and understand hundreds of languages including Chinese, English, French, Japanese, Korean, etc., improving the language adaptability of the technology and enabling it to be applied to teaching videos in most languages.
[0049] In terms of node extraction accuracy, the present invention constructs a node extraction model with appropriate granularity. Traditional methods for extracting nodes based on significant sentences (SSM) or keywords (KSM) have problems with inappropriate granularity. The SSM method directly divides sentences on the input text and uses sentence paragraphs as mind map nodes, with too coarse a granularity. The KSM method first uses techniques such as named entity recognition to extract keywords in the input text, and the resulting nodes often have too fine a granularity. In the teaching video understanding task, the granularity of the title nodes that meet the user's intentions is not fixed. The present invention uses a large language model for instruction fine-tuning to make the granularity of the nodes extracted at this stage conform to human intentions, so as to solve the problem of inaccurate mind map node extraction in traditional methods. Its advantages in the node extraction task lie first in its large amount of prior knowledge, which can accurately understand the syntactic and semantic relationships of the input text; in addition, the large model can perform node extraction with appropriate granularity according to natural language input in accordance with human intentions, avoiding the problems of too coarse or too fine granularity in traditional methods; finally, using prompts containing the original text and granularity, and benchmark mind map nodes annotated by humans to fine-tune the large model can enable the large model to learn the expressions of this specific downstream task of teaching video understanding, thereby improving the accuracy of node extraction. The nodes extracted by the node extraction model of the present invention have appropriate granularity and can accurately align with human intentions.
[0050] In terms of the accuracy of the hierarchical relationship, the present invention proposes a method for efficiently extracting visual modal features from teaching videos. This method is different from general video understanding methods and uses optical character recognition technology to extract text information from specific frames. Specifically, the extracted visual features include the titles, fonts, and position information of the text in the video presentation to guide the hierarchical relationship of nodes in the mind map generation process. The basis for adopting this visual feature extraction method is that the present invention observes the difference between teaching video understanding and general video understanding: in teaching videos, in addition to the teacher's voice explanation, the presentation in the video screen also contains a large amount of text content and hierarchical structure information, while traditional video understanding methods mainly focus on identifying people, objects, relationships, and actions in images, which is inconsistent with the characteristics of teaching videos. Therefore, this paper focuses on the feature extraction of text content, font, and position relationship to improve the accuracy of the hierarchical relationship organization of the mind map. Combining the extracted visual features with audio features to jointly guide the generation of the mind map solves the problem of incorrect hierarchical relationships in the mind map in traditional methods.
[0051] In terms of experimental indicators, the existing automated evaluation indicators for mind map generation only focus on the recall rate of the generated edge set in the benchmark edge set and do not pay attention to the overall hierarchical relationship of the mind map, which has a certain impact on the accuracy and reliability of the experimental results. The present invention proposes an automated evaluation indicator that combines the hierarchical relationship of the mind map. By adding the hierarchical and distance penalty factors of the edges, it can take into account both the recall rate of the edges and the hierarchical relationship of the edges, solve the problem that traditional methods only consider the recall rate of the edges and do not pay attention to the overall hierarchical relationship, and make the evaluation indicators for mind map generation more reliable.
[0052] In summary, the present invention has made technical improvements and innovations in terms of language transferability, node extraction accuracy, hierarchical relationship accuracy, and experimental indicator reliability, significantly improving the transferability of the "teaching video - mind map" method and the accuracy of generating mind maps, and making the evaluation indicators for automatically generating mind maps more reliable. Brief Description of the Drawings
[0053] Figure 1 is the overall flowchart of the present invention;
[0054] Figure 2 is the overall algorithm architecture diagram of the present invention;
[0055] Figure 3 is the pseudo code of the evaluation indicator algorithm of the present invention. Detailed Implementation Manner
[0056] The present invention will be further described below in conjunction with embodiments, but it should not be understood that the above-mentioned subject matter scope of the present invention is limited to the following embodiments. Without departing from the above-mentioned technical idea of the present invention, various substitutions and changes made according to common general knowledge and conventional means in the art should be included within the protection scope of the present invention.
[0057] Embodiment 1:
[0058] See Figures 1 to 3 , a mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation, including: a video acquisition module, a signal separation module, a text transcription module, a frame sampling module, an entity extraction module, a mind map node extraction module, a visual feature extraction module, a visual feature missing judgment module, a missing frame search module, and a mind map generation module.
[0059] The video acquisition module is used to acquire the original video signal.
[0060] The signal separation module is used to separate audio information and visual information from the original video signal.
[0061] The text transcription module is used to transcribe the audio information into text information.
[0062] The frame sampling module is used to uniformly sample the visual information to construct a video frame set.
[0063] The entity extraction module is used to extract entities and the relationships between entities from the text information.
[0064] The mind map node extraction module is used to extract mind map nodes from the text information.
[0065] The visual feature extraction module is used to extract visual features from the video frame set to obtain visual features.
[0066] The visual feature missing judgment module is used to judge whether the visual features extracted by the visual feature extraction module are complete. If so, the extracted visual features are input into the mind map generation module. If not, the extracted visual features are input into the missing frame search module.
[0067] The missing frame search module is used to select missing frames from the visual information, add the missing frames to the video frame set, and input the video frame set into the visual feature extraction module.
[0068] The mind map generation module is used to integrate entities and the relationships between entities, mind map nodes, and visual features to generate a mind map.
[0069] Embodiment 2:
[0070] A mind map generation system for teaching video summary based on multimodal retrieval enhanced generation. The main technical content is shown in Embodiment 1. Further, the text information is transcribed using the Whisper-NER model.
[0071] Embodiment 3:
[0072] A mind map generation system for teaching video summary based on multimodal retrieval enhanced generation. The main technical content is shown in any one of Embodiments 1 to 2. Further, the text information is as follows:
[0073] R speech =(S, σ) (1)
[0074] In the formula, S represents audio information, σ represents the text transcription method, and R speech represents text information.
[0075] Embodiment 4:
[0076] A mind map generation system for teaching video summary based on multimodal retrieval enhanced generation. The main technical content is shown in any one of Embodiments 1 to 3. Further, the entities and the relationships between entities are extracted using the Whisper-NER model.
[0077] The output of the Whisper-NER model is as follows:
[0078] y s =Decoder(y 1:s-1 , h, t) (2)
[0079] In the formula, s is the output sequence index, and t is the entity label. y s is the output sequence. Decoder is the decoder. y 1:s-1 is the set of the previous s - 1 output sequences.
[0080] Among them, the hidden state h is as follows:
[0081] h = Encoder(x) (3)
[0082] In the formula, x is the audio information. Encoder is the audio encoder.
[0083] The loss function of the Whisper-NER model is as follows:
[0084]
[0085] In the formula, is the loss function. n is the total number of output sequences. y* is the true sequence. y is the predicted output sequence to be minimized. P(y s =y s *|y1:s-1 The output sequence \(y = f(x, h, t)\) s and the true sequence \(y^*\) s The standard cross-entropy between them.
[0086] Example 5:
[0087] A mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation. The main technical content can be found in any one of Examples 1 to 4. Further, the relationships between entities include the syntactic and semantic relationships of the text.
[0088] Example 6:
[0089] A mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation. The main technical content can be found in any one of Examples 1 to 5. Further, the mind map nodes are extracted by the GLM4 model.
[0090] The GLM4 model performs adaptive granularity node extraction based on the input natural language, and adjusts the GLM4 model by combining the prompt words containing the original text and the granularity.
[0091] The mind map nodes are as follows:
[0092] N = (R speech , λ) (5)
[0093] In the formula, N is the mind map node. λ is the node extraction method. R speech represents the text information.
[0094] Example 7:
[0095] A mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation. The main technical content can be found in any one of Examples 1 to 6. Further, the visual features include the title in the video presentation, the content of the text, the font and position information of the text, and the hierarchical relationship of the mind map nodes.
[0096] Example 8:
[0097] A mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation. The main technical content can be found in any one of Examples 1 to 7. Further, whether the extracted visual features are complete is judged by the confidence score output by the GLM4 model.
[0098] If the confidence score output by the GLM4 model is greater than the preset threshold, it is judged that the extracted visual features are incomplete. If the confidence score output by the GLM4 model is less than or equal to the preset threshold, it is judged that the extracted visual features are complete.
[0099] The GLM4 model combines text information and the timeline for self-questioning. If the visual features are incomplete, it performs self-retrieval based on the visual features, text information, and timeline to select the missing frames.
[0100] Example 9:
[0101] For the mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement, the main technical content can be found in any one of Examples 1 to 8. Further, the visual features are extracted by an OCR model.
[0102] The visual features are as follows:
[0103] R img =(I, ρ), I = φ(video) (6)
[0104] In the formula, φ represents the frame selection method, I represents the set of video frames, ρ represents the visual extraction method, and the visual extraction method includes the optical character recognition method. R img represents the extracted visual features.
[0105] Example 10:
[0106] For the mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement, the main technical content can be found in any one of Examples 1 to 9. Further, the entities and the relationships between them, the mind map nodes, and the visual features are integrated by the GLM4 model.
[0107] The mind map is as follows:
[0108] M=(G, R img =(I, ρ), R speech =(S, σ), N=(R speech , λ)) (7)
[0109] In the formula, M is the generated mind map. G is the mind map generation method. φ represents the frame selection method, I represents the set of video frames, ρ represents the visual extraction method, and R img represents the extracted visual features. N is the mind map node. λ is the node extraction method. R speech represents the text information. S represents the audio information, σ represents the text transcription method, and R speech represents the text information.
[0110] Example 11:
[0111] See Figures 1 to 3, A mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation, including: a video acquisition module, a signal separation module, a text transcription module, a frame sampling module, an entity extraction module, a mind map node extraction module, a visual feature extraction module, a visual feature missing judgment module, a missing frame search module, and a mind map generation module.
[0112] The video acquisition module is used to downsample the original input video to a video with 1 FPS and obtain the original video signal.
[0113] The signal separation module is used to separate audio information and visual information from the original video signal.
[0114] The text transcription module is used to transcribe the audio information into text information.
[0115] The frame sampling module is used to uniformly sample the visual information and perform optical character recognition to construct a video frame set. For example, 1 frame is taken every 5 seconds and added to the frame set to be visually extracted.
[0116] The entity extraction module is used to extract entities and the relationships between entities from the text information.
[0117] The mind map node extraction module is used to extract mind map nodes from the text information.
[0118] The visual feature extraction module is used to extract visual features from the video frame set to obtain visual features.
[0119] The visual feature missing judgment module uses a proxy large model to perform iterative self-querying to judge whether the visual features extracted by the visual feature extraction module are complete. If so, the extracted visual features are input into the mind map generation module. If not, the extracted visual features are input into the missing frame search module.
[0120] The missing frame search module is used to select missing frames from the visual information using a multi-modal proxy model, add the missing frames to the video frame set, and input the video frame set into the visual feature extraction module.
[0121] The mind map generation module is used to integrate entities and the relationships between entities, mind map nodes, and visual features, fuse multi-modal information to enhance the generation ability of the central large model, and generate a mind map in mermaid format.
[0122] Example 12:
[0123] A mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation, the main technical content can be seen in Example 11. Further, the text information is transcribed using the Whisper-NER model.
[0124] In the audio processing part, "audio-text" transcription uses the pre-trained Whisper-NER model, which is an advanced open-source speech recognition and named entity extraction model. It can perform accurate and effective "audio-text" transcription and extract named entities and relationships in the text, providing information reference for audio modality in subsequent multimodal fusion.
[0125] Example 13:
[0126] A mind map generation system for teaching video summary based on multimodal retrieval enhanced generation. The main technical content can be found in any one of Examples 11 to 12. Further, the text information is as follows:
[0127] R speech =(S, σ) (1)
[0128] In the formula, S represents audio information, σ represents the text transcription method, and R speech represents text information.
[0129] Example 14:
[0130] A mind map generation system for teaching video summary based on multimodal retrieval enhanced generation. The main technical content can be found in any one of Examples 11 to 13. Further, the entities and the relationships between entities are extracted using the Whisper-NER model.
[0131] Let the audio input be x, and input x into the audio encoder to generate a series of hidden states h as follows:
[0132] h = Encoder(x) (2)
[0133] With a set of entity tags t = [t1, t2,..., t k (each t i tag represents a specific entity type, such as knowledge points or code) as a condition to regulate the decoding process, and then the decoder iteratively generates the output sequence y = [y1, y2,..., y n , which includes the transcribed text and entity tags. The representation of the output sequence is:
[0134] y t = Decoder(y 1:t-1 , h, t) (3)
[0135] Among them, y 1:t-1 is the set of outputs of all previous iteration steps, and each iteration step output y t contains the transcribed text and entity tags (a subset of t). For example: <person>Dani walks in the <location>national park. Among them <person>And <location>is an element in the entity tag set t, and the rest is the transcribed text.
[0136] The model is trained to minimize the standard cross-entropy loss between the predicted output sequence y and the true sequence y*, and the loss function is as follows:
[0137]
[0138] After training, the Whiper-NER model can extract entities and relationships between entities from audio.
[0139] Example 15:
[0140] A mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation. The main technical content can be found in any one of Examples 11 to 14. Further, the relationships between the entities include the syntactic and semantic relationships of the text.
[0141] Example 16:
[0142] A mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation. The main technical content can be found in any one of Examples 11 to 15. Further, the mind map nodes are extracted by the GLM4 model.
[0143] The GLM4 model performs adaptive granularity node extraction according to the input natural language, and adjusts the GLM4 model by combining the prompt words containing the original text and the granularity.
[0144] The extraction of mind map nodes uses GLM4. Its advantages in the node extraction task are firstly that its huge prior knowledge can accurately understand the syntactic and semantic relationships of the input text. In addition, the large model can perform node extraction with appropriate granularity according to the natural language input, avoiding the problems of too coarse or too fine granularity in traditional methods. Finally, using the prompt words containing the original text and the granularity, and the benchmark mind map nodes annotated by humans, to perform instruction fine-tuning on the large model can enable the large model to learn the expressions of this specific downstream task of teaching video understanding, thereby improving the accuracy of node extraction.
[0145] The mind map nodes are as follows:
[0146] N = (R speech , λ) (5)
[0147] In the formula, N is the mind map node. λ is the node extraction method. R speech represents the text information.
[0148] Example 17:
[0149] A mind map generation system for teaching video summary based on multi-modal retrieval enhanced generation. The main technical content can be found in any one of Embodiments 11 to 16. Further, the visual features include the title in the video presentation, the content of the text, the font and position information of the text, and the hierarchical relationship of the mind map nodes.
[0150] Embodiment 18:
[0151] A mind map generation system for teaching video summary based on multi-modal retrieval enhanced generation. The main technical content can be found in any one of Embodiments 11 to 17. Further, whether the extracted visual features are complete is judged by the confidence score output by the GLM4 model.
[0152] If the confidence score output by the GLM4 model is greater than the preset threshold, it is judged that the extracted visual features are incomplete. If the confidence score output by the GLM4 model is less than or equal to the preset threshold, it is judged that the extracted visual features are complete.
[0153] This confidence threshold is a decimal number in the range of [0, 1). For example, the confidence threshold for feature missing is set to 0.6 (indicating a 60% probability of feature missing). When the model outputs a confidence of 0.7 > 0.6, it is considered that feature missing has occurred; when the output confidence is 0.5 < 0.6, it is considered that feature missing has not occurred.
[0154] The GLM4 model combines text information and the timeline for self-questioning. If the visual features are incomplete, it performs self-retrieval according to the visual features, text information, and timeline to select the missing frames.
[0155] Whether visual features are missing is judged through the information integration of the proxy model GLM4 and self-questioning to ensure consistency. GLM4 first integrates text, timeline, and the currently collected visual features for information understanding. Subsequently, GLM4 generates multiple confidence scores for visual feature missing in parallel and samples the generated multiple confidence scores through a greedy decoding strategy to obtain a final self-consistent confidence output. According to the pre-set confidence threshold, if the confidence score output by the model is higher than the threshold, it is judged that there are missing visual features, otherwise it is judged that the visual features are complete.
[0156] After initial frame sampling and visual extraction, the extracted visual information is sent to the visual proxy large model, and GLM4 is also used as the proxy large model. The proxy large model will combine the transcribed audio text and the timeline for self-questioning. If visual information is missing, it performs self-retrieval according to the extracted visual information and the audio text timeline information, selects the missing frames, uses the OCR model to extract visual features, and sends them to the proxy large model for the next iteration; if the visual information has been completely obtained, the extracted visual information is sent to the multi-modal fusion step to guide the mind map generation.
[0157] Example 19:
[0158] A mind map generation system for summarizing teaching videos based on multi-modal retrieval enhanced generation. The main technical content can be found in any one of Examples 11 to 18. Further, the visual features are extracted by an OCR model.
[0159] The visual features are as follows:
[0160] R img =(I, ρ), I = φ(video) (6)
[0161] In the formula, φ represents the frame selection method, I represents the set of video frames, ρ represents the visual extraction method, and the visual extraction method includes the optical character recognition method. R img represents the extracted visual features.
[0162] Example 20:
[0163] A mind map generation system for summarizing teaching videos based on multi-modal retrieval enhanced generation. The main technical content can be found in any one of Examples 11 to 19. Further, the entities, the relationships between entities, the mind map nodes, and the visual features are integrated by the GLM4 model.
[0164] After extracting the features of the audio modality and the visual modality, it is necessary to align and fuse the features of the two modalities to guide the generation of the mind map. In this example, "audio-text" transcription is used for extracting the features of the audio modality, and the mind map nodes are also extracted from the text. For extracting the visual features in this example, an OCR model is used to obtain information such as the text content, font, and position in the image and output it in text form. In summary, the audio and visual modality information in the video has been aligned to the text modality after the feature extraction stage, and no additional alignment steps are required. Considering the powerful integration ability of the large language model for text modality data, in this example, the pre-trained large language model GLM4 is used to fuse the audio and visual features of the text modality, combined with the mind map nodes extracted in the first stage, to guide the generation of the mind map in the second stage.
[0165] The mind map is as follows:
[0166] M=(G, R img =(I, ρ), R speech =(S, σ), N=(R speech , λ)) (7)
[0167] In the formula, M is the generated mind map. G is the mind map generation method. φ represents the frame selection method, I represents the set of video frames, ρ represents the visual extraction method, R img Represents the extracted visual features. N is the mind map node. λ is the node extraction method. R speech Represents the text information. S represents the audio information, σ represents the text transcription method, R speech Represents the text information.
[0168] In addition, in terms of model evaluation, the existing automated mind map generation evaluation metrics only focus on the recall rate of the generated edge set in the benchmark edge set, and do not pay attention to the overall hierarchical relationship of the mind map, which has a certain impact on the accuracy and reliability of the experimental results. As Figure 3 shown, the present invention proposes an automated evaluation metric that combines the hierarchical relationship of the mind map. By adding the hierarchical and distance penalty factors of the edges, it solves the problem that the traditional method only considers the recall rate of the edges and does not pay attention to the overall hierarchical relationship, making the evaluation metric for mind map generation more reliable.
[0169] Example 21:
[0170] See Figures 1 to 3 , the mind map generation system for teaching video summary based on multi-modal retrieval enhancement, the main technical contents include:
[0171] Based on the technology of the mind map for teaching video summary generated by multi-modal retrieval enhancement in this embodiment, aiming at the problems existing in the existing methods, such as weak language transferability, the single-modal (text modality) method cannot meet the requirements of the video summary task, and the multi-modal method does not effectively utilize the information of the visual modality, improvements are made in aspects such as language transferability, node extraction accuracy, and hierarchical relationship organization. By using a large language model as the data integration center, relying on its rich multi-language prior knowledge to solve the language transfer problem; constructing a "video - mind map" generation architecture enhanced by multi-modal retrieval, accurately extracting the mind map nodes according to human intentions, and efficiently retrieving the visual information in the teaching video to enhance the generation ability of the large model, and then performing feature fusion of the audio and video modalities to ensure the accuracy of node extraction and the correctness of the hierarchical relationship.
[0172] As Figure 1 shown, it is the overall flowchart of this embodiment, and the specific steps are as follows:
[0173] 1) Downsample the original input video to a 1FPS video, input it into the model, and then perform steps 2) and 5).
[0174] 2) Separate the audio information from the video, perform text transcription, and then perform steps 3) and 4).
[0175] 3) Extract the named entities and relationships from the text transcribed from the audio and pass them to step 8).
[0176] 4) Use a large language model to extract the mind map nodes from the text transcribed from the audio and pass them to step 8).
[0177] 5) Separate the visual information from the video, perform initial uniform frame sampling and perform optical character recognition, and then proceed to step 6).
[0178] 6) Use the proxy large model for iterative self-querying. If the visual information is complete, pass it in and proceed to step 8); otherwise, proceed to step 7).
[0179] 7) The multi-modal proxy model selects the missing frames, adds them to the selected frame set, and returns to step 5).
[0180] 8) Integrate the mind map nodes, named entities and relationships, and visual information passed in from steps 3), 4), and 6), and combine with the audio text transcribed in step 2) to fuse multi-modal information to enhance the generation ability of the central large model, and generate a mind map in mermaid format.
[0181] As Figure 2 shown, it is the overall algorithm architecture diagram of this embodiment. Before performing multi-modal fusion in the final stage, this embodiment can be divided into two parts: audio processing and visual processing
[0182] In the audio processing part, the "audio-text" transcription uses the pre-trained Whisper-NER model, which is an advanced open-source speech recognition and named entity extraction model that can perform accurate and effective "audio-text" transcription and extract the named entities and relationships in the text, providing an information reference for audio modality for subsequent multi-modal fusion. The mind map node extraction uses GLM4. Its advantage in the node extraction task lies first in its vast prior knowledge, which can accurately understand the syntax and semantic relationships of the input text. In addition, the large model can perform node extraction with appropriate granularity according to the natural language input, following the human intention, avoiding the problems of overly coarse or overly fine granularity in traditional methods. Finally, using the prompt words containing the original text and granularity, and the benchmark mind map nodes annotated by humans to fine-tune the large model can enable the large model to learn the expressions of this specific downstream task of teaching video understanding, thereby improving the accuracy of node extraction. Among them, the audio subtitle extraction can be expressed by the following formula 1, where S represents the input audio, σ represents the audio extraction method, and Rspeech represents the extracted audio subtitle
[0183] R speech = (S, σ) (1)
[0184] And the node extraction can be expressed by the following formula 2, where λ represents the node extraction method
[0185] N = (R speech ,λ) (2)
[0186] In the visual processing part, the MinerU model is used to extract visual information. The extracted visual features include the title, font, and position information of the text in the video presentation, which are used to guide the hierarchical relationship of nodes in the mind map generation process. The initial frames are selected using the method of uniform frame sampling. For example, one frame is added to the set of frames to be visually extracted every 5 seconds. After initial frame sampling and visual extraction, the extracted visual information is sent to the visual proxy large model, and GLM4 is also used as the proxy large model. The proxy large model will conduct self-query in combination with the transcribed audio text and the timeline. If visual information is missing, it will perform self-retrieval based on the extracted visual information and the audio text timeline information, select the missing frames, extract visual features using the OCR model, and send them to the proxy large model for the next iteration; if the visual information has been completely obtained, the extracted visual information will be sent to the multi-modal fusion step to guide the mind map generation. Based on the characteristic that the variance between the number of pages of the presentation and the number of words per page in the teaching video (which can be regarded as proportional to the single-page presentation time) is relatively large, the method adopted in this embodiment can accurately select the unique frames close to the minimum, ensuring both the integrity of information and the efficiency of visual feature extraction. This processing part can be represented by the following formula 3, where φ represents the frame selection method, I is the set of frames selected from the input video, ρ represents the visual extraction method, and Rimg represents the visual information extracted by the visual extraction part:
[0187] R img =(I, ρ), I = φ(video) (3)
[0188] After extracting the features of the audio modality and the visual modality, it is necessary to align and fuse the features of the two modalities to guide the generation of the mind map. In this embodiment, "audio-text" transcription is used to extract the features of the audio modality, and the mind map nodes are also extracted from the text. In this embodiment, the OCR model is used to extract visual features, obtaining information such as the text content, font, and position in the image and outputting it in text form. In summary, the audio and visual modality information in the video has been aligned to the text modality after the feature extraction stage, and no additional alignment steps are required. Considering the powerful integration ability of the large language model for text modality data, in this embodiment, the pre-trained large language model GLM4 is used to fuse the audio and visual features of the text modality, and combined with the mind map nodes extracted in the first stage, it guides the generation of the mind map in the second stage.
[0189] The overall process of this embodiment can be represented by the following formula 4, where G represents the mind map generation method:
[0190] M=(G, R img =(I, ρ), R speech =(S, σ), N=(R speech , λ)) (4)
[0191] In addition, in terms of model evaluation, the existing automated mind map generation evaluation metrics only focus on the recall rate of the generated edge set in the benchmark edge set, and do not pay attention to the overall hierarchical relationship of the mind map, which has a certain impact on the accuracy and reliability of the experimental results. As Figure 3 shown, this embodiment proposes an automated evaluation metric that combines the hierarchical relationship of the mind map. By adding the hierarchical level and distance penalty factors of the edges, it solves the problem that only the recall rate of the edges is considered in the traditional method without paying attention to the overall hierarchical relationship, making the evaluation metric for mind map generation more reliable.< / location> < / person> < / location> < / person>
Claims
1. A mind map generation system for teaching video summary enhanced by multimodal retrieval and generation, characterized in that, Including: Video acquisition module, signal separation module, text transcription module, frame sampling module, entity extraction module, mind map node extraction module, visual feature extraction module, visual feature missing judgment module, missing frame search module, mind map generation module. The video acquisition module is used to acquire the original video signal; The signal separation module is used to separate the audio information and visual information from the original video signal; The text transcription module is used to transcribe the audio information into text information; The frame sampling module is used to uniformly sample the visual information to construct a video frame set; The entity extraction module is used to extract entities and the relationships between entities from the text information; The mind map node extraction module is used to extract mind map nodes from the text information; The visual feature extraction module is used to extract visual features from the video frame set to obtain visual features; The visual feature missing judgment module is used to judge whether the visual features extracted by the visual feature extraction module are complete. If so, the extracted visual features are input into the mind map generation module. If not, the extracted visual features are input into the missing frame search module; The missing frame search module is used to select missing frames from the visual information, add the missing frames to the video frame set, and input the video frame set into the visual feature extraction module; The mind map generation module is used to integrate entities and the relationships between entities, mind map nodes, and visual features to generate a mind map; 2. The mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation according to claim 1, wherein The text information is transcribed using the Whisper-NER model; 3. The teaching video summary mind map generation system based on multi-modal retrieval enhanced generation according to claim 1, wherein The text information is as follows: R speech = (S, σ) (1) Where S represents audio information, σ represents the text transcription method, and R speech represents text information.
4. The mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation according to claim 1, wherein The entities and the relationships between entities are extracted using the Whisper-NER model; The output of the Whisper-NER model is as follows: y s = Decoder(y 1:s-1 , h, t) (2) where s is the output sequence index, t is the entity label; y s is the output sequence; Decoder is the decoder; y 1:s-1 is the set of the previous s - 1 output sequences; Among them, the hidden state h is as follows: h = Encoder(x) (3) In the formula, x is the audio information; Encoder is the audio encoder; The loss function of the Whisper-NER model is as follows: Wherein, is the loss function; n is the total number of output sequences; y* is the true sequence; y is the minimized predicted output sequence; P(y s = y s *|y 1:s-1 , h, t) is the standard cross-entropy between the output sequence y s and the true sequence y s *.
5. The mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation according to claim 1, wherein The relationships between entities include the syntactic and semantic relationships of the text; 6. The mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation according to claim 1, wherein The mind map nodes are extracted by the GLM4 model; The GLM4 model performs adaptive granularity node extraction according to the input natural language, and adjusts the GLM4 model by combining the prompt words containing the original text and granularity; The mind map nodes are as follows: N = (R speech , λ) (5) Where N is the mind map node; λ is the node extraction method; R speech represents the text information.
7. The teaching video summary mind map generation system based on multi-modal retrieval enhanced generation according to claim 1, characterized in that, The visual features include the title in the video presentation, the content of the text, the font and position information of the text, and the hierarchical relationship of the mind map nodes; 8. The mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation according to claim 1, wherein Whether the extracted visual features are complete is judged by the confidence score output by the GLM4 model; If the confidence score output by the GLM4 model is greater than the preset threshold, it is judged that the extracted visual features are incomplete; If the confidence score output by the GLM4 model is less than or equal to the preset threshold, it is judged that the extracted visual features are complete; The GLM4 model combines the text information and the time axis for self-questioning. If the visual features are incomplete, it performs self-retrieval according to the visual features, text information, and time axis to select missing frames; 9. The mind map generation system for summarizing teaching videos based on multi-modal retrieval enhancement generation according to claim 1, characterized in that, The visual features are extracted by the OCR model; The visual features are as follows: R img = (I, ρ), I = φ(video) (6) Where φ represents the frame selection method, I represents the set of video frames, ρ represents the visual extraction method, and the visual extraction method includes an optical character recognition method, and R img represents the extracted visual features.
10. The teaching video summary mind map generation system based on multi-modal retrieval enhanced generation according to claim 1, wherein The entity, the relationships between entities, the mind map nodes, and the visual features are integrated through the GLM4 model; The mind map is as follows: M = (G, R img = (I, ρ), R speech = (S, σ), N = (R speech , λ)) (7) Wherein, M is the generated mind map; G is the mind map generation method; φ represents the frame selection method, I represents the video frame set, ρ represents the visual extraction method, and R img represents the extracted visual feature; N is the mind map node; λ is the node extraction method; R speech represents the text information; S represents the audio information, σ represents the text transcription method, and R speech represents the text information.
Citation Information
Patent Citations
Video processing method and video retrieval enhancement method for multi-modal large model
CN119339284A
AI-driven multilingual intelligent document abstract and mind map generation system
CN119494394A
Intelligent social media post generation method and system based on multi-modal architecture
CN119513425A
Multimodal heterogeneous feature fusion-based compact video event description method
WO2023050295A1