Teaching video generation method and device based on semantic knowledge base and electronic equipment

By constructing a hierarchical semantic knowledge base and digital human model, and generating structured slides and supporting speeches, the problems of imprigor and multimodal collaboration in the production of existing teaching videos are solved, and efficient and accurate teaching video generation is achieved.

CN120499463APending Publication Date: 2025-08-15PAZHOU LAB (HUANGPU) +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510762977.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing teaching video production process relies on manual recording and post-editing, which is difficult to ensure that the logic of the teaching content is rigorous and diverse. The existing automation tools cannot achieve integrated intelligent generation from "knowledge source" to "finished video products", and it is difficult to quickly adjust the output according to different learning situations and teaching styles.

Method used

A hierarchical semantic knowledge base is constructed based on multi-source data, and courses, chapters, concepts and relationships are organized through semantic knowledge graphs, structured slide content is generated, and supporting lectures are generated using language models, and teaching videos are generated based on digital human models.

Benefits of technology

The systematic organization, semantic correlation and structured expression of teaching content are realized, which significantly reduces the time cost of teaching video production, and ensures knowledge accuracy and teaching effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499463A_ABST
    Figure CN120499463A_ABST
Patent Text Reader

Abstract

The invention discloses a teaching video generation method and apparatus based on a semantic knowledge base, and an electronic device. The method comprises the steps of constructing a hierarchical semantic knowledge base based on multi-source data; the hierarchical semantic knowledge base comprises courses, chapters under the courses, concepts in the chapters and semantic relationships among the concepts; performing content extraction according to the concepts in the chapters in the hierarchical semantic knowledge base and the semantic relationship between the concepts, and generating structured slide content; based on the structured slide content, using a language large model to generate a matched lecture; generating voice according to the matched lecture, and driving the digital human model by using the voice; and the structured slide content is embedded into the demonstration scene of the digital human model in an animation mode to generate the teaching video, so that the knowledge accuracy and the teaching effect of the generated teaching video content are ensured while the time cost of teaching video production is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of educational technology and intelligent video synthesis, and specifically relates to a method, device and electronic equipment for generating teaching videos based on a semantic knowledge base. Background Art

[0002] With the rapid development of online education and intelligent teaching, instructional videos have become an important means of disseminating knowledge. They transcend time and space, allowing learners to access high-quality teaching resources anytime, anywhere. By integrating multiple formats, including text, images, audio, and video, they enhance learning immersion and effectiveness.

[0003] However, the existing teaching video production process mainly relies on manual recording and post-editing, involving multiple links such as lesson plan design, PPT (PowerPoint, presentation) production, lecture script writing, recording, editing, etc., which has some shortcomings:

[0004] Traditional production methods often struggle to ensure the logical rigor and diverse forms of teaching content: slides are sometimes disconnected from lectures, or diagrams don't match knowledge points well, impacting learning outcomes. Furthermore, current mainstream automation tools often focus on a single process, failing to achieve integrated intelligent generation from "knowledge source" to "finished video," making it difficult to quickly adjust output based on different learning situations and teaching styles. Furthermore, teaching videos involve multimodal content such as text, charts, illustrations, and live explanations. How to uniformly plan the generation of each modality while ensuring knowledge accuracy is a core difficulty that existing technologies have yet to fully address.

[0005] Therefore, how to provide a teaching video generation method that fully expresses knowledge and is multimodal and collaborative has become an important issue. Summary of the Invention

[0006] In order to solve the above problems existing in the prior art, the present invention provides a method, device and electronic equipment for generating teaching videos based on a semantic knowledge base.

[0007] The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0008] In a first aspect, the present invention provides a method for generating a teaching video based on a semantic knowledge base, the method comprising:

[0009] Constructing a hierarchical semantic knowledge base based on multi-source data; the hierarchical semantic knowledge base includes courses, chapters under the courses, concepts in each chapter, and semantic relationships between concepts;

[0010] Extracting content based on the concepts in each chapter and the semantic relationships between the concepts in the hierarchical semantic knowledge base to generate structured slide content; and generating accompanying lecture notes based on the structured slide content using a large language model;

[0011] A voice is generated according to the supporting lecture notes, and the voice is used to drive the digital human model; and the structured slide content is embedded into the demonstration scene of the digital human model through animation to generate a teaching video.

[0012] Optionally, build a hierarchical semantic knowledge base based on multi-source data, including:

[0013] Extracting structured information from multi-source data; the structured information includes text content and voice content;

[0014] Using a BERT-based named entity recognition algorithm to identify concepts in the structured information, and combining it with a relation extraction model to extract semantic relationships between the concepts;

[0015] The concepts and semantic relationships are semantically organized hierarchically to construct a verifiable semantic knowledge graph and obtain a hierarchical semantic knowledge base.

[0016] Optionally, generating speech according to the accompanying lecture notes and using the speech to drive the digital human model includes:

[0017] The accompanying lecture notes are input into the text-to-speech system to generate speech, and the digital human is driven in the Unity platform to perform lip synchronization and gesture matching.

[0018] Optionally, content extraction is performed based on the concepts in each chapter and the semantic relationships between the concepts in the hierarchical semantic knowledge base to generate structured slide content, including:

[0019] Performing template matching based on each chapter in the hierarchical semantic knowledge base to obtain a matched template;

[0020] Filling the matched template with content using concepts in each chapter and semantic relationships between the concepts to obtain initial structured slide content;

[0021] A schematic diagram is generated using a multimodal macro model, and the schematic diagram is inserted into the initial structured slide content to obtain structured slide content.

[0022] Optionally, the multimodal large model includes DALL-E.

[0023] Optionally, generating supporting lecture notes based on the structured slide content using a large language model includes:

[0024] Based on the structured slide content, a draft of a supporting lecture is generated using a large language model;

[0025] The accompanying lecture draft is optimized through a style transfer operation to generate an accompanying lecture.

[0026] In a second aspect, the present invention provides a device for generating a teaching video based on a semantic knowledge base, the device comprising:

[0027] A construction module is used to construct a hierarchical semantic knowledge base based on multi-source data; the hierarchical semantic knowledge base includes courses, chapters under the courses, concepts in each chapter, and semantic relationships between concepts;

[0028] The first generation module is configured to extract content based on the concepts in each chapter and the semantic relationships between the concepts in the hierarchical semantic knowledge base to generate structured slide content; and generate supporting lecture notes based on the structured slide content using a large language model;

[0029] The second generation module is used to generate voice according to the supporting lecture notes, use the voice to drive the digital human model; and embed the structured slide content into the demonstration scene of the digital human model through animation to generate a teaching video.

[0030] In a third aspect, the present invention provides an electronic device comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0031] Memory for storing computer programs;

[0032] The processor is used to implement the method steps described in any of the above-mentioned methods for generating teaching videos based on a semantic knowledge base when executing the computer program stored in the memory.

[0033] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any of the above-mentioned methods for generating teaching videos based on a semantic knowledge base are implemented.

[0034] The present invention provides a method for generating teaching videos based on a semantic knowledge base, which constructs a hierarchical semantic knowledge base based on multi-source data, wherein the hierarchical semantic knowledge base includes courses, chapters under the courses, concepts in each chapter, and semantic relationships between concepts, thereby realizing the systematic organization, semantic association, and structured expression of knowledge. Structured slide content is generated based on the concepts in each chapter and the semantic relationships between concepts in the hierarchical semantic knowledge base, and supporting lecture notes are generated based on the structured slide content using a large language model, thereby ensuring the compatibility, logic, and coherence of the teaching content. Digital human technology is then used to combine supporting lecture notes and structured slides to generate multimodal collaborative teaching videos, which significantly reduces the time cost of teaching video production while ensuring the knowledge accuracy and teaching effect of the generated teaching video content.

[0035] The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 1 is a flow chart of a method for generating a teaching video based on a semantic knowledge base provided by an embodiment of the present invention;

[0037] Figure 2 Schematic diagram of the system structure of the method for generating teaching videos based on a semantic knowledge base provided by an embodiment of the present invention;

[0038] Figure 3 Schematic diagram of the construction architecture of the hierarchical semantic knowledge base provided by an embodiment of the present invention;

[0039] Figure 4 1 is a structural diagram of a device for generating teaching videos based on a semantic knowledge base provided by an embodiment of the present invention;

[0040] Figure 5 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The present invention will be further described in detail below with reference to specific examples, but the embodiments of the present invention are not limited thereto.

[0042] In order to solve the problems of insufficient knowledge expression and difficulty in multimodal collaboration in existing teaching video generation methods, the embodiment of the present invention provides a teaching video generation method based on a semantic knowledge base, see Figure 1 , Figure 1 This is a flow chart of a method for generating a teaching video based on a semantic knowledge base provided by an embodiment of the present invention, which specifically includes the following steps:

[0043] Step S101 : constructing a hierarchical semantic knowledge base based on multi-source data; the hierarchical semantic knowledge base includes courses, chapters under the courses, concepts in each chapter, and semantic relationships between concepts.

[0044] In the embodiment of the present invention, a hierarchical semantic knowledge base is established. Each knowledge unit consists of four parts: course, chapter, concept, and semantic relationship.

[0045] Course section: identifies the course name corresponding to the current knowledge base, and is used to classify the entire knowledge system by subject and discipline.

[0046] Chapter: Divide the course into several teaching chapters. Each chapter corresponds to the natural structure in the textbook and serves as the middle-level unit of knowledge clustering.

[0047] Concept section: Lists the core knowledge concepts of each chapter, usually including definitions, theorems, laws, formulas, experimental phenomena, etc., which reflect the smallest semantic unit of teaching content. Each concept records its keywords, explanatory text and the chapter to which it belongs;

[0048] Relationship part: Describes the semantic relationship between concepts in the form of triples (Concept1, Relation, Concept2), such as "belongs to", "includes", "depends on", "causality", etc.

[0049] In an embodiment of the present invention, through this hierarchical semantic knowledge structure, the teaching content can be expressed in depth layer by layer from the path of courses, chapters, concepts, and relationships, realizing the systematic organization, semantic association and structured expression of knowledge, supporting the automatic generation and semantic consistency verification of teaching content, and providing a clear and operational knowledge foundation for subsequent PPT generation, lecture script writing and video synthesis.

[0050] In an embodiment of the present invention, a hierarchical semantic knowledge base is constructed based on multi-source data, including:

[0051] Extract structured information from multi-source data; structured information includes text content and voice content;

[0052] A named entity recognition algorithm based on BERT (Bidirectional Encoder Representations from Transformers) is used to identify concepts in structured information, and a relation extraction model is used to extract semantic relationships between concepts.

[0053] Concepts and semantic relationships are organized semantically and hierarchically, and a verifiable semantic knowledge graph is constructed for verification to obtain a hierarchical semantic knowledge base.

[0054] In this embodiment of the present invention, the multi-source data can specifically be educational resources such as teaching materials, such as PDFs and Word documents, or instructional videos, including video frames, subtitles, and audio. Natural language processing technology and multimodal parsing tools are used to preprocess the input teaching materials. Text-based teaching materials, such as PDFs and Word documents, are segmented and divided into chapters, extracting structured content such as titles and text. Subtitle text and speech transcription information are extracted from instructional videos to separate usable language segments from visual segments.

[0055] BERT-based named entity recognition algorithms such as NER (Named Entity Recognition) are used to identify core concepts such as important terms, definitions, and rules in the course. Relation extraction models such as RE-BERT are combined to extract the semantic relationships between concepts and represent the logical structure of knowledge in the form of triples.

[0056] The extracted knowledge entities, i.e., concepts and semantic relationships, are semantically organized hierarchically according to the aforementioned "Course-Chapter-Concept-Relationship" to construct a verifiable semantic knowledge graph and obtain a hierarchical semantic knowledge base. The knowledge graph embedding model can also be used for consistency verification and redundancy filtering to ensure the accuracy and logic of the knowledge structure.

[0057] The following describes the specific process of extracting and integrating course knowledge from multi-source data to build a hierarchical semantic knowledge base:

[0058] See also Figure 2 and Figure 3 , Figure 2 1 is a schematic diagram of the system structure of the method for generating teaching videos based on a semantic knowledge base provided by an embodiment of the present invention. Figure 3 This is a schematic diagram of the construction architecture of the hierarchical semantic knowledge base provided by an embodiment of the present invention. First, preprocessing operations are performed on multi-source data to extract structured information, including video key frame extraction and textbook text analysis.

[0059] In an embodiment of the present invention, extracting a video key frame includes:

[0060]

[0061] Among them, R 视频 is the vector representation of the video, T×H×W×C is the length of the vector representation of the video, T is the number of video frames; the video resolution is H×W; C is the number of video channels; F keyRepresents a video key frame; V represents a video; KeyFrameExtractor is a video key frame extraction operation.

[0062] In an embodiment of the present invention, the textbook text parsing includes:

[0063] T struct =TextParser(textbook), T struct ={Chapter, Knowledge Point, Example};

[0064] Among them, T struct Represents the textbook text parsing content; TextParser represents the text parsing operation;

[0065] In an embodiment of the present invention, after obtaining structured information, a named entity recognition (NER) algorithm based on BERT is used to identify concepts in the structured information, and a relation extraction (RE) model is combined to extract semantic relationships between concepts.

[0066] Use named entity recognition algorithms to identify concepts in structured information, including:

[0067] E=NER(T struct ∪F key ),E={e1,e2,...,e n};

[0068] Among them, E represents entity, that is, the smallest semantic unit; n represents the total number of entities; NER represents the named entity recognition algorithm.

[0069] The relation extraction model is used to extract the semantic relationship between concepts, including:

[0070] R=RE(E),R={(e i ,r ij ,e j )};

[0071] Among them, R represents the semantic relationship between concepts; RE represents the relation extraction operation; e i represents the i-th entity; e j represents the jth entity; i = 1, 2, ..., n, j = 1, 2, ..., n, r ij Represents the semantic relationship between the i-th entity and the j-th entity, such as belongs to, contains.

[0072] The extracted triples, i.e., concept 1, semantic relationship, and concept 2, are checked for consistency using a knowledge graph embedding model, such as TransE (a translation distance model for knowledge graph embedding). The constructed knowledge graph is organized into a hierarchical structure, including four levels: courses, chapters, concepts, and relationships. The structure of the hierarchical semantic knowledge base is as follows:

[0073]

[0074]

[0075] In an embodiment of the present invention, the operation of verifying the rationality of semantic relationships through a knowledge graph embedding model, such as TransE, is specifically as follows:

[0076] Score(e i ,r ij ,e j )=||e i +r ij -e j ||2;

[0077] Among them, Score(e i ,r ij ,e j ) represents a triple (e i ,r ij ,e j ) score;

[0078] After getting the triple (e i ,r ij ,e j ), the triples with scores higher than a threshold τ are retained, where the threshold τ can be set by those skilled in the art based on experience.

[0079] Step S102 : extracting content based on the concepts in each chapter and the relationships between the concepts in the hierarchical semantic knowledge base to generate structured slide content; and generating supporting lecture notes based on the structured slide content using a large language model.

[0080] In an embodiment of the present invention, teaching content can be generated based on the hierarchical semantic knowledge base, where the teaching content includes PPT and supporting lecture notes.

[0081] In an embodiment of the present invention, content extraction is performed based on the concepts in each chapter and the semantic relationships between the concepts in the hierarchical semantic knowledge base to generate structured slide content, including:

[0082] Perform template matching based on each chapter in the hierarchical semantic knowledge base to obtain a matched template;

[0083] The matching template is filled with content using the concepts in each chapter and the semantic relationships between the concepts to obtain the initial structured slide content;

[0084] A schematic diagram is generated using the multimodal large model, and the schematic diagram is inserted into the initial structured slide content to obtain the structured slide content.

[0085] In an embodiment of the present invention, content extraction is performed based on the concepts in each chapter and the relationships between the concepts in the hierarchical semantic knowledge base to generate structured slide content, including:

[0086] Based on the concepts in each chapter and the relationships between concepts in the hierarchical semantic knowledge base, key teaching content is extracted, and structured slide content is automatically generated through the PPT template engine. In combination with multimodal large models such as DALL-E, visual materials such as schematic diagrams, formula diagrams or teaching illustrations are generated.

[0087] In an embodiment of the present invention, the specific process of extracting content based on the concepts in each chapter and the relationships between the concepts in the hierarchical semantic knowledge base to generate structured slide content is as follows:

[0088] First, perform template matching:

[0089] Template = TemplateSelector(section type), such as "Definition-Example-Summary";

[0090] Among them, Template represents the template matching result; TemplateSelector represents the template selection operation;

[0091] Fill in the content based on the template matching results:

[0092] PPT content =ContentFiller(Template,hierarchical semantic knowledge base);

[0093] Among them, PPT content Indicates the content filled in PPT; ContentFiller indicates the content filling operation;

[0094] Then, perform visualization optimization and use a large multimodal model such as DALL-E3 to generate a schematic diagram:

[0095] I visual =DALL-E3("Schematic Description Text");

[0096] Among them, I visual The generated diagram is represented by the text in the diagram, which is the content filled in the PPT. DALL-E3 is a model that can generate images based on text descriptions.

[0097] In an embodiment of the present invention, generating supporting lecture notes based on structured slide content using a large language model includes:

[0098] Using PPT content as input, a large language model, such as the GPT (Generative Pre-Trained Transformer) language model, is called to generate supporting lecture notes. The supporting lecture notes can include the order of explanation, concept explanation, teaching terms, etc., to ensure smooth language, clear logic, natural expression, and adapt to various teaching styles.

[0099] In an embodiment of the present invention, generating supporting lecture notes based on structured slide content using a large language model includes:

[0100] Based on the structured slide content, a large language model is used to generate a draft of the accompanying lecture notes;

[0101] The accompanying lecture draft is optimized through a style transfer operation to generate an accompanying lecture.

[0102] Specifically, the generation of supporting lecture notes based on the structured slide content using the language model includes:

[0103] Generate a draft of the accompanying lecture:

[0104] S draft =GPT-4 (“Generate a draft of a lecture based on PPT content”);

[0105] Among them, S draft Indicates the draft of the accompanying lecture; GPT-4 is a type of GPT language model;

[0106] Then, the language of the supporting lecture draft was optimized to obtain the final supporting lecture draft:

[0107] S final =StyleTransfer(S draft , teaching style);

[0108] Among them, S final Indicates the accompanying lecture notes; StyleTransfer indicates the style transfer operation;

[0109] In this embodiment of the present invention, a PPT template engine is used to generate clearly structured slides based on chapters and knowledge points. A multimodal generative model, such as DALL-E, is used to generate teaching diagrams and illustrations based on keywords and automatically insert them into the PPT. A GPT-like language model is used to generate accompanying lecture notes, ensuring that the presentation is natural and fluent, covers all knowledge points, and adapts to the teaching context.

[0110] Step S103: Generate voice according to the supporting lecture content, and use the voice to drive the digital human model; and embed the structured slide content into the demonstration scene of the digital human model through animation to generate a teaching video.

[0111] In an embodiment of the present invention, generating speech based on the accompanying lecture notes and driving the digital human model using the speech includes:

[0112] The accompanying lecture notes are input into the text-to-speech system to generate speech, and the digital human is driven on the Unity platform to perform lip synchronization and posture movement matching.

[0113] Specifically, the accompanying lecture text is input into a TTS (text-to-speech) system for speech synthesis. A digital human model is loaded into the Unity platform, and speech-driven lip movements and facial expressions are synchronized. PPT content is embedded into the digital human presentation scene through animation. Once all content is synchronized on the timeline, the FFmpeg tool is used to render and output the video, supporting automatic subtitle generation and synthesis.

[0114] In an embodiment of the present invention, generating speech according to the accompanying speech script includes:

[0115]

[0116] Among them, A voice Represents the generated speech; R 语音 Represents the vector representation of speech; N×16000 is the length corresponding to the vector representation of speech;

[0117] The motion generation of the digital human model includes:

[0118] M acticon =ActionGenerator(S final );

[0119] Among them, M acticon Indicates the action of the generated digital human model; ActionGenerator represents the action generation operation;

[0120] Embed structured slide content into the digital human model's demonstration scene through animation to generate a teaching video, including:

[0121] V final =VideoRenderer(A voice ,M action ,PPT content );

[0122] Among them, V final Represents a teaching video; VideoRenderer represents a video rendering operation.

[0123] In an embodiment of the present invention, a hierarchical semantic knowledge base is constructed based on multi-source data, wherein the hierarchical semantic knowledge base includes courses, chapters under the courses, concepts in each chapter, and semantic relationships between concepts, thereby realizing systematic organization, semantic association, and structured expression of knowledge. Structured slide content is generated based on the concepts in each chapter and the semantic relationships between concepts in the hierarchical semantic knowledge base, and supporting lecture notes are generated based on the structured slide content using a large language model, thereby ensuring the compatibility, logic, and coherence of the teaching content. Digital human technology is then used to generate multimodal collaborative teaching videos in combination with supporting lecture notes and structured slides, which significantly reduces the time cost of producing teaching videos while ensuring the knowledge accuracy and teaching effect of the generated teaching video content.

[0124] In addition, in the embodiment of the present invention, the BERT-NER and TransE models can be implemented using PyTorch, and the hierarchical semantic knowledge base can be stored as a Neo4j graph database; automatic PPT typesetting can be achieved by calling the Microsoft PowerPoint API; the generation of supporting lecture notes can be achieved by integrating the GPT-4 interface of Hugging Face; the Unity engine is used to drive the digital human model, and FFmpeg is combined to synthesize teaching videos.

[0125] Based on the same inventive concept, the embodiment of the present invention also provides a teaching video generation device based on a semantic knowledge base, see Figure 4 , Figure 4 1 is a structural diagram of a teaching video generation device based on a semantic knowledge base provided by an embodiment of the present invention, the teaching video generation device comprising:

[0126] A construction module 401 is used to construct a hierarchical semantic knowledge base based on multi-source data; the hierarchical semantic knowledge base includes courses, chapters under the courses, concepts in each chapter, and semantic relationships between concepts;

[0127] The first generation module 402 is configured to extract content based on the concepts in each chapter and the semantic relationships between the concepts in the hierarchical semantic knowledge base to generate structured slide content; and generate accompanying lecture notes based on the structured slide content using a large language model;

[0128] The second generating module 403 is configured to generate speech based on the supporting lecture notes, drive the digital human model with the speech, and embed the structured slide content into the demonstration scene of the digital human model through animation to generate a teaching video.

[0129] In an embodiment of the present invention, a hierarchical semantic knowledge base is constructed based on multi-source data, wherein the hierarchical semantic knowledge base includes courses, chapters under the courses, concepts in each chapter, and semantic relationships between concepts, thereby realizing systematic organization, semantic association, and structured expression of knowledge. Structured slide content is generated based on the concepts in each chapter and the semantic relationships between concepts in the hierarchical semantic knowledge base, and supporting lecture notes are generated based on the structured slide content using a large language model, thereby ensuring the compatibility, logic, and coherence of the teaching content. Digital human technology is then used to generate multimodal collaborative teaching videos in combination with supporting lecture notes and structured slides, which significantly reduces the time cost of producing teaching videos while ensuring the knowledge accuracy and teaching effect of the generated teaching video content.

[0130] Optional, building blocks for:

[0131] Extracting structured information from multi-source data; the structured information includes text content and voice content;

[0132] Using a BERT-based named entity recognition algorithm to identify concepts in the structured information, and combining it with a relation extraction model to extract semantic relationships between the concepts;

[0133] The concepts and semantic relationships are semantically organized hierarchically to construct a verifiable semantic knowledge graph and obtain a hierarchical semantic knowledge base.

[0134] Optionally, a second generation module generates speech based on the accompanying lecture notes and drives the digital human model using the speech, including:

[0135] The accompanying lecture notes are input into the text-to-speech system to generate speech, and the digital human is driven in the Unity platform to perform lip synchronization and gesture matching.

[0136] Optionally, the first generation module extracts content based on the concepts in each chapter and the semantic relationships between the concepts in the hierarchical semantic knowledge base to generate structured slide content, including:

[0137] Performing template matching based on each chapter in the hierarchical semantic knowledge base to obtain a matched template;

[0138] Filling the matched template with content using concepts in each chapter and semantic relationships between the concepts to obtain initial structured slide content;

[0139] A schematic diagram is generated using a multimodal macro model, and the schematic diagram is inserted into the initial structured slide content to obtain structured slide content.

[0140] Optionally, the multimodal large model includes DALL-E.

[0141] Optionally, the first generation module generates supporting lecture notes based on the structured slide content using a large language model, including:

[0142] Based on the structured slide content, a draft of a supporting lecture is generated using a large language model;

[0143] The accompanying lecture draft is optimized through a style transfer operation to generate an accompanying lecture.

[0144] The embodiment of the present invention further provides an electronic device, such as Figure 5 As shown, it includes a processor 501, a communication interface 502, a memory 503 and a communication bus 504, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504.

[0145] Memory 503, used for storing computer programs;

[0146] The processor 501 is configured to implement any of the above-mentioned steps of the method for generating a teaching video based on a semantic knowledge base when executing the program stored in the memory 503 .

[0147] The communication bus mentioned in the electronic device mentioned above may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.

[0148] The communication interface is used for communication between the above electronic device and other devices.

[0149] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.

[0150] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.

[0151] The present invention also provides a computer-readable storage medium in which a computer program is stored, and when the computer program is executed by a processor, the method steps of any of the above-mentioned methods for generating teaching videos based on a semantic knowledge base are implemented.

[0152] Optionally, the computer-readable storage medium may be a non-volatile memory (NVM), such as at least one disk memory.

[0153] Optionally, the computer-readable storage medium may also be at least one storage device located away from the processor.

[0154] In another embodiment of the present invention, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute the method steps described in any of the above-mentioned methods for generating teaching videos based on a semantic knowledge base.

[0155] It should be noted that the terms "first," "second," and the like are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in sequences other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention.

[0156] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification.

[0157] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art can understand and implement other variations of the disclosed embodiments by viewing the drawings and the disclosed content. In the description of the present invention, the word "comprising" does not exclude other components or steps, "one" or "a" does not exclude multiple situations, and "multiple" means two or more, unless otherwise clearly and specifically defined. In addition, certain measures are recorded in different embodiments, but this does not mean that these measures cannot be combined to produce good results.

[0158] The method provided in the embodiments of the present invention can be applied to electronic devices. Specifically, the electronic devices can be desktop computers, portable computers, smart mobile terminals, servers, etc. This is not limited here; any electronic device that can implement the present invention falls within the scope of protection of the present invention.

[0159] As for the device / electronic device / storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0160] It should be noted that the device, electronic device and storage medium of the embodiments of the present invention are respectively the device, electronic device and storage medium that apply the above-mentioned method for generating teaching videos based on a semantic knowledge base. All embodiments of the above-mentioned method for generating teaching videos based on a semantic knowledge base are applicable to the device, electronic device and storage medium, and can achieve the same or similar beneficial effects.

[0161] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.

Claims

1. A method for generating teaching videos based on a semantic knowledge base, characterized in that: The teaching video generation method comprises: Constructing a hierarchical semantic knowledge base based on multi-source data; the hierarchical semantic knowledge base includes courses, chapters under the courses, concepts in each chapter, and semantic relationships between concepts; Extracting content based on the concepts in each chapter and the semantic relationships between the concepts in the hierarchical semantic knowledge base to generate structured slide content; and generating accompanying lecture notes based on the structured slide content using a large language model; A voice is generated according to the supporting lecture notes, and the voice is used to drive the digital human model; and the structured slide content is embedded into the demonstration scene of the digital human model through animation to generate a teaching video.

2. The teaching video generation method according to claim 1, characterized in that: Build a hierarchical semantic knowledge base based on multi-source data, including: Extracting structured information from multi-source data; the structured information includes text content and voice content; Using a BERT-based named entity recognition algorithm to identify concepts in the structured information, and combining it with a relation extraction model to extract semantic relationships between the concepts; The concepts and semantic relationships are semantically organized hierarchically to construct a verifiable semantic knowledge graph and obtain a hierarchical semantic knowledge base.

3. The teaching video generation method according to claim 1, characterized in that: Generating speech according to the accompanying lecture notes and driving the digital human model using the speech, including: The accompanying lecture notes are input into the text-to-speech system to generate speech, and the digital human is driven in the Unity platform to perform lip synchronization and gesture matching.

4. The teaching video generation method according to claim 1, characterized in that: Content is extracted based on the concepts in each chapter and the semantic relationships between the concepts in the hierarchical semantic knowledge base to generate structured slide content, including: Performing template matching based on each chapter in the hierarchical semantic knowledge base to obtain a matched template; Filling the matched template with content using concepts in each chapter and semantic relationships between the concepts to obtain initial structured slide content; A schematic diagram is generated using a multimodal macro model, and the schematic diagram is inserted into the initial structured slide content to obtain structured slide content.

5. The teaching video generation method according to claim 4, characterized in that: The multimodal large model includes DALL-E.

6. The teaching video generation method according to claim 1, characterized in that: Generate supporting lecture notes based on the structured slide content using a large language model, including: Based on the structured slide content, a draft of a supporting lecture is generated using a large language model; The accompanying lecture draft is optimized through a style transfer operation to generate an accompanying lecture.

7. A device for generating teaching videos based on a semantic knowledge base, characterized in that: The teaching video generating device comprises: A construction module is used to construct a hierarchical semantic knowledge base based on multi-source data; the hierarchical semantic knowledge base includes courses, chapters under the courses, concepts in each chapter, and semantic relationships between concepts; The first generation module is configured to extract content based on the concepts in each chapter and the semantic relationships between the concepts in the hierarchical semantic knowledge base to generate structured slide content; and generate supporting lecture notes based on the structured slide content using a large language model; The second generation module is used to generate voice according to the supporting lecture notes, use the voice to drive the digital human model; and embed the structured slide content into the demonstration scene of the digital human model through animation to generate a teaching video.

8. The teaching video generating device according to claim 7, characterized in that: The building blocks are specifically used for: Extract structured information from multi-source data; the structured information includes text content and voice content; use a BERT-based named entity recognition algorithm to identify concepts in the structured information, and combine it with a relationship extraction model to extract semantic relationships between the concepts; organize the concepts and semantic relationships in a semantic hierarchical manner, construct a verifiable semantic knowledge graph, and obtain a hierarchical semantic knowledge base.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; The processor is used to implement the method for generating a teaching video based on a semantic knowledge base as described in any one of claims 1 to 6 when executing a computer program stored in the memory.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method for generating a teaching video based on a semantic knowledge base according to any one of claims 1 to 6 is implemented.