Intelligent teaching video generation system fused with multi-modal large model and method thereof
Through the intelligent teaching video generation system integrating multimodal large models, the existing teaching video generation efficiency is solved and the problems of inappropriate multimodal content splicing are realized, and realistic and personalized teaching videos are automatically generated, which improves teaching quality and learning experience.
Patent Information
- Application Number
- CN202510817839.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-02
Smart Images

Figure CN120583293A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video generation, and in particular to an intelligent teaching video generation system and method integrating a multimodal large model. Background Art
[0002] In virtual classrooms and online education, the production of instructional videos has long relied on manual labor. Teachers spend a significant amount of time writing instructional content, recording audio lectures, and manually integrating text, subtitles, and audio using specialized software. This method of generating instructional videos is inefficient, with the production cycle for a single course video typically taking several days and high labor costs. While some automated tools can assist in generating single-modal content, the resulting content is rigid in form, lacking semantic connections between elements such as text, audio, and images, and requiring manual secondary adjustments. This makes end-to-end automated production difficult, and has become a core bottleneck restricting the large-scale output of educational content.
[0003] With the expansion of artificial intelligence applications, technologies such as natural language generation and speech synthesis have gradually penetrated into the production of educational content. For example, a lesson plan generation system based on deep learning can quickly output structured teaching text, and a speech synthesis engine can convert text into anthropomorphic explanation audio. However, existing technologies mostly focus on the independent generation of single-modal content. The coordination of content from different modalities still relies on manually preset rules or fixed templates for mechanical splicing. This leads to problems such as misalignment between speech rhythm and animation rhythm, and disconnection between graphic presentation and explanation logic. This seriously weakens the coherence and interactive experience of teaching videos, reducing teaching effectiveness. Summary of the Invention
[0004] In order to solve the technical problem that existing teaching videos, due to the mechanical splicing of multimodal content, lead to disconnected audio and video, chaotic logic, and seriously affect the teaching coherence and interactive experience, the purpose of the present invention is to provide an intelligent teaching video generation system and method that integrates a multimodal large model. The technical solutions adopted are as follows:
[0005] One embodiment of the present invention provides an intelligent teaching video generation system integrating a multimodal large model, the system comprising:
[0006] The input module is used to receive the course name input by the user and send it to the teaching content generation module;
[0007] A teaching content generation module is used to call the target DeepSeek model to generate teaching content based on the received course name and send it to the speech synthesis module. The teaching content includes knowledge point explanation, example analysis, exercises and knowledge summary;
[0008] The speech synthesis module is used to call the CosyVoice model to convert the received teaching content into speech content consistent with the timbre of the audio uploaded by the user, and send it to the animation generation module and the video synthesis module;
[0009] An animation generation module is used to generate animation content using an echomimic model based on the received voice content and the character image photo submitted by the user. The animation content includes a virtual teacher image and a dynamic explanation animation, and sends it to the video synthesis module. The dynamic explanation animation includes lip synchronization, expression changes, and hand gestures.
[0010] The video synthesis module is used to fuse the received voice content and the received animation content to generate a teaching video.
[0011] Furthermore, the target DeepSeek model adopts a hybrid expert architecture, a multi-head potential attention mechanism, a load balancing strategy without auxiliary loss, and multi-word prediction when generating the teaching content, and supports FP8 mixed precision training.
[0012] Furthermore, the hybrid expert architecture decomposes the target DeepSeek model into multiple expert networks and dynamically selects the optimal expert for calculation on each input.
[0013] Furthermore, the CosyVoice model includes:
[0014] A speech tagging unit, configured to generate speech tags corresponding to the input teaching content using an autoregressive transformer;
[0015] A speech coding unit, configured to convert a continuous speech signal into a discrete coding representation using a speech quantization coding technique;
[0016] an acoustic adjustment unit, configured to adjust and optimize the generated acoustic features using a flow matching technique;
[0017] A spectrum reconstruction unit, configured to reconstruct a Mel spectrum from the generated speech markers using an ODE-based diffusion model;
[0018] The speech synthesis unit is used to synthesize the final speech waveform from the Mel spectrum using a HiFTNet-based vocoder.
[0019] Furthermore, the speech synthesis module further includes:
[0020] A speech emotion generation unit, used to adjust the emotion and rhythm of the generated speech according to the text content or user instructions;
[0021] A voice timbre customization unit, used to imitate the voice characteristics of a user-specified speaker through fine-tuning technology;
[0022] The speech mixing unit is used to perform sound feature interpolation processing between two or more people.
[0023] Furthermore, the Echomimic model adopts APDH strategy, audio diffusion technology, HPA head local attention and PhD Loss stage denoising loss when generating the animation content, and supports multilingual applications.
[0024] Furthermore, the system further comprises:
[0025] User interaction module, used to support the management of historical generation teaching videos;
[0026] A storage module is used to store generated teaching content, voice content, animation content and teaching videos;
[0027] The output module is used to output the generated teaching video to the user terminal.
[0028] Another embodiment of the present invention provides a method for generating intelligent teaching videos integrating a multimodal large model, the method comprising:
[0029] Get the course name entered by the user;
[0030] Invoke the target DeepSeek model according to the course title to generate teaching content, wherein the teaching content includes explanation of knowledge points, analysis of example questions, exercises, and knowledge summary;
[0031] Call the CosyVoice model to convert the teaching content into voice content consistent with the timbre of the audio uploaded by the user;
[0032] Generate animation content using an Echomimic model based on the voice content and a user-submitted character image photo. The animation content includes a virtual teacher image and a dynamic explanation animation. The dynamic explanation animation includes lip synchronization, facial expression changes, and hand gestures.
[0033] The voice content and the animation content are integrated to obtain a teaching video, which is then pushed to a user terminal.
[0034] The present invention has the following beneficial effects:
[0035] The present invention realizes the full-scale automatic generation of teaching videos. From inputting the course name to outputting the complete teaching video, teachers do not need to manually write teaching content, record voice and create animations, which greatly saves teachers' time and energy, allowing teachers to have more time for teaching research, student guidance and curriculum innovation, and is conducive to optimizing the teaching process and improving the overall teaching quality.
[0036] The present invention supports the customization of personalized teaching videos. In terms of voice, the teaching style can be selected according to the teaching content, and the emotion, intonation, and speaking speed of the voice can be adjusted. For example, when explaining key knowledge, an emphatic tone can be used, and when introducing cases, a lively tone can be adopted to enhance the appeal and appeal of teaching. In terms of the virtual teacher image, users can upload different photos or choose preset images, and can also customize the amplitude of movements, the intensity of expressions, etc. For example, a lively teacher image can be paired with a larger amplitude of movement, while a calm teacher image can be more restrained in movements, which is convenient for meeting the needs of different teaching styles and scenarios. Personalized teaching video customization creates a learning environment for students that is more in line with their own learning habits, improves the learning experience and participation, promotes the development of personalized learning, and helps students better absorb knowledge and improve learning outcomes.
[0037] This invention utilizes advanced models such as DeepSeek, CosyVoice, and Echomimic to comprehensively guarantee the quality of teaching videos. The DeepSeek model, based on its unique architecture and powerful learning capabilities, generates accurate and comprehensive teaching content, covering key knowledge points, example analysis, exercises, and summaries, with clear logic and focused emphasis. The CosyVoice model synthesizes natural and fluent speech, highly recreating the rhythm and timbre of human speech, effectively avoiding a mechanical feel and providing students with a more comfortable auditory experience and enhanced concentration during the learning process. The Echomimic model generates realistic and vivid animations, with the virtual teacher's mouth shape, expressions, and gestures synchronized with the speech, enhancing the intuitiveness and fun of teaching, helping students better understand and master knowledge, and improving teaching effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0039] Figure 1 A structural diagram of an intelligent teaching video generation system integrating a multimodal large model provided by one embodiment of the present invention;
[0040] Figure 2 A flowchart of a method for generating intelligent teaching videos integrating a multimodal large model provided in another embodiment of the present invention;
[0041] Figure 3 This is a flowchart of the CosyVoice technology principle;
[0042] Figure 4This is an overall flow chart of the intelligent teaching assistance system in an embodiment of the present invention. DETAILED DESCRIPTION
[0043] To further illustrate the technical means and effects employed by the present invention to achieve its intended objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementations, structures, features, and effects of the technical solutions proposed by the present invention. In the following description, references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0044] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0045] The application scenarios targeted by the present invention may be:
[0046] In recent years, breakthroughs in multimodal large models and generative AI (Artificial Intelligence) technology have provided a new technical path for cross-modal content collaboration. To overcome the technical problems of existing teaching videos, which suffer from the disconnection between audio and video and logical confusion caused by the mechanical splicing of multimodal content, seriously affecting the coherence of teaching and the interactive experience, this paper proposes a cross-modal alignment algorithm to achieve semantic-level association and dynamic matching of text, speech, image and other elements. Specifically, it uses a large text language model, a speech synthesis model, and a video rendering model to achieve a closed-loop text-speech-video generation process.
[0047] Furthermore, the present invention has broad application prospects. In the field of online education, rich teaching resources can attract more students, enhance the competitiveness of the platform, and promote the development of the online education industry. In the field of corporate training, it can quickly generate standardized, high-quality training courses, reduce training costs, and improve training efficiency and effectiveness. In the field of content creation, it can provide efficient solutions for the production of educational short videos and audiobooks, enriching content formats and meeting the diverse needs of the market.
[0048] One embodiment of the present invention provides an intelligent teaching video generation system integrating a multi-modal large model, the structure of the system is shown in FIG. Figure 1 As shown, including:
[0049] The input module 101 is used to receive the course name input by the user and send it to the teaching content generation module.
[0050] In this embodiment, the input module serves as the data source entry for the intelligent teaching video generation system, providing a simple and easy-to-use system operation interface. Users can enter the course title, upload reference audio and images, and select teaching video generation parameters such as voice style, virtual teacher image, and video resolution. Course titles can be as detailed as "High School Mathematics - Fundamentals of Functions" or "Junior High School Physics - Buoyancy."
[0051] The input module signal is connected to the teaching content generation module, and the course name received by the input module is sent to the teaching content generation module.
[0052] The teaching content generation module 102 is used to call the target DeepSeek model to generate teaching content according to the received course name and send it to the speech synthesis module.
[0053] Here, the teaching content is to complete a structured lesson plan, which includes explanation of knowledge points, analysis of example questions, exercises and knowledge summary. The teaching content generation module supports docking with external teaching resource libraries, integrating text, charts, and multimedia materials to provide rich materials for subsequent teaching activities.
[0054] In the teaching content generation module, the target DeepSeek model is called with the help of the API (Application Programming Interface) key to automatically generate well-structured teaching content based on the course name and teaching objectives.
[0055] The target DeepSeek model adopts a hybrid expert architecture, a multi-head potential attention mechanism, a load balancing strategy without auxiliary loss, and multi-word prediction when generating teaching content, and supports FP8 mixed precision training.
[0056] First, a Mixture of Experts (MoE) architecture is adopted, which helps reduce computing resource consumption by decomposing the model into multiple expert networks and dynamically selecting the most appropriate expert for calculation on each input.
[0057] Secondly, a multi-head potential attention mechanism is adopted to reduce the key-value cache requirements in the reasoning process through low-rank joint compression, thereby improving the reasoning efficiency.
[0058] Then, a load balancing strategy without auxiliary loss is adopted to solve the expert load imbalance problem by dynamically adjusting the routing bias.
[0059] Next, a multi-token prediction (MTP) training objective is adopted, allowing the model to predict multiple tokens in a single forward propagation, in order to improve the speed and coherence of text generation, and enhance training efficiency and model performance.
[0060] Finally, FP8 mixed precision training is supported to reduce the GPU memory requirements and storage bandwidth pressure during training.
[0061] It should be noted that the teaching content in this embodiment is stored in JSON or Markdown format for subsequent modules to read. In addition, before generating teaching content, the teaching generation module also supports custom parameter settings, such as teaching level, teaching style, and language style, to enhance personalized output capabilities. Among them, the teaching level can be high school or junior high school, the teaching style can be basic or advanced, and the language style can be formal or colloquial.
[0062] The speech synthesis module 103 is used to call the CosyVoice model to convert the received teaching content into speech content consistent with the timbre of the audio uploaded by the user, and send it to the animation generation module and the video synthesis module.
[0063] Here, the speech synthesis module is based on the speech synthesis model of CosyVoice, which can convert teaching content into natural speech. The flowchart of CosyVoice technology principle is as follows: Figure 3 shown.
[0064] In this embodiment, the CosyVoice model includes:
[0065] The speech tagging unit 1031 is used to generate speech tags corresponding to the input teaching content using an autoregressive transformer, generating speech with realistic prosody and timbre. The speech tagging process typically involves a large number of parameters and a complex network structure to learn and simulate the complexity of human speech.
[0066] The speech coding unit 1032 uses speech quantization coding techniques to convert the continuous speech signal into a discrete coded representation. During speech coding, the original analog speech signal is converted into a series of digital values that represent the corresponding speech. The discrete code can capture key speech characteristics such as volume, pitch, and timbre.
[0067] It is worth mentioning that quantization coding is the first step in speech signal processing, which can provide a basis for subsequent processing and analysis.
[0068] Acoustic adjustment unit 1033 is used to adjust and optimize the generated acoustic features using stream matching technology. Stream matching is an optimization technology used in speech synthesis. It is typically used to adjust and optimize generated acoustic features, such as Mel spectra, to ensure that they are more natural, smooth, and more accurately reflect the characteristics of the original speech signal.
[0069] It should be noted that acoustic adjustment, as a step in the Mel spectrum reconstruction process, ensures that the generated Mel spectrum is closer to natural speech in terms of dynamic range and frequency domain characteristics, providing high-quality input for the final speech synthesis.
[0070] The spectrum reconstruction unit 1034 is configured to reconstruct a Mel spectrum from the generated speech markers using an ODE-based diffusion model. The Mel spectrum is a frequency domain representation of a speech signal that maps the frequency content of the sound to the Mel scale of human auditory perception.
[0071] The speech synthesis unit 1035 uses a HiFTNet-based vocoder to synthesize the final speech waveform from the mel spectrum. It uses the Transformer architecture from deep learning to generate highly realistic speech. In CosyVoice, the vocoder receives the mel spectrum from the previous stage as input and then generates a continuous speech waveform.
[0072] At this point, the speech synthesis in the teaching video can be completed. Of course, in order to improve the personalized features of the teaching video, the speech synthesis module also includes:
[0073] The speech emotion generation unit 1036 is used to adjust the emotion and rhythm of the generated speech according to the text content or user instructions.
[0074] In this embodiment, CosyVoice supports multi-language speech generation and, through fine-grained control technology, can adjust the emotion and rhythm of the generated speech according to the text content or user instructions, including but not limited to the simulation of emotional states such as happiness, sadness, and anger.
[0075] Of course, it also includes fine-grained control of emotion and rhythm. The CosyVoice model can adjust the emotional color and rhythmic characteristics of the speech output according to the emotional tags or instructions in the text. The fine-grained control capability of specific emotions and rhythm is achieved by introducing emotional annotation data during the model training process.
[0076] The voice timbre customization unit 1037 is used to imitate the voice characteristics of the user-specified speaker through fine-tuning technology.
[0077] The voice mixing unit 1038 is used to perform sound feature interpolation processing between two or more people to create an intermediate sound effect.
[0078] As an exemplary embodiment, the execution steps of the speech generation module include:
[0079] First, obtain the text of the teaching content and the optional reference voice, that is, the teacher's audio sample;
[0080] Extract acoustic features from teacher audio samples to obtain personalized timbre vectors, which include pitch, rhythm, and prosody.
[0081] The text of the teaching content is encoded into a sequence of speech tags and passed into the speech generation network for modeling;
[0082] Reconstruct high-quality Mel spectrum using a diffusion model based on ODE (Ordinary Differential Equation);
[0083] The Mel spectrum is converted into a complete speech waveform through the HiFTNet vocoder (High-Fidelity Transformer Network);
[0084] In the process of generating voice content, multi-language pronunciation (Chinese, English) and emotion control (happy, calm, emphasized, contemplative, etc.) are supported and can be specified through interface parameters.
[0085] It should be noted that the output voice file format is WAV (Waveform Audio File Format) or MP3 (Moving Picture Experts Group Audio Layer III), and the sound quality supports a 44.1kHz sampling rate. It has the advantages of high naturalness, no distortion, and low latency, making it suitable for teaching videos. In addition, the speech synthesis system also has a speech quality assessment and optimization function, which can monitor the clarity and naturalness of the synthesized speech in real time and dynamically adjust the synthesis parameters.
[0086] The animation generation module 104 is configured to generate animation content using an Echomimic model based on the received voice content and the character image photo submitted by the user.
[0087] Here, the Echomimic model belongs to a dynamic generation framework for multimodal collaboration, which can generate dynamic teaching videos from static images. To meet diverse needs, the animation generation module also provides editable functions for parameters such as movement amplitude, expression intensity, and background blur.
[0088] The Echomimic model uses the APDH (Audio-Pose Dynamic Harmonization) strategy, audio diffusion technology, HPA (Head Partial Attention) and PhD Loss (Phase-specific Denoising Loss) phase denoising loss when generating animation content, and supports multilingual applications.
[0089] Pose sampling in the APDH strategy: By gradually reducing the reliance on gesture videos, the model focuses more on the natural integration of audio and body movements. For example, gesture videos are mandatory in the early stages of training, and then gradually weakened, ultimately achieving full-body movements driven solely by audio.
[0090] Audio Diffusion: Expands the audio signal from lip sync to the entire body, for example, adjusting the head tilt angle and arm swing amplitude according to the voice emotion.
[0091] HPA Head Partial Attention: Addressing the scarcity of half-body data, this algorithm integrates close-up head data, such as facial key points, to enhance the subtlety of facial expressions. For example, when generating a smile, it not only synchronizes lip movements but also simulates micro-expressions such as wrinkles at the corners of the eyes.
[0092] Staged denoising loss: The generation process is divided into three stages: the pose-dominant stage, which quickly generates human silhouettes and basic movements such as standing or waving; the detail-dominant stage, which refines local features such as clothing textures and finger joints; and the quality-dominant stage, which optimizes lighting effects and color consistency.
[0093] It's worth noting that the animation generation achieved through the Echomimic model represents a technological breakthrough. End-to-end generation eliminates the need for intermediate skeletal binding or keyframe setup, going directly from input to video output to reduce manual intervention. Multi-language support is available, adapting to both Chinese and English speech, for example, automatically adjusting lip movement and gestures for Chinese pronunciation. Furthermore, the Audio Driven acceleration model achieves a 10x speedup on V100 GPUs, reducing the time required to generate 240 frames from 7 minutes to 50 seconds.
[0094] In the Echomimic model, reference images support any person's photo, such as ID photos or daily photos, and the model automatically extracts facial features and body proportions; audio clips support WAV / MP3 formats, up to 20 seconds long, to drive lip synchronization and emotional expression; gesture videos can provide hand movement references, such as gestures during explanations and speeches; the output content can have a frame rate of 24-30fps, including full-body movements such as head rotation, shoulder movement, and gestures.
[0095] For example, the comparison results of typical effects of Echomimic and traditional painting tools are shown in Table 1:
[0096] Table 1
[0097]
[0098] As shown in Table 1, EchomimicV2 achieves a 98.7% accuracy rate for Chinese lip synchronization, far surpassing the cumbersome process of traditional tools that rely on manual keyframe adjustments. EchomimicV2 uses AI to directly generate movements that are close to the naturalness of real people, while traditional tools require manual splicing from a preset movement library, lacking flexibility and realism. EchomimicV2 produces a 10-second video in just one minute, making it over a hundred times more efficient than traditional tools. EchomimicV2 uses end-to-end AI generation, avoiding the problems of manual parameter adjustment and movement library limitations in traditional processes. EchomimicV2 naturally supports multiple languages and multiple scenarios, while traditional tools require repeated configuration for different needs. Therefore, EchomimicV2 comprehensively surpasses traditional animation tools in terms of accuracy, efficiency, and cost, representing a technological leap from "handmade" to "intelligent generation." It is particularly suitable for fields such as education and short videos that require fast, high-quality output.
[0099] As an exemplary embodiment, the implementation process of generating animation content includes:
[0100] User-uploaded images of people, such as real-life photos or cartoon-style images;
[0101] The system extracts facial key points, body proportions and other features of the task image as a skeleton reference;
[0102] The system calls the generated voice audio and inputs it together with the image into the Echomimic generation model;
[0103] The Echomimic generative model optionally supports gesture videos as input, enabling further customization of hand movements during explanations.
[0104] The output is full-body dynamic video frames, with resolutions supporting 720P or 1080P and frame rates adjustable from 24fps to 30fps.
[0105] The Echomimic generative model boasts high inference efficiency, accurate lip synchronization, and natural movements. It generates videos approximately 1 minute / 10 seconds long and is compatible with conventional consumer-grade GPUs. Furthermore, voice-driven animation generation is achieved through the following key technologies:
[0106] Audio diffusion strategy, which synchronously maps speech features such as rhythm and emotion to movements such as mouth shape, head movement, and gestures;
[0107] Posture sampling and dynamic fusion to simulate teaching gestures, such as pointing, drawing, and explaining movements;
[0108] The local attention mechanism of the head refines facial expression details, including eyebrow raising, blinking, and changes in the corners of the mouth;
[0109] A staged denoising generation strategy is used, with the generation process divided into three stages: posture, detail, and lighting effects, to enhance the naturalness of movements.
[0110] The application scenarios of the animation generation module may include:
[0111] Virtual anchors and live broadcasts can support real-time interaction and multi-language switching. Real-time interaction refers to the use of voice recognition APIs to enable audience questions and real-time responses from virtual anchors, such as product introductions in e-commerce live broadcasts. Multi-language switching supports seamless switching between Chinese and English, such as cross-border e-commerce platforms that can quickly generate product promotion videos in different languages.
[0112] Education and training, which can support course production and virtual experiments. Course production means that teachers only need to record audio to generate explanation animations, reducing course production costs. Virtual experiments refer to the generation of virtual anatomy animations in medical education, matching voice explanations of human anatomy.
[0113] Film, television, and games, it can support low-cost animation and film and television special effects. Low-cost animation means that independent game developers can quickly generate NPC (Non-Player Character) dialogue animations, replacing the traditional motion capture process; film and television special effects means generating background characters to reduce the cost of extras.
[0114] Enterprise services can support intelligent customer service and product demonstrations. Intelligent customer service refers to the deployment of virtual customer service representatives in industries like banking and insurance to answer common questions through voice interaction. Product demonstrations involve the generation of animated tutorials for product usage, such as car manufacturers using virtual characters to explain functional operations.
[0115] The video synthesis module 105 is used to generate a teaching video by fusing the received voice content and animation content.
[0116] In this embodiment, the video synthesis module is used to integrate voice, animation and teaching auxiliary materials into the final teaching video. The output video supports export in multiple formats, which is convenient for playback on different platforms. It also supports post-editing operations such as video editing and splicing to meet the user's personalized needs.
[0117] As an exemplary embodiment, the video synthesis module performs the following steps:
[0118] Align the speech generated by CosyVoice with the action frame sequence generated by Echomimic at the frame level. The action frame sequence corresponds to the animation content.
[0119] The system supports superimposing static materials such as teaching PPT, blackboard animation, chart demonstration, etc. on the screen to enhance the effect of information transmission;
[0120] Users can set the background, virtual teacher size and position, subtitle style, etc. in the editing panel;
[0121] Use a multi-threaded video encoder to combine all elements into a standard format video, and support local download or automatic upload to educational platforms, content creation platforms, etc.
[0122] Video output supports 1080P30fps teaching videos. The characters in the video have natural expressions, the voice and mouth shape are highly synchronized, the gestures are coordinated with the teaching rhythm, and the overall presentation is highly professional and friendly.
[0123] An intelligent teaching video generation system integrating a multimodal large model, further comprising:
[0124] User interaction module, used to support the management of historically generated teaching videos, including search, play, download, delete and other operations;
[0125] A storage module is used to store generated teaching content, voice content, animation content and teaching videos;
[0126] The output module is used to output the generated teaching video to the user terminal.
[0127] Regarding the system architecture and deployment environment: the backend is built based on Python and FastAPI, and the front end uses Vue+ElementUI, with a user-friendly overall interface; model reasoning is based on the CUDA GPU environment, and the recommended hardware is: NVIDIA RTX4090 (24GB video memory), 64GB RAM, and 1TB SSD; the model can be deployed locally or called in the cloud, and RESTful API (Representative State Transfer) is provided to support system integration; the system supports user identity management, video history management, and video personalized template management. Among them, the overall flow chart of the intelligent teaching assistance system is as follows Figure 4 shown.
[0128] Another embodiment of the present invention provides a method for generating intelligent teaching videos by integrating a multimodal large model. The method is shown in the flowchart of the steps. Figure 2 Shown, including:
[0129] S201, obtaining the course name input by the user.
[0130] S202, calling the target DeepSeek model to generate teaching content according to the course title, the teaching content includes knowledge point explanation, example analysis, exercises and knowledge summary.
[0131] S203, calling the CosyVoice model to convert the teaching content into voice content that is consistent with the timbre of the audio uploaded by the user.
[0132] S204, based on the voice content and the character image photo submitted by the user, an Echomimic model is used to generate animation content, the animation content includes a virtual teacher image and a dynamic explanation animation, and the dynamic explanation animation includes lip synchronization, expression changes, and hand gestures.
[0133] S205: Fusing the voice content and the animation content to obtain a teaching video, and pushing it to the user terminal.
[0134] It should be noted that the implementation process of each step of an intelligent teaching video generation method that integrates a multimodal large model can refer to the content of the above-mentioned intelligent teaching video generation system that integrates a multimodal large model, and there is no need to repeat the same content.
[0135] In summary, the present invention can integrate DeepSeek, CosyVoice and Echomimic models, realize the automation of the entire process from teaching content generation to video output, and significantly reduce the time and energy invested by teachers in making teaching videos. The system supports voice emotion control, virtual teacher image customization and movement adjustment, and can generate personalized teaching videos according to different teaching scenarios and user needs. By utilizing DeepSeek's intelligent content generation capabilities, CosyVoice's high-quality speech synthesis technology and Echomimic's realistic animation generation technology, the professionalism and vividness of the teaching videos are ensured, and the teaching effect is effectively improved. In addition, the present invention also has good scalability, and can easily access new models and technologies to continuously improve system performance and functions.
[0136] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. An intelligent teaching video generation system integrating multimodal large models, characterized by: The system comprises: The input module is used to receive the course name input by the user and send it to the teaching content generation module; A teaching content generation module is used to call the target DeepSeek model to generate teaching content based on the received course title and send it to the speech synthesis module. The teaching content includes knowledge point explanation, example analysis, exercises and knowledge summary; The speech synthesis module is used to call the CosyVoice model to convert the received teaching content into speech content consistent with the timbre of the audio uploaded by the user, and send it to the animation generation module and the video synthesis module; An animation generation module is used to generate animation content using an echomimic model based on the received voice content and the character image photo submitted by the user. The animation content includes a virtual teacher image and a dynamic explanation animation, and sends it to the video synthesis module. The dynamic explanation animation includes lip synchronization, expression changes, and hand gestures. The video synthesis module is used to fuse the received voice content and the received animation content to generate a teaching video.
2. The intelligent teaching video generation system integrating multimodal large models according to claim 1 is characterized in that: The target DeepSeek model adopts a hybrid expert architecture, a multi-head potential attention mechanism, a load balancing strategy without auxiliary loss, and multi-word prediction when generating the teaching content, and supports FP8 mixed precision training.
3. The intelligent teaching video generation system integrating multimodal large models according to claim 2 is characterized in that: The hybrid expert architecture decomposes the target DeepSeek model into multiple expert networks and dynamically selects the optimal expert for calculation on each input.
4. The intelligent teaching video generation system integrating multimodal large models according to claim 1 is characterized in that: The CosyVoice model includes: A speech tagging unit, configured to generate speech tags corresponding to the input teaching content using an autoregressive transformer; A speech coding unit, configured to convert a continuous speech signal into a discrete coding representation using a speech quantization coding technique; an acoustic adjustment unit, configured to adjust and optimize the generated acoustic features using a flow matching technique; A spectrum reconstruction unit, configured to reconstruct a Mel spectrum from the generated speech markers using an ODE-based diffusion model; The speech synthesis unit is used to synthesize the final speech waveform from the Mel spectrum using a HiFTNet-based vocoder.
5. The intelligent teaching video generation system integrating multimodal large models according to claim 4 is characterized in that: The speech synthesis module also includes: A speech emotion generation unit, used to adjust the emotion and rhythm of the generated speech according to the text content or user instructions; A voice timbre customization unit, used to imitate the voice characteristics of a user-specified speaker through fine-tuning technology; The speech mixing unit is used to perform sound feature interpolation processing between two or more people.
6. The intelligent teaching video generation system integrating multimodal large models according to claim 1 is characterized in that: The Echomimic model adopts the APDH strategy, audio diffusion technology, HPA head local attention and PhD Loss stage denoising loss when generating the animation content, and supports multilingual applications.
7. The intelligent teaching video generation system integrating multimodal large models according to claim 1 is characterized in that: The system further comprises: User interaction module, used to support the management of historical generation teaching videos; A storage module is used to store generated teaching content, voice content, animation content and teaching videos; The output module is used to output the generated teaching video to the user terminal.
8. A method for generating intelligent teaching videos integrating multimodal large models, characterized in that: The method comprises: Get the course name entered by the user; Invoke the target DeepSeek model according to the course title to generate teaching content, wherein the teaching content includes explanation of knowledge points, analysis of example questions, exercises, and knowledge summary; Call the CosyVoice model to convert the teaching content into voice content consistent with the timbre of the audio uploaded by the user; Generate animation content using an Echomimic model based on the voice content and a user-submitted character image photo. The animation content includes a virtual teacher image and a dynamic explanation animation. The dynamic explanation animation includes lip synchronization, facial expression changes, and hand gestures. The voice content and the animation content are integrated to obtain a teaching video, which is then pushed to a user terminal.
Citation Information
Cited By
Teaching video generation method and device, equipment, storage medium and program product
CN120956989A