Method and terminal for automatic generation of a lecture video

CN122595992APending Publication Date: 2026-08-18FUJIAN TIANQUAN EDUCATION TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510170871.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

数字人的肢体动作和实际讲解的内容及内容的位置无关,且场景的切换大多遵循已有的模版套路,与内容无法有效结合

Benefits of technology

[0006]The beneficial effects of this invention are as follows: The method and terminal for automatically generating lecture videos of this invention, combined with the Big Prophecy model and RAG technology, intelligently adjusts the PPT content to make the explanation more in line with semantic logic; achieves dynamic animation matching, and through the precise matching of animation parameter files with PPT content and voice, achieves perfect synchronization of actions, scenes and explanation content, significantly improving the digital human presentation effect; and effectively improves teaching quality and interactivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122595992A_ABST
    Figure CN122595992A_ABST
Patent Text Reader

Abstract

The present application relates to the field of teaching video generation, in particular to a kind of teaching video automatic generation method and terminal, and the content extraction is carried out to target PPT file, and content optimization is carried out in combination with large language model technology and RAG technology;Utilize third party large language model to learn digital person animation module parameter rule, generate first edition animation parameter file and new speech script based on the optimized target PPT file;Transcribe voice file according to speech script using TTS technology, synchronously generate timestamp and subtitle file, and integrate first edition animation parameter file and speech script to generate second edition animation parameter file;Build the mapping relationship of speech script and the page in target PPT file, generate image frame, voice file generates mouth shape animation file, and render to obtain first animation file;Final animation rendering is completed by first animation file in combination with second edition animation parameter file, and teaching video is output;It can effectively improve teaching quality and interactivity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of lecture video generation, and in particular to a method and terminal for automatically generating lecture videos. Background Technology

[0002] With the rapid development of information technology, the technology of automatically generating videos using digital human technology combined with PowerPoint presentations has become quite mature. However, this technology still has many shortcomings: The presentation lacks interactivity between the digital figures and the PowerPoint content. The digital figures' gestures are unrelated to the actual content being presented or its location, and scene transitions mostly follow pre-existing templates, failing to effectively integrate with the content. This dull and unengaging presentation, especially in educational settings, significantly diminishes student engagement. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method and terminal for automatically generating lecture videos, which can achieve higher quality lecture videos automatically generated.

[0004] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A method for automatically generating lecture videos, comprising the following steps: S1. Extract content from the target PPT file and optimize the content using large language modeling and RAG technologies; S2. Use a third-party large language model to learn the parameter rules of the digital human animation module, and generate the first version of the animation parameter file and a new speech based on the optimized target PPT file; S3. Using TTS technology, transcribe the speech into an audio file based on the speech text, simultaneously generate timestamp and subtitle files, and integrate them with the first version of the animation parameter file and the speech text to generate a second version of the animation parameter file; S4. Construct a mapping relationship between the speech manuscript and the pages in the target PPT file, generate image frames, generate lip-sync animation files from the audio files, and render to obtain the first animation file; S5. The final animation rendering is completed by combining the first animation file with the second version of the animation parameter file, and the teaching video is output.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A terminal for automatically generating lecture videos includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above-described method for automatically generating lecture videos.

[0006] The beneficial effects of this invention are as follows: The method and terminal for automatically generating lecture videos of this invention, combined with the Big Prophecy model and RAG technology, intelligently adjusts the PPT content to make the explanation more in line with semantic logic; achieves dynamic animation matching, and through the precise matching of animation parameter files with PPT content and voice, achieves perfect synchronization of actions, scenes and explanation content, significantly improving the digital human presentation effect; and effectively improves teaching quality and interactivity. Attached Figure Description

[0007] Figure 1 This is a simplified flowchart of a method for automatically generating lecture videos according to an embodiment of the present invention; Figure 2 This is a structural diagram of a terminal for automatically generating lecture videos according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating a method for automatically generating lecture videos according to an embodiment of the present invention. Label Explanation: 1. A terminal for automatically generating lecture videos; 2. A processor; 3. A memory. Detailed Implementation

[0008] To explain in detail the technical content, objectives, and effects of the present invention, the following description is provided in conjunction with the embodiments and accompanying drawings.

[0009] Please refer to Figure 1 as well as Figure 3 A method for automatically generating lecture videos, comprising the following steps: S1. Extract content from the target PPT file and optimize the content using large language modeling and RAG technologies; S2. Use a third-party large language model to learn the parameter rules of the digital human animation module, and generate the first version of the animation parameter file and a new speech based on the optimized target PPT file; S3. Using TTS technology, transcribe the speech into an audio file based on the speech text, simultaneously generate timestamp and subtitle files, and integrate them with the first version of the animation parameter file and the speech text to generate a second version of the animation parameter file; S4. Construct a mapping relationship between the speech manuscript and the pages in the target PPT file, generate image frames, generate lip-sync animation files from the audio files, and render to obtain the first animation file; S5. The final animation rendering is completed by combining the first animation file with the second version of the animation parameter file, and the teaching video is output.

[0010] As can be seen from the above description, the beneficial effects of the present invention are as follows: The method and terminal for automatically generating lecture videos of the present invention, combined with the Big Prophecy model and RAG technology, intelligently adjusts the PPT content to make the explanation more in line with semantic logic; achieves dynamic animation matching, and through the precise matching of animation parameter files with PPT content and voice, achieves perfect synchronization of actions, scenes and explanation content, significantly improving the digital human presentation effect; and effectively improves teaching quality and interactivity.

[0011] Furthermore, the steps between step S1 and step S2 include: S11. Analyze the optimized target PPT file using the Big Prophet model and generate layout suggestions; S12. Optimize the layout of the target PPT according to the layout suggestions.

[0012] As described above, the layout of the target PPT file has also been optimized, which not only ensures the accurate transmission of information, but also incorporates a large-scale model for in-depth understanding and optimization of the content, making the PPT content more concise and insightful.

[0013] Furthermore, the typesetting suggestions include a first type of typesetting suggestions and a second type of typesetting suggestions; The first type of layout suggestion is to edit and replace the content of the target PPT file, including text highlighting, font bolding and emphasis, and extracting text to generate a comparison table; The second type of layout suggestion includes copying the relevant PPT and then editing the content or adjusting the page numbers to achieve content linkage.

[0014] As described above, layout optimization not only involves editing and replacing text content to ensure accurate information delivery, but also further editing and optimizing PPT content. By switching between different frames, it guides the audience to focus on key information, enhancing the attractiveness and readability of the PPT.

[0015] Furthermore, step S1 specifically includes: Extract the document content and speaker notes from the target PPT file, match the speaker notes with the document content one-to-one, and form structured content that can be recognized by large models; The structured content is optimized using large language modeling and RAG techniques to achieve content optimization of the target PPT file.

[0016] As described above, this application extracts speaker notes and then maps them one-to-one with the original content in the PPT document to form structured content that can be recognized by large models, so that speaker notes can be processed and analyzed together with the main content.

[0017] Further, step S2 includes: The parameter rules of the digital human animation module are learned using a third-party large language model. Based on the optimized target PPT file, digital human and animation scene are assigned, and action parameters of word particles are assigned to each PPT page to generate the first version of the animation parameter file.

[0018] As described above, by generating the first parameter animation file, a 2D mapping between the PPT and the animated character, as well as alignment with the audio file, is achieved.

[0019] Furthermore, step S3 specifically includes: Using TTS technology, an audio file is transcribed from the speech transcript, and timestamps and subtitle files are generated simultaneously. Combining the timestamp file and the first version of animation parameters, the audio file, subtitle file, and target PPT file are aligned according to the PPT page number, content, and timestamp. The animation effects corresponding to the PPT content are then assigned to specific time nodes to obtain the second version of the animation parameter file.

[0020] As described above, during the speech transcription process, a timestamp file containing detailed records of the start and end times of each speech segment is generated simultaneously. Based on this timestamp file, sentence-level subtitle files are further extracted. In addition to subtitle functionality, these subtitle files also provide accurate time references for lip-syncing of digital humans, ensuring the smoothness and coordination of the teaching animation in both visual and auditory aspects.

[0021] Furthermore, the specific steps of transcribing the speech transcript into an audio file using TTS technology are as follows: Based on the content of the target PPT file, parameters such as language, pronunciation, and intonation are selected, and the speech text is transcribed into an audio file according to the selected parameters.

[0022] As described above, the transcription of audio files requires the selection of language, speech, and intonation parameters based on the content of the target PPT file, so that the audio file is more compatible with the target PPT.

[0023] Furthermore, step S4 specifically includes: Extract page information from the target PPT file, construct a mapping relationship between the page information and the text content of the presentation, and generate image frames of the scene digital human based on the mapping relationship; Based on the audio file, a lip-sync animation file is generated using a lip-sync generation service; The first animation file is obtained by combining the lip-sync animation file and the image frames generated based on the mapping relationship.

[0024] As described above, based on the mapping relationship between page information and the text content of the speech manuscript, combined with the timestamp of the audio file and the lip-sync animation file, the teaching animation is made smooth and coordinated in both visual and auditory senses.

[0025] Furthermore, step S5 specifically includes: Place the second version of the animation parameter file and the first version of the animation file into the player to drive the character to complete the animation effects rendering, including scene switching, prop use, special effects addition and switching; The generated animation is then rendered and output as the required teaching video file.

[0026] As described above, the first animation file has completed the mapping between the PPT and the animated character, and has been aligned with the audio file. The second version of the animation parameter file and the first version of the animation file are placed in the player to drive the character to complete the animation effect rendering, such as scene switching, prop usage, and the addition and switching of special effects.

[0027] Please refer to Figure 2 A terminal for automatically generating lecture videos includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the above-described method for automatically generating lecture videos.

[0028] The present invention provides a method and terminal for automatically generating lecture videos, which is applicable to the automatic generation of lecture videos in various scenarios such as education and training.

[0029] Please refer to Figure 1 and Figure 3 Embodiment 1 of the present invention is as follows: A method for automatically generating lecture videos, comprising the following steps: S1. Extract content from the target PPT file and optimize the content using large language modeling and RAG technologies; Step S1 is as follows: Extract the document content and speaker notes from the target PPT file, match the speaker notes with the document content one-to-one, and form structured content that can be recognized by large models; The structured content is optimized using large language modeling and RAG techniques to achieve content optimization of the target PPT file.

[0030] In this embodiment, a data cleaning module is constructed to address the difficulty of directly processing PPT files using existing large-scale modeling techniques. This module combines multiple advanced technologies to achieve efficient and accurate conversion of PPT files, with particular attention paid to the preservation and utilization of speaker notes.

[0031] A large model module was constructed to optimize PPT content and guide digital human motion, effectively combining the two. The PPT content was revised and improved, and its layout optimized. Furthermore, corresponding animation parameter files were generated based on the PPT content to drive subsequent digital human technology.

[0032] The content of the PowerPoint presentation is optimized and adjusted using RAG technology. Based on the main content and outline of the presentation, appropriate explanations are automatically generated, including summaries, expansions, or questions about complex concepts, enhancing the interactivity of the lecture videos. Furthermore, semantic analysis is performed on each slide of the presentation to generate accurate and natural vocabulary and smooth transition sentences, enabling the digital human to better "explain" each page.

[0033] First, extract the speaker's notes and the specific content of each slide from the PPT document, and then match the extracted content with the PPT page numbers one by one. Finally, use RAG technology to process the extracted PPT content.

[0034] Traditional document conversion methods often overlook speaker notes in PowerPoint presentations, which are crucial for effective teaching. To preserve and utilize this text, a speaker notes extraction step is specifically included. This step automatically identifies and extracts speaker notes from the PowerPoint document, then maps them one-to-one with the original content, creating structured content that can be recognized by large-scale models. In this way, the speaker notes can be processed and analyzed alongside the main content in subsequent steps.

[0035] By combining large language model technology and RAG technology, the PPT content is optimized. The content of each PPT page is optimized, and while ensuring the professionalism and accuracy of the PPT content, the diversity and interest of the PPT content are appropriately increased.

[0036] While traditional large language models can optimize content to some extent, they often fail to converge effectively within a specified range to meet user needs when faced with specialized problems. To address this issue, RAG technology was introduced. RAG technology enhances the specialized capabilities of large models, enabling them to extract relevant information and perform optimization processing more accurately when dealing with specialized problems.

[0037] S11. Analyze the optimized target PPT file using the Big Prophet model and generate layout suggestions; The typesetting suggestions include a first type of typesetting suggestion and a second type of typesetting suggestion; The first type of layout suggestion is to edit and replace the content of the target PPT file, including text highlighting, font bolding and emphasis, and extracting text to generate a comparison table; The second type of layout suggestion includes copying the relevant PPT slides and then editing the content or adjusting the page numbers to achieve content linkage; S12. Optimize the layout of the target PPT according to the layout suggestions.

[0038] In this embodiment, the large model module uses a third-party large language model to analyze the optimized PPT and provides appropriate layout suggestions based on the PPT content. These suggestions can be divided into two main categories based on subsequent operations: the first category only requires editing the PPT content; the second category requires not only editing the content but also adding or adjusting page numbers. For example, bolding and highlighting key content falls under the first category. Gradual display of multiple content items and linking related content fall under the second category. Specific layout operations will be further processed in the subsequent PPT layout optimization module.

[0039] The PPT layout optimization module further refines the PPT based on content optimization suggestions, layout optimization suggestions, and animation parameter guidance suggestions from the overall model. This includes: Based on the layout suggestions from the large language model, for the first type of layout suggestion, it is only necessary to edit the relevant PPT content according to the layout suggestions and replace the original content. Adopting the first type of layout suggestion from the large language model, further editing and replacement of the text content of the original PPT page by page is required. This includes, but is not limited to, text highlighting and font bolding. This step not only ensures the accurate transmission of information but also incorporates the large model's deep understanding and optimization of the content, making the PPT content more concise and insightful.

[0040] Redesigning the layout: Follow layout suggestions to visually optimize your PPT. This includes highlighting and bolding key content, and generating tables for comparisons.

[0041] Innovative Presentation Effects: Based on the second type of layout suggestion, appropriately copy relevant PPT slides before editing the content, or adjust the PPT page numbers to achieve content synergy. For example, when the layout suggestion is to display specific content progressively, first copy the related PPT slides, and then edit the content of the copied PPT slides according to the actual structure of the content, thereby corresponding to different animation frames.

[0042] Following the second type of layout suggestions provided by the large model, the PPT content was further edited and optimized. The aim was to guide the audience to focus on key information and enhance the attractiveness and readability of the PPT by switching between different frames.

[0043] S2. Use a third-party large language model to learn the parameter rules of the digital human animation module, and generate the first version of the animation parameter file and a new speech based on the optimized target PPT file.

[0044] In this embodiment, a new speech manuscript is generated by combining the content of the optimized PPT with a large language model and RAG technology, which is then used to generate an audio file.

[0045] Step S2 includes: The parameter rules of the digital human animation module are learned using a third-party large language model. Based on the optimized target PPT file, digital human and animation scene are assigned, and action parameters of word particles are assigned to each PPT page to generate the first version of the animation parameter file.

[0046] In this embodiment, a large model is used to learn animation demonstration rules and parameters, including the digital human's performance style, voice, tone, and usage scenarios. Combined with optimized PPT content, precise animation parameter guidance suggestions down to the word level are generated, and a corresponding animation parameter document is output.

[0047] That is, by using a third-party big language and big model to learn the parameter rules unique to the digital human animation module, and combining them with the optimized PPT content, not only is the digital human and animation scene allocation completed, but also the first round of action parameter allocation for word particles on each PPT page is carried out to obtain the first version of the animation parameter file.

[0048] S3. Using TTS technology, transcribe the speech into an audio file based on the speech text, simultaneously generate timestamp and subtitle files, and integrate them with the first version of the animation parameter file and the speech text to generate a second version of the animation parameter file; Step S3 is as follows: Using TTS technology, an audio file is transcribed from the speech transcript, and timestamps and subtitle files are generated simultaneously. The specific steps for transcribing the speech transcript into an audio file using TTS technology are as follows: Based on the content of the target PPT file, parameters such as language, pronunciation, and intonation are selected, and the speech text is transcribed into an audio file according to the selected parameters.

[0049] Combining the timestamp file and the first version of animation parameters, the audio file, subtitle file, and target PPT file are aligned according to the PPT page number, content, and timestamp. The animation effects corresponding to the PPT content are then assigned to specific time nodes to obtain the second version of the animation parameter file.

[0050] In this embodiment, appropriate language, voice, tone, and other parameters are selected based on the content of the PPT, and the optimized presentation text is transcribed into an audio file. A TTS (Text-to-Speech) module is set up, which, combined with TTS technology, provides a variety of natural voice options (such as different tones, speaking speeds, and emotional expressions), and automatically adjusts the voice style and tone according to the content. For example, technical PPTs can choose a calm and professional tone, while humanities PPTs can choose a more vivid and expressive voice.

[0051] During the speech transcription process, timestamp files detailing the start and end times of each speech segment were generated simultaneously. Sentence-level subtitle files were then extracted from these timestamp files. In addition to their subtitle function, these subtitle files provided precise time references for synchronized lip-syncing with the digital human, ensuring both visual and auditory fluency and harmony in the teaching animation.

[0052] An animation parameter allocation module was set up. Combining the timestamp file and the first version of animation parameters, the module aligns the page numbers, content, and timestamp file of the PPT with each other, and assigns the corresponding animation effects on the PPT content to specific time nodes, resulting in a second version of the animation parameter file for subsequent animation rendering.

[0053] S4. Construct a mapping relationship between the speech manuscript and the pages in the target PPT file, generate image frames, generate lip-sync animation files from the audio files, and render to obtain the first animation file; Step S4 is as follows: Extract page information from the target PPT file, construct a mapping relationship between the page information and the text content of the presentation, and generate image frames of the scene digital human based on the mapping relationship; Based on the audio file, a lip-sync animation file is generated using a lip-sync generation service; The first animation file is obtained by combining the lip-sync animation file and the image frames generated based on the mapping relationship.

[0054] In this embodiment, an animation generation module is set up to generate a digital human presentation animation by combining high-quality input data (edited PPT, audio, subtitle files, and motion parameter files): The process involves uploading a PowerPoint presentation, extracting page information from the PPT file, constructing a 2D mapping between the presentation text and the PPT pages, and adding this mapping to a playback template to generate image frames. A 3D lip-sync animation file is then generated using the audio file via a lip-sync generation service. This lip-sync animation file, along with the image frames mapping the PPT content and the scene's digital human, is then rendered in a player to obtain the first animation file. In this first animation file, the 2D mapping between the PPT and the animated character is complete, and it is aligned with the audio file.

[0055] S5. The final animation rendering is completed by combining the first animation file with the second version of the animation parameter file, and the teaching video is output. Step S5 is as follows: Place the second version of the animation parameter file and the first version of the animation file into the player to drive the character to complete the animation effects rendering, including scene switching, prop use, special effects addition and switching; The generated animation is then rendered and output as the required teaching video file.

[0056] In this embodiment, the second version of the animation parameter file and the first version of the animation file are placed in the player to drive the character to complete the animation effect rendering, such as scene switching, prop usage, and the addition and switching of special effects. The generated animation is then rendered and output as the required video file. Please refer to Figure 2 Embodiment two of the present invention is as follows: A terminal 1 for automatically generating lecture videos includes a processor 2, a memory 3, and a computer program stored in the memory 3 and executable on the processor 2. When the processor 2 executes the computer program, it implements the steps of the method for automatically generating lecture videos described in Embodiment 1 above.

[0057] In summary, the method and terminal for automatically generating lecture videos provided by this invention, combined with the Big Prophecy model and RAG technology, intelligently adjusts the PPT content to make the explanation more consistent with semantic logic; achieves dynamic animation matching, and through precise matching of animation parameter files with PPT content and voice, achieves perfect synchronization of actions, scenes and explanation content, significantly improving the digital human presentation effect; and effectively enhances teaching quality and interactivity.

[0058] This invention applies technologies such as large-scale models and digital humans to generate lecture videos from PPT presentations, offering significant advantages in production efficiency, content quality, presentation effects, and application scenarios. It enhances the quality and efficiency of teaching and knowledge dissemination, bringing new development opportunities to the education and training industries. 1. Highly Automated Production Process: The patented technology automates the entire process from PPT content optimization to animation rendering. As the document states, "The entire process from PPT file to digital human teaching video is automated, covering text optimization, voice generation, animation matching and rendering. High-quality teaching videos can be generated without human intervention," greatly reducing manual operation, saving labor costs, and significantly improving video production efficiency.

[0059] 2. Deeply Optimize Teaching Content: Leveraging large-scale models and RAG technology, deeply optimize PPT content, including semantic analysis and interaction design, automatically generating logically consistent explanations and transitional statements. This enhances the interactivity of teaching content, improves teaching quality, and makes knowledge transfer more effective, aligning with educational needs. For example, "Deeply optimize PPT content, including content adjustments, semantic analysis, and interaction design. Automatically generate semantically logical explanations and natural transitional statements."

[0060] 3. Precisely match animation with content: By using animation parameter files, the animation parameters and audio timestamps are precisely aligned to ensure that the digital human's actions, scenes and PPT presentation content are perfectly synchronized, improving the digital human's presentation effect, making the teaching video more attractive and immersive, and avoiding the disconnect between content and presentation. "It achieves alignment between animation parameters and audio timestamps, generates dynamic and fine-grained motion rendering, and ensures that the digital human presentation and PPT content are completely synchronized."

[0061] 4. High-quality output videos: The final generated teaching videos are natural and smooth, with vivid and interesting scene transitions, which can meet the needs of various scenarios such as education and training, bring users a high-quality experience, help knowledge dissemination and learning, and adapt to the needs of knowledge explanation and training in different fields. "The video output is natural and smooth, and the scene transitions are vivid and interesting, suitable for various scenarios such as education and training, and improves the user experience."

[0062] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent modifications made based on the content of the present invention specification and drawings, or direct or indirect applications in related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for automatically generating lecture videos, characterized in that, Including the following steps: S1. Extract content from the target PPT file and optimize the content using large language modeling and RAG technologies; S2. Use a third-party large language model to learn the parameter rules of the digital human animation module, and generate the first version of the animation parameter file and a new speech based on the optimized target PPT file; S3. Using TTS technology, transcribe the speech into an audio file based on the speech text, simultaneously generate timestamp and subtitle files, and integrate them with the first version of the animation parameter file and the speech text to generate a second version of the animation parameter file; S4. Construct a mapping relationship between the speech manuscript and the pages in the target PPT file, generate image frames, generate lip-sync animation files from the audio files, and render to obtain the first animation file; S5. The final animation rendering is completed by combining the first animation file with the second version of the animation parameter file, and the teaching video is output.

2. The method for automatically generating lecture videos according to claim 1, characterized in that, The steps between step S1 and step S2 include: S11. Analyze the optimized target PPT file using the Big Prophet model and generate layout suggestions; S12. Optimize the layout of the target PPT according to the layout suggestions.

3. The method for automatically generating lecture videos according to claim 2, characterized in that, The typesetting suggestions include a first type of typesetting suggestion and a second type of typesetting suggestion; The first type of layout suggestion is to edit and replace the content of the target PPT file, including text highlighting, font bolding and emphasis, and extracting text to generate a comparison table; The second type of layout suggestion includes copying the relevant PPT and then editing the content or adjusting the page numbers to achieve content linkage.

4. The method for automatically generating lecture videos according to claim 1, characterized in that, Step S1 is as follows: Extract the document content and speaker notes from the target PPT file, match the speaker notes with the document content one-to-one, and form structured content that can be recognized by large models; The structured content is optimized using large language modeling and RAG techniques to achieve content optimization of the target PPT file.

5. The method for automatically generating lecture videos according to claim 1, characterized in that, Step S2 includes: The parameter rules of the digital human animation module are learned using a third-party large language model. Based on the optimized target PPT file, digital human and animation scene are assigned, and action parameters of word particles are assigned to each PPT page to generate the first version of the animation parameter file.

6. The method for automatically generating lecture videos according to claim 1, characterized in that, Step S3 is as follows: Using TTS technology, an audio file is transcribed from the speech transcript, and timestamps and subtitle files are generated simultaneously. Combining the timestamp file and the first version of animation parameters, the audio file, subtitle file, and target PPT file are aligned according to the PPT page number, content, and timestamp. The animation effects corresponding to the PPT content are then assigned to specific time nodes to obtain the second version of the animation parameter file.

7. A method for automatically generating lecture videos according to claim 1 or 6, characterized in that, The specific steps for transcribing the speech transcript into an audio file using TTS technology are as follows: Based on the content of the target PPT file, parameters such as language, pronunciation, and intonation are selected, and the speech text is transcribed into an audio file according to the selected parameters.

8. The method for automatically generating lecture videos according to claim 1, characterized in that, Step S4 is as follows: Extract page information from the target PPT file, construct a mapping relationship between the page information and the text content of the presentation, and generate image frames of the scene digital human based on the mapping relationship; Based on the audio file, a lip-sync animation file is generated using a lip-sync generation service; The first animation file is obtained by combining the lip-sync animation file and the image frames generated based on the mapping relationship.

9. The method for automatically generating lecture videos according to claim 1, characterized in that, Step S5 is as follows: Place the second version of the animation parameter file and the first version of the animation file into the player to drive the character to complete the animation effects rendering, including scene switching, prop use, special effects addition and switching; The generated animation is then rendered and output as the required teaching video file.

10. A terminal for automatically generating lecture videos, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps in the method for automatically generating teaching videos as described in any one of claims 1-9.