Digital human course generation method and device based on large model and storage medium

CN122655705APending Publication Date: 2026-08-28SHENHUA TRAINING CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610485474.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-14
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0004]有鉴于此,本申请实施例提供了一种基于大模型的数字人课程生成方法、装置及存储介质,以解决现有技术存在的制课流程繁琐、内容难结构化生成、视频课件生成自动化程度低的问题

Benefits of technology

通过获取课程生成请求,并对课程参考素材执行内容解析以生成课程上下文,其中,课程生成请求包括课程目标信息以及课程参考素材中的一种或多种;基于课程目标信息和课程上下文调用大语言模型生成课程大纲,并生成与课程大纲对应的逐页讲解脚本;接收用户对课程大纲和逐页讲解脚本的编辑结果并更新课程大纲和逐页讲解脚本;获取模板标识并基于模板标识从模板库确定课件模板,根据更新后的课程大纲和逐页讲解脚本生成课件文件,并将逐页讲解脚本按页写入课件文件的页级备注信息;对课件文件执行分页渲染以得到各页课件图像,并从页级备注信息读取与各页对应的脚本片段,建立课件图像与脚本片段的页级关联关系;获取数字人配置参数和音色配置参数,基于脚本片段生成对应的语音片段,并基于数字人配置参数对脚本片段执行口型对齐以生成数字人视频片段;基于页级关联关系将各页课件图像、语音片段及数字人视频片段按页序进行时序合成,输出授课视频课件。本申请能够提高制课自动化与生成效率、降低课件与视频制作成本、提升课件内容一致性与可复用性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122655705A_ABST
    Figure CN122655705A_ABST
Patent Text Reader

Abstract

The application provides a digital human course generation method and device based on a large model and a storage medium. The method comprises: performing content analysis on a course reference material to generate a course context; calling a large language model to generate a course outline and a page-by-page explanation script; obtaining a template identifier and determining a courseware template from a template library based on the template identifier, generating a courseware file according to the updated course outline and the page-by-page explanation script; performing page rendering on the courseware file to obtain page courseware images, and establishing a page-level association relationship between the courseware images and the script segments; generating corresponding voice segments based on the script segments, and generating digital human video segments; based on the page-level association relationship, the page courseware images, the voice segments and the digital human video segments are time-synchronously combined according to the page sequence, and a teaching video courseware is output. The application can improve the course generation automation and efficiency, reduce the courseware and video production cost, and improve the courseware content consistency and reusability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus and storage medium for generating digital human courses based on a large model. Background Technology

[0002] With the advancement of enterprise digital transformation and the development of artificial intelligence, virtual reality, and big data analytics technologies, online learning platforms are gradually becoming important vehicles for enterprises to conduct training management, knowledge dissemination, and experience accumulation. Existing online learning platforms typically provide functions such as course resource management, learning task publishing, learning process recording, examination and assessment, and statistical reports. Some platforms have introduced templated courseware creation, content recommendation, and basic profile analysis to support course development and training operations.

[0003] However, in practical applications, the course content production chain remains primarily manual. The processes of courseware arrangement, script writing, audio and video recording, and post-production editing are fragmented and rely on multiple tools, resulting in a complex and repetitive workflow that struggles to adapt to the rapid iteration needs of different roles and topics. Simultaneously, the platform's intelligence level is low, lacking the ability to understand, extract, and structure course objectives and reference materials, making it difficult to create reusable course outlines, page-by-page explanation scripts, and courseware content. Furthermore, its data analysis and application capabilities are limited, making it difficult to transform the content and learning process data collected by the platform into executable results that directly drive course production and operational adjustments. To meet the needs of enterprises for personalized courseware styles, template systems, and teaching presentation formats, customized development and project delivery are typically required, leading to high development costs, long cycles, and complex maintenance. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method, apparatus and storage medium for generating digital human courses based on a large model, in order to solve the problems of cumbersome course production process, difficulty in generating structured content and low degree of automation in video courseware generation in the prior art.

[0005] A first aspect of this application provides a method for generating digital human courses based on a large model, comprising: obtaining a course generation request and performing content parsing on course reference materials to generate a course context, wherein the course generation request includes one or more of course objective information and course reference materials; generating a course outline by calling a large language model based on the course objective information and the course context, and generating a page-by-page explanation script corresponding to the course outline; receiving the user's editing results on the course outline and the page-by-page explanation script and updating the course outline and the page-by-page explanation script; obtaining a template identifier and determining a courseware template from a template library based on the template identifier, and generating a courseware template according to the updated course outline and the courseware template; and generating a courseware template based on the updated courseware template. The process involves generating courseware files from the outline and page-by-page explanation script, and writing the explanation script page by page into the page-level annotation information of the courseware files. The courseware files are then rendered page by page to obtain the courseware images for each page. The script fragments corresponding to each page are read from the page-level annotation information, establishing a page-level association between the courseware images and the script fragments. Digital human configuration parameters and timbre configuration parameters are obtained, and corresponding audio fragments are generated based on the script fragments. Lip alignment is then performed on the script fragments based on the digital human configuration parameters to generate digital human video fragments. Based on the page-level association, the courseware images, audio fragments, and digital human video fragments for each page are sequentially synthesized according to page order to output the teaching video courseware.

[0006] A second aspect of this application provides a digital human course generation device based on a large model, comprising: an acquisition module, configured to acquire a course generation request and perform content parsing on course reference materials to generate a course context, wherein the course generation request includes one or more of course objective information and course reference materials; a generation module, configured to generate a course outline based on the course objective information and the course context by calling a large language model, and generate a page-by-page explanation script corresponding to the course outline; receive user editing results of the course outline and page-by-page explanation script and update the course outline and page-by-page explanation script; and a determination module, configured to acquire a template identifier and determine a courseware template from a template library based on the template identifier, and determine the courseware template according to the updated course outline. The system generates courseware files using a syllabus and page-by-page explanation script, and writes the explanation script page by page into the page-level annotation information of the courseware files. The rendering module performs page-by-page rendering of the courseware files to obtain the courseware images for each page, and reads the script fragments corresponding to each page from the page-level annotation information, establishing a page-level association between the courseware images and script fragments. The configuration module obtains digital human configuration parameters and timbre configuration parameters, generates corresponding speech fragments based on the script fragments, and performs lip-syncing on the script fragments based on the digital human configuration parameters to generate digital human video fragments. The synthesis module synthesizes the courseware images, speech fragments, and digital human video fragments in a time sequence according to the page-level association, outputting the teaching video courseware.

[0007] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0008] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0009] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: The process involves: obtaining a course generation request and parsing the course reference materials to generate a course context; the course generation request including course objective information and one or more course reference materials; generating a course syllabus based on the course objective information and course context using a large language model, and generating a corresponding page-by-page explanation script; receiving and updating the course syllabus and page-by-page explanation script based on user edits; obtaining a template identifier and determining a courseware template from a template library based on the template identifier; and generating courseware files based on the updated course syllabus and page-by-page explanation script. The process involves writing the page-by-page explanation script into the page-level annotation information of the courseware file; performing paginated rendering on the courseware file to obtain the courseware images for each page, and reading the script fragments corresponding to each page from the page-level annotation information to establish a page-level association between the courseware images and script fragments; obtaining digital human configuration parameters and voice configuration parameters, generating corresponding audio fragments based on the script fragments, and performing lip-syncing on the script fragments based on the digital human configuration parameters to generate digital human video fragments; and sequentially synthesizing the courseware images, audio fragments, and digital human video fragments of each page according to the page-level association to output the teaching video courseware. This application can improve the automation and generation efficiency of courseware production, reduce the cost of courseware and video production, and enhance the consistency and reusability of courseware content. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart illustrating the fusion-based training method based on performance feedback provided in an embodiment of this application. Figure 2 This is a schematic diagram of the structure of the fusion training device based on performance feedback provided in an embodiment of this application; Figure 3This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0013] With the development and application of technologies such as artificial intelligence and virtual reality, online learning platforms are rapidly evolving. The rapid development of technologies like artificial intelligence and big data has provided more possibilities and opportunities for learning and training businesses, offering new impetus for building learning organizations and creating a new pattern of knowledge co-creation and sharing across the entire enterprise. However, current online learning platforms suffer from problems such as a single operational model, low level of intelligence, insufficient data analysis and application capabilities, and high costs and time commitments for customized development.

[0014] In view of the problems existing in the prior art, this application provides a course generation method for an online learning platform. This application utilizes the contextual understanding capabilities of AI large-scale models to generate course outlines, create PPT courses, and produce digital human video courseware. Specifically, it supports the use of PPT and the instructor's script / voice to drive digital human narration during course production, synthesizing courseware PPTs and other materials to quickly generate teaching videos. Leveraging AI technology, it eliminates the cumbersome and tedious course production process, enabling corporate instructors and managers to customize digital avatars to accelerate course development. It also helps companies quickly achieve corporate knowledge accumulation, thereby overcoming limitations imposed by environment and equipment, resolving cost and supply-demand imbalances in course content production, and rapidly building and applying the latest AI-generated high-quality content to meet the diverse needs of instructors in creating courseware during teaching activities.

[0015] The following section provides a detailed explanation of the process for creating courseware using AI, using real-world examples. Specifically, it may include the following steps: Step 1. Input AI Course Objectives Input the course's target information, including the course title, writing style, and specific objectives. Also, provide reference materials such as PPTs and videos. The AI-powered course will then consider the actual content and application scenarios of these materials to generate a course outline.

[0016] Step 2. AI automatically generates and modifies the outline. AI-generated courseware automatically identifies, extracts, and processes content from course objectives and reference materials, generating an AI outline and a verbatim transcript of the explanations in the notes for each PPT slide. Users can then edit and modify the outline. The process includes: 1) selecting a PPT template (company-owned templates are supported; the platform has dozens of built-in PPT templates for users to choose from); 2) generating a detailed PPT and a verbatim transcript based on the selected template, course objectives, outline, and materials, and intelligently polishing the transcript; and 3) previewing and downloading the AI-generated PPT after successful generation.

[0017] In step 2, implementing PPT generation functionality based on a large language model within the AI ​​training platform will greatly simplify the content creation process and improve the efficiency of training material production. Large language models possess powerful natural language understanding and generation capabilities, enabling them to generate structured PPT content based on input text or topics. The core of this technical approach lies in how to efficiently convert natural language descriptions into multimedia content suitable for presentation.

[0018] First, users input a brief description, key points, or outline related to the PPT topic through the platform. The large language model receives this input and generates corresponding document outlines and text content. During this process, the model automatically identifies the chapter structure, generates titles and paragraph content based on the input, and extracts the most essential information for each slide.

[0019] The platform will integrate a PPT template system, which allows users to select suitable PPT templates to generate PPT presentations.

[0020] Finally, the platform should provide a user-friendly editing interface, allowing users to preview and modify the initial draft online to ensure the final PPT meets their specific needs and preferences. This technical approach, by combining the natural language processing capabilities of a large language model with an automated process for multimedia content generation, provides users with an efficient and flexible PPT production solution, significantly reducing the time cost of manual intervention while improving content quality and consistency.

[0021] In another branch of step 2, the PPT generation capability needs to support two scenarios. One scenario is when there are no materials available, relying solely on the administrator's input of the title and course objectives to generate a PPT outline using large language model capabilities. After appropriate editing of the PPT outline, a complete PPT file is generated. The other scenario is PPT generation based on existing materials. The material-based scenario is the more frequently used scenario in this project. By parsing and recognizing existing materials, the knowledge uploaded by the administrator is converted into a PPT.

[0022] By leveraging the capabilities of generative artificial intelligence, the text content is summarized, an outline is extracted, and it is converted into the format required to generate a PPT. Then, using the placeholder editing technology of PPT master slides, the content is filled into the PPT to generate a beautiful and easy-to-use result.

[0023] Step 3. Select to explain the AI ​​digital human. Choose a suitable digital avatar (gender, style, etc.); position and size. The platform has a variety of built-in digital avatar styles; you can choose a suitable digital avatar, adjusting its gender, style, and pose. Digital avatar adjustments include adjusting its page position and size. Step 4. Generate AI video, which allows you to change patterns and add interactive actions during the PPT presentation; The features include: Template selection: Users can choose existing templates to create courseware, improving course production efficiency; Multiple theme options are available; Video editing: Supports adding / copying / deleting / switching scene pages; Supports setting transitions between scene pages; Supports image, text, and video materials with effect settings; AI voice-over / subtitles; Supports multi-character voice-over dialogue; AI subtitle function generates subtitles with one click; Supports recording, voice-over, and AI intelligent voice-over team collaboration; Team creators and administrators manage all materials and categorize them.

[0024] Step 4 introduces a PPT-to-video conversion function based on intelligent media services into the AI ​​learning platform. This further enhances the dissemination and diversity of content, meeting the needs of different learning scenarios. Intelligent media services combine technologies such as natural language processing, speech synthesis, and image processing to convert static PPT content into dynamic video formats, thereby improving the user's learning experience.

[0025] This scenario uses Python as the scripting language to connect the business processes. The first step in the technical approach is to extract the text and image content from the PowerPoint presentation into image information, preparing for image-based video generation. This is done by using Python to call the LibreOffice command-line tool to export the PowerPoint presentation as images. The images are then uploaded to the cloud platform OSS, and their URLs are obtained.

[0026] The presentation script is read from the PPT notes (the generated PPT will automatically fill in the presentation notes, and manual editing is also supported). Based on the PPT page number association, image addresses are matched one-to-one with presentation segments.

[0027] The digital human model list interface and voice model list interface are provided by the AI ​​platform. After selecting the appropriate digital human and voice model numbers, the final video is generated by calling the AI ​​platform's video generation and editing services.

[0028] The technologies used include: Text-to-Speech (TTS) technology: Each page of the speech is a text segment. Based on a text-to-speech model, inference is performed to convert the text content into speech. This technology relies on a model trained on a large amount of speech and corresponding text data. This project selects an existing model and does not retrain it.

[0029] Digital Human Technology: Based on an AI platform, digital human technology uses LipSync technology to lip-sync script content and generate digital human video clips.

[0030] Then, through cloud-based intelligent media services, the digital human video, audio clips, and PPT slide images are combined into a video using video editing. The underlying technology for this compositing is FFmpeg.

[0031] The technical solution of this application will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0032] Figure 1 This is a flowchart illustrating the digital human course generation method based on a large model provided in an embodiment of this application. For example... Figure 1 As shown, this digital human course generation method based on a large model can specifically include: S101, Obtain a course generation request and perform content parsing on the course reference materials to generate a course context, wherein the course generation request includes course objective information and one or more of the course reference materials; S102: Based on the course objective information and course context, call the large language model to generate a course outline and generate a page-by-page explanation script corresponding to the course outline; receive the user's editing results on the course outline and page-by-page explanation script and update the course outline and page-by-page explanation script; S103, obtain the template identifier and determine the courseware template from the template library based on the template identifier, generate the courseware file according to the updated course outline and page-by-page explanation script, and write the page-by-page explanation script into the page-level remarks information of the courseware file; S104, Perform paginated rendering on the courseware file to obtain courseware images for each page, and read the script fragments corresponding to each page from the page-level notes information to establish the page-level association between the courseware images and the script fragments; S105: Obtain digital human configuration parameters and timbre configuration parameters, generate corresponding speech segments based on script segments, and perform lip-syncing on script segments based on digital human configuration parameters to generate digital human video segments. S106, based on page-level association, sequentially synthesize the courseware images, audio segments, and digital human video segments of each page according to page order, and output the teaching video courseware.

[0033] In some embodiments, obtaining a course generation request and performing content parsing on course reference materials to generate course context includes: Receive course reference materials uploaded or specified by users, and determine the material type and structure information of the course reference materials; Perform content extraction processing on course reference materials to obtain text content, key page information, and media content description information corresponding to the course reference materials; The text content, page key information, and media content description information are regularized and structured to generate structured material content, which includes a set of key themes and source tag information associated with the set of key themes. The course context is constructed based on structured material content and course objective information. The course context includes a set of context fragments used by the large language model to generate the course outline and structured index information corresponding to the set of context fragments.

[0034] Specifically, in this embodiment, the platform first obtains a course generation request. The course generation request is submitted by the administrator on the "AI Course Objective Input" interface and includes at least the course title, writing style, and specific course objectives. For example, the course title could be "Required Course for Safe Production Upon Entry," the writing style could be "corporate training tone, clear and organized," and the specific course objectives could include "covering safety precautions upon entry, common risk points, and emergency response procedures." In addition to the course objective information, the course generation request may also include course reference materials. These reference materials can be uploaded by the administrator or specified from the platform's material library. The reference materials can be in the form of PPT files, video files, or one or more existing course documents, used to ensure the actual content and application scenarios of the AI-powered intelligent course reference materials are relevant.

[0035] Furthermore, upon receiving the course reference materials, the platform determines the material type and structure information. For example, when the course reference material is a PPT file, the platform identifies it as "slideshow material" and parses it to obtain material structure information such as page number, layout identifier for each page, number of text boxes, number of image objects, and whether page-level notes exist. When the course reference material is a video file, the platform identifies it as "video material" and parses it to obtain structural information such as video duration, resolution, whether an audio track exists, and keyframe extraction interval. When the course reference material includes both PPT and video, the platform generates corresponding structural descriptions for each and records the relationships between the materials in the material list, such as recording whether the video is a screen recording of the PPT, for subsequent content extraction and alignment.

[0036] Furthermore, after determining the type and structure of the materials, the platform performs content extraction processing on the course reference materials to obtain text content, key page information, and media content descriptions. For PPT materials, the platform extracts the content of each page, obtaining text content such as page titles, body paragraphs, and bulleted lists, and extracts key page information by combining the hierarchical and layout structure of the text within the page. For example, for the "Safety Instructions for Entering the Factory" page, the platform can extract key page information such as "Wear personal protective equipment," "Flammable and explosive materials are prohibited," and "Comply with hot work permits," while retaining the page identifier and position index within the page during extraction.

[0037] For video materials, the platform can perform speech-to-text transcription on the video audio track to generate text content aligned with the timeline, and perform image content recognition on keyframes or extracted frames to generate media content description information. For example, the platform can identify descriptions such as "demonstrating the wearing of a safety helmet," "steps for using fire-fighting equipment," and "emergency assembly point diagram" from the video, and associate each description with its corresponding time period markers for storage. For images or charts contained in PowerPoint presentations, the platform can also perform image content recognition to generate media content description information, such as recognizing "safety evacuation route map" as "factory area evacuation route diagram and assembly point markers."

[0038] Furthermore, after obtaining the text content, page key information, and media content description information, the platform performs regularization and structuring processing on the above information to generate structured material content. Regularization processing may include removing noisy content such as headers and footers, standardizing terminology, merging duplicate key points, and aggregating cross-page information by topic. Structuring processing is used to organize the extracted results into a set of structured entries that can be consumed by the model. In this embodiment, the structured material content includes at least a set of key topics and source tagging information associated with the set of key topics. The set of key topics is used to express reusable knowledge points and key facts in the material, and the source tagging information is used to identify which course reference material, page, or time segment each key topic comes from. For example, the key topic "Hot work requires approval" can correspond to the source tagging information "Material A (PPT): Page 5", and the key topic "Four-step method for fire extinguisher" can correspond to the source tagging information "Material B (Video): 00:03:10-00:03:40", so that the source of the key points can be traced in the subsequent generation process and the administrator can verify it.

[0039] Furthermore, based on the structured content and course objective information, the platform constructs a course context. The course context serves to carry the input information required for the subsequent generation of the course syllabus by the large language model. In this embodiment, it includes at least a set of context fragments and corresponding structured index information. The set of context fragments can consist of several readable content segments. Each fragment contains a set of key themes, a brief explanatory text corresponding to the key themes, and a necessary media description summary, providing sufficient background information for the large language model to understand the course topics. The structured index information describes the organization and retrieval entry points of the fragments, and includes at least fragment identifiers, fragment topic tags, a source tag citation list, and association tags with the course objective information.

[0040] For example, the platform can divide the course context into multiple context fragments such as "factory entry behavior code fragment", "high-risk operation fragment", and "emergency response fragment", and generate structured index information for each fragment. This allows the platform to select the most relevant set of context fragments to input into the model when calling the large language model to generate the course outline, or to generate a structured framework of empty material context based on the course objective information when the administrator only inputs the title and objective and there is no material. This ensures that the input structure remains consistent.

[0041] Through the above processing, this embodiment realizes the process of starting from a course generation request, performing material type identification, structural analysis, content extraction, and structural organization on course reference materials, and constructing a course context based on this. This course context can carry the main themes and media descriptions in the materials in a searchable and referable structured form, providing a stable input foundation for subsequent generation of course outlines and page-by-page explanation scripts based on a large language model. The technical effect of this embodiment is to improve the efficiency of automated organization of course reference materials into course context, reduce the time cost of manually sorting materials and extracting outlines, and improve the traceability and consistency of the source of materials for subsequent generated content.

[0042] In some embodiments, a course syllabus is generated based on course objective information and course context by invoking a large language model, and a page-by-page explanation script corresponding to the course syllabus is generated, including: The course objective information and course context are encapsulated into the generation input of the large language model. The generation input includes the course title, writing style, specific course objectives, and a set of context fragments corresponding to the course reference materials. The large language model is invoked to perform structured generation on the generated input, and the output is a course outline corresponding to the course topic. The course outline includes a sequence of chapter titles and a set of key points associated with each chapter title. Based on the course outline, the page-level structure of the courseware is determined, and page-by-page explanation scripts corresponding to the page-level structure of the courseware are generated. The page-by-page explanation scripts include the page title, the summary of the key points of the page, and the explanation text corresponding to each page. The generated page-by-page explanation script is processed to ensure that the page titles and summary of key points in the page-by-page explanation script correspond to the chapter titles and key point sets in the course outline.

[0043] Specifically, the platform first encapsulates the course objective information and course context into the input for generating a large language model. The course objective information is filled in by the administrator on the platform interface and includes at least the course title, writing style, and specific course objectives. For example, the administrator might fill in the course title as "Required Course for Safe Production Upon Entry," select "corporate training tone, clear and organized" as the writing style, and fill in the specific course objectives as "covering safety precautions upon entry, common risk points, emergency response procedures, and providing key operational points for typical scenarios." The course context comes from the parsed results of the course reference materials and includes at least a set of context fragments and the corresponding structured index information.

[0044] For example, in some examples, the context fragment set may contain three types of fragments: "factory entry behavior code fragments," "high-risk operation fragments," and "emergency response fragments." Each fragment contains a set of key themes, a media content description summary, and source tagging information. The source tagging information indicates which page of the PPT or which time period of the video the key themes originate from. When encapsulating the generated input, the platform writes the course objective information as a global constraint field into the generated input, writes the context fragment set into the context field of the generated input according to the fragment identifier, and writes the structured index information into the index field of the generated input. This allows the large language model to locate and reference the key themes of the corresponding fragments by index during generation.

[0045] During the process of encapsulating and generating input, the platform can also organize the input order and weight of the context fragment set according to writing style and specific course objectives. For example, when the specific course objective emphasizes "emergency response procedures," the platform can place the "emergency response fragment" at the beginning of the context fragment set and mark it as a priority fragment in the index field to guide the model to prioritize extracting key points such as evacuation assembly points, alarm procedures, and fire extinguisher usage steps from that fragment.

[0046] In another example, when the administrator does not upload course reference materials and only enters the course title and specific course objectives, the platform still generates an empty material context frame according to the same generated input structure and marks the generated input as "no material mode" so that the model can complete the initial generation of the course outline based on the course objective information, thus satisfying the scenario of generating an outline without materials.

[0047] Furthermore, after completing the encapsulation of the generated input, the platform calls a large language model to perform structured generation on the generated input, outputting a course outline corresponding to the course topic. In this embodiment, the course outline includes at least a sequence of chapter titles and a set of key points associated with each chapter title. For example, the platform can obtain the chapter title sequence "Course Introduction and Learning Objectives", "Basic Safety Norms for Entering the Plant", "Common Risks and Prevention Points", "Emergency Response and Reporting Procedures", and "Summary and Self-Test Tips", and generate a set of key points such as "Key Points for Hot Work Approval", "Key Points for Fall Prevention in Work at Heights", and "Key Points for Temporary Power Use and Inspection" under the "Key Points for Emergency Response and Reporting Procedures" chapter; and generate key points such as "Alarm and Information Reporting Path", "Initial Fire Response Steps", and "Emergency Evacuation and Assembly Point Confirmation" under the "Emergency Response and Reporting Procedures" chapter. The platform can also retain references to source tag information in the generated course outline. For example, it can associate the key points of the "four-step fire extinguisher method" with the source tag of "video footage 00:03:10-00:03:40," allowing administrators to locate the corresponding footage segment during subsequent verification. This association method does not require explicit citation text to be displayed externally, but the platform can internally store it as evidence of key outline points in the outline data structure.

[0048] Furthermore, based on the generated course outline, the platform further determines the page-level structure of the courseware and generates a page-by-page explanation script corresponding to the page-level structure. The page-level structure of the courseware is used to describe the page order and page theme of splitting the course outline into several pages of courseware. In this embodiment, the platform can determine the page-level splitting rules based on the sequence of chapter titles and the set of key points, for example, mapping "chapter titles" to separator pages or chapter pages, mapping "key points" to content pages, and aggregating no more than a preset number of key points in each content page to control the information density within the page. For example, the "Basic Safety Specifications for Entering the Factory" chapter can be split into 3 content pages, covering "Labor Protection Equipment and Entry Inspection", "Factory Area Access and Prohibitions", and "Work Permits and On-site Signage" respectively; the "Emergency Response and Reporting Process" chapter can be split into 2 content pages, covering "Alarm and Reporting Process" and "Evacuation and Assembly Points" respectively.

[0049] When generating the page-by-page presentation script, the platform generates a page title, a summary of key points, and presentation text for each page. The page title summarizes the page's theme, the summary lists the core items, and the presentation text forms a verbatim script that can be directly read aloud. For example, on the "Initial Fire Response" page, the summary might include "Confirm the fire source type," "Select the appropriate fire extinguisher," "Pull-in, pull-out, and press-out operation," and "Evacuation and recheck." The presentation text is organized into coherent paragraphs in a corporate training tone and can embed verbal prompts for images or diagrams, such as "Please observe the safety pin position of the fire extinguisher in the diagram on the right," to match the course production method of using PPT and instructor scripts to drive digital human presentations.

[0050] To ensure structural consistency between the page-by-page explanation script and the course outline, the platform performs script standardization on the generated page-by-page explanation script, ensuring that the page titles and summary of key points in the page-by-page explanation script correspond to the chapter titles and key point sets in the course outline. In this embodiment, script standardization may include title normalization, key point consistency verification, and page-level merging and adjustment. Title normalization unifies the naming style of page titles to match the chapter titles, avoiding structural breaks caused by synonyms with different names. Key point consistency verification compares the page summary of key points with the set of key point items in the course outline page by page. If an item not found in the course outline is detected in the page summary, it is marked as a candidate for new key point and a second generation is triggered, or the process reverts to the course outline editing stage. If a key point in the course outline is not covered by any page, page-level merging and adjustment is triggered, inserting the missing key point into the best-matching page or generating a new page.

[0051] For example, if the course outline contains the point "Emergency Assembly Point Confirmation" but the page-by-page explanation script does not cover it, the platform can add this item to the page summary of the "Evacuation and Assembly Point" page and add verbal instructions for assembly point confirmation in the explanation text; if the page-by-page explanation script adds the item "Drill Sign-in and Recording" but it is not included in the course outline, the platform can add this item back as a new point in the "Summary and Self-Test Tips" section of the course outline, or prompt the administrator to confirm it in the course outline editing interface before proceeding to the subsequent PPT generation.

[0052] In this embodiment, after generating the course outline and page-by-page explanation script, the platform presents the results to the administrator in an editable format. The administrator can delete or modify the course outline, and manually revise or trigger intelligent polishing of the page-by-page explanation script to meet the user's course production workflow, allowing for deletion and modification of the outline and intelligent polishing of the script. After receiving the administrator's editing results, the platform uses the updated course outline and updated page-by-page explanation script as the sole valid input for selecting templates and generating PPT courseware files. This ensures that the content of subsequent courseware pages is consistent with the notes in the lecture script, facilitating the subsequent reading of the lecture script from the courseware notes and driving the digital human to narrate and generate the teaching video.

[0053] This embodiment encapsulates course objective information and course context into a unified structure as input and calls a large language model to perform structured generation. This enables the course outline and page-by-page explanation script to be automatically extracted from course reference materials and maintain a consistent, editable, and verifiable page-level structure. This reduces the workload of manually writing outlines and word-by-word lectures, improves the efficiency and consistency of course content generation, and provides a stable script source for subsequent digital human video synthesis driven by courseware annotation scripts.

[0054] In some embodiments, a template identifier is obtained and a courseware template is determined from a template library based on the template identifier. A courseware file is generated according to the updated course syllabus and page-by-page explanation script. The page-by-page explanation script is written page-by-page into the page-level notes information of the courseware file, including: Obtain the template identifier selected by the user, and read the courseware template corresponding to the template identifier from the template library based on the template identifier. The courseware template includes a set of page layouts and a placeholder structure corresponding to the set of page layouts. The updated course outline is mapped to page-level content planning, which includes the page title and the set of key points for each page. Page-level content planning and placeholder structure are used to generate page fill data, and the page fill data is written into the placeholder structure to generate the content of each page, thereby generating the courseware file; The page-by-page explanation script will be written into the page-level notes information of the courseware file. The page-level notes information and the page content of the corresponding page will share the same page identifier, so that the page content of each page in the courseware file can be linked to the corresponding script fragment at the page level.

[0055] Specifically, the platform first obtains the template identifier selected by the user. This template identifier is selected by the administrator in the platform's "Template Selection" interface. The selected template can be either a company-owned template or a template built into the platform. For example, if the administrator selects a company-branded template for "Mandatory Safety Production Training Course for Factory Entry," the template identifier could correspond to the "Company Blue and White Standard Training Template." Based on the template identifier, the platform reads the corresponding courseware template from the template library. The courseware template includes at least a set of page layouts and a placeholder structure corresponding to that set.

[0056] In some examples, the page layout set describes the available page types and layouts in the template, such as cover page, table of contents page, chapter page, summary page, image and text page, and summary page; the placeholder structure describes the positions and attributes of placeholders that can be filled in each page type, such as title placeholders, body summary placeholders, image placeholders, and footer placeholders. When reading the courseware template, the platform can also simultaneously read basic elements related to the company brand in the template, such as color scheme, font, header and footer information, and company logo graphics, to ensure that the generated courseware is consistent with the company's style.

[0057] Furthermore, after reading the courseware template, the platform maps the updated course outline to page-level content planning. The updated course outline comes from the administrator's deletions and modifications to the course outline in the previous embodiment, and the platform uses this outline as the only valid outline version for mapping. Page-level content planning includes at least the page title and set of key points for each page, and reflects the page order and page type selection results.

[0058] For example, the platform maps the "Course Introduction and Learning Objectives" section of the course syllabus to a table of contents or an introductory page, and maps "Basic Safety Norms for Entering the Factory," "Common Risks and Protection Points," and "Emergency Response and Reporting Procedures" to a combination of chapter pages and key point pages. It also generates a page title for each page, such as "Labor Protection and Inspection for Entering the Factory," "Factory Area Access Rules and Prohibitions," "Initial Fire Response Procedures," and "Emergency Evacuation and Assembly Points." Correspondingly, the platform generates a set of key points for each page. For example, the key points set for the "Initial Fire Response Procedures" page might include "Confirming the Fire Source Type," "Selecting the Appropriate Fire Extinguisher," "Plugging, Inserting, and Pressing Operation," and "Evacuation and Review." During this mapping process, the platform can select appropriate page types for different pages based on the template's page layout set. For example, it might select a chapter page layout for the beginning of a chapter, a key point page layout for pages with many items, and a graphic page layout for pages containing illustrations, to maintain a stable layout for the generated courseware.

[0059] Furthermore, based on page-level content planning and placeholder structure, the platform generates page fill data and writes it into the placeholder structure to generate page content, thereby generating courseware files. The page fill data describes the content objects to be written into the placeholders and their attributes, including at least title text, key point text, and image references corresponding to the media content description information. In this embodiment, for title placeholders, the platform writes the page title into the corresponding placeholder; for main text key point placeholders, the platform writes the set of page key points in order of entries, and can perform entry grouping or generate continuation pages when the number of entries exceeds the placeholder capacity; for image placeholders, the platform can select the image object that best matches the key point of the page based on the source tag information and media content description information in the course context, for example, associating the "Emergency Evacuation and Assembly Point" page with the "Evacuation Route Map" image in the materials, or associating the "Four-Step Fire Extinguisher Method" page with the "Fire Extinguisher Operation Diagram" in the materials.

[0060] In some examples, when the course reference material is a PowerPoint presentation, the platform can directly reuse image objects from the presentation and create references by page identifier; when the course reference material is a video, the platform can generate image objects based on keyframe extraction results and create references. After completing the writing of each placeholder, the platform generates a courseware file containing multiple pages of content and records the courseware file as a courseware artifact associated with this course generation task.

[0061] While generating the courseware file, the platform writes the page-by-page explanation script into the page-level annotation information of the courseware file. The page-by-page explanation script comes from the generation and editing results of the previous embodiment. The platform writes the explanation text corresponding to each page as a script fragment into the page-level annotation information of that page, so that the page-level annotation information and the page content of the corresponding page share the same page identifier.

[0062] For example, in some examples, the platform includes the corresponding explanatory text paragraphs in the notes of the "Initial Fire Response Steps" page and the "Factory Area Access Rules and Prohibitions" page, maintaining semantic and sequential consistency between the script fragments in the notes and the set of key points on the page. This provides a direct data foundation for the subsequent video generation process, which involves reading the presentation script from the PPT notes, associating image addresses and presentation fragments page by page, and driving the digital human's narration. The platform also provides preview and download access after the courseware file is successfully generated, and the preview interface allows administrators to perform secondary proofreading and minor editing of the page content and note scripts to ensure that the final courseware meets the needs of corporate lecturers in creating courseware for teaching activities.

[0063] In an optional implementation, when writing page-level notes, the platform can also simultaneously write structured auxiliary tags for page-level scripts for subsequent segmentation and duration estimation during video generation. For example, segmentation tags can be inserted within the notes according to paragraph boundaries, or the index range of page key points corresponding to each explanatory text segment can be recorded. However, these auxiliary tags do not change the binding relationship between the page identifier and the note script; the same page identifier is still used as the unified key for subsequent material association and video compositing.

[0064] This embodiment reads a courseware template containing placeholder structures based on a template identifier, maps the updated course outline to page-level content planning, and then performs placeholder filling to generate courseware files, ensuring that the courseware page content can be stably formatted according to the template. At the same time, the page-by-page explanation script is written into page-level notes and shares the same page identifier with the page content, thereby establishing a page-level correspondence between page content and script fragments within the courseware file. This reduces the workload of manual typesetting and page-by-page writing of lecture notes, improves courseware generation efficiency, and provides a consistent data interface for subsequent generation and temporal synthesis of digital human video fragments driven by the notes script.

[0065] In some embodiments, paginated rendering is performed on the courseware file to obtain courseware images for each page, and script fragments corresponding to each page are read from page-level annotation information to establish a page-level association between courseware images and script fragments, including: Perform paginated export processing on the courseware file, render the content of each page of the courseware file into a courseware image corresponding to the page identifier, and generate accessible image address information for each courseware image; Read the script fragments corresponding to each page identifier from the page-level notes information of the courseware file, and perform segmentation and normalization processing on the script fragments to generate page-level script units; Page-level associations are established based on page identifiers. These associations are used to indicate the one-to-one correspondence between the image address information of the courseware image corresponding to each page identifier and the page-level script unit. Page-level relationships are written into the course generation context or index data associated with the courseware file, so that subsequent speech segments and digital human video segments can be generated based on page-level relationships and temporal synthesis can be performed.

[0066] Specifically, the platform first performs pagination export processing on the courseware file, rendering the content of each page into a courseware image corresponding to the page identifier, and generating accessible image address information for each courseware image. Upon receiving the "Generate AI Video" instruction, the platform reads the courseware file from the product storage of the course generation task, determines the number of pages in the courseware file, and generates a set of page identifiers. Page identifiers can be generated using a combination of "courseware file identifier + page number" to ensure that page identifiers are globally unique within the same course generation task.

[0067] The platform then triggers a paginated export process, rendering each page's content as a static image file. The resolution and format of these image files can be configured according to platform presets to meet the video synthesis clarity requirements. After paginated export is complete, the platform uploads each image file to the object storage service and returns the image address information for each file. For example, for the "Mandatory Course for Safe Production Upon Entry" courseware, the platform generates page identifiers P1 to P12 and obtains 12 image address information entries for each. These image address information can be the accessible path or access token returned by the object storage service.

[0068] Furthermore, after obtaining the courseware image and image address information, the platform reads the script fragments corresponding to each page identifier from the page-level annotation information of the courseware file, and performs segmentation and normalization processing on the script fragments to generate page-level script units. For example, the platform traverses the page-level annotation information of the courseware file by page identifier, and reads the explanatory text in the annotation of each page as a script fragment. Since the annotation script may contain multiple explanatory contents, voice prompts, or inconsistent punctuation, the platform performs segmentation and normalization processing on the script fragments to form page-level script units. Segmentation and normalization processing may include splitting the script by natural paragraphs, periods, or preset delimiters, removing empty paragraphs and repeated whitespace characters, unifying the expression of voice prompts, and maintaining the script order consistent with the order of page key points. For example, on the "Initial Fire Handling Steps" page corresponding to page identifier P6, the annotation script fragment may contain continuous explanatory text of four key points. The platform normalizes it into a page-level script unit, which can further contain several script sub-segments, each script sub-segment corresponding to the voice content of one key point, thereby facilitating subsequent speech synthesis by segment or by page.

[0069] Furthermore, after preparing the courseware images and page-level script units, the platform establishes page-level associations based on page identifiers. These page-level associations indicate the one-to-one correspondence between the image address information of the courseware image corresponding to each page identifier and the page-level script unit. For example, the platform uses the page identifier as the association key to construct an association record for each page. Each association record includes at least the page identifier, image address information, and page-level script unit. For instance, association record R6 may include page identifier P6, image address information URL6, and page-level script unit S6; association record R7 may include page identifier P7, image address information URL7, and page-level script unit S7, thus forming a stable mapping from page identifier to image + script. If a courseware page has an anomaly of no notes or empty notes, the platform can set the page-level script unit corresponding to that page identifier to empty and mark it as pending completion, or return to the courseware editing interface to prompt the administrator to supplement the explanation script for that page, ensuring the completeness of subsequent digital human broadcast content.

[0070] Furthermore, after establishing page-level relationships, the platform writes these relationships into the course generation context or into the index data associated with the courseware files, so that subsequent audio clips and digital human video clips can be generated based on these page-level relationships and temporal synthesis can be performed. For example, the platform can add a "page-level material index" field to the course generation context of this course generation task, and write the association records of each page into this field in page order; or it can store the page-level relationships separately as an index data object associated with the courseware file identifier, and generate an index identifier for this index data object.

[0071] Subsequently, when calling the AI ​​platform's timbre model list interface and digital human model list interface and initiating video generation and editing service calls, the platform can extract image address information and page-level script units page by page based on the page-level material index. First, it generates audio segments associated with page identifiers, then generates digital human video segments associated with page identifiers. Finally, it inputs the images, audio, and digital human videos under the same page identifier as the composite materials of the same page into the video editing and compositing process, thereby realizing the implementation route of compositing digital human videos, audio segments, and PPT page images into a video.

[0072] This embodiment obtains image address information by performing paginated export and upload of courseware files. At the same time, it reads and organizes page-level script units from page-level notes, establishes a one-to-one correspondence between courseware images and script fragments using page identifiers, and writes them into the course generation context or index data. This unifies the static courseware content and the spoken script into a searchable set of page-level materials, reducing the reliance on manual page-by-page organization and alignment in the video generation process, and improving the automation and processing efficiency of subsequent speech synthesis, digital human video generation, and temporal synthesis processes.

[0073] In some embodiments, obtaining digital human configuration parameters and timbre configuration parameters, generating corresponding speech segments based on script segments, and performing lip-syncing on the script segments based on the digital human configuration parameters to generate digital human video segments, including: Obtain the digital human configuration parameters and timbre configuration parameters selected by the user. The digital human configuration parameters include the digital human image identifier and the presentation parameters associated with the digital human image identifier. The presentation parameters include page position parameters and size parameters. The timbre configuration parameters include the timbre model identifier. The script calls a text-to-speech model to generate speech segments, and generates speech tag information associated with page identifiers for the speech segments; Based on the digital human image identifier, the digital human generation model is called to generate digital human video segments corresponding to the script segments. Based on the speech segments, the lip-sync processing of the digital human video segments is performed to align the lip-sync sequence of the digital human video segments with the speech timing of the speech segments. Voice segments, voice tag information, and digital human video segments are associated with page identifiers and stored for subsequent time-series synthesis according to page order.

[0074] Specifically, the platform first obtains the configuration parameters of the digital human and the voice configuration parameters selected by the user. On the "Digital Human Selection" interface, the platform displays a list of digital human models and a list of voice models provided by the cloud-based AI platform. The administrator selects the digital human image and voice model based on the course type and company style. For example, for "Mandatory Course for Safe Production Entry," the administrator selects the "Standard Training Female Instructor" digital human image and the "Mandarin Female Voice A" voice model. The platform writes the digital human image identifier of the selected digital human image and the voice model identifier of the selected voice model into the configuration data for this course generation task, and further receives the digital human's presentation parameters, which include at least page position parameters and size parameters. For example, the page position parameter may indicate that the digital human's overlay area in the courseware image is located in the lower right corner, and the size parameter may indicate that the proportion of the digital human's screen width is a preset proportion, thus ensuring that the digital human does not obscure key areas of the page.

[0075] Furthermore, after obtaining the digital human configuration parameters and timbre configuration parameters, the platform calls a text-to-speech model based on the script fragment to generate a speech fragment, and generates speech tag information associated with page identifiers for the speech fragment. The script fragment comes from the page-level association relationship established in the previous embodiment. The platform reads page-level script units as text input for the speech to be synthesized according to page identifiers. For example, the platform reads script fragment S6 for page identifier P6 and script fragment S7 for page identifier P7. The platform calls a cloud-based text-to-speech service or a locally deployed text-to-speech model based on the timbre model identifier to generate a speech fragment.

[0076] In some examples, to ensure that audio segments can be accurately referenced in subsequent synthesis processes, the platform generates speech tagging information for each audio segment. This speech tagging information includes at least a page identifier, a speech segment identifier, and speech duration information; it may also include basic parameters such as the speech sampling rate and audio format if necessary. For example, speech tagging information A6 includes a page identifier P6 and a speech segment identifier Audio6, and records the duration of Audio6 as 18.5 seconds; speech tagging information A7 includes a page identifier P7 and a speech segment identifier Audio7, and records the duration of Audio7 as 16.2 seconds. The platform can upload the audio segments to an object storage service and obtain the speech address information, or save them in the platform's media resource library for subsequent video editing services to reference.

[0077] Furthermore, after generating the audio segments, the platform uses the digital human avatar identifier to call the digital human generation model to generate digital human video segments corresponding to the script segments. It then performs lip-syncing processing on the digital human video segments based on the audio segments to ensure that the lip-sync sequence of the digital human video segments aligns with the audio sequence of the audio segments. For example, the platform encapsulates the digital human avatar identifier, script segment identifier or script text content, audio segment address information, and presentation parameters into a digital human generation request, which calls the cloud-based digital human service to generate the corresponding page's digital human video segments. The digital human generation request can be made page-by-page to ensure that the digital human video segments on each page strictly correspond to the script and audio on that page.

[0078] Lip alignment can be performed by the cloud-based digital human service during the generation process, or the audio segment can be input as synchronization on the platform side and aligned by the LipSync module. For example, for page identifier P6, the platform uses Audio6 as the reference audio input for lip alignment to generate a digital human video segment Video6, synchronizing the lip movements of the digital human in Video6 with the rhythm of the audio in Audio6; for page identifier P7, the platform uses Audio7 as the reference input to generate Video7. To maintain consistency between different pages, the platform reuses the same digital human identifier and the same set of page position and size parameters when generating the digital human for each page, ensuring a stable superposition of the digital human and facilitating direct overlay onto the courseware image background during subsequent compositing.

[0079] In an optional implementation of this embodiment, the platform can also perform segment-level splicing or segment-level alignment of the audio segment and the digital human video segment based on the segmentation and normalization results of the script segment. For example, if the script segment S6 corresponding to page identifier P6 contains four script sub-segments, the platform can choose to first generate an audio sub-segment and a corresponding lip-sync video sub-segment for each script sub-segment, and then splice them in the order of the sub-segments to form Video6 and Audio6, thereby improving the stability of long scripts in lip-sync and pause rhythm. However, regardless of whether page-level or segment-level generation is used, the platform maintains the one-to-one association between the final product and the page identifier, so that subsequent temporal synthesis can be performed according to the page order.

[0080] Furthermore, after generating the audio segments and digital human video segments, the platform associates and stores the audio segments, audio tag information, and digital human video segments with page identifiers for subsequent sequential synthesis according to page order. For example, the platform adds a "page-level audio and video index" field to the media index of the course generation task, recording the page identifier, audio segment identifier or audio address information, audio duration information, digital human video segment identifier or video address information, and presentation parameters for each page.

[0081] For example, page-level audio / video index I6 records Audio6 and Video6 corresponding to P6 and their presentation parameters, and page-level audio / video index I7 records Audio7 and Video7 corresponding to P7 and their presentation parameters. When the platform subsequently calls the cloud-based video editing and compositing service, it can directly retrieve and use the corresponding courseware image address information, audio address information, and digital human video address information by page identifier, thereby realizing the link connection of compositing digital human video, audio clips, and PPT page images into video.

[0082] This embodiment acquires and solidifies digital human configuration parameters and voice configuration parameters on the platform side, generates voice segments associated with page identifiers based on page-level script fragments, and uses the voice segments to drive the lip-syncing generation of digital human video segments. This ensures that the voice and digital human visuals are synchronized and reusable at the page level, reducing the involvement of manual recording and video shooting and editing, improving the automation and consistency of digital human teaching material generation, and providing a searchable set of page-level audio and video materials for subsequent sequential synthesis of teaching video courseware.

[0083] In some embodiments, based on page-level association, the courseware images, audio clips, and digital human video clips from each page are sequentially synthesized according to page order to output teaching video courseware, including: Based on page-level association, determine the set of courseware images, audio clips, and digital human video clips corresponding to each page identifier, and determine the page sequence information corresponding to each page identifier; For any page identifier, the playback period of the current page is determined based on the duration information of the audio segment. During the playback period, the courseware image is used as the background image, the digital human video segment is superimposed on the background image according to the digital human configuration parameters, and the audio segment is used as the audio track of the current page. The composite segments corresponding to each page are sequentially spliced ​​together to generate the target video timeline according to the page sequence information, and transition segments or transition parameters are inserted between the composite segments of adjacent pages to generate teaching video courseware. Write the lecture video courseware to the storage medium and generate a courseware output identifier associated with the lecture video courseware for the platform to preview or download.

[0084] Specifically, the platform first determines the set of courseware images, audio clips, and digital human video clips corresponding to each page identifier based on page-level association relationships, and then determines the page order information corresponding to each page identifier. The page-level association relationships are written into the course generation context or index data by the aforementioned embodiments, and at least include a one-to-one correspondence between page identifiers, image address information of courseware images, and page-level script units. After generating audio clips and digital human video clips, the platform also associates and stores audio clip identifiers, audio duration information, and digital human video clip identifiers with page identifiers. Therefore, the platform can aggregate a set of materials containing "courseware image address information + audio clip address information + digital human video clip address information" under the same page identifier, and determine the page order information based on the page order or page identifier sequence number when the courseware file was generated.

[0085] For example, in some examples, the courseware "Required Course for Factory Safety" has 12 pages. The platform determines the page order information as P1 to P12 and constructs a material set M1 to M12 for each page. Among them, M6 contains URL6, Audio6, and Video6, and M7 contains URL7, Audio7, and Video7. This page order information and the material set together constitute the input basis for subsequent time-series synthesis.

[0086] Furthermore, after obtaining the page order and material set, the platform determines the playback period for any given page identifier based on the duration of the audio segment. During this playback period, the courseware image is used as the background image, the digital human video segment is overlaid onto the background image according to the digital human configuration parameters, and the audio segment is used as the audio track for the current page. For example, the platform reads the audio duration information for each page and uses it as the baseline length for the playback period of that page. For instance, if Audio6 has a duration of 18.5 seconds, the platform sets the playback period for page identifier P6 to 18.5 seconds, or adds a preset buffer duration to this to ensure stable page transitions. During this playback period, the platform displays the courseware image corresponding to URL6 as a fixed background image and overlays the digital human image from Video6 onto the specified position of the background image according to the digital human configuration parameters.

[0087] The digital human configuration parameters have been fixed in the aforementioned embodiments, including page position and size parameters. For example, the digital human is placed in the lower right corner with a size that is a preset proportion of the screen width, thereby avoiding obscuring key areas of the page. The platform also uses Audio6 as the audio track for this page to write the audio channel into the synthesized segment, ensuring that the digital human's lip movements are consistent with the audio playback sequence. The same method is used to generate synthesized segments page by page for other pages, ensuring that the synthesized segments on each page have consistent rules in terms of screen composition and audio input.

[0088] In an optional implementation of this embodiment, the platform may further add a subtitle track or a key point prompt track within each page of the synthesized segment. This subtitle track can be generated by triggering script or audio segments. For example, the platform may perform speech-to-text time alignment processing on Audio6 to obtain a subtitle timeline and overlay the subtitles onto the bottom area of ​​the synthesized segment; or it may extract the page's key points into key point prompts that appear according to time periods. This optional implementation does not change the page-level association relationship or the main synthesis chain; it only serves as an additional track output for the synthesized segment.

[0089] Furthermore, after generating the composite segments for each page, the platform sequentially splices the corresponding composite segments from each page according to the page order information to generate the target video timeline. Transition segments or transition parameters are inserted between composite segments on adjacent pages to generate the teaching video courseware. For example, the platform sequentially splices the composite segments according to the page order P1 to P12 to form a continuous video timeline, and inserts preset transitions between adjacent segments, such as fade-in / fade-out, sliding transitions, etc.; or only transition parameters are inserted, and the cloud-based intelligent media service generates the transition effect according to the parameters during rendering output. For example, a 0.4-second fade-in / fade-out transition is inserted between P5 and P6 to avoid abrupt cuts. For operations performed by the administrator in the video editing interface, such as deleting or copying a page, the platform can update the page order information and timeline splicing order accordingly, triggering timeline generation and rendering output again, thereby completing the editable course production process.

[0090] Furthermore, after constructing the target video timeline, the platform writes the lecture video courseware to the storage medium and generates a courseware output identifier associated with the courseware for platform preview or download. Specifically, the platform calls the cloud-based video editing and export service to encode the target video timeline, obtaining the final video file, and uploads the final video file to object storage or the platform's media resource library. The platform generates a courseware output identifier for this video file, which may include a course generation task identifier, a video version marker, and access address information, used to enable "preview," "download," and subsequent course publishing and learning task distribution within the platform. For example, administrators can click play on the "preview" interface to watch the lecture video courseware presented by the digital human. If adjustments to a page's explanation text or transition effects are needed, they can return to the upstream steps to modify the notes script or edit the timeline and regenerate the output version.

[0091] This embodiment aggregates courseware images, audio segments, and digital human video segments based on page-level association relationships. It uses audio duration to drive the determination of page-level playback time periods and overlays the materials for synthesis. Then, it splices the materials according to page order and inserts transitions to generate the target video timeline. This achieves automated synthesis and output from PPT pagination images and annotation scripts to teaching video courseware, reducing manual editing and multi-tool alignment operations, improving the efficiency and consistency of teaching video generation, and supporting platform-side preview, download, and subsequent iterative generation.

[0092] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0093] Figure 2 This is a schematic diagram of the structure of the digital human course generation device based on a large model provided in an embodiment of this application. Figure 2 As shown, the digital human course generation device based on a large model includes: The acquisition module 201 is used to acquire a course generation request and perform content parsing on the course reference materials to generate a course context. The course generation request includes course objective information and one or more of the course reference materials. The generation module 202 is used to generate a course outline based on the course objective information and course context by calling the large language model, and to generate a page-by-page explanation script corresponding to the course outline; it also receives the user's editing results on the course outline and page-by-page explanation script and updates the course outline and page-by-page explanation script. Module 203 is used to obtain the template identifier and determine the courseware template from the template library based on the template identifier. It generates a courseware file according to the updated course outline and page-by-page explanation script, and writes the page-by-page explanation script into the page-level remarks information of the courseware file. The rendering module 204 is used to perform paginated rendering on the courseware file to obtain the courseware images for each page, and to read the script fragments corresponding to each page from the page-level notes information to establish the page-level association between the courseware images and the script fragments; The configuration module 205 is used to obtain digital human configuration parameters and timbre configuration parameters, generate corresponding speech segments based on script segments, and perform lip-syncing on script segments based on digital human configuration parameters to generate digital human video segments. The synthesis module 206 is used to synthesize the courseware images, audio segments and digital human video segments of each page in a time sequence according to the page-level association relationship, and output the teaching video courseware.

[0094] In some embodiments, Figure 2The acquisition module 201 receives course reference materials uploaded or specified by the user and determines the material type and structure information of the course reference materials; it performs content extraction processing on the course reference materials to obtain the text content, page key information, and media content description information corresponding to the course reference materials; it performs regularization and structuring processing on the text content, page key information, and media content description information to generate structured material content, wherein the structured material content includes a set of thematic key points and source tag information associated with the set of thematic key points; it constructs a course context based on the structured material content and course objective information, the course context including a set of context fragments used by the large language model to generate the course outline and structured index information corresponding to the set of context fragments.

[0095] In some embodiments, Figure 2 The generation module 202 encapsulates the course objective information and course context into the generation input of the large language model. The generation input includes the course title, writing style, specific course objectives, and a set of context fragments corresponding to the course reference materials. The large language model is invoked to perform structured generation on the generation input, outputting a course outline corresponding to the course theme. The course outline includes a sequence of chapter titles and a set of key points associated with each chapter title. Based on the course outline, the page-level structure of the courseware is determined, and a page-by-page explanation script corresponding to the page-level structure of the courseware is generated. The page-by-page explanation script includes a page title, a summary of key points for each page, and explanation text for each page. The generated page-by-page explanation script is subjected to script normalization processing to ensure that the page titles and page summaries of the page-by-page explanation script maintain a correspondence with the chapter titles and key point sets in the course outline.

[0096] In some embodiments, Figure 2 The determination module 203 obtains the template identifier selected by the user, and reads the courseware template corresponding to the template identifier from the template library. The courseware template includes a set of page layouts and a placeholder structure corresponding to the set of page layouts. The updated course outline is mapped to a page-level content plan, which includes the page title and a set of page highlights for each page. Page fill data is generated based on the page-level content plan and the placeholder structure, and the page fill data is written into the placeholder structure to generate the page content for each page, thereby generating the courseware file. The page-by-page explanation script is written into the page-level notes information of the courseware file, where the page-level notes information and the page content of the corresponding page share the same page identifier, so that the page content of each page in the courseware file and the corresponding script fragment are established in a page-level correspondence.

[0097] In some embodiments, Figure 2The rendering module 204 performs pagination export processing on the courseware file, rendering the content of each page of the courseware file into a courseware image corresponding to the page identifier, and generating accessible image address information for each courseware image; it reads the script fragments corresponding to each page identifier from the page-level comments information of the courseware file, and performs segmentation and normalization processing on the script fragments to generate page-level script units; it establishes page-level association relationships based on the page identifiers, which are used to indicate the one-to-one correspondence between the image address information of the courseware image corresponding to each page identifier and the page-level script units; and it writes the page-level association relationships into the course generation context or the index data associated with the courseware file, so that subsequent generation of audio segments and digital human video segments based on the page-level association relationships and the execution of temporal synthesis can be performed.

[0098] In some embodiments, Figure 2 The configuration module 205 obtains the digital human configuration parameters and voice configuration parameters selected by the user. The digital human configuration parameters include a digital human image identifier and presentation parameters associated with the digital human image identifier. The presentation parameters include page position parameters and size parameters. The voice configuration parameters include a voice model identifier. Based on the script fragment, the module calls a text-to-speech model to generate a speech fragment and generates speech mark information associated with the page identifier for the speech fragment. Based on the digital human image identifier, the module calls a digital human generation model to generate a digital human video fragment corresponding to the script fragment. Based on the speech fragment, the module performs lip-syncing processing on the digital human video fragment to align the lip-sync sequence of the digital human video fragment with the speech timing of the speech fragment. The module stores the speech fragment, speech mark information, and digital human video fragment associated with the page identifier for subsequent timing synthesis according to page order.

[0099] In some embodiments, Figure 2 The synthesis module 206 determines the set of courseware images, audio segments, and digital human video segments corresponding to each page identifier based on page-level association relationships, and determines the page sequence information corresponding to each page identifier; for any page identifier, it determines the playback period of the current page based on the duration information of the audio segment, and within the playback period, the courseware image is used as the background image, the digital human video segment is superimposed on the background image according to the digital human configuration parameters, and the audio segment is used as the audio track of the current page; according to the page sequence information, the synthesis segments corresponding to each page are sequentially spliced ​​to generate the target video timeline, and transition segments or transition parameters are inserted between the synthesis segments of adjacent pages to generate the teaching video courseware; the teaching video courseware is written to the storage medium, and a courseware output identifier associated with the teaching video courseware is generated for platform preview or download.

[0100] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0101] Figure 3 This is a schematic diagram of the electronic device 3 provided in an embodiment of this application. Figure 3 As shown, the electronic device 3 of this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, it implements the steps in the various method embodiments described above. Alternatively, when the processor 301 executes the computer program 303, it implements the functions of each module / unit in the various device embodiments described above.

[0102] Electronic device 3 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 3 may include, but is not limited to, processor 301 and memory 302. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown, or different components.

[0103] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0104] The memory 302 can be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. The memory 302 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 3. The memory 302 can also include both internal and external storage units of the electronic device 3. The memory 302 is used to store computer programs and other programs and data required by the electronic device.

[0105] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0106] If integrated modules / units are implemented as software functional units and sold or used as independent products, they can be stored in a readable storage medium (e.g., a computer-readable storage medium). Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which may be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0107] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for generating digital human courses based on a large model, characterized in that, include: Obtain a course generation request and perform content parsing on the course reference materials to generate a course context, wherein the course generation request includes course objective information and one or more of the course reference materials; Based on the course objective information and course context, the system calls the large language model to generate a course outline and generates a page-by-page explanation script corresponding to the course outline; it receives the user's editing results on the course outline and page-by-page explanation script and updates the course outline and page-by-page explanation script. Obtain the template identifier and determine the courseware template from the template library based on the template identifier. Generate the courseware file according to the updated course outline and page-by-page explanation script, and write the page-by-page explanation script into the page-level notes information of the courseware file. Perform paginated rendering on the courseware file to obtain the courseware images for each page, and read the script fragments corresponding to each page from the page-level notes information to establish the page-level association between the courseware images and the script fragments; Obtain digital human configuration parameters and voice configuration parameters, generate corresponding speech segments based on script segments, and perform lip-syncing on script segments based on digital human configuration parameters to generate digital human video segments; Based on the page-level association, the courseware images, audio clips, and digital human video clips on each page are sequentially synthesized according to page order to output the teaching video courseware.

2. The method according to claim 1, characterized in that, The process of obtaining the course generation request and performing content parsing on the course reference materials to generate course context includes: Receive course reference materials uploaded or specified by users, and determine the material type and structure information of the course reference materials; The course reference materials are subjected to content extraction processing to obtain the text content, page key information and media content description information corresponding to the course reference materials; The text content, page key information, and media content description information are normalized and structured to generate structured material content, wherein the structured material content includes a set of key themes and source tag information associated with the set of key themes; A course context is constructed based on the structured material content and course objective information. The course context includes a set of context fragments for generating a course outline using a large language model, as well as structured index information corresponding to the set of context fragments.

3. The method according to claim 1, characterized in that, The process of generating a course syllabus based on the course objective information and course context by calling a large language model, and generating a page-by-page explanation script corresponding to the course syllabus, includes: The course objective information and course context are encapsulated into the generation input of a large language model. The generation input includes the course title, writing style, specific course objectives, and a set of context fragments corresponding to the course reference materials. The large language model is invoked to perform structured generation on the generated input, and a course outline corresponding to the course topic is output. The course outline includes a sequence of chapter titles and a set of key points associated with each chapter title. Based on the course outline, the page-level structure of the courseware is determined, and a page-by-page explanation script corresponding to the page-level structure of the courseware is generated. The page-by-page explanation script includes the page title, the summary of the key points of the page, and the explanation text corresponding to each page. The generated page-by-page explanation script is processed to ensure that the page titles and summary of key points in the page-by-page explanation script correspond to the chapter titles and key point items in the course outline.

4. The method according to claim 1, characterized in that, The process involves obtaining a template identifier and determining a courseware template from the template library based on that identifier, generating a courseware file according to the updated course outline and page-by-page explanation script, and writing the page-by-page explanation script into the page-level notes information of the courseware file, including: Obtain the template identifier selected by the user, and read the courseware template corresponding to the template identifier from the template library based on the template identifier. The courseware template includes a set of page layouts and a placeholder structure corresponding to the set of page layouts. The updated course outline is mapped to a page-level content plan, which includes the page title and the set of key points for each page. Page filling data is generated based on the page-level content planning and placeholder structure, and the page filling data is written into the placeholder structure to generate the content of each page, thereby generating the courseware file; The page-by-page explanation script is written into the page-level notes information of the courseware file, wherein the page-level notes information and the page content of the corresponding page share the same page identifier, so as to establish a page-level correspondence between the page content of each page in the courseware file and the corresponding script fragment.

5. The method according to claim 1, characterized in that, The process of performing paginated rendering on the courseware file to obtain courseware images for each page, and reading the script fragments corresponding to each page from the page-level annotation information to establish a page-level association between the courseware images and the script fragments includes: Perform paginated export processing on the courseware file, render the content of each page of the courseware file into a courseware image corresponding to the page identifier, and generate accessible image address information for each courseware image; Read the script fragments corresponding to each page identifier from the page-level notes information of the courseware file, and perform segmentation and normalization processing on the script fragments to generate page-level script units; A page-level association relationship is established based on the page identifier. The page-level association relationship is used to indicate the one-to-one correspondence between the image address information of the courseware image corresponding to each page identifier and the page-level script unit. The page-level association is written into the course generation context or the index data associated with the courseware file, so that subsequent speech segments and digital human video segments can be generated based on the page-level association and time-series synthesis can be performed.

6. The method according to claim 1, characterized in that, The steps of obtaining digital human configuration parameters and timbre configuration parameters, generating corresponding speech segments based on script segments, and performing lip-syncing on the script segments based on the digital human configuration parameters to generate digital human video segments include: Obtain the digital human configuration parameters and timbre configuration parameters selected by the user. The digital human configuration parameters include the digital human image identifier and the presentation parameters associated with the digital human image identifier. The presentation parameters include page position parameters and size parameters. The timbre configuration parameters include the timbre model identifier. Based on the script fragment, a text-to-speech model is invoked to generate a speech fragment, and speech tag information associated with the page identifier is generated for the speech fragment; Based on the digital human image identifier, the digital human generation model is invoked to generate a digital human video segment corresponding to the script segment, and lip-syncing processing is performed on the digital human video segment based on the speech segment so that the lip-sync sequence of the digital human video segment is aligned with the speech timing of the speech segment. The speech segments, speech tag information, and digital human video segments are associated with the page identifier and stored for subsequent time-series synthesis according to page order.

7. The method according to claim 1, characterized in that, The process involves sequentially synthesizing the courseware images, audio clips, and digital human video clips from each page according to the page-level association relationship, and outputting the teaching video courseware, including: Based on the page-level association relationship, determine the set of courseware images, audio clips, and digital human video clips corresponding to each page identifier, and determine the page sequence information corresponding to each page identifier; For any page identifier, the playback period of the current page is determined based on the duration information of the audio segment, and the courseware image is used as the background image during the playback period. The digital human video segment is superimposed on the background image according to the digital human configuration parameters, and the audio segment is used as the audio track of the current page. According to the page sequence information, the corresponding composite segments of each page are sequentially spliced ​​to generate the target video timeline, and transition segments or transition parameters are inserted between the composite segments of adjacent pages to generate teaching video courseware; The teaching video courseware is written to a storage medium, and a courseware output identifier associated with the teaching video courseware is generated for the platform to preview or download.

8. A digital human course generation device based on a large model, characterized in that, include: The acquisition module is used to acquire a course generation request and perform content parsing on the course reference materials to generate a course context. The course generation request includes one or more of the course target information and course reference materials. The generation module is used to generate a course outline based on the course objective information and course context by calling a large language model, and to generate a page-by-page explanation script corresponding to the course outline; it also receives the user's editing results on the course outline and page-by-page explanation script and updates the course outline and page-by-page explanation script. The determination module is used to obtain the template identifier and determine the courseware template from the template library based on the template identifier. It generates courseware files according to the updated course outline and page-by-page explanation script, and writes the page-by-page explanation script into the page-level remarks information of the courseware file. The rendering module is used to perform paginated rendering of the courseware file to obtain the courseware images for each page, and to read the script fragments corresponding to each page from the page-level notes information to establish the page-level association between the courseware images and the script fragments; The configuration module is used to obtain digital human configuration parameters and timbre configuration parameters, generate corresponding speech segments based on script segments, and perform lip-syncing on script segments based on digital human configuration parameters to generate digital human video segments. The synthesis module is used to synthesize the courseware images, audio segments, and digital human video segments of each page in a time sequence according to the page-level association relationship, and output the teaching video courseware.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.