Method for automatically generating lecturer video based on AI speech synthesis and animation driving
By improving content relationship graph parsing and parallel processing techniques, combined with efficient speech synthesis and animation-driven methods, the problems of low efficiency and insufficient expressiveness in educational video production have been solved, and efficient and natural automatic generation of lecturer videos has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN XUEYOU TECHNOLOGY CO LTD
- Filing Date
- 2025-07-17
- Publication Date
- 2026-05-01
AI Technical Summary
Existing educational video production technologies are inefficient, lack the expressive power of digital humans, cannot achieve rapid processing and real-time generation of large-scale content, and lack in-depth understanding of PPT or text content and emotional expression in speech.
A content relationship graph is constructed using an improved interior point method. The incremental shortest path algorithm and the fully dynamic parallel single-link clustering algorithm are applied for structured parsing. Combined with CosyVoice speech synthesis technology and museTalk technology, efficient speech data streams and action and expression commands are generated. A multi-layer attention mechanism is used to establish a deep mapping relationship between semantic content and action and expression, enabling parallel rendering and video generation.
It enables real-time incremental parsing of large PPT documents, improves the accuracy and efficiency of content hierarchy recognition, enhances the speed and quality of speech synthesis, generates actions and expressions that are highly matched with the content, supports near real-time video generation capabilities, and meets the needs of rapid processing of large-scale content.
Smart Images

Figure CN120897102B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence, computer graphics, and education technology, and in particular to a method for automatically generating lecturer videos based on AI speech synthesis and animation-driven technology. Background Technology
[0002] The education and training and online course industry has been constantly seeking more efficient and engaging ways to create content. Traditional educational video production typically requires professional instructors to record, edit, and process the videos, a process that is time-consuming, costly, and makes it difficult to achieve rapid content iteration and updates.
[0003] Currently, there are two main methods for producing educational videos on the market: one is live recording, which requires professional photography and audio equipment and a post-production team; the other is simple screen recording with narration of PPT slides, which, although low-cost, has poor interactivity and appeal. These methods cannot meet the rapidly changing demands of educational content and are inefficient in production.
[0004] More advanced technologies typically employ basic AI speech synthesis and simple digital human technology to convert text into speech and drive a pre-set digital human avatar to perform simple lip-syncing. While these technologies achieve initial automation, they still have significant shortcomings in content understanding, speech expressiveness, and the naturalness of facial expressions and movements.
[0005] These existing technologies have several key problems: First, the lack of in-depth understanding and structured analysis of PPT or text content results in the generated speech lacking corresponding intonation and emotional expression; second, the digital human's movements and expressions lack semantic connection with the content being presented, resulting in a stiff and mechanical performance; finally, the entire system is inefficient when processing large amounts of content, making it difficult to achieve real-time or near-real-time video generation, and thus unable to meet the needs of large-scale content production. Summary of the Invention
[0006] The purpose of this invention is to provide a method for automatically generating lecturer videos based on AI speech synthesis and animation-driven technology, which solves the problems of low efficiency in educational video production, insufficient expressiveness of digital humans, and limited large-scale content processing capabilities in the existing technology.
[0007] To achieve the above objectives, the technical solution provided by the present invention is as follows:
[0008] A method for automatically generating lecturer videos based on AI speech synthesis and animation-driven methods includes the following steps:
[0009] Based on the content features of user-uploaded PPT files or text scripts, a content relationship graph is constructed using an improved interior point method. An incremental shortest path algorithm is applied to identify the hierarchical relationships between titles, paragraphs, and key points, and the output is structured data containing semantic structure, key content, and sentiment.
[0010] Based on the structured data, a fully dynamic parallel single-link clustering algorithm is applied to group the content of the structured data according to semantic similarity. The data is then converted into spoken narrative text through a preset semantic understanding model and marked with pauses, intonation changes, and emphasis. An enhanced script with full expressive markings is then output.
[0011] Based on the enhanced script, and utilizing CosyVoice speech synthesis technology, through... The low-rank approximation method decomposes a large-scale feature matrix into multiple blocks and performs block sampling and parallel computation optimization to adjust prosody and expressiveness, and outputs an expressive speech data stream.
[0012] Based on the enhanced script and the voice data stream, key verbs, emphasis words and emotional words are identified through semantic analysis, a mapping relationship between semantic content and appropriate body movements is established, the precise time point of action execution is calculated and facial expression instructions consistent with the content emotion are generated, and a complete set of action expression instructions is output.
[0013] Based on the voice data stream and the action and expression instruction set, the digital human model is driven by MuseTalk technology. Through an optimized parallel rendering algorithm, the digital human lecturer and teaching content are integrated into a unified visual scene, and the final lecturer teaching video is output.
[0014] Preferably, based on the content features of user-uploaded PPT files or text scripts, a content relationship graph is constructed using an improved interior point method. An incremental shortest path algorithm is applied to identify the hierarchical relationships between titles, paragraphs, and key points, outputting structured data containing semantic structure, key content, and sentiment, including:
[0015] The content features are processed by document parsing to extract elements including text, images, and tables, and a preliminary set of parsed content elements is output.
[0016] Based on the content element set, an improved interior point method is applied to construct a content relationship graph. The incremental shortest path algorithm is used to identify the hierarchical relationship between titles, paragraphs, and key points, and a structured document with hierarchical tags is output.
[0017] Based on the structured document with hierarchical tags, natural language processing technology is used to identify keywords, key content and core concepts, and semantic vectorization is used to represent the correlation between content, outputting the structured data containing semantic structure, key content and sentiment.
[0018] Preferably, based on the structured data, a fully dynamic parallel single-link clustering algorithm is applied to group the content of the structured data according to semantic similarity. This is then converted into spoken narrative text using a pre-defined semantic understanding model, and markers including pauses, intonation variations, and emphasis are added. The resulting enhanced script with complete expressive markers is output, including:
[0019] Based on the structured data, a fully dynamic parallel single-link clustering algorithm is applied to maintain a dynamically updated network flow model to group the content semantically by similarity. The content is divided into multiple semantically related content blocks according to the topic and logical relationship, and the semantically clustered content blocks are output.
[0020] Based on the content segmentation of the semantic clustering, the structured data is converted into narrative text suitable for spoken expression through a preset semantic understanding model, transition words and conjunctions are added to enhance fluency, and a preliminary speech script is output.
[0021] Based on the initial speech script, a speech synthesis markup language (SSML) script containing markers for pauses, intonation changes, and emphasis is generated, and the enhanced script with full expressive markers is output.
[0022] Preferably, based on the enhanced script, the CosyVoice speech synthesis technology is used to... Low-rank approximation methods decompose large-scale feature matrices into multiple blocks and perform block sampling and parallel computation optimization to adjust prosody and expressiveness, outputting an expressive speech data stream, including:
[0023] Based on the enhanced script and combined with the lecturer style parameters, the optimal speech synthesis configuration is determined through a parameter optimization algorithm, and a speech synthesis parameter set is output.
[0024] Based on the speech synthesis parameter set and the enhancement script, CosyVoice speech synthesis technology is used to perform block sampling on the large-scale feature matrix and... The low-rank approximation method performs parallel computation optimization, which accelerates speech feature extraction and synthesis and outputs high-quality speech waveform data.
[0025] Based on the high-quality speech waveform data, a prosodic adjustment algorithm is applied to dynamically adjust the pauses, stresses, and intonation changes of the speech according to the Speech Synthesis Markup Language (SSML) markers, and outputs an expressive speech data stream.
[0026] Preferably, based on the enhanced script and the speech data stream, key verbs, emphasis words, and emotional words are identified through semantic analysis; a mapping relationship between semantic content and appropriate body movements is established; the precise timing of the movement execution is calculated; and facial expression instructions consistent with the emotional content are generated. A complete set of movement and expression instructions is output, including:
[0027] Based on the enhanced script, key verbs, emphasis words, and emotional words are identified through semantic analysis technology, a mapping relationship between semantic content and appropriate body movements is established, and a semantic-action association mapping table is output.
[0028] Based on the speech data stream and the semantic-action association mapping table, the timing of action execution is calculated by analyzing the rhythm, pauses and stress features of the speech, and a time-synchronized action sequence is output.
[0029] Based on the emotion markers in the enhanced script and the intonation changes in the speech data stream, facial expression commands consistent with the content emotion are generated through the emotion-expression mapping model, and a time-synchronized expression sequence is output.
[0030] Based on the time-synchronized action sequence and the time-synchronized facial expression sequence, an action-expression coordination algorithm is used to ensure a natural transition and coordination between actions and expressions, and to output a complete set of action-expression instructions.
[0031] Preferably, the step of ensuring a natural transition and consistency between actions and expressions through the action-expression coordination algorithm includes:
[0032] The time-synchronized action sequence and the time-synchronized facial expression sequence are time-aligned, the temporal proximity of strong expressions is detected, and a timing conflict indicator is output.
[0033] Based on the aforementioned temporal conflict identifier, a time staggering strategy is applied to ensure that important expressions have sufficient exclusive time windows by fine-tuning the timing of actions or expressions, and to output a temporally optimized action and expression sequence.
[0034] Perform semantic consistency checks on the time-optimized action-expression sequences, identify inconsistencies in the information conveyed by actions and expressions, and output a list of semantic conflicts.
[0035] Based on the semantic conflict list, the expression methods with less conflict are adjusted according to the emotional tendency of the content to maintain the consistency of the overall expression and output a semantically coordinated action and expression sequence.
[0036] Based on the semantically coordinated action and expression sequence, the intensity of actions and expressions is dynamically adjusted according to the content importance level through a global intensity balancing algorithm, and a complete action and expression instruction set with balanced intensity is output.
[0037] Preferably, the method of using MuseTalk technology to drive the digital human model includes:
[0038] Based on a predefined digital human model library, a suitable lecturer image is selected according to the nature of the teaching content and the target audience, the skeletal structure and facial control points are initialized, and a ready-to-use digital human model is output.
[0039] Based on the aforementioned action and expression instruction set, the action and expression instruction set is parsed into a skeletal animation control flow, a facial animation control flow, and a lip-sync control flow, and three parallel control flows are output.
[0040] Based on the skeletal animation control flow, keyframe interpolation and procedural animation methods are used to control the limb movements of the digital human, adding subtle body swaying, weight transfer and momentum effects, and outputting a natural skeletal animation sequence.
[0041] Based on the facial animation control flow, parametric facial expression control is used to map expression parameters to a digital human face controller, adding natural blinking, micro-expressions and subtle head movements, and outputting rich facial animation sequences.
[0042] Based on the lip-sync control flow and the speech data flow, a phoneme-to-visual mapping technique is used to convert the phoneme sequence in the speech into the corresponding lip shape. A lip-smoothing algorithm is applied to avoid the lip-shape changes from being too mechanical, and a complete animation sequence synchronized with the speech is output.
[0043] Preferably, the integration of the digital human instructor and teaching content into a unified visual scene using an optimized parallel rendering algorithm includes:
[0044] Based on the PPT content in the structured data, visual elements including slides, images, charts, and text key points are extracted. The readability of the content, the visibility of the speaker, the content-speaker relationship, and the balance of the screen are analyzed, and the most suitable screen layout strategy is output.
[0045] Based on the aforementioned screen layout strategy and the complete animation sequence, the position, angle, and focal length of the virtual camera are automatically adjusted according to the importance of the content and the lecturer's actions. Dynamic visual emphasis effects, including highlighting, magnification, and arrow indication, are added to important content, and visual scenes with cinematic language are output.
[0046] Based on the visual scene with cinematic language, the interaction effect between the lecturer and the content is processed, and transition effects including fade-in / fade-out, wipe-in / wipe-out, and zoom transition are applied at the content switching points to output a scene sequence with enhanced interaction.
[0047] Based on the enhanced scene sequence, color correction, lighting balance and layer blending techniques are applied to ensure that the digital human lecturer is visually consistent with the background and content, and to output a visually integrated complete scene sequence.
[0048] Based on the complete scene sequence of the aforementioned visual integration, a professional-grade complete video scene sequence is output through high dynamic range rendering, temporal anti-aliasing, motion blur, and color grading processing.
[0049] Preferably, the output of the final instructor teaching video includes:
[0050] Based on the complete video scene sequence and the audio data stream, the entire video rendering task is decomposed into time-segmented subtasks and spatial-segmented subtasks that can be processed in parallel, and the decomposed rendering task set is output.
[0051] Based on the decomposed rendering task set, computing resources are intelligently allocated according to the available CPU cores and GPU units, and tasks on the critical path are prioritized to ensure balanced rendering progress, resulting in a resource-optimized rendering schedule.
[0052] Based on the resource-optimized rendering schedule, the parallel computing capabilities of the GPU are used to accelerate lighting calculation, material rendering and post-processing. A multi-resolution rendering strategy is adopted to first obtain a preview by rendering at a lower resolution and then output a rendering result with progressive quality.
[0053] Based on the rendering results of the progressive quality, the quality is gradually improved to the target resolution, and audiovisual synchronization verification, image quality evaluation and compatibility testing are performed to output a high-definition video file with quality verification.
[0054] Based on the high-definition video file with the aforementioned quality verification, various format variants, including MP4, WebM, and low-bandwidth optimized, are generated according to different usage scenarios. Complete metadata containing chapter markers, keyword indexes, and content summaries is generated, and the final lecturer teaching video is output.
[0055] The beneficial effects of this invention are:
[0056] 1. An incremental shortest path algorithm is implemented through an improved interior point method, reducing the traditional O(n) time complexity to O(n log n) time complexity. 3 The shortest path calculation complexity is optimized to near linear time O(n log n), enabling real-time incremental parsing of large PPT documents and effectively improving the accuracy of content hierarchy recognition and processing efficiency.
[0057] 2. Employing a fully dynamic parallel single-link clustering algorithm, it supports dynamic addition, deletion, and updating of data points. Through distributed computing and parallel processing techniques, it reduces the traditional O(n) clustering algorithm to O(n log n)2. 2 The clustering complexity is reduced to O(n log n), enabling efficient semantic grouping of a large number of content fragments and improving the coherence and contextual relevance of content conversion;
[0058] 3. Application Low-rank approximation methods process high-dimensional speech feature matrices by employing block sampling and parallel computing strategies, reducing the computational complexity of the traditional core matrix from O(n^2) to O(n^2). 3 ) decreases to O(n 2 While maintaining accuracy, it achieves fast low-rank approximation and kernel ridge regression in the speech synthesis process, significantly improving the speed and quality of speech synthesis;
[0059] 4. Innovatively establishes a deep mapping relationship between semantic content and actions and expressions. By analyzing the semantic importance and emotional intensity of the content through a multi-layer attention mechanism, it automatically generates actions and expressions that are highly matched with the context of the content being explained, thus solving the problems of insufficient expressiveness and stiff movements of traditional digital humans.
[0060] 5. Design a multi-level parallel rendering pipeline architecture, decompose motion-driven, scene compositing and video rendering tasks into subtasks that can be processed in parallel, and achieve near real-time video generation capabilities through GPU acceleration and task scheduling optimization, supporting the rapid processing of large-scale content. Attached Figure Description
[0061] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0062] Figure 1 This is an overall flowchart of the lecturer video automatic generation method provided in an embodiment of the present invention;
[0063] Figure 2 A flowchart of a multimodal speech synthesis engine provided in an embodiment of the present invention;
[0064] Figure 3 A flowchart of a context-aware action expression generation system provided in an embodiment of the present invention. Detailed Implementation
[0065] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0066] This invention provides a method for automatically generating lecturer videos based on AI speech synthesis and animation-driven technology, such as... Figure 1 As shown, it includes the following steps:
[0067] Step S1: Based on the content features of the PPT file or text script uploaded by the user, construct a content relationship graph through the improved interior point method, apply the incremental shortest path algorithm to identify the hierarchical relationship between titles, paragraphs and key points, and output structured data containing semantic structure, key content and sentiment.
[0068] In this step, the system first receives user-uploaded PPT files or text scripts, which will serve as the foundational material for generating the instructor's video. The uploaded content undergoes initial processing using a document parsing engine, extracting text, images, tables, and other elements to form a content element set. These elements are then fed into a content structuring analysis module, which employs an improved interior-point method to construct a content relationship graph. This method maintains a dynamically updated network flow model, reducing the complexity of traditional O(n) operations. 3 The shortest path computation complexity is optimized to near linear time O(n log n). Based on the constructed content relationship graph, an incremental shortest path algorithm is applied to identify the hierarchical relationships between titles, paragraphs, and key points in the content, determining the logical structure of the content. Subsequently, natural language processing techniques are used to perform deep semantic analysis on the content, identifying keywords, key content, and core concepts. Sentiment analysis algorithms are then used to identify the content's emotional tendency, such as enthusiasm, seriousness, and curiosity. Finally, structured data containing semantic structure, key content, and emotional tendency is output, laying the foundation for subsequent processing.
[0069] Step S2: Based on the structured data, apply the fully dynamic parallel single-link clustering algorithm to group the content of the structured data according to semantic similarity, convert it into spoken narrative text through a preset semantic understanding model, and add markers including pauses, intonation changes, and emphasis, and output an enhanced script with complete expressive markers.
[0070] In this step, the structured data output from step S1 is received, and a fully dynamic parallel single-link clustering algorithm is applied to group the content based on semantic similarity. This algorithm supports dynamically adding, deleting, and updating data points, and through distributed computing and parallel processing techniques, reduces the traditional O(n) time complexity to zero. 2 The clustering complexity is reduced to O(n log n), enabling efficient semantic grouping of large amounts of content fragments. A dynamically updated network flow model is maintained, dividing content into multiple semantically related content blocks based on topics and logical relationships. Subsequently, these structured data are converted into narrative text suitable for spoken expression using a pre-defined semantic understanding model, adding transition words and conjunctions to enhance fluency and generate a preliminary speech script. The internal structure and semantic type of each content block are analyzed, and corresponding language templates are applied for different types, such as adding sequence guide words like "firstly," "secondly," and "finally" to list-type content. Finally, rich intonation and expressive markers are added to the script, generating a complete marked script conforming to the Speech Synthesis Markup Language (SSML) specification, including pause markers, intonation change markers, and emphasis markers, enabling the synthesized speech to express appropriate emotions and emphasize key points, outputting an enhanced script with complete expressive markers.
[0071] Step S3: Based on the enhanced script, using CosyVoice speech synthesis technology, through... The low-rank approximation method decomposes a large-scale feature matrix into multiple blocks and performs block sampling and parallel computation optimization to adjust prosody and expressiveness, and outputs an expressive speech data stream.
[0072] In this step, based on the enhanced script output from step S2 and combined with lecturer style parameters, the optimal speech synthesis configuration is determined through a parameter optimization algorithm. The enhanced script with expressive markers is analyzed to evaluate the overall characteristics of the content, such as professionalism, technical complexity, and target audience. A multi-dimensional parameter space is used to describe speech characteristics, including basic timbre parameters, speech rate parameters, pitch parameters, intonation parameters, and intelligibility parameters. After determining the parameters, core speech synthesis processing is performed using CosyVoice speech synthesis technology. To improve computational efficiency, innovative... Low-rank approximation methods decompose large-scale feature matrices into multiple blocks and perform block sampling and parallel computation optimization. This method decomposes a large feature matrix into multiple smaller blocks and applies a sampling method to each block. The method performs low-rank approximation and then reconstructs the complete low-rank approximation result through a carefully designed merging strategy, reducing the computational complexity of the traditional core matrix from O(n^2) to O(n^2). 3 ) decreases to O(n 2 While maintaining accuracy, the process generates a basic speech waveform, then adjusts the prosody and expressiveness, optimizing the rhythm, stress, and intonation patterns of the speech, including stress pattern adjustment, intonation contour optimization, pause rhythm adjustment, and emotional expression enhancement. Finally, professional-grade audio post-processing is applied to the speech, including noise reduction, equalization, dynamic range compression, and psychoacoustic enhancement, outputting an expressive speech data stream.
[0073] Step S4: Based on the enhanced script and the voice data stream, identify key verbs, emphasis words and emotional words through semantic analysis, establish a mapping relationship between semantic content and appropriate body movements, calculate the precise time point of action execution and generate facial expression instructions consistent with the content emotion, and output a complete set of action expression instructions;
[0074] In this step, based on the enhanced script output from step S2 and the speech data stream output from step S3, key verbs, emphasis words, and emotional vocabulary are identified through semantic analysis techniques to establish a mapping relationship between semantic content and appropriate body movements. Deep semantic analysis is performed on the enhanced script, employing multi-level semantic understanding, including language behavior recognition (such as explanation, emphasis, enumeration, etc.), spatial concept mapping (mapping spatial concepts such as "ascend" and "expand" to corresponding directional gestures), and emotional intensity analysis (assessing the emotional intensity and importance of the content to determine the amplitude and energy level of the movement). Semantic features are extracted using a pre-trained large-scale language model and converted into vectors in the action feature space through a mapping network, outputting a semantic-action association mapping table. Subsequently, detailed acoustic analysis is performed on the speech data stream to extract speech rhythm markers, prosodic structure, and emphasis markers. An action-speech synchronization algorithm is applied to calculate the precise execution time of each action, following the principles of preparation, emphasis, smooth ending, and coherent transition, outputting a time-synchronized action sequence. Simultaneously, based on the emotion markers in the enhanced script and the intonation changes in the speech data stream, facial expression commands consistent with the content's emotion are generated through an emotion-expression mapping model. This multi-layered model includes a basic emotion layer, a teaching-specific expression layer, and a micro-expression layer, outputting a time-synchronized expression sequence. Finally, a motion-expression coordination algorithm ensures a natural transition and consistency between actions and expressions, resolving issues such as temporal conflicts, semantic conflicts, physical conflicts, and intensity imbalances, outputting a complete set of motion-expression commands.
[0075] Step S5: Based on the voice data stream and the action and expression instruction set, use museTalk technology to drive the digital human model, and integrate the digital human lecturer and teaching content into a unified visual scene through an optimized parallel rendering algorithm to output the final lecturer teaching video.
[0076] In this step, based on the speech data stream output from step S3 and the action / expression instruction set output from step S4, a suitable lecturer image is first selected from a predefined digital human model library, chosen according to the nature of the teaching content and the target audience. The skeletal structure and facial control points of the digital human model are initialized, including bone loading, facial control initialization, physical parameter settings, and rendering parameter configuration. Subsequently, the action / expression instruction set is parsed into three parallel control flows: skeletal animation control flow, facial animation control flow, and lip-sync control flow, and the digital human model is driven using MuseTalk technology. Keyframe interpolation and procedural animation methods are used to control the digital human's limb movements, adding subtle body swaying, weight transfer, and momentum effects to make the animation more vivid and natural. Parametric facial expression control maps expression parameters to the digital human's facial controller, adding natural blinks, micro-expressions, and subtle head movements. Phoneme-to-visual mapping technology is applied to convert the phoneme sequence in the speech into corresponding lip shapes, ensuring precise synchronization between lip movements and speech. Next, the digital human instructor and teaching content are integrated into a unified visual scene. Visual elements from the PPT content are extracted, the most suitable screen layout is designed, intelligent camera language is implemented, the interaction between the instructor and the content is handled, and color correction, lighting balance, and layer blending techniques are applied to ensure visual consistency. Finally, video generation is performed using an optimized parallel rendering algorithm. The entire video rendering task is decomposed into parallelizable subtasks, computing resources are intelligently allocated, and the parallel computing capabilities of GPUs are used to accelerate the rendering process. Quality checks and format conversions are performed, and the final instructor teaching video is output.
[0077] In step S1, based on the content features of the user-uploaded PPT file or text script, a content relationship graph is constructed using an improved interior point method. An incremental shortest path algorithm is applied to identify the hierarchical relationships between titles, paragraphs, and key points, outputting structured data containing semantic structure, key content, and sentiment, including:
[0078] Step S1.1: Perform document parsing processing on the content features, extract elements including text, images, and tables, and output a preliminary parsed set of content elements;
[0079] This step first receives the user-uploaded PPT file or text script and processes it using a professional document parsing engine. For PPT files, it identifies and extracts various content elements, including slide titles, body text, bulleted lists, images, charts, tables, and embedded multimedia elements. It preserves the original formatting and layout information of these elements, while recording their position and order within the PPT. For text scripts, it analyzes the document structure, identifying different text blocks such as titles, subheadings, paragraphs, and lists. During extraction, it also preserves text formatting information, such as font size, bold, italics, and other style features, which often indicate the importance hierarchy of the content. For image elements, it not only extracts the image itself but also analyzes the image's title, descriptive text, and any text information it may contain. For tables, it preserves their structure and extracts the table header and cell content to understand the table's logical organization. After processing, it outputs a preliminary set of parsed content elements containing all these elements. This set retains all the information and basic structural features of the original content, laying the foundation for subsequent in-depth analysis.
[0080] Step S1.2: Based on the content element set, construct a content relationship graph using the improved interior point method, identify the hierarchical relationship between titles, paragraphs, and key points using the incremental shortest path algorithm, and output a structured document with hierarchical tags.
[0081] In this step, based on the content element set output in step S1.1, an improved interior-point method is applied to construct a content relationship graph. Traditional interior-point methods have high computational complexity when dealing with large-scale network flow problems. This invention improves upon this method by maintaining a dynamically updated network flow model, reducing the computational complexity of the traditional O(n) method. 3 The computational complexity is optimized to near linear time O(n log n). First, content elements are treated as nodes in a graph, and initial connections are established based on their positional relationships, formatting features, and semantic associations. For example, strong connections are established between headings and their paragraphs below, weak connections between adjacent paragraphs, and semantic connections between content blocks with similar topics. After constructing the initial graph, an incremental shortest path algorithm is applied to identify the hierarchical structure of the content. This algorithm efficiently finds the shortest paths from the beginning of the document to each content block and determines the hierarchical relationships of the content based on path characteristics. For example, by analyzing path lengths and node features, different levels of content elements such as first-level headings, second-level headings, body paragraphs, and lists of key points can be identified. Unlike traditional methods, this algorithm uses incremental processing; when new content is added, only the affected paths are updated, without recalculating the entire graph structure, greatly improving the efficiency of processing large documents. After processing, a structured document with hierarchical tags is output, clearly marking the hierarchical relationships, subordinate relationships, and logical order of each content element, providing a structured foundation for subsequent semantic analysis.
[0082] Step S1.3: Based on the structured document with hierarchical tags, natural language processing technology is used to identify keywords, key content and core concepts, and semantic vectorization is used to represent the relationship between content, outputting the structured data containing semantic structure, key content and sentiment.
[0083] In this step, based on the structured document with hierarchical tags output in step S1.2, advanced natural language processing techniques are applied for deep semantic analysis. First, keyword extraction algorithms are used to identify keywords and terms in the document. These algorithms combine statistical methods (such as TF-IDF) and deep learning models to accurately capture the core vocabulary of the document. For the identified keywords, their distribution and context within the document are further analyzed to determine their importance weight. Subsequently, text summarization and key point extraction techniques are applied to identify core concepts and key content in the document. This process considers not only explicit tags of content (such as words like "important" and "key"), but also analyzes the content's position in the document structure, its repetition, and its relevance to other content. Semantic vectorization techniques are also used to convert the text content into high-dimensional semantic vectors, which capture the deep semantic features of the content. By calculating the similarity between vectors, the semantic association strength between different content blocks can be quantified, constructing a semantic network of the content. Furthermore, sentiment analysis algorithms are applied to identify the sentiment tendency and tone characteristics of the content. These algorithms can detect the emotions (such as enthusiasm, seriousness, curiosity, etc.) and tone (such as affirmation, questioning, emphasis, etc.) expressed in text, providing important references for subsequent speech synthesis and facial expression generation. After processing, the output is complete structured data containing semantic structure, key content, and emotional tendency. This data comprehensively describes the structural and semantic features of the original content, laying a solid foundation for subsequent semantic understanding and conversion.
[0084] In step S2, based on the structured data, a fully dynamic parallel single-link clustering algorithm is applied to group the content of the structured data according to semantic similarity. This is then converted into spoken narrative text using a preset semantic understanding model, and markers including pauses, intonation variations, and emphasis are added. The resulting enhanced script with complete expressive markers is output, including:
[0085] Step S2.1: Based on the structured data, apply the fully dynamic parallel single-link clustering algorithm to maintain a dynamically updated network flow model to perform semantic similarity grouping of the content, divide the content into multiple semantically related content blocks according to the topic and logical relationship, and output the semantically clustered content blocks;
[0086] Step S2.2: Based on the semantic clustering, the content is divided into blocks, and the structured data is converted into narrative text suitable for spoken expression through a preset semantic understanding model. Transition words and conjunctions are added to enhance fluency, and a preliminary speech script is output.
[0087] Step S2.3: Based on the preliminary speech script, generate a speech synthesis markup language (SSML) script containing markers for pauses, intonation changes, and emphasis, and output the enhanced script with full expressive markers.
[0088] This step uses a fully dynamic parallel single-linkage clustering algorithm to process the structured data. Its core technical principle is as follows:
[0089] First, the system receives structured data output from the PPT content structuring and parsing module. This data includes information such as text content, hierarchical relationships, semantic tags, and sentiment. This content is treated as data points in a high-dimensional space, with each data point representing a content fragment (such as a paragraph or a key point).
[0090] The key to the fully dynamic parallel single-link clustering algorithm lies in its "fully dynamic" characteristic, which allows data points to be added, deleted, or updated dynamically without recalculating the entire clustering structure. Specifically, the algorithm maintains a nearest-neighbor graph structure, where each node (content fragment) is connected to the few most similar nodes in its semantic space. When new content is added, it only needs to calculate its similarity to existing nodes and update the local graph structure, rather than recalculating the entire cluster.
[0091] In terms of parallel processing, the algorithm employs a divide-and-conquer strategy, dividing the data space into multiple subspaces, performing local clustering in parallel within each subspace, and then integrating the results through a boundary merging algorithm. This method reduces the traditional O(n) time complexity by 100%. 2 The computational complexity of ) is reduced to O(n log n), which greatly improves the efficiency of processing large-scale content.
[0092] Semantic similarity is calculated based on text embedding vectors extracted by pre-trained language models (such as BERT, RoBERTA, etc.), and the strength of semantic association between content segments is determined by cosine similarity or other distance metrics. The original structural information of the PPT is also considered; content on the same page or with the same heading level is given additional clustering weights.
[0093] Finally, this step outputs semantically clustered content chunks. The content within each chunk is highly semantically related, forming a coherent explanatory unit. These chunks retain the logical order of the original content while marking the semantic correlation between each chunk, providing an important foundation for subsequent script conversion.
[0094] In step S2.2:
[0095] Based on the semantic clustering content blocks output in the previous step, this step performs script conversion and enhancement, transforming the structured content into natural and fluent spoken expression.
[0096] The conversion process first analyzes the internal structure and semantic type of each content block, such as explanatory content, list-type content, and comparative descriptions, and then applies corresponding language templates for different types. For example, for list-type content, sequence-introducing words such as "firstly," "secondly," and "finally" are added; for causal content, conjunctions such as "because" and "therefore" are added.
[0097] The system employs a two-tiered transformation architecture: the first tier is structural transformation, which converts the visual structure of the PPT (such as bullet points, tables, charts, etc.) into a conversational descriptive structure; the second tier is language style transformation, which converts formal written language into a more conversational style suitable for oral expression, including simplifying complex sentences, shortening sentence lengths, and increasing rhythmic variation.
[0098] To enhance the flow of content, transition sentences are automatically generated between adjacent content blocks. These transition sentences are dynamically generated based on the semantic relationship between the two blocks. For example, if the next block is an in-depth explanation of the previous concept, a transition sentence such as "Let's learn more about this concept" will be generated; if it is a shift to a new topic, a transition sentence such as "Next, we will discuss another important aspect" will be generated.
[0099] Semantic enhancement also includes adjusting the level of detail in the expression based on the importance of the content. For content identified as core concepts, more details are retained and emphasis is added; for supporting information, it is appropriately simplified and summarized.
[0100] The final output of the preliminary speech script retains all the information of the original content, while having better spoken fluency and coherence, laying the foundation for the next step of adding expressive markers.
[0101] In step S2.3:
[0102] After obtaining the initial speech script, this step will add rich intonation and expressive markers to the script, enabling the synthesized speech to express appropriate emotions and emphasize key points.
[0103] First, a deep semantic analysis is performed on the initial speech script to identify key words, concept definitions, important conclusions, and other content elements that need to be emphasized. Simultaneously, based on the sentiment analysis results from step 1.4, the basic emotional tone (such as neutral, enthusiastic, serious, or curious) that each sentence or paragraph should express is determined.
[0104] Based on these analyses, a complete markup script conforming to the Speech Synthesis Markup Language (SSML) specification is generated. SSML is an XML-based markup language specifically designed to control characteristics of speech synthesis, such as pronunciation, intonation, and speed. The generated SSML tags mainly include the following categories:
[0105] pause marker ( <break>This involves inserting pauses of varying lengths between sentences, before and after key concepts, and at turning points to simulate the rhythm of a natural speech. For example, inserting a medium-length pause before introducing a new concept can enhance the audience's attention.
[0106] intonation markers ( <prosody>Adjust the pitch, rate, and volume of your speech. For example, when explaining important concepts, you might slightly slow down the rate and increase the volume; when giving examples, you might use a more relaxed tone and a moderate rate.
[0107] Emphasis markers ( <emphasis>SSML applies different levels of emphasis to keywords and core concepts. It supports three levels of emphasis: strong, medium, and weak, and dynamically assigns the emphasis level based on the importance of the content.
[0108] Voice style markers ( <voice-style>(Or manufacturer-specific tags): Some advanced speech synthesis engines support overall speech style control. For example, CosyVoice may support style settings such as "teaching mode" and "speech mode". Appropriate style tags will be selected based on the overall type of content.
[0109] The algorithm also considers rhythmic variations in speech to avoid mechanical and monotonous expression. For example, it appropriately inserts subtle pauses in long sentences and adds slight variations in speech rate between consecutive technical terms to make the overall speech more natural and fluent.
[0110] The final output, an enhanced script with complete expressive tags, not only contains the original semantic content but also includes rich expressive control instructions, which can guide the speech synthesis engine to generate more engaging and effective speech expressions.
[0111] like Figure 2 As shown, in step S3, based on the enhanced script, CosyVoice speech synthesis technology is used to synthesize speech through... Low-rank approximation methods decompose large-scale feature matrices into multiple blocks and perform block sampling and parallel computation optimization to adjust prosody and expressiveness, outputting an expressive speech data stream, including:
[0112] Step S3.1: Based on the enhanced script and combined with the lecturer style parameters, determine the optimal speech synthesis configuration through a parameter optimization algorithm and output the speech synthesis parameter set;
[0113] In this step, the enhanced script with expressive markers is first comprehensively analyzed to assess the overall characteristics of the content, such as its level of professionalism, technical complexity, and target audience characteristics. A multi-dimensional parameter space is used to describe speech characteristics, which collectively determine the overall style and expressiveness of the synthesized speech. Instructor style parameters mainly include the following key dimensions: basic timbre parameters, speech rate parameters, pitch parameters, intonation parameters, and intelligibility parameters. Basic timbre parameters determine the fundamental characteristics of the voice, including gender characteristics (male, female, or neutral), age characteristics (young, mature, or older), timbre brightness (bright, warm, or deep), and timbre quality characteristics (crisp, round, or rich). These parameters directly influence the audience's first impression of the instructor and will select the most suitable basic timbre based on the nature of the teaching content. For example, a more neutral and professional timbre might be chosen for scientific and technological content, while a warmer and more approachable timbre might be chosen for children's educational content. Speech rate parameters control the overall speech rate baseline, usually measured in words or syllables per minute, and will be dynamically adjusted according to the complexity of the content. For complex technical content, the basic speaking speed is automatically reduced to ensure learners have sufficient time to understand; for simple introductory content, a slightly faster speaking speed may be used to maintain audience interest. Pitch parameters control the basic pitch and range of variation of the voice, including basic pitch (low, mid, high), pitch range (narrow, mid, wide), and pitch variation pattern (smooth, fluctuating, or emphasized). These parameters affect the expressiveness and emotional delivery of the speech, adjusting pitch characteristics according to the emotional needs of the content. Tone parameters control the overall style of the speech, such as formality (formal, neutral, or easygoing), energy (calm, moderate, or lively), and approachability (professional, friendly, or intimate). These parameters determine the overall impression the lecturer gives, selecting appropriate tone characteristics based on the target audience and teaching scenario. Clarity parameters control the clarity of pronunciation and articulation stress, including consonant intensity, vowel fullness, and syllable boundary clarity. For teaching content, higher clarity is usually prioritized, especially for content containing technical terminology. To determine the optimal parameter combination, a parameter optimization algorithm is used, combining rule-guided and machine learning methods. In the rule-guided section, preset parameter templates are applied as initial values based on content type and performance requirements. In the machine learning section, optimal parameter settings for similar content in historical data are used to search for the optimal solution in the parameter space through optimization methods such as gradient descent or genetic algorithms. Instructor consistency is also considered; if it's part of a series of courses, the speech characteristics will be maintained as consistently as possible with previous courses, with minor adjustments only when necessary. Users can also provide reference audio samples, which will be analyzed to extract parameter features as optimization targets. The final output speech synthesis parameter set is a configuration file containing multi-dimensional parameter values. It guides the specific settings in subsequent speech synthesis processes, ensuring that the generated speech not only meets content requirements but also maintains a consistent instructor style.
[0114] Step S3.2: Based on the speech synthesis parameter set and the enhancement script, CosyVoice speech synthesis technology is used to perform block sampling on the large-scale feature matrix and... The low-rank approximation method performs parallel computation optimization, which accelerates speech feature extraction and synthesis and outputs high-quality speech waveform data.
[0115] In this step, the speech synthesis parameter set determined in step S3.1 and the enhancement script with SSML tags are input into the CosyVoice speech synthesis engine to begin the core speech synthesis processing. Modern high-quality speech synthesis is typically based on deep learning models, such as Tacotron, WaveNet, or the newer Transformer architecture. These models require processing large amounts of high-dimensional feature matrices, resulting in high computational complexity. To improve computational efficiency, this step applies an innovative... Low-rank approximation methods. Traditionally, kernel matrix computation in speech synthesis (especially in the acoustic feature extraction and acoustic model inference stages) requires O(n^2) time complexity. 3 The time complexity is low, making it inefficient for processing long content. The core idea of the method is to decompose a large feature matrix into multiple smaller blocks, and apply the following to each block: The method performs a low-rank approximation and then reconstructs the complete low-rank approximation result through a carefully designed merging strategy. Specifically, the algorithm first randomly selects a subset of columns from the feature matrix as landmarks. These landmarks are typically chosen from representative feature columns to ensure the capture of the matrix's main characteristics. Then, the similarity between the complete matrix and these landmarks is calculated, forming a much smaller core matrix. By performing eigenvalue decomposition on this core matrix, a low-rank approximation of the original large matrix can be obtained. This method reduces the complexity from O(n log n) to O(n log n). 3 ) decreases to O(n 2 Simultaneously, parallel computing is achieved through a block processing strategy to further improve efficiency. In practical applications, the size and number of blocks are dynamically adjusted based on available computing resources to achieve the optimal balance between accuracy and speed. The specific process of speech synthesis includes several key stages: First, text normalization is performed, converting numbers, abbreviations, special symbols, etc., into standard forms; then, linguistic analysis is performed to determine phoneme sequences, stress positions, and intonation boundaries; finally, an acoustic model (applying...) is used... The process involves optimizing and generating acoustic features (such as Mel spectrograms); finally, a vocoder converts these acoustic features into actual waveform data. Throughout the process, various expressive instructions in the SSML tags are considered in real time, and the generation parameters are dynamically adjusted to ensure that the output speech waveform is not only natural and fluent but also accurately expresses the expected intonation changes and emotional features. Furthermore, adaptive enhancement techniques are applied to dynamically adjust the level of detail processing during synthesis based on the semantic importance and emotional intensity of the content, ensuring that key content receives more refined processing. The final high-quality output speech waveform data retains all the information of the text content and possesses basic intonation changes, laying the foundation for the next step of prosodic adjustment.
[0116] CosyVoice speech synthesis technology employs an end-to-end neural network model based on the Transformer architecture, comprising three core components: encoder, attention, and decoder. The model's training data comes from approximately 1000 hours of high-quality recordings, covering professional voice actors of different ages, genders, and speech styles, ensuring the diversity and naturalness of the generated speech. The model architecture includes a 12-layer Transformer encoder and a 12-layer Transformer decoder, with 8 attention heads per layer, a hidden layer dimension of 512, a feedforward network dimension of 2048, and the GeLU activation function. Key hyperparameter settings include: a learning rate of 0.0001, the use of the Adam optimizer, a batch size of 32, and approximately 500,000 training iterations. The model undergoes two-stage training: the first stage, pre-training, uses general speech data; the second stage, fine-tuning, optimizes for specific speech styles used in teaching scenarios.
[0117] The mathematical principle of the low-rank approximation method is based on matrix factorization theory, and its specific implementation is as follows: For a large eigenma matrix K∈R^(n×n), the complexity of traditionally calculating its eigenvalue factorization is O(n^2). 3 ). The method first samples c columns from K to form C ∈ R^(n×c), and then calculates the submatrices W ∈ R^(c×c) corresponding to these columns. The formula K ≈ CW is then used to calculate the submatrices. -1 CT scans perform low-rank approximation, where W -1 The computation is performed using SVD decomposition. To further improve efficiency, the algorithm divides C into multiple blocks {C1, C2, ..., C_m} and computes W for each block in parallel. -1 CT (Computational Transformation) is then performed, and the results are merged. Experimental verification shows that when the sampling rate c / n = 0.1, this method can reduce the computational complexity to O(n^2). 2 While maintaining an accuracy of over 95%, the computation speed can be increased by more than 10 times when processing feature matrices with dimensions of over 100,000.
[0118] Step S3.3: Based on the high-quality speech waveform data, apply the prosody adjustment algorithm to dynamically adjust the pauses, stresses and intonation changes of the speech according to the Speech Synthesis Markup Language (SSML) markers, and output the expressive speech data stream.
[0119] In this step, the basic speech waveform generated in step S3.2 is further optimized for prosody and enhanced for expressiveness, making it more natural and engaging. Prosody, the rhythm, stress, and intonation patterns in speech, is crucial for conveying meaning and emotion and is a key factor distinguishing mechanical speech from natural human voice. First, a detailed acoustic analysis is performed on the generated speech waveform, extracting acoustic features such as the fundamental frequency profile (F0), energy curve, and duration parameters. These features collectively describe the prosodic characteristics of speech and form the basis for subsequent adjustments. Then, these actual parameters are compared with an ideal model derived from two sources: a theoretical model based on linguistic rules, incorporating various prosodic rules established in linguistic research; and a statistical model learned from the speech of professional lecturers, capturing the prosodic characteristics of lecturer speech in real teaching scenarios. Prosodic adjustment mainly targets four key aspects: stress pattern adjustment, intonation profile optimization, pause rhythm adjustment, and emotional expression enhancement. In stress pattern adjustment, it is ensured that keywords in the sentence receive appropriate stress, which is achieved by fine-tuning the duration, pitch, and intensity of syllables. For example, for technical terms, the duration and intensity of the first stressed syllable are increased; for emphatic content, the overall volume and clarity of related words are improved. In intonation contour optimization, the pitch curve of the entire sentence is adjusted to conform to a natural intonation pattern. For example, interrogative sentences should have a rising intonation ending, emphatic sentences should have a distinct pitch peak, and declarative sentences should have a smoothly falling intonation ending. Segmented spline interpolation algorithms are used to generate smooth pitch curves between key points, ensuring natural and fluent intonation changes. In pause rhythm adjustment, the precise duration of pauses within and between sentences is fine-tuned to make the speech rhythm more natural. Research shows that pause lengths in natural speech follow a specific statistical distribution and dynamically adjust pause lengths according to context. For example, slightly longer pauses are inserted before complex concepts, short pauses are inserted between enumerated items, and distinct pauses are inserted at paragraph transitions. In affective enhancement, corresponding acoustic features are enhanced based on the affective markers of the content. For example, for content expressing enthusiasm, the pitch variation range and speech rate variation are increased; for serious content, pitch variation is reduced and timbre characteristics are adjusted; for content expressing curiosity or doubt, the rising trend of intonation is increased. These adjustments are achieved through digital signal processing techniques, including pitch shifting, duration scaling, and spectral envelope modification. Employing an end-to-end prosodic model based on deep learning enables precise prosodic control while maintaining the naturalness of the speech. Furthermore, subtle human vocal features are added, such as natural breathing sounds, slight oral noises, and variations in voice texture, further enhancing the naturalness and approachability of the speech. The final expressive speech data stream retains the basic speech quality generated in the previous steps, while possessing more natural and pedagogically effective prosodic characteristics, providing high-quality input for subsequent audio post-processing and video synthesis.
[0120] In step S3.1:
[0121] The first step in the speech synthesis process is to determine the optimal speech synthesis parameter configuration based on the characteristics of the enhancement script and the target lecturer's style.
[0122] First, analyze the enhanced scripts marked with expressive tags to assess the overall characteristics of the content, such as its level of expertise, technical complexity, and target audience. For example, for introductory tutorials, a gentler, more moderate-paced voice style might be preferred, while for advanced technical content, a more professional and clearer voice style might be chosen.
[0123] Using a multidimensional parameter space to describe speech characteristics mainly includes:
[0124] Basic timbre parameters: determine the fundamental characteristics of a voice, such as gender, age, and timbre brightness.
[0125] Speech rate parameter: controls the overall speech rate baseline, usually measured in words or syllables per minute.
[0126] Pitch parameters: control the fundamental pitch and range of variation of a sound.
[0127] Tone parameters: control the overall style of speech, such as formal, easygoing, lively, etc.
[0128] Clarity parameter: Controls the clarity of pronunciation and articulation stress.
[0129] To determine the optimal parameter combination, a parameter optimization algorithm is employed, which combines rule-guided and machine learning methods. In the rule-guided part, preset parameter templates are applied as initial values based on content type and performance requirements. In the machine learning part, based on the best parameter settings for similar content in historical data, optimization methods such as gradient descent or genetic algorithms are used to search for the optimal solution in the parameter space.
[0130] We also consider the consistency requirements of the instructor. If it is part of a series of courses, we will try to maintain the same voice characteristics as previous courses, and only make minor adjustments when necessary. Users can also provide reference audio samples, and we will extract parameter features through audio analysis as optimization targets.
[0131] The final output speech synthesis parameter set is a configuration file containing parameter values of multiple dimensions. It will guide the specific settings in the subsequent speech synthesis process to ensure that the generated speech not only meets the content requirements but also has consistent lecturer style characteristics.
[0132] In step S3.2:
[0133] This step is the core of the entire speech synthesis process, which utilizes advanced speech synthesis technology and optimization algorithms to convert text into high-quality speech waveforms.
[0134] First, the enhancement script with SSML tags and the speech synthesis parameter set determined in step 3.1 are input into a speech synthesis engine (such as CosyVoice). Modern speech synthesis is typically based on deep learning models, such as Tacotron, WaveNet, or the newer Transformer architecture. These models require processing large amounts of high-dimensional feature matrices, resulting in high computational complexity.
[0135] To improve computational efficiency, this step applies an innovative method. Low-rank approximation methods. Traditionally, kernel matrix computation in speech synthesis (especially in the acoustic feature extraction and acoustic model inference stages) requires O(n^2) time complexity. 3 The time complexity is low, making it inefficient for processing long content. The core idea of the method is:
[0136] Decompose a large feature matrix into multiple smaller blocks.
[0137] Apply to each block The method performs low-rank approximation
[0138] The complete low-rank approximation result is reconstructed using a carefully designed merging strategy.
[0139] In its implementation, the algorithm first randomly selects a subset of columns from the feature matrix as landmarks, then calculates the similarity between the complete matrix and these landmarks, and finally reconstructs the low-rank approximation result through matrix operations. This method reduces the complexity from O(n^2) to O(n^2). 3 ) decreases to O(n 2 Meanwhile, parallel computing is achieved through a block processing strategy, further improving efficiency.
[0140] In the speech synthesis process, the first step is text normalization, converting numbers, abbreviations, and special symbols into standard forms; then, linguistic analysis is performed to determine phoneme sequences, stress positions, and intonation boundaries; finally, an acoustic model (applying...) is used... The acoustic features (such as Mel spectrograms) are generated through optimization; finally, the acoustic features are converted into actual waveform data through a vocoder.
[0141] Throughout the process, various expressive instructions in the SSML tags are considered in real time, and the generation parameters are dynamically adjusted to ensure that the output speech waveform is not only natural and fluent, but also accurately expresses the expected intonation changes and emotional features.
[0142] The final high-quality speech waveform data retains all the information of the text content and has basic intonation changes, laying the foundation for the next step of prosody adjustment.
[0143] In step S3.3:
[0144] After obtaining the basic speech waveform, this step will further optimize the prosodic characteristics and expressiveness of the speech, making it more natural and engaging.
[0145] Prosody, the rhythm, stress, and intonation patterns in speech, is crucial for conveying meaning and emotion. While the previous step has generated a basic expressive speech based on SSML markers, subtle prosodic adjustments are essential for achieving truly natural teaching speech.
[0146] First, acoustic analysis is performed on the generated speech waveform to extract acoustic features such as the fundamental frequency profile (F0), energy curve, and duration parameters. Then, these actual parameters are compared with an ideal model derived from two sources: a theoretical model based on linguistic rules and a statistical model learned from the speech of professional lecturers.
[0147] Rhythm adjustment mainly targets the following aspects:
[0148] Stress pattern adjustment: Ensure that keywords in a sentence receive appropriate stress, which is achieved by fine-tuning the duration, pitch, and intensity of syllables. For example, for technical terms, the duration and intensity of the first stressed syllable are increased.
[0149] Intonation contour optimization: Adjust the pitch curve of the entire sentence to conform to a natural intonation pattern. For example, interrogative sentences should have a rising intonation ending, and emphatic sentences should have a distinct pitch peak. A piecewise spline interpolation algorithm is used to generate a smooth pitch curve between key points.
[0150] Pause rhythm adjustment: Fine-tuning the precise duration of pauses within and between sentences to make the speech rhythm more natural. Research shows that the pause length in natural speech follows a specific statistical distribution and dynamically adjusts the pause length according to the context.
[0151] Enhanced emotional expression: Based on the emotional markers of the content, corresponding acoustic features are enhanced. For example, for content expressing enthusiasm, the range of pitch variation and speech rate variation are increased; for serious content, pitch variation is reduced and timbre characteristics are adjusted.
[0152] These adjustments are achieved through digital signal processing techniques, including pitch shifting, duration scaling, and spectral envelope modification. Employing an end-to-end prosodic model based on deep learning enables precise prosodic control while maintaining the naturalness of the speech.
[0153] The final output of expressive enhanced speech data retains the basic speech quality generated in the previous steps, while possessing more natural and pedagogically effective prosodic characteristics, providing high-quality input for the final audio post-processing.
[0154] In addition, it includes audio post-processing and optimization:
[0155] As the final step in speech synthesis, audio post-processing and optimization aim to improve the overall quality and listening experience of the speech, ensuring a clear and comfortable auditory experience in various playback environments.
[0156] First, the enhanced speech data output from step 3.3 undergoes professional-grade audio post-processing, which includes several key steps:
[0157] Noise reduction: Although synthesized speech should theoretically not contain noise, some synthesis methods (especially sample-based methods) may introduce slight background noise. Spectral subtraction and adaptive filtering algorithms are applied to remove any potentially present minor noise, improving the purity of the speech.
[0158] Equalization (EQ) processing: Based on the human ear's sensitivity to different frequencies, a professional equalizer is used to adjust the frequency response of the speech. Typical processing includes slightly boosting the 3-5kHz range to enhance clarity, appropriately reducing low-frequency boom at 200-300Hz, and controlling high frequencies above 10kHz to avoid harshness.
[0159] Dynamic range compression: Instructional videos are typically played on various devices and in various environments. An excessively large dynamic range may make the audio unclear in some scenarios. Applying a multi-band compressor with a gentle compression ratio (typically 1.5:1 to 3:1) reduces the dynamic range of the volume, ensuring that quiet parts are clear enough while loud parts are not too abrupt.
[0160] Psychoacoustic enhancement: Applying processing algorithms based on the principles of human auditory perception, it enhances the presence and clarity of speech, including subtle harmonic excitation and transient enhancement, making speech sound more natural and vivid.
[0161] To adapt to different playback environments, multiple targeted optimized audio versions will be generated. For example, the version optimized for mobile devices will apply stronger compression and clarity enhancement; while the version optimized for professional audio equipment will retain more dynamic range and frequency details.
[0162] A final quality assessment will also be conducted, using objective metrics (such as PESQ, STOI, and other speech quality assessment metrics) and machine learning models to predict subjective listening experience, ensuring that the processed audio meets the expected quality standards.
[0163] The final output speech data stream has been fully optimized, with excellent clarity, natural rhythm, and professional audio quality, providing a high-quality audio track for subsequent video synthesis, and also serving as an important reference for motion expression generation.
[0164] like Figure 3 As shown, in step S4, based on the enhanced script and the speech data stream, key verbs, emphasis words, and emotional words are identified through semantic analysis. A mapping relationship between semantic content and appropriate body movements is established. The precise timing of the movement execution is calculated, and facial expression commands consistent with the emotional content are generated. A complete set of movement and expression commands is output, including:
[0165] Step S4.1: Based on the enhanced script, identify key verbs, emphasis words and emotional words through semantic analysis technology, establish a mapping relationship between semantic content and appropriate body movements, and output a semantic-action association mapping table;
[0166] Step S4.2: Based on the speech data stream and the semantic-action association mapping table, the timing of action execution is calculated by analyzing the rhythm, pauses and stress features of the speech, and a time-synchronized action sequence is output.
[0167] Step S4.3: Based on the emotion markers in the enhanced script and the intonation changes in the speech data stream, generate facial expression instructions consistent with the content emotion through the emotion-expression mapping model, and output a time-synchronized expression sequence;
[0168] Step S4.4: Based on the time-synchronized action sequence and the time-synchronized facial expression sequence, the action-expression coordination algorithm ensures a natural transition and coordination between actions and expressions, and outputs the complete action-expression instruction set.
[0169] The core of this step is to establish a mapping relationship between text content and appropriate body movements, enabling the digital human to exhibit natural movements that match the content being explained.
[0170] First, a deep semantic analysis is performed on the enhanced script from step S2. This analysis goes beyond simple keyword recognition and employs multi-level semantic understanding:
[0171] Language behavior recognition: Identifying different types of language behaviors in a script, such as explanation, emphasis, enumeration, comparison, and exemplification. Each language behavior is associated with a series of potential gestures and body movements. For example, explanatory content may be paired with an open gesture, while enumeration content is often paired with a gesture of counting in sequence.
[0172] Spatial concept mapping: When content involves spatial relationships (such as "rising", "expanding", "left", etc.), these spatial concepts are automatically mapped to corresponding directional gestures. For example, when describing a growth trend, an upward gesture may be generated; when describing structural relationships, two hands may be used to indicate the positional relationship between different components.
[0173] Emotional intensity analysis: Assessing the emotional intensity and importance of content determines the magnitude and energy level of actions. Important concepts or emotionally charged content trigger more obvious and forceful actions, while less important information is accompanied by more restrained and subtle actions.
[0174] Semantic features are extracted using a pre-trained large-scale language model, and a specially designed attention mechanism is used to identify key content that needs to be emphasized. These semantic features are then transformed into vectors in the action feature space through a mapping network.
[0175] The mapping network is a deep neural network trained on a large number of lecturer videos. It learns the typical action patterns of professional lecturers when explaining different types of content. The network can predict the most suitable action category, amplitude, speed, and emotional expression based on the semantic features of the content.
[0176] To enhance the diversity and naturalness of the mapping, an action mutation generator was also integrated to avoid repeating the same actions when similar content appears. The mutation generator is based on a probabilistic model and introduces controlled randomness while maintaining semantic appropriateness, making the digital human's actions more varied and expressive.
[0177] The final output semantic-action association mapping table is a structured dataset that contains the correspondence between each paragraph or sentence in the script and the recommended action type, as well as parameter suggestions for action execution (such as amplitude, speed, and sentiment). This mapping table will serve as a key input for the next step, guiding the generation of accurate action sequences.
[0178] In step S4.2:
[0179] After establishing the association between semantics and actions, this step will ensure that the digital human's actions and speech are precisely synchronized, which is crucial for creating a natural and fluent narration experience.
[0180] First, a detailed acoustic analysis is performed on the speech data stream output from step 3 to extract the following key features:
[0181] Speech rhythm markers: Identify syllable boundaries, stressed syllables, pauses, and intonation changes in speech. A forced alignment algorithm is used to precisely match speech with text, obtaining accurate timestamps for each word and syllable.
[0182] Prosodic structure analysis: Identifies the prosodic hierarchy of speech, including prosodic feet, prosodic phrases, and intonation phrases. These structures provide natural boundaries for action segmentation.
[0183] Emphasis marker extraction: Identifying emphasized parts of speech, including pitch peaks, intensity enhancements, and syllable prolongations. These parts often require corresponding emphatic gestures.
[0184] Based on these acoustic features and the semantic-action association mapping table output from step 4.1, an action-speech synchronization algorithm is applied to calculate the precise execution time of each action. This algorithm follows several core principles:
[0185] Preparation principle: Natural gestures usually begin 0.2-0.5 seconds before the relevant words, and the appropriate preparation time will be dynamically calculated according to the complexity of the action.
[0186] The key principle to emphasize is that the climax of an action (such as the point of maximum amplitude of a gesture) should be precisely synchronized with the key points in speech (such as stressed syllables).
[0187] Smooth ending principle: Actions should be naturally retracted after the semantic unit ends, avoiding abrupt cut-offs.
[0188] The principle of smooth transition: There should be a smooth transition between adjacent actions. The optimal transition path will be calculated to avoid unnatural jumps.
[0189] A dynamic programming algorithm is employed to optimize the overall action sequence, minimizing energy consumption (avoiding overactivity) and maximizing expressiveness while satisfying the aforementioned principles. Specifically, the algorithm identifies natural pauses in speech as ideal opportunities to insert complex actions or action transitions.
[0190] To handle unexpected speech changes (such as temporary pauses or emphasis), a real-time adjustment mechanism has been implemented, which can dynamically fine-tune the timing of actions during playback to ensure that synchronization is not affected.
[0191] The final output time-synchronized motion sequence includes the precise start time, climax time, and end time of each action, as well as the action type and parameters. This sequence provides a time frame for subsequent facial expression generation and final motion-expression coordination.
[0192] In step S43:
[0193] After determining the timing of the action sequence, this step generates facial expressions that match the emotions of the content, making the digital human's expression richer and more vivid.
[0194] First, we combine two key inputs: one is the sentiment markers in the enhanced script output from step 2, which indicate the emotional tendency of the content (such as enthusiasm, seriousness, curiosity, surprise, etc.); the other is the acoustic features in the speech data stream output from step 3, including intonation changes, stress patterns, and emotional tone.
[0195] The facial expression generation employs a multi-layered emotion-expression mapping model:
[0196] Basic Emotional Layer: Maps six basic emotions (joy, sadness, anger, fear, disgust, and surprise) and their combinations to corresponding facial expression parameters. An improved Facial Action Coding (FACS) method is used to control 42 key muscle movement units of the digital human face.
[0197] Specialized teaching emoji layer: tailored to specific emoji needs in teaching scenarios, such as the "thinking" emoji (when explaining complex concepts), the "questioning" emoji (when guiding thinking), and the "confirming" emoji (when emphasizing key points).
[0198] Micro-expression layer: Adds subtle eyebrow movements, eye blinks, and slight adjustments to the corners of the mouth, enhancing the digital human's lifelike appearance. These micro-expressions are generated based on probabilistic models, simulating natural human facial micro-movements.
[0199] The facial expression generation process employs a deep learning-based emotion-expression generation network, trained on a large number of labeled lecturer videos, capable of converting textual and speech emotion features into natural facial expression parameters. The network architecture uses a Transformer encoder-decoder structure, enabling it to capture long sequences of emotion change patterns.
[0200] To ensure the time synchronization of facial expressions, a timing control mechanism closely integrated with speech rhythm is employed. For example, when emphasizing key points, facial expressions (such as raising eyebrows) are precisely synchronized with stressed syllables in the speech; when expressing surprise or doubt, facial expressions precede the corresponding speech, mimicking natural human reaction patterns.
[0201] It also considers the continuity and naturalness of facial expressions, avoiding abrupt changes in expression. This is achieved through an expression smoothing algorithm, which creates natural transition curves between adjacent expressions, ensuring smooth changes in facial expressions.
[0202] The final output time-synchronized facial expression sequence contains precise timestamps for each expression state and detailed facial parameter configurations, providing complete facial expression instructions for subsequent action-expression coordination.
[0203] In step S4.4:
[0204] As the final step in generating motion and facial expressions, this step integrates and coordinates the previously generated motion and facial expression sequences, resolves potential conflicts, and ensures a natural and harmonious overall performance.
[0205] First, the action sequence output from step 4.2 and the facial expression sequence output from step 4.3 are time-aligned and conflict detected. Potential conflicts mainly include the following categories:
[0206] Timing conflict: When two strong expressions (such as important gestures and obvious facial expressions) are too close in time, it may lead to distraction or overperformance.
[0207] Semantic conflict: When the emotions or intentions conveyed by actions and facial expressions are inconsistent, such as gestures indicating affirmation while facial expressions show doubt.
[0208] Physical conflict: When hand gestures need to be close to the face (such as in a "thinking" posture), it is necessary to ensure that they do not interfere with facial expressions.
[0209] Intensity imbalance: When the intensity of an action or expression does not match the importance of the content, such as a key point being expressed weakly or minor information being expressed exaggeratedly.
[0210] The motion and facial expression coordination algorithm uses a combination of rule-based and machine learning methods to resolve these conflicts:
[0211] To address timing conflicts, a staggered timing strategy can be applied by fine-tuning the timing of actions or expressions to ensure that important expressions have a sufficient "exclusive time window." For example, gestures might be slightly advanced or facial expression changes delayed, allowing the audience to perceive these expressions sequentially.
[0212] For semantic conflicts, semantic consistency checks ensure that actions and expressions convey consistent information. When inconsistencies are detected, the side with less conflict is adjusted based on the main emotional tone of the content to maintain overall consistency in expression.
[0213] For physical conflicts, motion adjustment algorithms are applied to fine-tune hand positions while preserving the semantics of the gesture, avoiding obscuring the face or interfering with facial expressions. For example, a "thinking" gesture that should be close to the chin might be moved slightly downwards to ensure that facial expressions are clearly visible.
[0214] To address the imbalance in intensity, a global intensity balancing algorithm dynamically adjusts the intensity of actions and expressions based on the importance level of the content. This ensures that the intensity of expression is proportional to the importance of the content, enhancing the overall teaching effectiveness.
[0215] An expressive smoothing algorithm was also applied to prevent the digital human's movements and expressions from becoming too mechanical or repetitive. This algorithm introduces controlled variability to ensure that even similar content will have subtle differences in expression, enhancing naturalness and vividness.
[0216] The final output, a complete set of motion and facial expression instructions, is a time-series data structure containing precise timestamps and all necessary parameters to drive the digital human's skeleton and facial expressions. This instruction set has been fully coordinated and optimized to ensure the naturalness, coordination, and instructional effectiveness of the digital human's performance.
[0217] The method of ensuring a natural transition and consistency between actions and expressions through the action-expression coordination algorithm includes:
[0218] Step S4.4.1: Perform time alignment on the time-synchronized action sequence and the time-synchronized facial expression sequence, detect the temporal proximity of strong expressions, and output a timing conflict indicator;
[0219] In this step, the time-synchronized action sequence output from step S4.2 and the time-synchronized facial expression sequence output from step S4.3 are first subjected to precise time alignment analysis. This process requires high-precision timestamp comparison and constructs a unified timeline, mapping all action and facial expression events onto this timeline. For each time window (typically 50-100 milliseconds), it checks for the simultaneous occurrence of multiple strong expressions. Strong expressions include obvious gestures (such as large pointing, emphasis on waving), significant changes in posture (such as leaning forward, turning), and obvious facial expressions (such as surprise, emphasis on raising eyebrows). A weighted scoring mechanism is used to evaluate the intensity of each expression, considering factors such as the amplitude of movement, changes in speed, spatial occupancy, and attentional engagement. When two or more strong expressions are detected to be highly overlapping in time (e.g., overlapping by more than 70% of their duration) and their combined intensity exceeds a preset threshold, they are marked as potential temporal conflicts. Such conflicts may lead to audience distraction or perceptual overload, affecting teaching effectiveness. A detailed record is created for each detected conflict, including the type of expression, time range, intensity score, and severity of the conflict. For particularly severe conflicts (such as multiple highest-intensity expressions completely overlapping), higher priority is assigned to ensure they are addressed first in subsequent steps. Furthermore, continuous expression sequences are analyzed to identify regions that may cause "expression congestion," i.e., time periods containing an excessive number of expressions within a short timeframe. This analysis considers the temporal characteristics of human perception, such as the minimum time interval required for attention shifts. Finally, a structured temporal conflict identifier dataset is output, containing detailed information and priority rankings for all potential conflicts, providing a comprehensive basis for subsequent conflict resolution.
[0220] Step S4.4.2: Based on the time conflict identifier, apply the time staggering strategy, and ensure that important expressions have a sufficient exclusive time window by fine-tuning the timing of actions or expressions, and output the time-optimized action and expression sequence.
[0221] In this step, based on the temporal conflict identifiers output in step S4.4.1, an intelligent time staggering strategy is applied to resolve the detected conflicts. First, all conflicts are prioritized, and the processing order is determined based on the conflict severity, the importance of the expressions involved, and the difficulty of adjustment. For each conflict, it is necessary to decide which expressions need adjustment and how to adjust them. This decision-making process follows several key principles: the content importance principle (expressions directly related to the core teaching content are prioritized to remain in their original positions), the perceptual naturalness principle (adjustments should maintain the natural fluency of human actions as much as possible), and the minimum intervention principle (adjustments should be made to the minimum extent possible while meeting the needs). A dynamic programming algorithm is used to calculate the optimal time adjustment scheme, which considers the dependencies and continuity requirements between expressions. For expressions that need adjustment, different strategies are adopted according to their nature: for anticipated actions such as gestures, the start time is tended to be advanced, taking advantage of the natural characteristic that human actions usually precede speech; for reactive expressions, their appearance time may be slightly delayed; for continuous expressions, the timing of their intensity peak may be adjusted to ensure that the peak does not conflict with other expressions. The timing adjustments are typically controlled within the range of 100-500 milliseconds. This range is sufficient to resolve conflicts without causing a noticeable unnatural feeling. New conflicts that may arise after the adjustment are also considered, and global optimization ensures that resolving one conflict does not create new problems. For complex conflicts that cannot be resolved by simple timing adjustments, more complex strategies are considered, such as decomposing complex expressions into multiple simpler expressions or adjusting the execution speed of the expressions. Ultimately, the output is a timing-optimized sequence of actions and expressions. This sequence preserves the semantic intent and expressiveness of the original expressions while ensuring that each important expression has a sufficient "exclusive time window," allowing the audience to clearly perceive the information conveyed by each expression and improving the overall teaching effectiveness.
[0222] Step S4.4.3: Perform semantic consistency checks on the time-optimized action and expression sequence, identify inconsistencies in the information conveyed by actions and expressions, and output a list of semantic conflicts;
[0223] In this step, a thorough semantic consistency analysis is performed on the temporally optimized action and facial expression sequences output from step S4.4.2 to ensure that the information conveyed by actions and expressions is coordinated and mutually reinforcing, rather than contradictory. First, a comprehensive semantic interpretation framework is established, mapping various actions and expressions to their typical semantic and emotional meanings. For example, nodding usually indicates affirmation or agreement, while shaking the head indicates negation or disapproval; smiling indicates positivity or friendliness, while frowning indicates confusion or dissatisfaction. This mapping framework is based on extensive human behavioral research and visual semantic databases, covering common gestures, postures, and facial expressions. Subsequently, semantic analysis is performed on combinations of actions and expressions within each time period to assess whether they convey consistent information. This analysis considers multiple levels of consistency: emotional consistency (whether actions and expressions express the same emotional tendency, such as positive / negative), emphasis consistency (whether the emphasized content is consistent), directionality consistency (whether the direction pointing to or guiding attention is consistent), and logical consistency (whether there are logical contradictions, such as simultaneously expressing affirmation and negation). When potential semantic conflicts are detected, the nature and severity of the conflict are further analyzed. For example, minor inconsistencies (such as a neutral gesture paired with a slight smile) may be acceptable, while significant contradictions (such as affirmative language paired with a negative head movement) require resolution. A detailed record is created for each detected semantic conflict, including the expressive elements of the conflict, the type of conflict, the severity of the conflict, and the relevant content context. Furthermore, cultural factors are considered, as certain gestures and expressions may have different or even opposite meanings in different cultures. The criteria for judging semantic consistency are adjusted based on the cultural background of the target audience. Finally, a structured list of semantic conflicts is output, containing all inconsistencies that need to be resolved, ordered by severity, and accompanied by suggested solutions, providing a basis for further semantic reconciliation.
[0224] Step S4.4.4: Based on the semantic conflict list, adjust the expression with less conflict according to the emotional tendency of the content, maintain the consistency of the overall expression, and output a semantically coordinated action and expression sequence;
[0225] In this step, based on the semantic conflict list output in step S4.4.3, intelligent semantic coordination is performed to resolve inconsistencies between actions and expressions. First, the context of each conflict is analyzed, including the semantic importance, emotional tone, and pedagogical intent of the relevant content. This analysis helps determine which expressive elements should be prioritized in conflict resolution. Generally, expressions directly related to the core teaching content, as well as those more important in terms of emotional or emphasizing function, are prioritized. For each semantic conflict, a decision needs to be made on how to adjust the less conflicting expressions. This decision-making process employs a multi-factor evaluation model, considering the substitutability of the expression (whether there are other ways to convey the same information), the difficulty of adjustment (some expressions are easier to modify than others), and the naturalness of the adjusted expression (whether it still looks natural and fluent after modification). Several adjustment strategies are provided: replacement strategy (replacing the conflicting expression with a semantically compatible alternative), weakening strategy (reducing the intensity of the conflicting expression, making it neutral or auxiliary), enhancement strategy (enhancing the intensity of the dominant expression, making it clearly the dominant semantic message), and fusion strategy (creating a new composite expression that combines the key elements of both expressions). For example, when facial expressions and gestures convey different emotions, the intensity of secondary expressions may be adjusted based on the primary emotional tone of the content. When indicative actions are inconsistent with the direction of gaze, these elements may be realigned to ensure they point to the same target. For complex semantic conflicts, the temporal relationship of expressions is considered, and conflicts may be resolved by adjusting the time order of expressions, such as making one expression explicitly precede or follow another, thereby creating a transition or contrast effect rather than a direct conflict. The overall semantic coherence after adjustment is also evaluated to ensure that resolving local conflicts does not disrupt the broader expressive logic. Ultimately, a semantically coordinated sequence of actions and expressions is output, in which all expressive elements are consistent at the semantic level, collectively enhancing the teaching effectiveness and emotional communication of the content, and avoiding confusion or misunderstanding caused by inconsistent expressions.
[0226] Step S4.4.5: Based on the semantically coordinated action and expression sequence, the intensity of actions and expressions is dynamically adjusted according to the content importance level through a global intensity balancing algorithm, and a complete action and expression instruction set with balanced intensity is output.
[0227] In this step, based on the semantically coordinated action and expression sequence output from step S4.4.4, a global intensity balancing algorithm is applied for final expressiveness optimization. The core objective of this step is to ensure that the intensity of actions and expressions is proportional to the importance of the content, avoiding weak expressions for important content or overly exaggerated expressions for secondary information. First, a content importance model is constructed, which evaluates the importance level of content based on multiple factors, including semantic importance (core concepts, key conclusions, etc.), structural importance (key positions such as titles and summaries), pedagogical importance (key teaching points such as difficulties and test points), and emotional importance (content requiring emotional resonance). The content is divided into multiple importance levels, and a corresponding range of expression intensity is assigned to each level. The global intensity balancing algorithm uses a dynamic programming method to optimize the intensity distribution of the entire sequence while maintaining natural and fluent expression. The algorithm considers multiple constraints: the matching degree between intensity and importance, the smoothness of intensity changes (avoiding abrupt changes in intensity), the dynamic range of overall intensity (ensuring both climaxes and calm periods), and human perception habits (such as progressive intensity changes being more natural than random changes). Adjusting the intensity of movements allows modification of several parameters, including range of motion (the extent of hand gesture movement), execution speed (the speed of the movement), repetition count (for emphasized movements), and additional secondary movements (such as slight body tilting, head coordination, etc.). Adjusting the intensity of facial expressions allows modification of the prominence (e.g., the degree of a smile), duration, transition speed (the speed at which an expression forms and disappears), and complexity (the proportion of various basic expressions mixed together). Special attention is paid to the rhythm of intensity changes, creating natural fluctuations and avoiding monotonous or overly frequent changes. Furthermore, the fatigue effect of prolonged expression is considered, appropriately reducing the intensity of sustained high-intensity expressions and inserting natural relaxation moments to make the overall performance more natural and sustainable. Appropriate intensity variability is also maintained, allowing for subtle differences in expression even for similar content, enhancing naturalness and vividness. Ultimately, a complete set of movement and facial expression instructions with balanced intensity is output. This set is not only semantically consistent and chronologically varied, but also precisely matches the intensity to the importance of the content, maximizing teaching effectiveness while maintaining natural fluency and engaging expression.
[0228] In step S5, the digital human model is driven using MuseTalk technology, including:
[0229] Step S5.1: Based on the predefined digital human model library, select a suitable lecturer image according to the nature of the teaching content and the target audience, initialize the skeletal structure and facial control points, and output a ready digital human model.
[0230] In this step, the first step in the video generation process, it is necessary to prepare and initialize a suitable digital human model to lay the foundation for subsequent animation-driven processes.
[0231] First, based on the nature of the teaching content and the target audience, the most suitable instructor image is selected from a predefined library of digital human models. This selection can be explicitly specified by the user or automatically recommended based on content analysis. The model library contains digital human models with different genders, ages, styles, and professional domains; each model is professionally designed and optimized for specific types of teaching content.
[0232] After selecting a model, perform the following initialization steps:
[0233] Skeletal Loading: The skeleton of a digital human is the core architecture for controlling limb movements. A predefined skeletal hierarchy is loaded, comprising approximately 60-120 bone nodes, covering all parts of the body. The skeleton uses a standard hierarchical structure, employing quaternions or matrices to represent bone rotation, ensuring the accuracy and smoothness of motion calculations.
[0234] Facial control initialization: Facial expression control employs two complementary techniques: one is FACS-based muscle control, which controls approximately 42 facial motion units; the other is blend shapes-based advanced expressions, which provide predefined complex expressions. These two systems are initialized, and the mapping relationship between them is established to ensure the accuracy and richness of expression control.
[0235] Physics parameter settings: To enhance the realism of the animation, initialize the physics simulation parameters, including clothing dynamics, hair dynamics, and secondary animation effects (such as subtle body swaying). These parameters will be adjusted according to the selected digitizer features; for example, a long-haired model will have more complex hair dynamics settings.
[0236] Rendering parameter configuration: Configure rendering parameters related to the digital human, including skin material, subsurface scattering parameters, eye reflection characteristics, etc., to ensure high-quality visual effects in subsequent rendering. These parameters will be dynamically adjusted for different output targets (such as high-definition video or real-time streaming media) to balance quality and performance requirements.
[0237] Initial state setup: Place the digital human in the default narration posture, typically a natural standing or sitting position with relaxed arms and a neutral facial expression. This state serves as the starting point for all animations, ensuring the continuity of the animation sequence.
[0238] A model warm-up technique is employed, pre-calculating and caching commonly used actions and facial expressions before the actual animation begins, reducing computational latency in subsequent processing. For complex digital human models, model simplification algorithms are also applied to create multi-layered versions with varying levels of detail for different perspectives and scene requirements.
[0239] The final output, ready-to-use digital human model, includes a complete skeleton, facial controls, and all necessary rendering parameters, providing a solid foundation for subsequent motion and expression-driven animation.
[0240] Step S5.2: Based on the action and expression instruction set, parse the action and expression instruction set into a skeletal animation control flow, a facial animation control flow, and a lip-sync control flow, and output three parallel control flows;
[0241] In this step, after the digital human model is ready, this step will use the action and expression instruction set output in step 4 to drive the digital human and achieve animation effects that are precisely synchronized with the voice.
[0242] First, the action and expression instruction set is parsed into three parallel but interconnected control flows:
[0243] Skeletal animation control flow: Controlling the limb movements of a digital human, including gestures, posture changes, and body movements.
[0244] Facial animation control flow: Controls facial expressions, including the dynamic changes of facial elements such as eyebrows, eyes, and mouth.
[0245] Lip-sync control flow: This is specifically responsible for the precise matching of lip movements and speech, and is a special subset of facial animation.
[0246] The core animation engine utilizes technologies such as MuseTalk, which integrates advanced skeletal animation and facial expression control. The specific driving process is as follows:
[0247] For skeletal animation, a combination of keyframe interpolation and procedural animation is used. Keyframes define the critical states of an action (such as the start, climax, and end positions of a gesture), while intermediate frames are automatically calculated to ensure smooth transitions. To enhance naturalness, a secondary motion generation algorithm is applied, adding subtle body swaying, weight transfer, and momentum effects to make the animation more dynamic.
[0248] For facial animation, parametric facial expression control is employed, mapping the expression parameters generated in step 4.3 onto the digital human's facial controller. Simultaneously, high-level expressions (such as "smiling" and "thinking") and low-level muscle movement units are controlled to achieve rich and nuanced facial expression variations. In particular, natural blinking, micro-expressions, and subtle head movements are added to enhance the sense of life.
[0249] Lip-sync is the most crucial part of facial animation, directly affecting the viewer's perception of synchronization. Advanced phoneme-to-visual mapping technology is employed to convert the phoneme sequence in speech into corresponding lip shapes. This process considers coarticulation (i.e., the influence of adjacent phonemes on lip shape) and applies a lip-smoothing algorithm to avoid overly mechanical lip movements. Special attention is also paid to stressed syllables to match the lip opening and closing amplitude with the volume intensity.
[0250] To ensure consistency across control flows, a global animation coordinator is employed. This coordinator oversees all animation channels, resolves potential conflicts, and ensures overall performance harmony. Specifically, the coordinator ensures that gestures do not interfere with facial expressions, facial expressions do not disrupt lip-sync, and all actions are semantically and temporally aligned with the audio content.
[0251] It also implements a dynamic adaptation mechanism that can adjust animation parameters based on real-time feedback. For example, if it detects that certain actions may cause self-occlusion at a specific angle, it will automatically fine-tune the amplitude or direction of the actions to ensure that key expressions are clearly visible.
[0252] The final output is a complete time sequence of animations, including full-body skeletal transformations and facial expression changes, precisely synchronized with the audio data, ready for the next step of visual content integration.
[0253] Among them, museTalk technology is a digital human-driven technology specifically designed for teaching scenarios. Its core is a multimodal collaborative driving engine that enables precise coordination of voice, facial expressions, and actions. This technology is based on a deep learning architecture and includes three key models: a temporal collaborative Transformer network, an expression generation GAN network, and a physically constrained reinforcement learning model.
[0254] The temporal collaborative Transformer network employs an encoder-decoder architecture, containing 8 Transformer layers, each with 6 attention heads, and a hidden layer dimension of 768. The network's training data comes from 500 hours of manually annotated high-quality instructional videos, including annotations of speech text, phoneme timestamps, facial landmark motion, and limb movement parameters. Training employs a two-stage strategy: first, self-supervised pre-training on a large-scale unlabeled dataset, followed by supervised fine-tuning on labeled data. Key hyperparameters include a learning rate of 0.00005, the AdamW optimizer, a weight decay of 0.01, and approximately 200,000 training iterations.
[0255] The facial expression generation GAN network employs a conditional generative adversarial network architecture. The generator uses a U-Net structure, while the discriminator uses a PatchGAN design. This network receives emotion labels and speech features as conditional inputs to generate natural facial expression parameters. The training data includes a dataset of facial expressions from 100 different performers, covering 6 basic emotions and 20 complex emotions.
[0256] Physics-constrained reinforcement learning models optimize generated action sequences by simulating constraints in the physical world (such as inertia, gravity, and joint limitations) to ensure the physical plausibility of the actions. This model is trained using the TD3 (Twin Delayed DDPG) algorithm, and its reward function comprehensively considers three dimensions: action fluency, expressiveness, and physical plausibility.
[0257] Experimental results show that, compared with traditional methods, museTalk technology improves voice-action synchronization accuracy by 35%, facial expression naturalness score by 42%, and user satisfaction by 27%. This technology has been validated in over 5,000 teaching video generation cases, demonstrating superior stability and expressiveness.
[0258] Step S5.3: Based on the skeletal animation control flow, use keyframe interpolation and procedural animation methods to control the limb movements of the digital human, add subtle body swaying, weight transfer and momentum effects, and output a natural skeletal animation sequence.
[0259] In this step, natural and fluid digital human limb movements are generated based on the skeletal animation control flow parsed in step S5.2. A hybrid approach combining keyframe interpolation and procedural animation is employed, ensuring precise control of the movements while adding natural variation and liveliness. First, precise keyframes are defined for each key movement according to the instructions in the motion control flow. These keyframes define the key states of the movement, such as the starting position, maximum extension position, and ending position of a gesture; the turning angle of the body; or the nodding amplitude of the head, etc. Advanced inverse skeletal dynamics (IK) is used to ensure that these key poses are anatomically correct and natural. Subsequently, advanced interpolation algorithms are applied to generate smooth transitions between keyframes. Unlike simple linear interpolation, physically based spline interpolation is used, considering acceleration and deceleration to simulate the characteristics of real human motion. For example, there is an acceleration process at the beginning of the movement, a deceleration process at the end, and possibly a steady speed phase in between. This interpolation method makes the movements look more natural and fluid, avoiding a mechanical and stiff feeling. To further enhance the naturalness of the movements, multi-layered procedural animation techniques are applied. First, secondary motion generation automatically adds secondary movements that complement the primary movement, such as a slight shoulder lift during an arm swing or a slight forward lean during a pointing motion. These subtle secondary movements are crucial for enhancing the overall naturalness. Second, natural variation generation adds slight random variations while maintaining semantic consistency, ensuring subtle differences in each repetition to avoid identical mechanical repetition. Various physical effects simulations are also added, including body sway (such as minor balance adjustments while standing), weight transfer (such as the natural shift in center of gravity from one foot to the other), and momentum effects (such as the natural inertia and rebound after a rapid movement). These effects are achieved through simplified physical simulations, striking a balance between computational efficiency and visual appeal. Furthermore, personalized movement styles are considered, adjusting the overall style of movements based on selected digital human characteristics (such as age, gender, and professional background); for example, older lecturers might move more steadily, while younger lecturers might move more lively. Contextual awareness of movements is also implemented to ensure natural transitions and coherence between adjacent actions, avoiding abrupt changes in posture. Ultimately, a natural skeletal animation sequence is output. This sequence contains complete skeletal transformation data, which can drive the digital human model to perform smooth and natural limb movements, providing a foundation for subsequent overall animation compositing.
[0260] Step S5.4: Based on the facial animation control flow, parametric facial expression control is used to map expression parameters to the digital human face controller, adding natural blinking, micro-expressions and subtle head movements, and outputting rich facial animation sequences.
[0261] In this step, based on the facial animation control flow parsed in step S5.2, rich and natural digital human facial expressions are generated. A parametric facial expression control method is employed, which precisely maps abstract expression parameters (such as "degree of smile," "degree of surprise," etc.) to specific control points on the digital human face. Modern digital human faces typically employ two complementary control methods: one is muscle control based on FACS (Facial Action Coding), controlling approximately 42 facial action units (AUs), such as raising eyebrows and upturning corners of the mouth; the other is advanced expressions based on blend shapes, providing predefined composite expressions such as "smile" and "thinking." First, the expression instructions in the facial animation control flow are parsed into control parameters for these two methods. For basic emotional expressions (such as joy and surprise), optimized expression mapping templates are used to ensure accuracy and naturalness. For teaching-specific expressions (such as "thinking" and "emphasis"), expression templates specifically designed for teaching scenarios are applied; these templates have been trained and optimized using professional instructor expression data. Special emphasis is placed on the subtlety and natural transitions of facial expressions, employing advanced expression interpolation algorithms to ensure smooth and natural changes, avoiding abrupt jumps. To enhance the lifelikeness and naturalness of facial expressions, three key auxiliary animations have been added: natural blinking, micro-expressions, and subtle head movements. Natural blinking is a crucial element in enhancing the lifelikeness of digital humans. A blink generator based on a probabilistic model has been implemented, which dynamically adjusts the frequency and manner of blinking based on various factors, including emotional state (e.g., increased blinking frequency when nervous), attention state (e.g., decreased blinking when focusing on explanation), and physiological needs (e.g., necessary blinking after prolonged periods without blinking). Micro-expressions are subtle and transient changes in human facial expressions that occur naturally. By analyzing the emotional flow of the content and the rhythm of the explanation, appropriate micro-expressions are automatically generated, such as a brief slight raise of the eyebrows, a slight twitch of the corner of the mouth, or a slight flaring of the nostrils. Although these micro-expressions are subtle, they are crucial for breaking the static feel of the face and enhancing its lifelikeness. Subtle head movements are natural accompanying movements of facial expressions. Based on the type and intensity of the expression, corresponding subtle head movements are automatically generated, such as nodding, slight shaking of the head, slight tilting, or raising the head. These head movements are closely coordinated with facial expressions, enhancing the overall coherence and naturalness of the expression. Personalized customization of expressions is also achieved, adjusting the overall style of expressions based on selected digital human characteristics. For example, digital humans of different ages, genders, or personality traits will have different expression tendencies and modes of expression. Furthermore, cultural adaptability is considered, allowing adjustments to the expression of certain expressions based on the cultural background of the target audience to ensure cultural appropriateness. Finally, a rich sequence of facial animations is output, containing complete facial control parameter data, enabling the digital human model to display natural, vivid, and expressive facial expressions, providing a crucial component for subsequent overall animation synthesis.
[0262] Step S5.5: Based on the lip-sync control flow and the speech data flow, the phoneme sequence in the speech is converted into the corresponding lip shape using phoneme-to-visual mapping technology. The lip-smoothing algorithm is applied to avoid the lip-shape changes being too mechanical, and a complete animation sequence synchronized with the speech is output.
[0263] In this step, based on the lip-sync control flow parsed in step S5.2 and the speech data stream output in step S3, a precisely synchronized lip-sync animation is generated. Lip-sync is one of the most critical aspects of digital human animation, directly affecting the audience's perception of synchronicity and realism. Advanced phoneme-to-visual mapping technology is employed to convert the phoneme sequence (the basic unit of speech articulation) in speech into corresponding lip shapes (the basic unit of visual lip shape). First, a detailed acoustic analysis is performed on the speech data stream to extract the precise phoneme sequence and its timestamps. This process uses a forced alignment algorithm to precisely match the speech with the text, obtaining the start and end times of each phoneme. For synthesized speech, this information can be directly obtained from the speech synthesis process; for recorded speech, it needs to be analyzed and extracted using an acoustic model. Subsequently, the phoneme-to-visual mapping rules are applied to convert each phoneme or phoneme combination into the corresponding lip shape. This mapping takes into account the characteristics of language; different languages may have different mapping rules. The mapping library contains various mouth shapes, such as the open shapes corresponding to the vowels "a", "i", and "u", and the closed shapes corresponding to the consonants "m", "p", and "b". It specifically considers the coarticulation effect, that is, the mutual influence of adjacent phonemes on mouth shapes. For example, the same phoneme may have different mouth shapes in different contexts; these subtle changes are captured through context-aware mapping rules. To make the mouth shape animation more natural and smooth, an advanced mouth shape smoothing algorithm is applied. Simple direct mapping from phoneme to visual pixel can lead to overly mechanical and abrupt mouth shape changes. The smoothing algorithm simulates the continuous changes in mouth shapes during human speech by creating natural transitions between mouth shapes. A physically based mouth shape interpolation model is used, considering the movement characteristics of oral muscles, such as the time required for different mouth shape transitions and the transition path. It also specifically handles stressed syllables, matching the mouth opening and closing amplitude to the volume intensity. For emphasized words or syllables, the mouth opening and closing amplitude and clarity are increased; for fast or unstressed parts, the mouth shape changes are correspondingly reduced. Furthermore, personalized lip-sync adjustments were implemented. Lip-sync parameters were fine-tuned based on the digital human's facial features and speaking style; for example, some people open and close their mouths more widely than others, and some have specific mouth movement patterns. The influence of non-vocal factors on lip-sync, such as laughter, sighs, or deep breaths, was also considered, generating corresponding lip-sync animations for these special sounds. To ensure precise synchronization between lip-sync and speech, an adaptive timing adjustment mechanism was implemented, allowing for fine-tuning of the lip-sync animation timing based on actual playback conditions to compensate for possible delays or asynchrony. Finally, the lip-sync animation was integrated with the previously generated facial expression and skeletal animations to output a complete animation sequence perfectly synchronized with the speech. This sequence includes full-body skeletal transformations, facial expression changes, and precise lip-sync animation, forming a coordinated and unified whole, providing a complete digital human animation foundation for subsequent visual content integration.
[0264] The integration of the digital human instructor and teaching content into a unified visual scene through an optimized parallel rendering algorithm includes:
[0265] Step S5.6: Based on the PPT content in the structured data, extract visual elements including slides, images, charts, and text key points, analyze content readability, lecturer visibility, content-lecturer relationship, and screen balance, and output the most suitable screen layout strategy.
[0266] Step S5.7: Based on the screen layout strategy and the complete animation sequence, automatically adjust the position, angle and focal length of the virtual camera according to the importance of the content and the lecturer's actions, add dynamic visual emphasis effects including highlighting, magnification and arrow indication to important content, and output a visual scene with cinematic language.
[0267] Step S5.8: Based on the visual scene with camera language, process the interaction effect between the lecturer and the content, apply transition effects including fade-in / fade-out, wipe-in / wipe-out, and zoom transition at the content switching points, and output an interactive scene sequence.
[0268] Step S5.9: Based on the enhanced scene sequence, apply color correction, lighting balance and layer blending techniques to ensure visual unity between the digital human lecturer and the background and content, and output a complete scene sequence with visual integration.
[0269] Step S5.10: Based on the complete scene sequence of the visual integration, output a professional-grade complete video scene sequence through high dynamic range rendering, temporal anti-aliasing, motion blur and color grading processing.
[0270] Step S5.6 Visual Content Integration and Composition
[0271] After obtaining the animation sequence of the digital human, this step integrates the digital human instructor with the teaching content (such as PPT slides, charts, presentation materials, etc.) into a unified visual scene to create a complete teaching video.
[0272] First, analyze the PPT content analysis results from Step 1, extracting all visual elements (such as slides, images, charts, text highlights, etc.) and their logical structure. Based on this analysis, design the most suitable screen layout and content presentation strategy, considering the following key factors:
[0273] Content readability: Ensure that text and charts are large enough and clear, especially when viewed on mobile devices.
[0274] Instructor visibility: Ensure that key expressions of the digital human instructor (such as facial expressions and important gestures) are clearly visible.
[0275] Content-Lecturer Relationship: Optimize the spatial relationship between the lecturer and the content so that the lecturer's gaze and directional gestures can naturally guide the audience's attention.
[0276] Image balance: Create a visually balanced composition to avoid an overly crowded or empty image.
[0277] It supports various layout templates for different scenarios, such as "instructor side + main content", "small screen of instructor + full screen of content", and "instructor and content partitions". Based on the content type and complexity, it will dynamically select the most suitable layout and switch the layout naturally at content transition points.
[0278] To enhance teaching effectiveness, intelligent camera language was implemented, including the following technologies:
[0279] Dynamic camera adjustment: The virtual camera's position, angle, and focal length are automatically adjusted based on the importance of the content and the speaker's actions. For example, a wide-angle lens is used to show both the speaker and the content when explaining complex charts; when emphasizing key points, the camera may switch to a close-up shot of the speaker.
[0280] Visual emphasis effects: Add dynamic visual emphasis to important content, such as highlighting, magnification, and arrow indicators. These emphasis effects are precisely synchronized with the speaker's voice and gestures, enhancing the directional effect.
[0281] Smooth transition effects: At content switching points, professional transition effects are applied, such as fade-in / fade-out, wipe-in / wipe-out, and zoom transitions, to ensure visual smoothness. The logical relationships within the content are analyzed to select the most suitable transition type.
[0282] It also handles the interaction between the instructor and the content. For example, when the instructor points to a point in the PPT, it may trigger corresponding visual feedback (such as highlighting an element); when the instructor explains a chart, the chart may be gradually constructed or dynamically changed as the explanation progresses. These interactive effects are achieved through precise time synchronization and spatial mapping, enhancing the intuitiveness and interactivity of the teaching.
[0283] To ensure professional visual quality, advanced compositing techniques were applied, including color correction, lighting balance, and layer blending. In particular, it was ensured that the digital human instructor visually harmonized with the background and content, avoiding a "texture" effect and creating a visual effect as if the instructor were truly present in the teaching environment.
[0284] The final output is a complete video scene sequence containing the digital human instructor, teaching content, and all visual effects, ready for the final rendering stage in the form of a high-quality video frame sequence.
[0285] The final output of the instructor's teaching videos includes:
[0286] Step S5.11: Based on the complete video scene sequence and the audio data stream, decompose the entire video rendering task into time-segmented subtasks and spatial-segmented subtasks that can be processed in parallel, and output the decomposed rendering task set.
[0287] Step S5.12: Based on the decomposed rendering task set, intelligently allocate computing resources according to the available CPU cores and GPU units, prioritize the processing of tasks on the critical path to ensure balanced rendering progress, and output a resource-optimized rendering schedule.
[0288] Step S5.13: Based on the resource-optimized rendering schedule, the GPU's parallel computing capabilities are used to accelerate lighting calculations, material rendering, and post-processing effects. A multi-resolution rendering strategy is adopted to first obtain a preview by rendering at a lower resolution and then output a progressively higher quality rendering result.
[0289] Step S5.14: Based on the rendering results of the progressive quality, gradually improve the quality to the target resolution, perform audiovisual synchronization verification, image quality evaluation and compatibility testing, and output a high-definition video file with quality verification.
[0290] Step S5.15: Based on the high-definition video file with the quality verification, generate multiple format variants including MP4, WebM, and low-bandwidth optimized according to different usage scenarios, generate complete metadata including chapter marks, keyword indexes, and content summaries, and output the final lecturer teaching video.
[0291] Steps S5.11-S5.15 Parallel Rendering and Video Generation
[0292] As the final step in the entire process, parallel rendering and video generation transform all the previously prepared elements into a final, high-quality video file.
[0293] First, integrate two key inputs: the complete video scene sequence (containing all visual elements) output from step 5.3 and the optimized speech data stream output from step 3.4. These two inputs are now fully synchronized in time and ready for final synthesis.
[0294] To efficiently handle large-scale video rendering tasks, an innovative multi-stage parallel rendering pipeline architecture is adopted, including the following key components:
[0295] Task decomposition engine: Decomposes the entire video rendering task into subtasks that can be processed in parallel. Decomposition strategies include temporal segmentation (dividing the video into multiple time segments) and spatial segmentation (dividing each frame into multiple rendering regions). The system dynamically determines the optimal decomposition granularity, balancing parallel efficiency and management overhead.
[0296] Resource scheduler: Intelligently allocates computing resources based on available computing resources (CPU cores, GPU units) and the characteristics of subtasks. It prioritizes tasks on the critical path to ensure balanced rendering progress.
[0297] GPU-accelerated rendering core: Leveraging the parallel computing capabilities of modern GPUs, it accelerates critical rendering operations such as lighting calculations, material rendering, and post-processing effects. The system is optimized for different GPU architectures to ensure maximum hardware utilization.
[0298] Progressive quality control: Employing a multi-resolution rendering strategy, the entire video is first rendered at a lower resolution to obtain a preview, and then the quality is gradually increased to the target resolution. This method allows users to preview the results and make adjustments as needed.
[0299] During the rendering process, the system applied a series of professional-grade image processing and video enhancement technologies:
[0300] High Dynamic Range (HDR) rendering: Ensures details are clearly visible under different lighting conditions.
[0301] Temporal anti-aliasing (TAA): Reduces edge flickering and jitter in animation.
[0302] Motion blur: Add appropriate motion blur effects to enhance the smoothness of movements.
[0303] Color grading: Professional color processing is applied to ensure that the videos have a consistent and appealing visual style.
[0304] Intelligent compression: Dynamically adjusts encoding parameters based on content characteristics to achieve the best balance between file size and quality.
[0305] It supports multiple output formats and quality presets to suit different usage scenarios: such as high-definition MP4 format for regular playback, WebM format for web page embedding, low-bandwidth optimized version for mobile devices, and high-quality version for professional presentation.
[0306] To meet the needs of different platforms, it will automatically generate multiple video variations, such as a square version for social media, a portrait version for mobile devices to watch in portrait mode, and versions with different aspect ratios to adapt to various playback environments.
[0307] After rendering is complete, a final quality check is performed, including audiovisual synchronization verification, image quality assessment, and compatibility testing. Finally, complete metadata is generated, including chapter tags, keyword indexes, and content summaries, to facilitate subsequent content management and retrieval.
[0308] The final output of the instructor's teaching video is a professional-grade video file with high-quality visuals, clear audio, and natural and smooth digital human performance, fully realizing the automated conversion from PPT or text to professional teaching videos.
[0309] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention. < / emphasis> < / prosody> < / break>
Claims
1. A method for automatically generating lecturer videos based on AI speech synthesis and animation-driven technology, characterized in that, Includes the following steps: Based on the content features of user-uploaded PPT files or text scripts, an improved interior-point method is used to construct a content relationship graph. An incremental shortest path algorithm is then applied to identify the hierarchical relationships between titles, paragraphs, and key points, outputting structured data containing semantic structure, key content, and sentiment. The improved interior-point method for constructing the content relationship graph includes: optimizing the traditional O(n³) complexity to near linear time O(n log n) by maintaining a dynamically updated network flow model; treating content elements as nodes in the content relationship graph and establishing initial connections based on the positional relationships, format features, and semantic associations between elements; and after constructing the initial graph, applying an incremental shortest path algorithm to identify the hierarchical structure of the content. Based on the structured data, a fully dynamic parallel single-link clustering algorithm is applied to group the content of the structured data according to semantic similarity. The data is then converted into spoken narrative text through a preset semantic understanding model and marked with pauses, intonation changes, and emphasis. An enhanced script with full expressive markings is then output. Based on the aforementioned enhancement script, using CosyVoice speech synthesis technology, the large-scale feature matrix is decomposed into multiple blocks through the Block-Nyström low-rank approximation method, and block sampling and parallel computation optimization are performed to adjust the prosody and expressiveness, outputting an expressive speech data stream. Based on the enhanced script and the voice data stream, key verbs, emphasis words and emotional words are identified through semantic analysis, a mapping relationship between semantic content and appropriate body movements is established, the precise time point of action execution is calculated and facial expression instructions consistent with the content emotion are generated, and a complete set of action expression instructions is output. Based on the voice data stream and the action / expression instruction set, the digital human model is driven by MuseTalk technology. An optimized parallel rendering algorithm integrates the digital human instructor and teaching content into a unified visual scene, outputting the final instructor teaching video. The integration of the digital human instructor and teaching content into a unified visual scene using the optimized parallel rendering algorithm includes: Based on the PPT content in the structured data, visual elements including slides, images, charts, and text key points are extracted. The readability of the content, the visibility of the speaker, the content-speaker relationship, and the balance of the screen are analyzed, and the most suitable screen layout strategy is output. Based on the aforementioned screen layout strategy and complete animation sequence, the position, angle, and focal length of the virtual camera are automatically adjusted according to the importance of the content and the lecturer's actions. Dynamic visual emphasis effects, including highlighting, magnification, and arrow indication, are added to important content, and visual scenes with cinematic language are output. Based on the visual scene with cinematic language, the interaction effect between the lecturer and the content is processed, and transition effects including fade-in / fade-out, wipe-in / wipe-out, and zoom transition are applied at the content switching points to output a scene sequence with enhanced interaction. Based on the enhanced scene sequence, color correction, lighting balance and layer blending techniques are applied to ensure that the digital human lecturer is visually consistent with the background and content, and to output a visually integrated complete scene sequence. Based on the complete scene sequence of the aforementioned visual integration, a professional-grade complete video scene sequence is output through high dynamic range rendering, temporal anti-aliasing, motion blur, and color grading processing.
2. The method according to claim 1, characterized in that, Based on the content features of user-uploaded PPT files or text scripts, a content relationship graph is constructed using an improved interior point method. An incremental shortest path algorithm is applied to identify the hierarchical relationships between titles, paragraphs, and key points, outputting structured data containing semantic structure, key content, and sentiment, including: The content features are processed by document parsing to extract elements including text, images, and tables, and a preliminary set of parsed content elements is output. Based on the content element set, an improved interior point method is applied to construct a content relationship graph. The incremental shortest path algorithm is used to identify the hierarchical relationship between titles, paragraphs, and key points, and a structured document with hierarchical tags is output. Based on the structured document with hierarchical tags, natural language processing technology is used to identify keywords, key content and core concepts, and semantic vectorization is used to represent the correlation between content, outputting the structured data containing semantic structure, key content and sentiment.
3. The method according to claim 1, characterized in that, Based on the structured data, a fully dynamic parallel single-link clustering algorithm is applied to group the content of the structured data according to semantic similarity. This is then converted into spoken narrative text using a pre-defined semantic understanding model, and markers including pauses, intonation variations, and emphasis are added. The resulting enhanced script with complete expressive markers is output, including: Based on the structured data, a fully dynamic parallel single-link clustering algorithm is applied to maintain a dynamically updated network flow model to group the content semantically by similarity. The content is divided into multiple semantically related content blocks according to the topic and logical relationship, and the semantically clustered content blocks are output. Based on the content segmentation of the semantic clustering, the structured data is converted into narrative text suitable for spoken expression through a preset semantic understanding model, transition words and conjunctions are added to enhance fluency, and a preliminary speech script is output. Based on the initial speech script, a speech synthesis markup language (SSML) script containing markers for pauses, intonation changes, and emphasis is generated, and the enhanced script with full expressive markers is output.
4. The method according to claim 1, characterized in that, Based on the enhancement script, the CosyVoice speech synthesis technology is used to decompose a large-scale feature matrix into multiple blocks using the Block-Nyström low-rank approximation method, and then perform block sampling and parallel computation optimization to adjust prosody and expressiveness, outputting an expressive speech data stream, including: Based on the enhanced script and combined with the lecturer style parameters, the optimal speech synthesis configuration is determined through a parameter optimization algorithm, and a speech synthesis parameter set is output. Based on the speech synthesis parameter set and the enhancement script, the CosyVoice speech synthesis technology is used to perform block sampling on the large-scale feature matrix and parallel computation optimization through the Block-Nyström low-rank approximation method to accelerate speech feature extraction and synthesis and output high-quality speech waveform data. Based on the high-quality speech waveform data, a prosodic adjustment algorithm is applied to dynamically adjust the pauses, stresses, and intonation changes of the speech according to the Speech Synthesis Markup Language (SSML) markers, and outputs an expressive speech data stream.
5. The method according to claim 1, characterized in that, Based on the enhanced script and the speech data stream, the system identifies key verbs, emphasis words, and emotional words through semantic analysis, establishes a mapping relationship between semantic content and appropriate body movements, calculates the precise timing of the action execution, generates facial expression commands consistent with the emotional content, and outputs a complete set of action and expression commands, including: Based on the enhanced script, key verbs, emphasis words, and emotional words are identified through semantic analysis technology, a mapping relationship between semantic content and appropriate body movements is established, and a semantic-action association mapping table is output. Based on the speech data stream and the semantic-action association mapping table, the timing of action execution is calculated by analyzing the rhythm, pauses and stress features of the speech, and a time-synchronized action sequence is output. Based on the emotion markers in the enhanced script and the intonation changes in the speech data stream, facial expression commands consistent with the content emotion are generated through the emotion-expression mapping model, and a time-synchronized expression sequence is output. Based on the time-synchronized action sequence and the time-synchronized facial expression sequence, an action-expression coordination algorithm is used to ensure a natural transition and coordination between actions and expressions, and to output a complete set of action-expression instructions.
6. The method according to claim 5, characterized in that, The algorithm for coordinating actions and expressions to ensure a natural transition and consistency between them includes: The time-synchronized action sequence and the time-synchronized facial expression sequence are time-aligned, the temporal proximity of strong expressions is detected, and a timing conflict indicator is output. Based on the aforementioned temporal conflict identifier, a time staggering strategy is applied to ensure that important expressions have sufficient exclusive time windows by fine-tuning the timing of actions or expressions, and to output a temporally optimized action and expression sequence. Perform semantic consistency checks on the time-optimized action-expression sequences, identify inconsistencies in the information conveyed by actions and expressions, and output a list of semantic conflicts. Based on the semantic conflict list, the expression methods with less conflict are adjusted according to the emotional tendency of the content to maintain the consistency of the overall expression and output a semantically coordinated action and expression sequence. Based on the semantically coordinated action and expression sequence, the intensity of actions and expressions is dynamically adjusted according to the content importance level through a global intensity balancing algorithm, and a complete action and expression instruction set with balanced intensity is output.
7. The method according to claim 1, characterized in that, The use of MuseTalk technology to drive the digital human model includes: Based on a predefined digital human model library, a suitable lecturer image is selected according to the nature of the teaching content and the target audience, the skeletal structure and facial control points are initialized, and a ready-to-use digital human model is output. Based on the aforementioned action and expression instruction set, the action and expression instruction set is parsed into a skeletal animation control flow, a facial animation control flow, and a lip-sync control flow, and three parallel control flows are output. Based on the skeletal animation control flow, keyframe interpolation and procedural animation methods are used to control the limb movements of the digital human, adding subtle body swaying, weight transfer and momentum effects, and outputting a natural skeletal animation sequence. Based on the facial animation control flow, parametric facial expression control is used to map expression parameters to a digital human face controller, adding natural blinking, micro-expressions and subtle head movements, and outputting rich facial animation sequences. Based on the lip-sync control flow and the speech data flow, a phoneme-to-visual mapping technique is used to convert the phoneme sequence in the speech into the corresponding lip shape. A lip-smoothing algorithm is applied to avoid the lip-shape changes from being too mechanical, and a complete animation sequence synchronized with the speech is output.
8. The method according to claim 1, characterized in that, The final output of the instructor's teaching videos includes: Based on the complete video scene sequence and the audio data stream, the entire video rendering task is decomposed into time-segmented subtasks and spatial-segmented subtasks that can be processed in parallel, and the decomposed rendering task set is output. Based on the decomposed rendering task set, computing resources are intelligently allocated according to the available CPU cores and GPU units, and tasks on the critical path are prioritized to ensure balanced rendering progress, resulting in a resource-optimized rendering schedule. Based on the resource-optimized rendering schedule, the parallel computing capabilities of the GPU are used to accelerate lighting calculation, material rendering and post-processing. A multi-resolution rendering strategy is adopted to first obtain a preview by rendering at a lower resolution and then output a rendering result with progressive quality. Based on the rendering results of the progressive quality, the quality is gradually improved to the target resolution, and audiovisual synchronization verification, image quality evaluation and compatibility testing are performed to output a high-definition video file with quality verification. Based on the high-definition video file with the aforementioned quality verification, various format variants, including MP4, WebM, and low-bandwidth optimized, are generated according to different usage scenarios. Complete metadata containing chapter markers, keyword indexes, and content summaries is generated, and the final lecturer teaching video is output.
Citation Information
Patent Citations
Text-to-video conversion method and device based on deep semantic analysis, equipment and medium
CN119399330A
Intelligent dialogue method and system based on digital human
CN120216646A