Automatic lecturer video generation method based on AI speech synthesis and animation driving
By improving the content relationship graph and semantic grouping algorithm, combined with efficient speech synthesis and animation-driven technology, the problems of low efficiency and insufficient expressiveness in educational video production are solved, enabling rapid understanding of PPT content and high-quality video generation.
Patent Information
- Application Number
- CN202510989167.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-07-17
AI Technical Summary
Existing educational video production technologies are inefficient, lack the expressive power of digital humans, cannot achieve rapid processing and real-time generation of large-scale content, and lack in-depth understanding of PPT or text content and emotional expression in speech.
A content relationship graph is constructed using an improved interior point method. An incremental shortest path algorithm is applied to identify hierarchical relationships. A fully dynamic parallel single-link clustering algorithm is used for semantic grouping. CosyVoice speech synthesis technology is used for prosodic adjustment. A mapping relationship between semantic content and action/expression is established. MuseTalk technology is used to drive a digital human model for video generation.
It enables real-time incremental parsing of large PPT documents, improves the accuracy and efficiency of content hierarchy recognition, enhances the speed and quality of speech synthesis, generates actions and expressions that highly match the content being presented, supports near real-time video generation capabilities, and meets the needs of rapid processing of large-scale content.
Smart Images

Figure CN120897102A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the fields of artificial intelligence, computer graphics and educational technology, and particularly relates to a lecturer video automatic generation method based on AI speech synthesis and animation driving. BACKGROUND
[0002] The field of education training and online courses has been seeking more efficient and attractive content production methods. Traditional education video production usually requires professional lecturers to record, edit and process, which is time-consuming, costly and difficult to achieve rapid iteration and update of content.
[0003] There are two common methods of education video production in the market: one is real person recording, which requires professional camera equipment, recording equipment and post-production team; the other is simple PPT screen recording with voiceover, which has low production cost but poor interactivity and attractiveness. These methods cannot meet the current rapid changing needs of education content and are inefficient.
[0004] More advanced technologies usually use basic AI speech synthesis and simple digital human technology to convert text content into speech and drive a pre-set digital human image for simple lip synchronization. This technology realizes preliminary automation, but still has obvious shortcomings in content understanding, speech expressiveness and naturalness of action and expression.
[0005] These existing technologies have several key problems: first, they lack deep understanding and structured analysis of PPT or text content, resulting in lack of corresponding intonation changes and emotional expression in generated speech; second, the actions and expressions of digital humans lack semantic association with the content of the explanation, appearing stiff and mechanical; finally, the entire system is inefficient in processing large amounts of content, making it difficult to achieve real-time or near real-time video generation, and unable to meet the needs of large-scale content production. SUMMARY
[0006] The purpose of the present application is to provide a lecturer video automatic generation method based on AI speech synthesis and animation driving, which solves the problems of low production efficiency, insufficient digital human expressiveness and limited large-scale content processing capacity in existing technologies.
[0007] To achieve the above purpose, the technical solution provided by the present application is:
[0008] A lecturer video automatic generation method based on AI speech synthesis and animation driving, comprising the following steps:
[0009] Based on the content features of the user-uploaded PPT file or text script, a content relationship graph is constructed by an improved interior point method, and an incremental shortest path algorithm is applied to identify the hierarchical relationship between the title, paragraphs, and main points, and output structured data containing semantic structure, key content, and emotional tendency.
[0010] Based on the structured data, a full-dynamic parallel single-link clustering algorithm is applied to group the content of the structured data based on semantic similarity, and a pre-set semantic understanding model is used to convert the content into narrative text in spoken language and add markers including pauses, tone changes, and emphasis, and output an enhanced script with complete expressiveness markers.
[0011] Based on the enhanced script, CosyVoice speech synthesis technology is used to A low-rank approximation method is used to decompose a large-scale feature matrix into multiple blocks and perform block sampling and parallel computing optimization, adjust the prosody and expressiveness, and output expressive voice data stream.
[0012] Based on the enhanced script and the voice data stream, key verbs, emphasized words, and emotional words are identified through semantic analysis, a mapping relationship between semantic content and appropriate body movements is established, the precise time points of movement execution are calculated, and facial expression instructions consistent with the content emotion are generated, and a complete set of action expression instructions is output.
[0013] Based on the voice data stream and the action expression instruction set, a museTalk technology is used to drive a digital human model, and an optimized parallel rendering algorithm is used to integrate the digital human lecturer and the teaching content into a unified visual scene, and output the final lecturer teaching video.
[0014] Preferably, based on the content features of the user-uploaded PPT file or text script, a content relationship graph is constructed by an improved interior point method, and an incremental shortest path algorithm is applied to identify the hierarchical relationship between the title, paragraphs, and main points, and output structured data containing semantic structure, key content, and emotional tendency, including:
[0015] The content features are processed for document analysis, and elements including text, images, and tables are extracted, and a set of content elements after preliminary analysis is output.
[0016] Based on the set of content elements, an improved interior point method is used to construct a content relationship graph, and an incremental shortest path algorithm is used to identify the hierarchical relationship between the title, paragraphs, and main points, and output a structured document with hierarchical markers.
[0017] Based on the structured document with hierarchical markers, natural language processing technology is used to identify key words, key content, and core concepts, and semantic vectorization is used to represent the relevance between content, and the structured data containing semantic structure, key content, and emotional tendency is output.
[0018] Preferably, based on the structured data, a full-dynamic parallel single-link clustering algorithm is applied to semantically similar group the content of the structured data, converted into narrative text suitable for spoken expression by a pre-set semantic understanding model and added with markers including pauses, intonation changes, emphasis, output an enhanced script with complete expressiveness markers, including:
[0019] Based on the structured data, a full-dynamic parallel single-link clustering algorithm is applied to maintain a dynamically updated network flow model to semantically similar group the content, and the content is divided into multiple semantically associated content blocks according to the theme and logical relationship, output the semantically clustered content blocks;
[0020] Based on the semantically clustered content blocks, the structured data is converted into narrative text suitable for spoken expression by a pre-set semantic understanding model, and transitional words and conjunctions are added to enhance fluency, output a preliminary voice script;
[0021] Based on the preliminary voice script, a Speech Synthesis Markup Language (SSML) script containing markers for pauses, intonation changes, and emphasis is generated, and the enhanced script with complete expressiveness markers is output.
[0022] Preferably, based on the enhanced script, the CosyVoice speech synthesis technology is used to The low-rank approximation method decomposes a large-scale feature matrix into multiple blocks and performs block sampling and parallel computation optimization, adjusts the prosody and expressiveness, and outputs expressive voice data streams, including:
[0023] Based on the enhanced script, the lecturer style parameters are combined, and the best speech synthesis configuration is determined by a parameter optimization algorithm, output a set of speech synthesis parameters;
[0024] Based on the set of speech synthesis parameters and the enhanced script, the CosyVoice speech synthesis technology is used to block sample a large-scale feature matrix and perform parallel computation optimization by The low-rank approximation method accelerates speech feature extraction and synthesis, and outputs high-quality speech waveform data;
[0025] Based on the high-quality speech waveform data, a prosody adjustment algorithm is applied to dynamically adjust the pauses, stress, and intonation changes of the speech according to the Speech Synthesis Markup Language (SSML) markers, and output the expressive voice data stream.
[0026] Preferably, based on the enhanced script and the speech data stream, key verbs, emphasis words, and sentiment lexicons are identified through semantic analysis, a mapping relationship between semantic content and appropriate body movements is established, the precise time points of movement execution are calculated, and facial expression instructions consistent with the sentiment of the content are generated, and a complete set of movement-expression instructions is output, including:
[0027] Based on the enhanced script, key verbs, emphasis words, and sentiment lexicons are identified through semantic analysis techniques, a semantic-action correlation mapping table is established, and the semantic-action correlation mapping table is output.
[0028] Based on the speech data stream and the semantic-action correlation mapping table, the time points of movement execution are calculated by analyzing the rhythm, pause, and accent characteristics of the speech, and a time-synchronized movement sequence is output.
[0029] Based on the sentiment markers in the enhanced script and the intonation changes in the speech data stream, facial expression instructions consistent with the sentiment of the content are generated through a sentiment-expression mapping model, and a time-synchronized expression sequence is output.
[0030] Based on the time-synchronized movement sequence and the time-synchronized expression sequence, natural transitions and coordinated consistency between movements and expressions are ensured through a movement-expression coordination algorithm, and a complete set of movement-expression instructions is output.
[0031] Preferably, the natural transitions and coordinated consistency between movements and expressions through the movement-expression coordination algorithm include:
[0032] Time alignment is performed on the time-synchronized movement sequence and the time-synchronized expression sequence, the proximity of strong expressions in time is detected, and a timing conflict identifier is output.
[0033] Based on the timing conflict identifier, a time staggering strategy is applied, the time of movements or expressions is fine-tuned to ensure that important expressions have sufficient exclusive time windows, and a timing-optimized movement-expression sequence is output.
[0034] The timing-optimized movement-expression sequence is checked for semantic consistency, inconsistencies in the information conveyed by movements and expressions are identified, and a semantic conflict list is output.
[0035] Based on the semantic conflict list, the expression method with less conflict is adjusted according to the sentiment tendency of the content to maintain the consistency of the overall expression, and a semantically coordinated movement-expression sequence is output.
[0036] Based on the semantically coordinated movement-expression sequence, the strength of movements and expressions is dynamically adjusted according to the importance level of the content through a global strength balancing algorithm, and a strength-balanced complete movement-expression instruction set is output.
[0037] Preferably, the driving digital human model with museTalk technology comprises:
[0038] Based on the pre-defined digital human model library, select the appropriate lecturer image according to the nature of the teaching content and the target audience, initialize the skeleton structure and facial control points, and output the ready digital human model;
[0039] Based on the action expression instruction set, the action expression instruction set is parsed into skeleton animation control flow, facial animation control flow and mouth shape synchronization control flow, and three parallel control flows are output;
[0040] Based on the skeleton animation control flow, use key frame interpolation and programmatic animation method to control the body movement of digital human, add subtle body sway, weight transfer and momentum effect, output natural skeleton animation sequence;
[0041] Based on the facial animation control flow, adopt parameterized facial expression control to map expression parameters to digital human facial controller, add natural blinking, micro-expression and subtle head movement, output rich facial animation sequence;
[0042] Based on the mouth shape synchronization control flow and the speech data flow, adopt phoneme to viseme mapping technology to convert phoneme sequence in speech into corresponding mouth shape, apply mouth shape smoothing algorithm to avoid mechanical mouth shape change, output complete animation sequence synchronized with speech.
[0043] Preferably, the integration of digital human lecturer and teaching content into a unified visual scene through optimized parallel rendering algorithm comprises:
[0044] Based on the PPT content in the structured data, extract visual elements including slides, images, charts, text points, analyze content readability, lecturer visibility, content-lecturer relationship and picture balance, output the most suitable picture layout strategy;
[0045] Based on the picture layout strategy and the complete animation sequence, automatically adjust the position, angle and focal length of virtual camera according to the importance of content and lecturer action, add dynamic visual emphasis effects including highlight, zoom, arrow indication to important content, output visual scene with shot language;
[0046] Based on the visual scene with shot language, process the interaction effect between lecturer and content, apply transition effects including fade in and out, wipe in and out, zoom transition at content switching points, output scene sequence with enhanced interaction;
[0047] Based on the interactive enhanced scene sequence, apply color correction, light balance and layer mixing technology to ensure that the digital human lecturer is visually unified with the background and content, output the complete scene sequence with visual integration;
[0048] Based on the complete visual integrated scene sequence, through high dynamic range rendering, temporal anti-aliasing, motion blur and color grading processing, a professional complete video scene sequence is output.
[0049] Preferably, the output final lecturer teaching video includes:
[0050] Based on the complete video scene sequence and the voice data stream, the entire video rendering task is decomposed into time segmentation sub-tasks and space segmentation sub-tasks that can be processed in parallel, and a set of decomposed rendering tasks is output.
[0051] Based on the set of decomposed rendering tasks, the available CPU cores and GPU units are intelligently allocated for computing resources, the tasks on the critical path are preferentially processed to ensure balanced rendering progress, and a resource-optimized rendering schedule is output.
[0052] Based on the resource-optimized rendering schedule, the GPU parallel computing capability is used to accelerate lighting calculation, material rendering and post-processing, and a multi-resolution rendering strategy is adopted to first render at a lower resolution to obtain a preview, and a rendering result with progressive quality is output.
[0053] Based on the rendering result with progressive quality, the quality is gradually improved to the target resolution, audio-visual synchronization verification, picture quality evaluation and compatibility testing are performed, and a high-definition video file with quality verification is output.
[0054] Based on the high-definition video file with quality verification, various format variants including MP4, WebM and low-bandwidth optimization are generated according to different use scenarios, complete metadata containing chapter markers, keyword indexes and content summaries are generated, and the final lecturer teaching video is output.
[0055] The beneficial effects of the present application are:
[0056] 1. The incremental shortest path algorithm is realized by the improved interior point method, and the traditional shortest path calculation with O(n 3 ) complexity is optimized to nearly linear time O(n log n), realizing real-time incremental analysis of large PPT documents, effectively improving the recognition accuracy and processing efficiency of the content hierarchy;
[0057] 2. The full-dynamic parallel single-link clustering algorithm is adopted, which supports dynamic addition, deletion and update of data points, and through distributed computing and parallel processing technology, the traditional clustering complexity of O(n 2 ) is reduced to O(n log n), realizing efficient semantic grouping of a large number of content segments, and improving the coherence and context correlation of content conversion;
[0058] 3. Application The low-rank approximation method processes a high-dimensional speech feature matrix, and through block sampling and parallel computing strategy, the traditional core matrix calculation complexity is reduced from O(n 3 ) to O(n 2 ), while maintaining accuracy, realizing fast low-rank approximation and kernel ridge regression in the speech synthesis process, and significantly improving the speed and quality of speech synthesis;
[0059] 4. The deep mapping relationship between semantic content and action expression is innovatively established, the semantic importance and emotional intensity of the content are analyzed through a multi-layer attention mechanism, and the action and expression highly matched with the content context are automatically generated and explained, solving the problems of insufficient expressiveness and stiff action of traditional digital people;
[0060] 5. A multi-level parallel rendering pipeline architecture is designed, the action driving, scene synthesis and video rendering tasks are divided into parallel processing subtasks, GPU acceleration and task scheduling optimization are used, near real-time video generation capability is realized, and large-scale content can be quickly processed. BRIEF DESCRIPTION OF DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0062] Figure 1 The overall flowchart of the lecturer video automatic generation method provided by the embodiment of the present application;
[0063] Figure 2 The flowchart of the multi-modal speech synthesis engine provided by the embodiment of the present application;
[0064] Figure 3 The flowchart of the context-aware action expression generation system provided by the embodiment of the present application. DETAILED DESCRIPTION
[0065] The present application will be further described in detail below in combination with the drawings and specific embodiments.
[0066] The present application provides a lecturer video automatic generation method based on AI speech synthesis and animation driving, as shown in the following steps: Figure 1
[0067] Step S1, based on the content characteristics of the PPT file or text script uploaded by the user, a content relationship graph is constructed by an improved interior point method, an incremental shortest path algorithm is applied to identify the hierarchical relationship between the title, paragraphs and main points, and structured data containing semantic structure, key content and emotional tendency is outputted;
[0068] In this step, the user uploads a PPT file or a text script, which will serve as the basis for generating the lecturer video. The uploaded content is first processed by a document parsing engine, which extracts elements such as text, images, and tables, forming a set of content elements. These elements are then sent to a content structuring analysis module, which uses an improved interior point method to construct a content relationship graph. This method optimizes the traditional O(n 3 ) complexity of the shortest path calculation to near-linear time O(n log n) by maintaining a dynamically updated network flow model. Based on the constructed content relationship graph, an incremental shortest path algorithm is applied to identify the hierarchical relationships between titles, paragraphs, and main points in the content, determining the logical structure of the content. Subsequently, natural language processing techniques are used to perform deep semantic analysis of the content, identifying keywords, key points, and core concepts, and sentiment analysis algorithms are used to identify the emotional tendencies of the content, such as enthusiasm, seriousness, and curiosity. Finally, structured data containing semantic structure, key points, and emotional tendencies is output, laying the foundation for subsequent processing.
[0069] In step S2, based on the structured data, a full-dynamic parallel single-link clustering algorithm is applied to the content of the structured data for semantic similarity grouping. Through a pre-set semantic understanding model, the narrative text of spoken language expression is converted and marked with pauses, tone changes, and emphasis, outputting an enhanced script with complete expressiveness markers.
[0070] In this step, the structured data output from step S1 is received, and a full-dynamic parallel single-link clustering algorithm is applied to the content for semantic similarity grouping. This algorithm supports dynamic addition, deletion, and updating of data points, and through distributed computing and parallel processing techniques, the traditional O(n 2 ) clustering complexity is reduced to O(n log n), achieving efficient semantic grouping of a large number of content segments. A dynamically updated network flow model is maintained, and the content is divided into multiple semantically related content blocks based on themes and logical relationships. Subsequently, these structured data are converted into narrative text suitable for spoken language expression through a pre-set semantic understanding model, and transitional words and conjunctions are added to enhance fluency, generating a preliminary voice script. The internal structure and semantic type of each content block are analyzed, and corresponding language templates are applied to different types, such as adding sequence guide words "first", "second", "last" for list-type content. Finally, rich tone and expressiveness markers are added to the script to generate a complete marked script that meets the specifications of the Speech Synthesis Markup Language (SSML), including pause markers, tone change markers, and emphasis markers, so that the synthesized speech can express appropriate emotions and emphasize key points, outputting an enhanced script with complete expressiveness markers.
[0071] Step S3, based on the enhanced script, utilize CosyVoice speech synthesis technology to generate a high-quality voice stream through Low-rank approximation method decomposes large-scale feature matrix into multiple blocks and performs block sampling and parallel computing optimization, adjusts prosody and expressiveness, and outputs expressive speech data stream;
[0072] In this step, based on the enhanced script output in step S2, combined with the lecturer style parameters, the best speech synthesis configuration is determined by parameter optimization algorithm. Analyze the enhanced script with expressiveness markers, evaluate the overall characteristics of the content such as professional degree, technical complexity and target audience, and use multi-dimensional parameter space to describe speech characteristics, including basic timbre parameters, speech rate parameters, pitch parameters, tone parameters and clarity parameters. After determining the parameters, the core speech synthesis processing is carried out by using CosyVoice speech synthesis technology. To improve the computing efficiency, the innovative Low-rank approximation method decomposes large-scale feature matrix into multiple blocks and performs block sampling and parallel computing optimization. This method decomposes large feature matrix into multiple smaller blocks, applies Method for low-rank approximation to each block, and then reconstructs the complete low-rank approximation result through carefully designed merging strategy, reduces the traditional core matrix calculation complexity from O(n 3 ) to O(n 2 ), while maintaining accuracy. After generating the basic speech waveform, adjust the prosody and expressiveness, optimize the rhythm, stress and intonation pattern of the speech, including stress pattern adjustment, intonation contour optimization, pause rhythm adjustment and emotion expression enhancement. Finally, perform professional-level audio post-processing on the speech, including noise reduction processing, equalization processing, dynamic range compression and psychoacoustic enhancement, and output an expressive speech data stream.
[0073] Step S4, based on the enhanced script and the speech data stream, identify key verbs, emphasized words and emotional words through semantic analysis, establish the mapping relationship between semantic content and appropriate body movements, calculate the accurate time points of movement execution and generate facial expression instructions consistent with the content emotion, and output a complete set of movement expression instructions;
[0074] In this step, based on the enhanced script output in step S2 and the voice data stream output in step S3, key verbs, emphasis words and emotional vocabulary are identified through semantic analysis technology, and a mapping relationship between semantic content and appropriate body movements is established. Deep semantic analysis is performed on the enhanced script, and multi-level semantic understanding is adopted, including language behavior recognition (such as explanation, emphasis, enumeration, etc.), spatial concept mapping (mapping spatial concepts such as "rise" and "expand" to corresponding directional gestures), and emotional intensity analysis (evaluating the emotional intensity and importance of the content to determine the amplitude and energy level of the movements). Semantic features are extracted through a pre-trained large-scale language model, and are converted into vectors in the action feature space through a mapping network, outputting a semantic-action correlation mapping table. Subsequently, detailed acoustic analysis is performed on the voice data stream to extract speech rhythm markers, prosodic structure and emphasis markers, and an action-voice synchronization algorithm is applied to calculate the precise execution time points of each movement, following the preparation principle, emphasis principle, ending smoothing principle and coherent transition principle, to output a time-synchronized movement sequence. At the same time, based on the emotional markers in the enhanced script and the intonation changes in the voice data stream, a content-emotion-consistent facial expression instruction is generated through an emotion-expression mapping model, a multi-level emotion-expression mapping model is adopted, including a basic emotion layer, a teaching-specific expression layer and a micro-expression layer, and a time-synchronized expression sequence is output. Finally, a movement-expression coordination algorithm is used to ensure the natural transition and consistent coordination between movements and expressions, solving problems such as timing conflicts, semantic conflicts, physical conflicts and intensity imbalance, and outputting a complete set of movement-expression instructions.
[0075] Step S5, based on the voice data stream and the movement-expression instruction set, using museTalk technology to drive the digital human model, integrating the digital human lecturer and the teaching content into a unified visual scene through an optimized parallel rendering algorithm, and outputting the final lecturer teaching video.
[0076] In this step, based on the voice data stream output in step S3 and the action expression instruction set output in step S4, first, select a suitable lecturer image from the pre-defined digital human model library according to the nature of the teaching content and the target audience. Initialize the skeleton structure and facial control points of the digital human model, including skeleton loading, facial control initialization, physical parameter setting and rendering parameter configuration. Then, parse the action expression instruction set into three parallel control streams: skeleton animation control stream, facial animation control stream and lip synchronization control stream, and drive the digital human model using the museTalk technology. Use keyframe interpolation and procedural animation methods to control the body movements of the digital human, add subtle body sway, weight transfer and momentum effect to make the animation more lively and natural; use parameterized facial expression control to map expression parameters to digital human facial controllers, add natural blinking, micro-expression and subtle head movements; apply phoneme-to-viseme mapping technology to convert phoneme sequences in speech into corresponding lip shapes, ensuring accurate synchronization of lip movements with speech. Next, integrate the digital human lecturer and teaching content into a unified visual scene, extract visual elements from PPT content, design the most suitable picture layout, realize intelligent shot language, handle the interaction effect between lecturer and content, apply color correction, lighting balance and level mixing technology to ensure visual unity. Finally, generate the video through the optimized parallel rendering algorithm, divide the entire video rendering task into parallel processing subtasks, intelligently allocate computing resources, use GPU parallel computing capabilities to speed up the rendering process, perform quality check and format conversion, and output the final lecturer teaching video.
[0077] In step S1, based on the content features of the user uploaded PPT file or text script, a content relationship graph is constructed by an improved interior point method, and an incremental shortest path algorithm is applied to identify the hierarchical relationship between titles, paragraphs and main points. The structured data containing semantic structure, key content and emotional tendency is output, including:
[0078] Step S1.1, perform document parsing processing on the content features, extract elements including text, images and tables, and output a set of preliminary parsed content elements;
[0079] In this step, first receive the user uploaded PPT file or text script, and process it through professional document parsing engine. For PPT files, various content elements can be identified and extracted, including slide titles, body text, bulleted lists, images, charts, tables, and embedded multimedia elements, etc. The original format and layout information of these elements will be preserved, while their positional relationships and order in the PPT are recorded. For text scripts, the document structure will be analyzed to identify headings, subheadings, paragraphs, lists, and other different text blocks. During extraction, format information such as font size, bold, italic, and other style features that often imply the importance level of the content will also be preserved. For image elements, not only the images themselves will be extracted, but also the image titles, captions, and text information that may be contained within the images will be analyzed. For tables, their structure will be preserved while extracting table headers and cell contents, understanding the logical organization of the table. After processing, a preliminary parsed content element set containing all these elements is output, which preserves the full information and basic structural features of the original content, laying the foundation for subsequent in-depth analysis.
[0080] Step S1.2, based on the content element set, an improved interior point method is applied to construct a content relationship graph, and a hierarchical relationship between titles, paragraphs, and points is identified through an incremental shortest path algorithm, and a structured document with hierarchical labels is output;
[0081] In this step, based on the content element set output in step S1.1, an improved interior point method is applied to construct a content relationship graph. The traditional interior point method has high computational complexity when dealing with large-scale network flow problems, and the invention improves the method by maintaining a dynamically updated network flow model, optimizing the traditional O(n 3 ) complexity calculation to nearly linear time O(n logn). First, the content elements are regarded as nodes in the graph, and the initial connection is established based on the positional relationship between the elements, format features, and semantic association. For example, a strong connection is established between the title and the paragraph below it, a weak connection is established between adjacent paragraphs, and a semantic connection is established between content blocks with similar themes. After constructing the initial graph, an incremental shortest path algorithm is applied to identify the hierarchical structure of the content. This algorithm can efficiently find the shortest path from the beginning of the document to each content block, and determine the hierarchical relationship of the content according to the path characteristics. For example, by analyzing the path length and node characteristics, it can identify content elements at different levels, such as primary headings, secondary headings, body paragraphs, and point lists. Unlike traditional methods, this algorithm uses incremental processing, so when new content is added, only the affected paths need to be updated, rather than recalculating the entire graph structure, greatly improving the efficiency of processing large documents. After processing, a structured document with hierarchical labels is output, which clearly marks the hierarchical relationship, affiliation, and logical order of each content element, providing a structured foundation for subsequent semantic analysis.
[0082] Step S1.3, based on the structured document with hierarchical markers, utilize natural language processing techniques to identify keywords, key points and core concepts, represent the relevance between contents through semantic vectorization, output the structured data containing semantic structure, key points and sentiment orientation.
[0083] In this step, based on the structured document with hierarchical markers output in step S1.2, advanced natural language processing techniques are applied for deep semantic analysis. Firstly, keyword extraction algorithms are used to identify keywords and terms in the document, which combine statistical methods (such as TF-IDF) and deep learning models to accurately capture the core vocabulary of the document. For the identified keywords, further analysis of their distribution and context in the document is performed to determine their importance weight. Subsequently, text summarization and key extraction techniques are applied to identify core concepts and key points in the document. This process not only considers explicit markers of content (such as "important", "key" and other words), but also analyzes the location of content in the document structure, repetition and relevance to other content. Semantic vectorization techniques are also used to convert text content into high-dimensional semantic vectors, which can capture the deep semantic features of the content. By calculating the similarity between vectors, the semantic association strength between different content blocks can be quantified, and a semantic network of content can be constructed. In addition, sentiment analysis algorithms are applied to identify the sentiment orientation and tone features of the content. These algorithms can detect the expressed emotions (such as enthusiasm, seriousness, curiosity, etc.) and tone (such as affirmation, questioning, emphasis, etc.) in the text, providing important references for subsequent speech synthesis and expression generation. After processing, complete structured data containing semantic structure, key points and sentiment orientation are output, which comprehensively describe the structural and semantic features of the original content, laying a solid foundation for subsequent semantic understanding and conversion.
[0084] In step S2, based on the structured data, a full-dynamic parallel single-link clustering algorithm is applied to group the contents of the structured data according to semantic similarity, and a pre-set semantic understanding model is used to convert the contents into narrative text expressed in spoken language and add markers including pauses, tone changes and emphasis, output an enhanced script with complete expressiveness markers, including:
[0085] Step S2.1, based on the structured data, a full-dynamic parallel single-link clustering algorithm is applied to maintain a dynamically updated network flow model to group the contents according to semantic similarity, and the contents are divided into multiple semantically associated content blocks according to theme and logical relationship, output the semantically clustered content blocks;
[0086] Step S2.2, based on the semantic clustering of content blocks, converts the structured data into narrative text suitable for spoken expression through a pre-set semantic understanding model, adds transitional words and conjunctions to enhance fluency, and outputs a preliminary voice script;
[0087] Step S2.3, based on the preliminary voice script, generates a speech synthesis markup language (SSML) script containing pauses, tone changes, and emphasis markers, and outputs the enhanced script with complete expressive markers.
[0088] This step uses a full-dynamic parallel single-link clustering algorithm to process structured data, and its core technical principles are as follows:
[0089] First, receive the structured data output from the PPT content structured analysis module, which contains text content, hierarchical relationship, semantic markers, and sentiment orientation, etc. These contents are regarded as data points in high-dimensional space, and each data point represents a content segment (such as a paragraph or a point).
[0090] The key to the full-dynamic parallel single-link clustering algorithm is its "full-dynamic" feature, which allows dynamic addition, deletion or update of data points without re-computing the entire clustering structure. Specifically, the algorithm maintains a near-neighbor graph structure, and each node (content segment) is connected to the most similar nodes in its semantic space. When new content is added, only the similarity between it and existing nodes needs to be calculated, and the local graph structure needs to be updated, rather than re-computing the entire clustering.
[0091] In terms of parallel processing, the algorithm uses a divide-and-conquer strategy to divide the data space into multiple subspaces, perform local clustering in each subspace in parallel, and then integrate the results through a boundary merging algorithm. This method reduces the traditional O(n 2 ) computational complexity to O(n log n), greatly improving the efficiency of processing large-scale content.
[0092] The calculation of semantic similarity is based on the text embedding vectors extracted by pre-trained language models (such as BERT, RoBERTA, etc.), which determine the semantic association strength between content segments through cosine similarity or other distance measures. The original structural information of the PPT, such as content within the same page or content with the same title level, is also considered, giving additional clustering weight.
[0093] Finally, this step outputs semantic clustering of content blocks, and the content within each block is highly related in semantics, forming a coherent explanation unit. These blocks retain the logical order of the original content and mark the semantic association between blocks, providing an important foundation for subsequent script conversion.
[0094] In step S2.2:
[0095] Based on the semantic clustering of the content blocks from the previous step, this step performs script conversion and enhancement, transforming structured content into natural and fluent spoken language expressions.
[0096] The conversion process first analyzes the internal structure and semantic types of each content block, such as explanatory content, list-type content, comparative explanations, etc., and applies corresponding language templates for different types. For example, for list-type content, sequence guide words such as "first," "second," and "last" are added; for cause-and-effect content, conjunctions such as "because" and "therefore" are added.
[0097] A double-layer conversion architecture is adopted: the first layer is structural conversion, which converts the visual structure of PPT (such as bullets, tables, charts, etc.) into spoken language description structure; the second layer is language style conversion, which converts formal written language into a more suitable spoken language style for oral expression, including simplifying complex sentence patterns, shortening sentence length, and increasing rhythm changes.
[0098] To enhance the fluency of the content, transition sentences are automatically generated between adjacent content blocks, which are dynamically generated based on the semantic relationship between the two blocks. For example, if the next block is a deep explanation of the previous concept, transition sentences such as "Let's further understand this concept" are generated; if it is a shift to a new topic, transition sentences such as "Next, we will discuss another important aspect" are generated.
[0099] Semantic enhancement also includes adjusting the level of detail in expression according to the importance of the content. For content identified as core concepts, more details are retained and emphasis is added; for auxiliary explanation content, appropriate simplification and generalization are performed.
[0100] The final output of the preliminary speech script retains all the information of the original content, while having better oral fluency and coherence, laying the foundation for the next step of adding expressive markers.
[0101] In step S2.3:
[0102] After obtaining the preliminary speech script, this step adds rich intonation and expressive markers to the script, enabling the synthesized speech to express appropriate emotions and emphasize key points.
[0103] First, the preliminary speech script is subjected to in-depth semantic analysis to identify key words, concept definitions, important conclusions, and other content elements that need to be emphasized. At the same time, combined with the emotional inclination analysis results in step 1.4, the basic emotional tone (such as neutral, enthusiasm, seriousness, or curiosity) that each sentence or paragraph should express is determined.
[0104] Based on these analyses, a complete markup script conforming to the Speech Synthesis Markup Language (SSML) specification is generated. SSML is an XML-based markup language specifically designed to control the pronunciation, prosody, speed, and other characteristics of speech synthesis. The generated SSML markup mainly includes the following categories:
[0105] Pause markup ( <break>) : Insert pauses of different lengths at the end of sentences, before key concepts, at turning points, etc. to simulate the natural rhythm of speech. For example, insert a medium-length pause before introducing a new concept to heighten the listener's attention.
[0106] Pronunciation variation markers <prosody>) Adjusting the pitch, rate, and volume of the voice. For example, for important concept explanations, the rate can be slightly decreased and the volume increased; for example illustrations, a more relaxed tone and moderate rate can be used.
[0107] Emphasis markers <emphasis>) : Applying different levels of emphasis on keywords and core concepts. SSML supports three levels of emphasis: strong, medium, and weak, and dynamically assigns the level of emphasis according to the importance of the content.
[0108] Voice style tags <voice-style>Or vendor-specific markers): Some advanced TTS engines support overall voice style control, such as CosyVoice may support "teaching mode", "speech mode" and other style settings. Appropriate style markers will be selected according to the overall type of content.
[0109] The algorithm also considers the rhythm changes of the voice to avoid mechanical and monotonous expression. For example, appropriate micro-pauses are inserted in long sentences, and slight speed changes are added between consecutive technical terms to make the overall voice more natural and fluent.
[0110] The final output of the enhanced script with complete expressiveness markers not only contains the original semantic content, but also contains rich expressiveness control instructions, which can guide the TTS engine to generate more expressive and teaching-effective voice expression.
[0111] As shown in Figure 2 In step S3, based on the enhanced script, the CosyVoice TTS technology is used to Low-rank approximation method decomposes large-scale feature matrix into multiple blocks and performs block sampling and parallel computing optimization, adjusts prosody and expressiveness, and outputs expressive voice data stream, including:
[0112] Step S3.1, based on the enhanced script, combining the lecturer style parameters, determining the best TTS configuration through parameter optimization algorithm, outputting the TTS parameter set;
[0113] In this step, the enhanced script with expressiveness markers is first comprehensively analyzed to evaluate the overall characteristics of the content, such as the level of professionalism, technical complexity, and target audience features. A multi-dimensional parameter space is used to describe the voice characteristics, which collectively determine the overall style and expressiveness of the synthesized voice. The lecturer style parameters mainly include the following key dimensions: basic timbre parameters, speech rate parameters, pitch parameters, tone parameters, and clarity parameters. The basic timbre parameters determine the basic characteristics of the voice, including gender characteristics (male, female, or neutral), age characteristics (young, mature, or old), timbre brightness (bright, warm, or deep), and voice quality characteristics (clear, round, or thick). These parameters directly affect the listener's first impression of the lecturer's image, and the most suitable basic timbre will be selected according to the nature of the teaching content. For example, for technical content, a more neutral and professional timbre may be chosen, while for children's education content, a warmer and more friendly timbre may be chosen. The speech rate parameter controls the overall speech rate baseline, usually measured in words per minute or syllables per minute, and dynamically adjusts the speech rate according to the content complexity. For complex technical content, the basic speech rate will be automatically reduced to ensure that the learner has enough time to understand; for simple introductory content, a slightly faster speech rate may be used to maintain the interest of the audience. The pitch parameter controls the basic pitch and variation range of the voice, including the basic pitch (low, medium, high), the pitch variation range (narrow, medium, wide), and the pitch variation mode (smooth, fluctuating, or emphasized). These parameters affect the expressiveness and emotional communication ability of the voice, and the pitch characteristics will be adjusted according to the emotional needs of the content. The tone parameter controls the overall style of the voice, such as the degree of formality (formal, neutral, or casual), the degree of vitality (calm, moderate, or lively), and the degree of affinity (professional, friendly, or intimate). These parameters determine the overall feeling of the lecturer, and appropriate tone characteristics will be selected according to the target audience and teaching scenario. The clarity parameter controls the clarity and emphasis of pronunciation, including consonant intensity, vowel fullness, and syllable boundary clarity. For teaching content, a higher clarity is usually preferred, especially for content containing professional terms. In order to determine the optimal parameter combination, a parameter optimization algorithm is used, which combines rule guidance and machine learning methods. In the rule guidance part, according to the content type and performance requirements, the pre-set parameter template is used as the initial value; in the machine learning part, based on the best parameter settings of similar content in historical data, the optimal solution is searched in the parameter space through optimization methods such as gradient descent or genetic algorithm. The lecturer's consistency requirement is also considered, if it is part of a series of courses, the voice characteristics will be kept consistent with the previous courses as much as possible, and only minor adjustments will be made when necessary. Users can also provide reference audio samples, which will be analyzed to extract parameter features as optimization targets. The final output of the voice synthesis parameter set is a configuration file containing multiple dimension parameter values, which will guide the specific settings in the subsequent voice synthesis process, ensuring that the generated voice not only meets the content requirements, but also has consistent lecturer style features.
[0114] Step S3.2, based on the speech synthesis parameter set and the enhanced script, use CosyVoice speech synthesis technology to perform block sampling on the large-scale feature matrix and output high-quality speech waveform data through The low-rank approximation method performs parallel computation optimization, accelerates speech feature extraction and synthesis, and outputs high-quality speech waveform data;
[0115] In this step, the speech synthesis parameter set determined in step S3.1 and the enhanced script with SSML tags are input into the CosyVoice speech synthesis engine to start the core speech synthesis process. Modern high-quality speech synthesis is usually based on deep learning models such as Tacotron, WaveNet or newer Transformer architecture, which require processing large high-dimensional feature matrices with high computational complexity. To improve computational efficiency, this step applies an innovative low-rank approximation method. Traditionally, kernel matrix computation in speech synthesis (especially in acoustic feature extraction and acoustic model inference) requires O(n 3 ) time complexity, which is inefficient for long content processing. The core idea of the method is to decompose the large feature matrix into multiple smaller blocks, apply the method to each block for low-rank approximation, and then reconstruct the complete low-rank approximation result through a carefully designed merging strategy. Specifically, the algorithm first randomly selects a subset of columns of the feature matrix as landmarks, which are usually selected to represent the main characteristics of the matrix. Then, the similarity between the complete matrix and these landmarks is calculated to form a much smaller core matrix. By performing eigenvalue decomposition on this core matrix, a low-rank approximation of the original large matrix can be obtained. This method reduces the complexity from O(n 3 ) to O(n 2 ), and through the block processing strategy, parallel computation is realized, further improving efficiency. In practical applications, the size and number of blocks will be dynamically adjusted according to available computing resources to achieve the best balance between accuracy and speed. The specific process of speech synthesis includes multiple key stages: first, text normalization is performed to convert numbers, abbreviations, special symbols, etc. to standard form; then linguistic analysis is performed to determine the phoneme sequence, stress position and intonation boundary; then through the acoustic model (application of Optimization) to generate acoustic features such as Mel spectrograms; finally, the acoustic features are converted into actual waveform data through a vocoder. Throughout the process, various expressiveness instructions in the SSML tags are considered in real time, and the generation parameters are dynamically adjusted to ensure that the output speech waveform is not only natural and smooth, but also accurately expresses the intended intonation changes and emotional characteristics. In addition, adaptive enhancement techniques are applied to dynamically adjust the processing intensity during synthesis according to the semantic importance and emotional intensity of the content, ensuring that key content receives more detailed processing. The final output of high-quality speech waveform data retains all the information of the text content and has basic intonation changes, laying the foundation for the next step of prosody adjustment.
[0116] CosyVoice speech synthesis technology uses an end-to-end neural network model based on the Transformer architecture, which includes three core components: encoder-attention-decoder. The training data for this model comes from about 1000 hours of high-quality recorded data, covering professional voice actors of different ages, genders, and speech styles, ensuring the diversity and naturalness of the generated speech. The model architecture includes 12 layers of Transformer encoder and 12 layers of Transformer decoder, each with 8 attention heads, a hidden layer dimension of 512, a feedforward network dimension of 2048, and a GeLU activation function. Key hyperparameter settings include a learning rate of 0.0001, using the Adam optimizer, a batch size of 32, and about 500,000 training iterations. The model is trained in two stages: the first stage uses general speech data for pre-training, and the second stage optimizes the specific speech style for the teaching scenario.
[0117] The mathematical principle of the low-rank approximation method is based on matrix decomposition theory, and the specific implementation is as follows: for a large feature matrix K ∈ R^(n×n), the traditional calculation of its eigenvalue decomposition has a complexity of O(n 3 ). The method first samples c columns from K to form C ∈ R^(n×c), and then calculates the submatrix W ∈ R^(c×c) corresponding to these columns. The low-rank approximation is performed through the formula K ≈ CW -1 CT, where W -1 is calculated by SVD decomposition. To further improve efficiency, the algorithm divides C into multiple blocks {C1, C2,..., C_m}, and calculates the W -1 CT of each block in parallel, and finally combines the results. Experimental verification shows that when the sampling rate c / n = 0.1, this method can reduce the computational complexity to O(n 2 ), while maintaining an accuracy of more than 95%. When processing feature matrices with more than 100,000 dimensions, the calculation speed can be improved by more than 10 times.
[0118] Step S3.3, based on the high-quality speech waveform data, applying a prosody adjustment algorithm to dynamically adjust the pauses, stress and intonation variations of the speech according to the Speech Synthesis Markup Language (SSML) tags, outputting the expressive speech data stream.
[0119] In this step, further prosodic optimization and expressiveness enhancement are performed on the base speech waveform generated in step S3.2, making it more natural and engaging. Prosody is the rhythm, stress, and intonation pattern in speech, which is crucial for conveying meaning and emotion, and is a key factor in distinguishing between mechanical and natural human voices. First, a detailed acoustic analysis is performed on the generated speech waveform, extracting features such as fundamental frequency contour (F0), energy curve, and duration parameters. These features collectively describe the prosodic characteristics of the speech and serve as the basis for subsequent adjustments. Then, these actual parameters are compared with ideal models, which come from two sources: theoretical models based on linguistic rules, which incorporate various prosodic rules established in linguistic research, and statistical models learned from professional instructor speech, which capture the prosodic characteristics of instructor speech in real teaching scenarios. Prosodic adjustments focus on four key aspects: stress pattern adjustment, intonation contour optimization, pause rhythm adjustment, and emotion expression enhancement. In stress pattern adjustment, the appropriate stress is ensured for key words in a sentence, which is achieved by fine-tuning the duration, pitch, and intensity of syllables. For example, for technical terms, the duration and intensity of the first stressed syllable are increased; for emphasized content, the overall volume and clarity of relevant words are improved. In intonation contour optimization, the pitch curve of the entire sentence is adjusted to conform to natural intonation patterns. For example, questions should have a rising intonation at the end, emphasized sentences should have a clear pitch peak, and statements should have a smooth falling intonation at the end. A piecewise spline interpolation algorithm is used to generate a smooth pitch curve between key points, ensuring smooth and natural intonation changes. In pause rhythm adjustment, the precise duration of intra-sentence and inter-sentence pauses is fine-tuned to make the speech rhythm more natural. Research shows that pause lengths in natural speech follow specific statistical distributions and dynamically adjust pause lengths based on context. For example, slightly longer pauses are inserted before complex concepts, short pauses are inserted between listed items, and clear pauses are inserted at paragraph transitions. In emotion expression enhancement, the corresponding acoustic features are enhanced based on the emotional markers of the content. For example, for content expressing enthusiasm, the pitch variation range and speech rate variation are increased; for serious content, the pitch variation is reduced and the timbre characteristics are adjusted; for curious or questioning content, the rising trend of intonation is increased. These adjustments are achieved through digital signal processing techniques, including pitch shifting, duration stretching, and spectral envelope modification. An end-to-end prosody model based on deep learning is used to achieve precise prosodic control while maintaining the naturalness of the speech. In addition, subtle human voice features such as natural breathing sounds, slight oral cavity noises, and voice texture changes are added to further enhance the naturalness and affinity of the speech. The final output of the expressive speech data stream retains the quality of the base speech generated in the previous steps, while having more natural and more teaching-effective prosodic characteristics, providing high-quality input for subsequent audio post-processing and video synthesis.
[0120] In step S3.1:
[0121] In the first step of the voice synthesis process, the optimal voice synthesis parameter configuration needs to be determined based on the characteristics of the enhanced script and the target lecturer style.
[0122] Firstly, the enhanced script with expressive markers is analyzed to assess the overall characteristics of the content, such as professional level, technical complexity, and target audience. For example, for a beginner-level tutorial, a more moderate and moderate-paced voice style may be preferred; while for advanced technical content, a more professional and clearer voice style may be chosen.
[0123] A multi-dimensional parameter space is used to describe voice characteristics, mainly including:
[0124] Basic timbre parameters: determine the basic characteristics of the sound, such as gender characteristics, age characteristics, and timbre brightness
[0125] Speech rate parameters: control the overall speech rate baseline, usually measured in words per minute or phonemes
[0126] Pitch parameters: control the basic pitch and variation range of the voice
[0127] Tone parameters: control the overall style of the voice, such as formal, casual, lively, etc.
[0128] Clarity parameters: control the clarity of pronunciation and emphasis on words
[0129] In order to determine the optimal parameter combination, a parameter optimization algorithm is used, which combines rule guidance and machine learning methods. In the rule guidance part, according to the content type and performance requirements, the preset parameter template is used as the initial value; in the machine learning part, based on the best parameter settings of similar content in historical data, through gradient descent or genetic algorithm optimization methods, the optimal solution is searched in the parameter space.
[0130] The lecturer's consistency requirements are also considered, if it is part of a series of courses, the voice characteristics will be kept consistent with the previous courses as much as possible, and only necessary adjustments will be made. Users can also provide reference audio samples, which will be analyzed to extract parameter characteristics as optimization targets.
[0131] The final output of the voice synthesis parameter set is a configuration file containing multiple dimensional parameter values, which will guide the specific settings in the subsequent voice synthesis process, ensuring that the generated voice meets the content requirements and has consistent lecturer style characteristics.
[0132] In step S3.2:
[0133] This step is the core of the entire voice synthesis process, which will use advanced voice synthesis technology and optimization algorithms to convert text into high-quality voice waveforms.
[0134] First, the enhanced script with SSML tags and the set of voice synthesis parameters determined in step 3.1 are input into a voice synthesis engine (such as CosyVoice). Modern voice synthesis is usually based on deep learning models such as Tacotron, WaveNet, or newer Transformer architectures. These models require processing large high-dimensional feature matrices with high computational complexity.
[0135] To improve computational efficiency, this step applies an innovative low-rank approximation method. Traditionally, kernel matrix computation in voice synthesis (especially in the acoustic feature extraction and acoustic model inference stages) requires O(n 3 ) time complexity, which is inefficient for long content processing. The core idea of the method is:
[0136] decomposing a large feature matrix into multiple smaller blocks
[0137] applying the method for low-rank approximation to each block
[0138] reconstructing the complete low-rank approximation result through a carefully designed merging strategy
[0139] In specific implementation, the algorithm first randomly selects a subset of columns of the feature matrix as landmarks, then calculates the similarity between the complete matrix and these landmarks, and finally reconstructs the low-rank approximation result through matrix operations. This method reduces the complexity from O(n 3 ) to O(n 2 ), and through the block processing strategy, parallel computing is realized, further improving efficiency.
[0140] In the process of voice synthesis, first, text normalization is performed to convert numbers, abbreviations, special symbols, etc. into standard form; then linguistic analysis is performed to determine the phoneme sequence, stress position, and intonation boundary; then acoustic features (such as mel-spectrogram) are generated through an acoustic model (with optimization); finally, the acoustic features are converted into actual waveform data through a vocoder.
[0141] Throughout the process, various expressiveness instructions in the SSML tags are considered in real time, and the generated parameters are dynamically adjusted to ensure that the output voice waveform is not only natural and smooth, but also accurately expresses the expected intonation changes and emotional characteristics.
[0142] The final output of high-quality voice waveform data retains all the information of the text content and has basic intonation changes, laying the foundation for the next step of prosody adjustment.
[0143] In step S3.3:
[0144] After obtaining the base speech waveform, this step further optimizes the prosodic characteristics and expressiveness of the speech, making it more natural and engaging.
[0145] Prosody is the rhythm, stress, and intonation pattern in speech, which is crucial for conveying meaning and emotion. Although the previous step has generated speech with basic expressiveness based on SSML tags, subtle prosodic adjustments are essential to achieve truly natural instructional speech.
[0146] First, acoustic analysis is performed on the generated speech waveform to extract features such as fundamental frequency contour (F0), energy curve, and duration parameters. Then, these actual parameters are compared with ideal models from two sources: theoretical models based on linguistic rules and statistical models learned from professional instructor speech.
[0147] Prosodic adjustments focus on the following aspects:
[0148] Stress pattern adjustment: Ensures that key words in a sentence receive appropriate stress, achieved by fine-tuning the duration, pitch, and intensity of syllables. For example, technical terms may have increased duration and intensity of the first stressed syllable.
[0149] Intonation contour optimization: Adjusts the pitch curve of the entire sentence to conform to natural intonation patterns. For example, questions should have a rising intonation at the end, while emphasis sentences should have a clear pitch peak. A piecewise spline interpolation algorithm is used to generate smooth pitch curves between key points.
[0150] Pause rhythm adjustment: Fine-tunes the precise duration of intra-sentence and inter-sentence pauses to make the speech rhythm more natural. Research shows that pause lengths in natural speech follow specific statistical distributions, with dynamic adjustments based on context.
[0151] Emotional expressiveness enhancement: Enhances relevant acoustic features based on the emotional markers of the content. For example, for content expressing enthusiasm, the pitch variation range and speech rate variation are increased; for serious content, the pitch variation is reduced, and timbre characteristics are adjusted.
[0152] These adjustments are implemented through digital signal processing techniques, including pitch shifting, duration stretching, and spectral envelope modification. An end-to-end prosody model based on deep learning is used to achieve precise prosodic control while maintaining the naturalness of the speech.
[0153] The final output of the expressive enhanced speech data retains the quality of the base speech generated in previous steps, while possessing more natural and more effective teaching prosodic characteristics, providing high-quality input for the final audio post-processing.
[0154] Additionally, audio post-processing and optimization are included:
[0155] As the final step of speech synthesis, audio post-processing and optimization aim to enhance the overall quality and listening experience of the speech, ensuring clear and comfortable auditory perception in various playback environments.
[0156] Firstly, the enhanced speech data output from step 3.3 undergoes professional-level audio post-processing, which includes multiple key aspects:
[0157] Noise reduction: Although synthesized speech should theoretically contain no noise, some synthesis methods (especially sample-based methods) may introduce slight background noise. Spectral subtraction and adaptive filtering algorithms are applied to remove any potential minor noise, improving the purity of the speech.
[0158] Equalization (EQ) processing: Based on the sensitivity of the human ear to different frequencies, a professional equalizer is used to adjust the frequency response of the speech. Typical processing includes slightly boosting the 3-5 kHz range to enhance clarity, appropriately reducing the low-frequency rumble in the 200-300 Hz range, and controlling high frequencies above 10 kHz to avoid harshness.
[0159] Dynamic range compression: Educational videos are often played on various devices and in different environments, and excessive dynamic range may result in inaudibility in certain scenarios. A multi-band compressor is applied to reduce the dynamic range of the volume with a moderate compression ratio (usually 1.5:1 to 3:1), ensuring that quiet parts are clear enough, while loud parts are not overly jarring.
[0160] Psychoacoustic enhancement: Processing algorithms based on human auditory perception principles are applied to enhance the presence and clarity of the speech, including slight harmonic excitation and transient enhancement, making the speech sound more natural and lively.
[0161] To adapt to different playback environments, multiple targeted optimized audio versions are generated. For example, versions optimized for mobile devices apply stronger compression and clarity enhancement, while versions optimized for professional sound systems retain more dynamic range and frequency details.
[0162] Final quality assessment is also conducted, using objective indicators (such as PESQ, STOI, and other speech quality assessment indicators) and machine learning models to predict subjective listening experience, ensuring that the processed audio meets the expected quality standards.
[0163] The final output speech data stream is comprehensively optimized, with excellent clarity, natural rhythm performance, and professional audio quality, providing high-quality audio tracks for subsequent video synthesis, and also serving as an important reference for action expression generation.
[0164] As Figure 3 As shown, in step S4, based on the enhanced script and the speech data stream, key verbs, emphasis words, and emotional vocabulary are identified through semantic analysis, a mapping relationship between semantic content and appropriate body movements is established, the precise time points of movement execution are calculated, and facial expression instructions consistent with the content emotion are generated, and a complete set of movement expression instructions is output, including:
[0165] Step S4.1, based on the enhanced script, key verbs, emphasis words, and emotional vocabulary are identified through semantic analysis techniques, a mapping relationship between semantic content and appropriate body movements is established, and a semantic-action correlation mapping table is output;
[0166] Step S4.2, based on the speech data stream and the semantic-action correlation mapping table, the time points of movement execution are calculated by analyzing the rhythm, pause, and accent features of the speech, and a time-synchronized movement sequence is output;
[0167] Step S4.3, based on the emotional markers in the enhanced script and the intonation changes in the speech data stream, facial expression instructions consistent with the content emotion are generated through an emotion-expression mapping model, and a time-synchronized expression sequence is output;
[0168] Step S4.4, based on the time-synchronized movement sequence and the time-synchronized expression sequence, natural transitions and coordinated consistency between movements and expressions are ensured through movement-expression coordination algorithms, and a complete set of movement-expression instructions is output.
[0169] The core of this step is to establish a mapping relationship between text content and appropriate body movements, enabling the digital person to exhibit natural movements that match the content being explained.
[0170] First, the enhanced script from step S2 is subjected to in-depth semantic analysis, which goes beyond simple keyword recognition and employs multi-level semantic understanding:
[0171] Language behavior recognition: Different language behavior types in the script are identified, such as explanation, emphasis, enumeration, comparison, example, etc. Each language behavior is associated with a series of potential hand gestures and body movements. For example, explanatory content may be accompanied by open hand gestures, and enumeration content may be accompanied by sequential counting gestures.
[0172] Spatial concept mapping: When the content involves spatial relationships (such as "rise", "expand", "left side", etc.), these spatial concepts are automatically mapped to corresponding directional gestures. For example, when describing a growing trend, an upward gesture may be generated; when describing structural relationships, two hands may be used to represent the positional relationships between different components.
[0173] Emotional Intensity Analysis: Assess the emotional intensity and importance of the content to determine the amplitude and energy level of the actions. Important concepts or emotionally intense content will trigger more pronounced and forceful actions, while secondary information will be accompanied by more restrained and subtle movements.
[0174] Semantic features are extracted using a pre-trained large-scale language model, and a specially designed attention mechanism is used to identify key content that needs to be emphasized. These semantic features are then converted into vectors in the action feature space through a mapping network.
[0175] The mapping network is a deep neural network trained on a large number of instructor videos, which learns the typical action patterns of professional instructors when explaining different types of content. This network can predict the most suitable action category, amplitude, speed, and emotional expression based on the semantic features of the content.
[0176] To enhance the diversity and naturalness of the mapping, a motion variation generator is also integrated to avoid repeating the same actions when similar content appears. The variation generator is based on a probabilistic model that introduces controlled randomness while maintaining semantic appropriateness, making the digital person's action performance more varied and rich.
[0177] The final output semantic-action mapping table is a structured dataset that contains the correspondence between each paragraph or sentence in the script and the recommended action type, as well as parameter suggestions for action execution (such as amplitude, speed, emotional color, etc.). This mapping table will serve as a key input for the next step, guiding the generation of precise action sequences.
[0178] In step S4.2:
[0179] After establishing the association between semantics and actions, this step will ensure that the digital person's actions are accurately synchronized with the speech expression, which is crucial for creating a natural and smooth explanation experience.
[0180] First, perform a detailed acoustic analysis on the speech data stream output in step 3, extracting the following key features:
[0181] Speech rhythm tagging: Identify the syllable boundaries, stressed syllables, pause points, and intonation change points in the speech. Use forced alignment algorithms to accurately match the speech with the text, obtaining precise timestamps for each word and syllable.
[0182] Prosodic structure analysis: Identify the prosodic hierarchical structure of the speech, including prosodic feet, prosodic phrases, and intonational phrases. These structures provide natural boundaries for action segmentation.
[0183] Emphasis extraction: Identify emphasized parts in speech, including pitch peaks, intensity enhancements, and syllable lengthening. These parts often require corresponding emphasized actions.
[0184] Based on these acoustic features and the semantic-action association mapping table output in Step 4.1, apply an action-speech synchronization algorithm to calculate the precise execution time of each action. This algorithm follows several core principles:
[0185] Preparation principle: Natural gestures usually start 0.2-0.5 seconds before the relevant word, with a dynamically calculated appropriate preparation time based on action complexity.
[0186] Emphasis principle: The climax of the action (such as the maximum amplitude point of the gesture) should be precisely synchronized with the emphasis in the speech (such as the stressed syllable).
[0187] End smoothness principle: Actions should naturally retract after the end of a semantic unit, avoiding abrupt cuts.
[0188] Smooth transition principle: Maintain smooth transitions between adjacent actions, calculating the optimal transition path to avoid unnatural jumps.
[0189] Optimize the overall action sequence using dynamic programming algorithms, minimizing energy consumption (avoiding excessive activity) and maximizing expressiveness while meeting the above principles. In particular, the algorithm identifies natural pauses in speech as ideal opportunities for inserting complex actions or action transitions.
[0190] To handle unexpected speech changes (such as temporary added pauses or emphasis), implement a real-time adjustment mechanism that can dynamically fine-tune action times during playback, ensuring that synchronization is not affected.
[0191] The final output of the time-synchronized action sequence includes the precise start time, climax time, and end time of each action, as well as the type and parameters of the action. This sequence provides a temporal framework for subsequent expression generation and final action-expression coordination.
[0192] In step S43:
[0193] After determining the timing of the action sequence, this step will generate facial expressions that match the content emotions, making the digital person's expression more rich and lively.
[0194] First, combine two key inputs: one is the emotional markers in the enhanced script output in Step 2, which indicate the emotional tendency of the content (such as enthusiasm, seriousness, curiosity, surprise, etc.); the second is the acoustic features in the speech data stream output in Step 3, including intonation changes, stress patterns, and emotional intonation.
[0195] The expression generation adopts a multi-level emotion-expression mapping model:
[0196] Basic emotion layer: mapping six basic emotions (joy, sadness, anger, fear, disgust, surprise) and their combinations to corresponding facial expression parameters. Adopting improved Facial Action Coding System (FACS), controlling 42 key muscle action units of the digital human face.
[0197] Teaching-specific expression layer: special expression needs for teaching scenarios, such as "thinking" expression (when explaining complex concepts), "asking" expression (when guiding thinking), "confirming" expression (when emphasizing key points), etc.
[0198] Micro-expression layer: adding subtle eyebrow movements, eye blinks, mouth corner adjustments, and other natural expression changes to enhance the vitality of the digital human. These micro-expressions are generated based on a probabilistic model, simulating human natural facial micro-movements.
[0199] The expression generation process adopts a deep learning-based emotion-expression generation network trained on a large number of annotated lecturer videos, which can convert text emotion and speech emotion features into natural facial expression parameters. The network architecture adopts a Transformer encoder-decoder structure, which can capture long sequence emotion change patterns.
[0200] To ensure the temporal synchronization of expressions, a timing control mechanism closely integrated with speech prosody is adopted. For example, when emphasizing key points, expression changes (such as raising eyebrows) will be accurately synchronized with the stressed syllables in the speech; when expressing surprise or doubt, expression changes will be slightly earlier than the corresponding speech, simulating human natural reaction patterns.
[0201] The continuity and natural transition of expressions are also considered to avoid abrupt changes in expressions. This is achieved through an expression smoothing algorithm that creates natural transition curves between adjacent expressions, ensuring coherent changes in facial expressions.
[0202] The final output of the time-synchronized expression sequence contains precise timestamps and detailed facial parameter configurations for each expression state, providing complete expression instructions for subsequent motion-expression coordination.
[0203] In step S4.4:
[0204] As the last step of motion-expression generation, this step will integrate and coordinate the previously generated motion sequence and expression sequence, solve potential conflicts, and ensure the natural harmony of the overall performance.
[0205] First, the motion sequence output in step 4.2 and the expression sequence output in step 4.3 are time-aligned and conflict detected. Potential conflicts mainly include the following categories:
[0206] Timing conflict: When two strong expressions (such as important gestures and obvious facial expressions) are too close in time, it may cause distraction or overexpression.
[0207] Semantic conflict: When the emotions or intentions conveyed by actions and expressions are inconsistent, such as gestures indicating affirmation while facial expressions show doubt.
[0208] Physical conflict: When hand actions need to be close to the face (such as the "thinking" gesture), it is necessary to ensure that they do not interfere with facial expressions.
[0209] Intensity imbalance: When the intensity of actions or expressions does not match the importance of the content, such as key points with weak expressions or minor information with exaggerated expressions.
[0210] The action-expression coordination algorithm uses a combination of rule-based and machine learning methods to solve these conflicts:
[0211] For timing conflicts, the time staggering strategy is applied, which adjusts the timing of actions or expressions to ensure that important expressions have enough "exclusive time window". For example, gestures may be slightly ahead or facial expression changes may be delayed, allowing the audience to perceive these expressions in sequence.
[0212] For semantic conflicts, semantic consistency checks are used to ensure that actions and expressions convey consistent information. When inconsistencies are detected, the less conflicting side is adjusted according to the main emotional tendency of the content, maintaining the consistency of the overall expression.
[0213] For physical conflicts, the action adjustment algorithm is applied to fine-tune the hand position while maintaining the semantics of the gesture, avoiding blocking the face or interfering with facial expressions. For example, the "thinking" gesture that should be close to the chin may be slightly moved down to ensure that the facial expression is clearly visible.
[0214] For intensity imbalance, the global intensity balancing algorithm is applied to dynamically adjust the intensity of actions and expressions according to the importance hierarchy of the content. This ensures that the expression intensity is proportional to the importance of the content, enhancing the overall teaching effectiveness.
[0215] The expressiveness smoothing algorithm is also applied to avoid the digital person's actions and expressions being too mechanical or repetitive. This algorithm introduces controlled variability, ensuring that even similar content will have subtle expression differences, enhancing naturalness and liveliness.
[0216] The final output of the complete action-expression instruction set is a time series data structure that contains precise timestamps and all necessary parameters for driving the digital person's skeleton and facial expressions. This instruction set is fully coordinated and optimized to ensure that the digital person's performance is natural, coordinated and teaching effective.
[0217] The action-expression coordination algorithm ensures natural transition and consistent coordination between actions and expressions, including:
[0218] Step S4.4.1, time alignment of the time-synchronized action sequence and the time-synchronized expression sequence, detecting strong expressions in time proximity, outputting a time conflict identifier;
[0219] In this step, first, the time-synchronized action sequence output in step S4.2 and the time-synchronized expression sequence output in step S4.3 are precisely time-aligned and analyzed. This process requires high-precision timestamp comparison and will build a unified time axis to map all action and expression events onto this time axis. For each time window (usually 50-100 milliseconds), it is checked whether multiple strong expressions occur simultaneously. Strong expressions include obvious gestures (such as large-scale pointing, emphatic hand waving), significant posture changes (such as leaning forward, turning), and obvious facial expressions (such as surprise, emphatic eyebrow raising). A weighted scoring mechanism is used to evaluate the strength of each expression, taking into account factors such as action amplitude, speed change, space occupation, and attention attraction. When two or more strong expressions are highly overlapping in time (such as overlapping more than 70% of the duration), and their combined strength exceeds a pre-set threshold, it is marked as a potential time conflict. Such conflicts can cause audience distraction or perceptual overload, affecting teaching effectiveness. Detailed records will be created for each detected conflict, including the type of expression, time range, strength score, and conflict severity. For particularly serious conflicts (such as multiple highest intensity expressions completely overlapping), higher priority will be given to ensure priority processing in subsequent steps. In addition, continuous expression sequences are analyzed to identify areas that may cause "expression congestion", i.e., time periods containing too many expressions in a short period of time. This analysis takes into account the time characteristics of human perception, such as the minimum time interval required for attention switching. Finally, a structured time conflict identifier dataset is output, containing detailed information and priority ranking of all potential conflicts, providing a comprehensive basis for the next step of conflict resolution.
[0220] Step S4.4.2, based on the time conflict identifier, applying a time staggering strategy to ensure that important expressions have enough exclusive time windows, outputting a time-optimized action-expression sequence;
[0221] In this step, based on the temporal conflict identifiers output in step S4.4.1, an intelligent time staggering strategy is applied to resolve the detected conflicts. First, all conflicts are prioritized, and the processing order is determined based on the conflict severity, the importance of the expressions involved, and the difficulty of adjustment. For each conflict, it is necessary to decide which expressions need adjustment and how to adjust them. This decision-making process follows several key principles: the content importance principle (expressions directly related to the core teaching content are prioritized to remain in their original positions), the perceptual naturalness principle (adjustments should maintain the natural fluency of human actions as much as possible), and the minimum intervention principle (adjustments should be made to the minimum extent possible while meeting the needs). A dynamic programming algorithm is used to calculate the optimal time adjustment scheme, which considers the dependencies and continuity requirements between expressions. For expressions that need adjustment, different strategies are adopted according to their nature: for anticipated actions such as gestures, the start time is tended to be advanced, taking advantage of the natural characteristic that human actions usually precede speech; for reactive expressions, their appearance time may be slightly delayed; for continuous expressions, the timing of their intensity peak may be adjusted to ensure that the peak does not conflict with other expressions. The timing adjustments are typically controlled within the range of 100-500 milliseconds. This range is sufficient to resolve conflicts without causing a noticeable unnatural feeling. New conflicts that may arise after the adjustment are also considered, and global optimization ensures that resolving one conflict does not create new problems. For complex conflicts that cannot be resolved by simple timing adjustments, more complex strategies are considered, such as decomposing complex expressions into multiple simpler expressions or adjusting the execution speed of the expressions. Ultimately, the output is a timing-optimized sequence of actions and expressions. This sequence preserves the semantic intent and expressiveness of the original expressions while ensuring that each important expression has a sufficient "exclusive time window," allowing the audience to clearly perceive the information conveyed by each expression and improving the overall teaching effectiveness.
[0222] Step S4.4.3: Perform semantic consistency checks on the time-optimized action and expression sequence, identify inconsistencies in the information conveyed by actions and expressions, and output a list of semantic conflicts;
[0223] In this step, the action-expression sequence output from step S4.4.2 is subjected to an in-depth semantic consistency analysis to ensure that the information conveyed by the actions and expressions complement each other rather than contradict each other. First, a comprehensive semantic interpretation framework is established to map various actions and expressions to their typical conveyed semantic and emotional meanings. For example, nodding generally indicates affirmation or agreement, while shaking the head indicates negation or disagreement; smiling indicates positivity or friendliness, while furrowing the brow indicates confusion or dissatisfaction. This mapping framework is based on extensive human behavior studies and visual semantic databases, covering common hand gestures, postures, and facial expressions. Subsequently, the semantic analysis is performed on the combination of actions and expressions within each time segment to assess whether they convey consistent information. This analysis considers multiple levels of consistency: emotional consistency (whether the actions and expressions express the same emotional inclination, such as positive / negative), emphasis consistency (whether the emphasized content is consistent), indication consistency (whether the direction of pointing or directing attention is consistent), and logical consistency (whether there are logical contradictions, such as expressing both affirmation and negation at the same time). When potential semantic conflicts are detected, further analysis is performed on the nature and severity of the conflicts. For example, minor inconsistencies (such as a neutral hand gesture accompanied by a slight smile) may be acceptable, while obvious contradictions (such as affirmative verbal content accompanied by negative head movements) need to be resolved. Detailed records are created for each detected semantic conflict, including the conflicting expression elements, conflict type, conflict severity, and related content context. In addition, cultural background factors are considered, as some gestures and expressions may have different or even opposite meanings in different cultures. The judgment criteria for semantic consistency are adjusted according to the cultural background of the target audience. Finally, a structured semantic conflict list is output, containing all inconsistencies that need to be resolved, sorted by severity, and accompanied by resolution suggestions, providing the basis for the next step of semantic coordination.
[0224] Step S4.4.4, based on the semantic conflict list, adjusting the less conflicting expression methods according to the emotional inclination of the content, maintaining the consistency of the overall expression, and outputting a semantic coordinated action-expression sequence;
[0225] In this step, based on the semantic conflict list output in step S4.4.3, an intelligent semantic coordination process is performed to resolve the inconsistency between actions and expressions. First, the context of each conflict is analyzed, including the semantic importance of the relevant content, emotional tendency, and teaching intent. This analysis helps determine which expression elements should be prioritized in conflict resolution. Generally, expressions directly related to core teaching content, as well as those that are more important in terms of emotion or emphasis, are given priority. For each semantic conflict, a decision needs to be made on how to adjust the less dominant expression. This decision-making process uses a multi-factor evaluation model, considering the substitutability of expressions (whether there are other ways to convey the same information), adjustment difficulty (some expressions are easier to modify than others), and naturalness after adjustment (whether the modified expression still looks natural and smooth). Various adjustment strategies are provided: replacement strategy (replacing the conflicting expression with a semantically compatible alternative), weakening strategy (reducing the intensity of the conflicting expression to make it neutral or auxiliary), enhancement strategy (enhancing the intensity of the dominant expression to make it clearly dominant in semantic communication), and fusion strategy (creating a new composite expression that combines key elements of both expressions). For example, when facial expressions and hand gestures convey different emotions, the emotional intensity of the secondary expression can be adjusted according to the primary emotional tendency of the content; when the pointing action and gaze direction are inconsistent, these elements can be realigned to ensure they point to the same target. For complex semantic conflicts, the temporal relationship between expressions is considered, and conflicts can be resolved by adjusting the time sequence of expressions, such as making one expression clearly precede or follow the other, thus creating a transition or contrast effect instead of direct conflict. The overall semantic coherence after adjustment is also evaluated to ensure that resolving local conflicts does not disrupt the expression logic on a larger scale. Finally, a semantically coordinated action-expression sequence is output, in which all expression elements in the sequence are consistent in terms of semantics, enhancing the teaching effect and emotional communication of the content, and avoiding confusion or misunderstanding caused by inconsistent expressions.
[0226] Step S4.4.5, based on the semantically coordinated action-expression sequence, dynamically adjusts the intensity of actions and expressions according to the content importance hierarchy through a global intensity balancing algorithm, outputting a complete intensity-balanced action-expression instruction set.
[0227] In this step, based on the semantically coordinated action-expression sequence output in step S4.4.4, a global intensity balancing algorithm is applied for the final expressiveness optimization. The core goal of this step is to ensure that the intensity of actions and expressions is proportional to the importance of the content, avoiding weak expressions for important content or exaggerated expressions for secondary information. First, a content importance model is constructed, which assesses the importance level of content based on multiple factors, including semantic importance (core concepts, key conclusions, etc.), structural importance (key positions such as titles, summaries), teaching importance (difficult points, exam points, etc.), and emotional importance (content that requires emotional resonance). The content is divided into multiple importance levels, and each level is assigned a corresponding expression intensity range. The global intensity balancing algorithm uses dynamic programming to optimize the intensity distribution of the entire sequence while maintaining natural and smooth expression. The algorithm considers multiple constraints: the matching degree of intensity and importance, the smoothness of intensity changes (avoiding abrupt changes in intensity), the dynamic range of overall intensity (ensuring both climax and calm), and human perception habits (such as progressive intensity changes are generally more natural than random changes). For action intensity adjustment, multiple parameters can be modified, including action amplitude (range of hand movements), execution speed (speed of action), repetition times (for emphasis actions), and additional secondary actions (such as slight body tilt, head coordination, etc.). For expression intensity adjustment, the apparent degree of expression (such as the degree of smile), duration, transition speed (speed of expression formation and disappearance), and complexity (mixing ratio of multiple basic expressions) can be modified. Special attention is paid to the rhythm of intensity changes to create natural intensity fluctuations and avoid monotony or excessive frequency of intensity changes. In addition, the fatigue effect of long-term expression is also considered, and the intensity of continuous high-intensity expression is appropriately reduced, with natural relaxation moments inserted to make the overall performance more natural and sustainable. Appropriate intensity variability is also preserved, with subtle expression differences even for similar content, enhancing naturalness and liveliness. Finally, the complete action-expression instruction set with balanced intensity is output, which not only has semantic consistency, but also has a staggered timing and precise matching of content importance, maximizing teaching effectiveness while maintaining natural and lively expression.
[0228] In step S5, the museTalk technology is used to drive the digital human model, including:
[0229] Step S5.1, based on the pre-defined digital human model library, according to the teaching content nature and target audience, select the appropriate lecturer image, initialize the skeleton structure and facial control points, output the ready digital human model;
[0230] In this step, in the first step of the video generation process, the appropriate digital human model needs to be prepared and initialized, laying the foundation for subsequent animation driving.
[0231] First, based on the nature of the teaching content and the target audience, the most suitable lecturer avatar is selected from a predefined library of digital human models. This selection can be explicitly specified by the user or automatically recommended based on content analysis. The model library contains digital human models of different genders, ages, styles, and professional field characteristics, each designed and optimized for specific types of teaching content.
[0232] After selecting the model, the following initialization steps are performed:
[0233] Skeleton loading: The skeleton of the digital human is the core framework for controlling limb movements. A predefined skeleton hierarchy is loaded, including about 60-120 bone nodes covering all parts of the body. The skeleton uses a standard hierarchy, using quaternions or matrices to represent bone rotations, ensuring the accuracy and smoothness of movement calculations.
[0234] Facial control initialization: Facial expression control uses two complementary techniques: one is based on FACS muscle control, controlling about 42 facial action units; the other is based on advanced expressions using blend shapes, providing predefined composite expressions. Initialize both sets and establish the mapping between them to ensure the accuracy and richness of expression control.
[0235] Physical parameter settings: To enhance the realism of the animation, initialize the physical simulation parameters, including clothing dynamics, hair dynamics, and secondary animation effects (such as subtle body sways). These parameters will be adjusted according to the selected digital human characteristics, for example, long hair models will have more complex hair dynamics settings.
[0236] Rendering parameter configuration: Configure rendering parameters related to the digital human, including skin material, sub-surface scattering parameters, eye reflection characteristics, etc., to ensure high-quality visual effects in subsequent rendering. For different output targets (such as high-definition video or real-time streaming), these parameters will be dynamically adjusted to balance quality and performance requirements.
[0237] Initial state setting: Place the digital human in the default lecture posture, usually a natural standing or sitting posture, with relaxed hands and neutral facial expression. This state serves as the starting point for all animations, ensuring the coherence of animation sequences.
[0238] Model warm-up techniques are used to pre-calculate and cache commonly used movements and expression states before actual animation begins, reducing computational delays in subsequent processing. For complex digital human models, model simplification algorithms are also applied to create multi-level detail versions for different viewing angles and scene requirements.
[0239] The final output of the ready digital human model contains complete skeletons, facial controls, and all necessary rendering parameters, providing a solid foundation for subsequent motion expression driving.
[0240] Step S5.2, based on the motion expression instruction set, parse the motion expression instruction set into a skeletal animation control flow, a facial animation control flow, and a lip-sync control flow, output three parallel control flows;
[0241] In this step, after the digital human model is ready, this step will drive the digital human using the motion expression instruction set output in step 4, achieving accurate synchronization with the animation effect of the voice.
[0242] First, the motion expression instruction set is parsed into three parallel but interrelated control flows:
[0243] Skeletal animation control flow: control the body movements of the digital human, including gestures, posture changes, and body movements.
[0244] Facial animation control flow: control facial expressions, including dynamic changes in facial elements such as eyebrows, eyes, and mouth.
[0245] Lip-sync control flow: responsible for the precise matching of lip movements and speech, a special subset of facial animation.
[0246] Use museTalk and other technologies as the core animation driving engine, which integrates advanced skeletal animation and facial expression control. The specific driving process is as follows:
[0247] For skeletal animation, use a combination of keyframe interpolation and procedural animation. Keyframes define the key states of the action (such as the starting, climax, and ending positions of gestures), and automatically calculate intermediate frames to ensure smooth transitions. To enhance naturalness, a secondary action generation algorithm is applied to add subtle body swaying, weight transfer, and momentum effects, making the animation more lively.
[0248] For facial animation, use parameterized facial expression control to map the expression parameters generated in step 4.3 to the digital human's facial controllers. At the same time, control high-level expressions (such as "smile", "think") and low-level muscle action units to achieve rich and delicate expression changes. In particular, natural blinking, micro-expressions, and subtle head movements will be added to enhance the sense of life.
[0249] Mouth synchronization is the most critical part of facial animation, directly affecting the audience's perception of synchronization. Advanced phoneme-to-viseme mapping techniques are used to convert phoneme sequences in speech into corresponding mouth shape. This process takes into account the coarticulation effect (i.e., the influence of adjacent phonemes on mouth shape) and applies a mouth smoothing algorithm to avoid overly mechanical mouth changes. It also specially handles accented syllables to match the opening and closing amplitude of the mouth to the volume intensity.
[0250] To ensure coordination and consistency between control flows, a global animation coordinator is applied, which supervises all animation channels, resolves potential conflicts, and ensures the harmony of the overall performance. In particular, the coordinator ensures that gestures do not interfere with facial expressions, facial expressions do not disrupt mouth synchronization, and all movements are consistent with the semantic and temporal content of the speech.
[0251] A dynamic adaptation mechanism is also implemented, which can adjust animation parameters based on real-time feedback. For example, if it is detected that certain movements may cause self-occlusion at certain angles, the movement amplitude or direction will be automatically fine-tuned to ensure clear visibility of key expressions.
[0252] The final output of the original animation sequence is a complete time sequence containing full-body skeletal transformations and facial expression changes, precisely synchronized with speech data, ready for the next step of visual content integration.
[0253] Among them, museTalk technology is a digital human driving technology specially designed for teaching scenarios, and its core is a multi-modal collaborative driving engine that can realize the precise collaboration of speech, expression, and movement. This technology is based on a deep learning architecture and includes three key models: a timing coordination Transformer network, an expression generation GAN network, and a physical constraint reinforcement learning model.
[0254] The timing coordination Transformer network uses an encoder-decoder structure, containing 8 layers of Transformer blocks, 6 attention heads per layer, and a hidden layer dimension of 768. The network training data comes from 500 hours of high-quality teaching videos annotated by humans, including speech text, phoneme timestamps, facial key point movements, and body movement parameters. Training uses a two-stage strategy: first, pre-training on large-scale unlabeled data in a self-supervised manner, then fine-tuning on labeled data in a supervised manner. Key hyperparameters include: learning rate of 0.00005, using AdamW optimizer, weight decay of 0.01, and training iteration number of about 200,000 times.
[0255] The expression generation GAN network adopts a conditional generative adversarial network architecture, the generator uses a U-Net structure, and the discriminator adopts a PatchGAN design. The network receives emotion labels and speech features as conditional input and generates natural facial expression parameters. The training data contains a facial expression dataset from 100 different performers, covering 6 basic emotions and 20 composite emotions.
[0256] The physical constraint reinforcement learning model optimizes the generated action sequence by simulating the constraints of the physical world, such as inertia, gravity, joint limits, etc., to ensure the physical rationality of the action. The model uses the TD3 (Twin Delayed DDPG) algorithm for training, and the reward function considers three dimensions: action fluency, expressiveness, and physical rationality.
[0257] Experimental verification shows that compared with traditional methods, the museTalk technology improves the voice-action synchronization accuracy by 35%, the expression naturalness score by 42%, and the user satisfaction by 27%. The technology has been verified in more than 5000 teaching video generation cases, showing superior stability and expressiveness.
[0258] Step S5.3, based on the skeletal animation control flow, use keyframe interpolation and procedural animation methods to control the digital person's limb movement, add subtle body sway, weight transfer and momentum effect, output natural skeletal animation sequence;
[0259] In this step, based on the skeletal animation control flow parsed in step S5.2, natural and smooth digital human limb movements are generated. A hybrid approach combining keyframe interpolation and procedural animation is adopted, which ensures precise control of movements while adding natural variations and liveliness. First, according to the instructions in the movement control flow, precise keyframes are defined for each key movement. These keyframes define the key states of the movement, such as the starting position, maximum extension position, and ending position of a gesture; the turning angle of the body; or the nodding amplitude of the head, etc. Using advanced inverse kinematics (IK), it is ensured that these key poses are anatomically correct and natural. Subsequently, advanced interpolation algorithms are applied to generate smooth transitions between keyframes. Unlike simple linear interpolation, physics-based spline interpolation is used, considering acceleration and deceleration, simulating the characteristics of real human motion. For example, there is an acceleration process at the beginning of the movement, a deceleration process at the end, and possibly a steady speed stage in the middle. This interpolation method makes the movement look more natural and smooth, avoiding a mechanical and stiff feeling. To further enhance the naturalness of the movement, multi-level procedural animation techniques are applied. First, secondary movement generation automatically adds secondary movements that coordinate with the main movement, such as a slight lifting of the shoulders when waving the arms, or a slight leaning forward of the upper body when pointing. These secondary movements, although subtle, are crucial to enhancing the overall naturalness. Second, natural variation generation adds slight random variations while maintaining the semantic meaning of the movement, such as slight differences in each repeated movement, avoiding completely identical mechanical repetition. Various physical effect simulations are also added, including body sway (such as slight balance adjustment when standing), weight transfer (such as natural center of gravity changes from one foot to the other), and momentum effects (such as natural inertia and rebound after fast movements). These effects are achieved through simplified physical simulations, balancing computational efficiency and visual effects. In addition, personalized movement styles are considered, adjusting the overall style of the movement according to the selected digital human image characteristics (such as age, gender, professional background, etc.), such as older lecturers may have more stable movements, while younger lecturers may be more lively. Context-aware movement is also implemented to ensure natural transitions and coherence between adjacent movements, avoiding abrupt posture changes. Finally, a natural skeletal animation sequence is output, which contains complete skeletal transformation data and can drive the digital human model to perform smooth and natural limb movements, providing a foundation for subsequent overall animation synthesis.
[0260] Step S5.4, based on the face animation control flow, using parameterized facial expression control to map expression parameters to digital human face controllers, adding natural blinking, micro-expressions, and subtle head movements, outputting rich facial animation sequences;
[0261] In this step, based on the facial animation control flow parsed in step S5.2, rich and natural digital human facial expressions are generated. A parameterized facial expression control method is adopted, which accurately maps abstract expression parameters (such as "smile degree", "surprise degree", etc.) to specific control points of the digital human face. Modern digital human faces usually adopt two complementary controls: one is muscle control based on FACS (Facial Action Coding System), which controls about 42 facial action units (AUs), such as eyebrow raising, mouth corner lifting, etc.; the other is high-level expression based on blend shapes, which provides predefined composite expressions such as "smile", "thinking", etc. First, the expression instructions in the facial animation control flow are parsed into the control parameters of the two. For basic emotional expressions (such as joy, surprise, etc.), optimized expression mapping templates are used to ensure the accuracy and naturalness of the expressions. For teaching-specific expressions (such as "thinking", "emphasis", etc.), expression templates specially designed for teaching scenarios are applied, which are trained and optimized from professional instructor expression data. Special attention is paid to the delicacy and natural transition of expressions, and advanced expression interpolation algorithms are applied to ensure smooth and natural expression changes and avoid abrupt jumps. To enhance the life and naturalness of facial expressions, three key types of auxiliary animations are added: natural blinking, micro-expression, and subtle head movements. Natural blinking is an important element to enhance the life of digital humans, and a blinking generator based on a probability model is implemented, which dynamically adjusts the blinking frequency and manner considering various factors, including emotional state (such as increased blinking frequency when nervous), attention state (such as reduced blinking when concentrating on explanation), and physiological needs (such as necessary blinking after a long time without blinking). Micro-expression is a subtle and short expression change that naturally exists on human faces, and appropriate micro-expressions are automatically generated by analyzing the emotional flow and explanation rhythm of the content, such as short eyebrow lifting, slight mouth twitching, or slight nose expansion. These micro-expressions, although subtle, are crucial to breaking the static feeling of the face and enhancing the life. Subtle head movements are natural accompanying movements of facial expressions, and appropriate head movements are automatically generated according to the expression type and intensity, such as nodding, slight shaking, tilting, or lifting the head. These head movements are closely coordinated with facial expressions to enhance the overall feeling and naturalness of expression. Personalized customization of expressions is also implemented, adjusting the overall style of expressions according to the selected digital human characteristics, such as different age, gender, or personality characteristics of digital humans will have different expression tendencies and expression methods. In addition, cultural adaptability is also considered, and the expression method of some expressions can be adjusted according to the cultural background of the target audience to ensure the cultural appropriateness of the expressions. Finally, a rich facial animation sequence is output, which contains complete facial control parameter data and can drive the digital human model to exhibit natural, lively, and expressive facial expressions, providing an important component for subsequent overall animation synthesis.
[0262] Step S5.5, based on the lip-sync control stream and the speech data stream, convert the phoneme sequence in the speech into corresponding lip shapes using phoneme-to-viseme mapping techniques, apply a lip smoothing algorithm to avoid too mechanical lip changes, and output a complete animation sequence synchronized with the speech.
[0263] In this step, based on the lip-sync control stream parsed in step S5.2 and the speech data stream output in step S3, a precisely synchronized lip animation is generated. Lip synchronization is one of the most critical aspects of digital human animation, directly affecting the audience's perception of synchronization and realism. Advanced phoneme-to-viseme mapping techniques are employed to convert the phoneme sequence in speech (the basic pronunciation unit of speech) into the corresponding lip shape (the basic visual unit of lip shape). First, a detailed acoustic analysis of the speech data stream is performed to extract the accurate phoneme sequence and its timestamp. This process uses a forced alignment algorithm to accurately match the speech with the text, obtaining the start and end times of each phoneme. For synthetic speech, this information can be directly obtained from the speech synthesis process; for recorded speech, it needs to be analyzed and extracted through an acoustic model. Subsequently, phoneme-to-viseme mapping rules are applied to convert each phoneme or combination of phonemes into the corresponding lip shape. This mapping takes into account the characteristics of the language, and different languages may have different mapping rules. The mapping library contains various lip shapes, such as the open mouth shapes corresponding to the vowels "a", "i", "u", and the closed-lip shapes corresponding to the consonants "m", "p", "b", etc. Special consideration is given to the coarticulation effect, i.e., the mutual influence of adjacent phonemes on the lip shape. For example, the same phoneme may have different lip shapes in different contexts, and the mapping rules with context perception capture these subtle changes. To make the lip animation more natural and smooth, advanced lip smoothing algorithms are applied. Simple phoneme-to-viseme direct mapping may result in lip changes that are too mechanical and abrupt, and the smoothing algorithm simulates the continuous changes in lip shape when a person speaks by creating natural transitions between lip shapes. A physics-based lip interpolation model is used, taking into account the motion characteristics of the mouth muscles, such as the time required for different lip shape transitions and the transition path. Special handling is also given to accented syllables, matching the opening and closing amplitude of the lip shape to the volume intensity. For emphasized words or syllables, the opening and closing amplitude and clarity of the lip shape are increased; for fast or soft parts, the lip shape changes are correspondingly reduced. In addition, personalized adjustments of the lip shape are implemented, fine-tuning the lip parameters according to the facial features and speaking style of the digital human, such as some people speaking with a larger mouth opening amplitude, while others have a smaller one; some people have specific motion patterns in the corners of the mouth when speaking, etc. Non-speech factors that affect the lip shape are also considered, such as laughter, sighs, or deep breaths, and corresponding lip animations are generated for these special sounds. To ensure precise synchronization of the lip shape with the speech, an adaptive time adjustment mechanism is implemented, which can fine-tune the lip animation time according to the actual playback situation, compensating for possible delays or desynchronization. Finally, the lip animation is integrated with the previously generated facial expressions and skeletal animations to output a complete animation sequence perfectly synchronized with the speech. This sequence contains full-body skeletal transformations, facial expression changes, and precise lip animations, forming a coordinated and unified whole, providing a complete digital human animation foundation for subsequent visual content integration.
[0264] wherein the optimized parallel rendering algorithm integrates the digital presenter and teaching content into a unified visual scene, including:
[0265] Step S5.6, based on the PPT content in the structured data, extract visual elements including slides, images, charts, and text highlights, analyze content readability, presenter visibility, content-presenter relationship, and picture balance, and output the most suitable picture layout strategy;
[0266] Step S5.7, based on the picture layout strategy and the complete animation sequence, automatically adjust the position, angle, and focal length of the virtual camera according to the importance of the content and the action of the presenter, add dynamic visual emphasis effects including highlighting, zooming, and arrow pointing to important content, and output a visual scene with shot language;
[0267] Step S5.8, based on the visual scene with shot language, process the interaction effect between the presenter and the content, apply transition effects including fade-in, fade-out, and zoom transition at content switching points, and output an interaction-enhanced scene sequence;
[0268] Step S5.9, based on the interaction-enhanced scene sequence, apply color correction, lighting balance, and layer blending techniques to ensure that the digital presenter is visually unified with the background and content, and output a visually integrated complete scene sequence;
[0269] Step S5.10, based on the visually integrated complete scene sequence, output a professional complete video scene sequence through high dynamic range rendering, temporal anti-aliasing, motion blur, and color grading processing.
[0270] Step S5.6 Visual Content Integration and Synthesis
[0271] After obtaining the animation sequence of the digital person, this step integrates the digital presenter and teaching content (such as PPT slides, charts, demonstration materials, etc.) into a unified visual scene, creating a complete teaching video picture.
[0272] First, analyze the PPT content analysis results from Step 1, extract all visual elements (such as slides, images, charts, and text highlights) and their logical structure. Based on these analyses, design the most suitable picture layout and content display strategy, considering the following key factors:
[0273] Content readability: Ensure that text and charts are large enough and clear, especially when viewed on mobile devices.
[0274] Presenter visibility: Ensure that the key expressions of the digital presenter (such as facial expressions and important gestures) are clearly visible.
[0275] Content-lecturer relationship: Optimize the spatial relationship between the lecturer and the content, so that the lecturer's gaze direction and indicative gestures can naturally guide the audience's attention.
[0276] Picture balance: Create a visually balanced composition, avoiding overly crowded or empty frames.
[0277] Support multiple scene layout templates, such as "lecturer side + content main", "lecturer small picture + content full screen", "lecturer and content partition", etc. According to the content type and complexity, the most suitable layout will be dynamically selected, and the layout will be naturally switched at the content transition point.
[0278] To enhance teaching effectiveness, intelligent lens language is implemented, including the following technologies:
[0279] Dynamic lens adjustment: According to the importance of the content and the lecturer's actions, the position, angle and focal length of the virtual camera will be automatically adjusted. For example, when explaining a complex chart, a wide-angle lens is used to show the lecturer and the content at the same time; when emphasizing key points, the lecturer's close-up shot may be switched to.
[0280] Visual emphasis effect: Add dynamic visual emphasis to important content, such as highlighting, zooming, arrow pointing, etc. These emphasis effects are precisely synchronized with the lecturer's voice and gestures, enhancing the indication effect.
[0281] Smooth transition effect: At the content transition point, professional transition effects such as fade-in and fade-out, wipe-in and wipe-out, zoom transition, etc. are applied to ensure visual smoothness. The logical relationship of the content is analyzed, and the most suitable transition type is selected.
[0282] Also handle the interaction effect between the lecturer and the content, such as when the lecturer points to a certain point in the PPT, it may trigger the corresponding visual feedback (such as element highlighting); when the lecturer explains the chart, the chart may be gradually constructed or dynamically changed with the explanation progress. These interactive effects are realized through precise time synchronization and space mapping, enhancing the intuitiveness and interactivity of teaching.
[0283] To ensure professional visual quality, advanced synthesis techniques are applied, including color correction, light balance and level mixing. In particular, ensure that the digital lecturer is visually harmonious with the background and content, avoiding the "paste" feeling, and creating the visual effect of the lecturer's real presence in the teaching environment.
[0284] The final output of the complete video scene sequence includes the digital lecturer, teaching content and all visual effects, in the form of high-quality video frame sequence, ready for the final rendering stage.
[0285] The output final lecturer teaching video includes:
[0286] Step S5.11, based on the complete video scene sequence and the voice data stream, decompose the entire video rendering task into time-sliced and space-sliced subtasks that can be processed in parallel, output the decomposed rendering task set;
[0287] Step S5.12, based on the decomposed rendering task set, intelligently allocate computing resources according to available CPU cores and GPU units, prioritize tasks on critical paths to ensure balanced rendering progress, output resource-optimized rendering scheduling;
[0288] Step S5.13, based on the resource-optimized rendering scheduling, utilize GPU parallel computing capabilities to accelerate lighting calculation, material rendering, and post-processing effects, adopt a multi-resolution rendering strategy to first render at a lower resolution for a preview, output a rendering result with progressive quality;
[0289] Step S5.14, based on the rendering result with progressive quality, gradually improve the quality to the target resolution, perform audio-visual synchronization verification, picture quality evaluation, and compatibility testing, output a high-definition video file that has passed quality verification;
[0290] Step S5.15, based on the high-definition video file that has passed quality verification, generate multiple format variants including MP4, WebM, and low-bandwidth optimization according to different usage scenarios, generate complete metadata containing chapter markers, keyword indexes, and content summaries, output the final lecturer teaching video.
[0291] Parallel rendering and video generation
[0292] As the final step of the entire process, parallel rendering and video generation convert all the elements prepared in the previous steps into the final high-quality video file.
[0293] First, integrate the two key inputs: the complete video scene sequence output in step 5.3 (containing all visual elements) and the optimized voice data stream output in step 3.4. These two inputs are already fully synchronized in time, ready for final synthesis.
[0294] To efficiently handle large-scale video rendering tasks, an innovative multi-level parallel rendering pipeline architecture is adopted, including the following key components:
[0295] Task decomposition engine: decompose the entire video rendering task into subtasks that can be processed in parallel. Decomposition strategies include time slicing (dividing the video into multiple time periods) and space slicing (dividing each frame into multiple rendering regions). The system dynamically determines the optimal decomposition granularity, balancing parallel efficiency and management overhead.
[0296] Resource Scheduler: intelligently allocates computing resources (CPU cores, GPU units) based on their availability and the characteristics of the sub-tasks. Prioritizes tasks on critical paths to ensure balanced rendering progress.
[0297] GPU Accelerated Rendering Core: leverages the parallel computing capabilities of modern GPUs to accelerate critical rendering operations such as lighting calculations, material rendering, and post-effects. The system is optimized for different GPU architectures to maximize hardware utilization.
[0298] Progressive Quality Control: employs a multi-resolution rendering strategy, first rendering the entire video at a lower resolution for a preview, then progressively increasing the quality to the target resolution. This approach allows users to view the results early and make adjustments if necessary.
[0299] During rendering, the system applies a series of professional-grade image processing and video enhancement techniques:
[0300] High Dynamic Range (HDR) Rendering: ensures clear visibility of details under different lighting conditions.
[0301] Temporal Anti-Aliasing (TAA): reduces edge flickering and jitter in animations.
[0302] Motion Blur: adds appropriate motion blur effects to enhance the smoothness of actions.
[0303] Color Grading: applies professional color processing to ensure a consistent and appealing visual style for the video.
[0304] Intelligent Compression: dynamically adjusts encoding parameters based on content characteristics to achieve the best balance between file size and quality.
[0305] Supports multiple output formats and quality presets to adapt to different use scenarios: high-definition MP4 format for regular playback, WebM format for web embedding, low-bandwidth optimized versions for mobile devices, and high-quality versions for professional display.
[0306] To meet the needs of different platforms, various variants of the video are automatically generated, such as square versions for social media, vertical versions for mobile device viewing, and versions with different aspect ratios to adapt to various playback environments.
[0307] After rendering is complete, a final quality check is performed, including audio-visual synchronization verification, picture quality assessment, and compatibility testing. Finally, complete metadata is generated, including chapter markers, keyword indexes, and content summaries, to facilitate subsequent content management and retrieval.
[0308] The final output of the lecturer teaching video is a professional level video file with high quality visual performance, clear audio and natural and smooth digital human performance, which fully realizes the automatic conversion from PPT or text to professional teaching video.
[0309] The above merely describes the preferred embodiments of the present application and is not intended to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application. < / emphasis> < / prosody> < / break>
Claims
1. An AI-based speech synthesis and animation-driven lecturer video automatic generation method, characterized in that, The method comprises the following steps: Based on the content features of the user-uploaded PPT file or text script, a content relationship graph is constructed by an improved interior point method, an incremental shortest path algorithm is applied to identify the hierarchical relationship between titles, paragraphs, and main points, and structured data containing semantic structure, key content, and emotional tendency is outputted; Based on the structured data, a full-dynamic parallel single-link clustering algorithm is applied to group the content of the structured data according to semantic similarity, a preset semantic understanding model is used to convert the structured data into narrative text expressed in spoken language and add marks including pauses, tone changes, and emphasis, and an enhanced script with complete expressiveness marks is outputted; Based on the enhanced script, using CosyVoice speech synthesis technology, through The low-rank approximation method decomposes the large-scale feature matrix into multiple blocks and performs block sampling and parallel computing optimization, adjusts prosody and expressiveness, and outputs expressive speech data stream; Based on the enhanced script and the voice data stream, key verbs, emphasized words, and emotional words are identified through semantic analysis, a mapping relationship between semantic content and appropriate body movements is established, the accurate time points of movement execution are calculated, and facial expression instructions consistent with the content emotion are generated, and a complete set of action expression instructions is outputted; Based on the voice data stream and the set of action expression instructions, a digital human model is driven by using the museTalk technology, an optimized parallel rendering algorithm is used to integrate the digital human lecturer and the teaching content into a unified visual scene, and a final lecturer teaching video is outputted.
2. The method of claim 1, wherein, The method comprises the following steps: The content features are subjected to document analysis processing, elements including text, images, and tables are extracted, and a set of content elements subjected to preliminary analysis is outputted; Based on the set of content elements, an improved interior point method is applied to construct a content relationship graph, an incremental shortest path algorithm is used to identify the hierarchical relationship between titles, paragraphs, and main points, and a structured document with hierarchical marks is outputted; Based on the structured document with hierarchical marks, natural language processing technology is used to identify key words, key content, and core concepts, semantic vectors are used to represent the relevance between contents, and the structured data containing semantic structure, key content, and emotional tendency is outputted.
3. The method of claim 1, wherein, Based on the structured data, a full-dynamic parallel single-link clustering algorithm is applied to group the content of the structured data according to semantic similarity, a preset semantic understanding model is used to convert the structured data into narrative text expressed in spoken language and add marks including pauses, tone changes, and emphasis, and an enhanced script with complete expressiveness marks is outputted; Based on the structured data, a full-dynamic parallel single-link clustering algorithm is applied to maintain a dynamically updated network stream model to group the content according to semantic similarity, the content is divided into multiple semantically associated content blocks according to themes and logical relationships, and the content blocks subjected to semantic clustering are outputted; Based on the content blocks subjected to semantic clustering, the structured data is converted into narrative text suitable for spoken language expression by using a preset semantic understanding model, transitional words and conjunctions are added to enhance fluency, and a preliminary voice script is outputted; Based on the preliminary speech script, a speech synthesis markup language (SSML) script containing pauses, tone changes, and emphasis markers is generated, and the enhanced script with complete expressiveness markers is output.
4. The method of claim 1, wherein, The CosyVoice speech synthesis technology is used to synthesize the speech according to the enhanced script The low-rank approximation method decomposes a large-scale feature matrix into multiple blocks, performs block sampling and parallel computing optimization, adjusts prosody and expressiveness, and outputs an expressive speech data stream, comprising: Based on the enhanced script, the best speech synthesis configuration is determined through a parameter optimization algorithm combined with the lecturer style parameters, and a set of speech synthesis parameters is output. Based on the speech synthesis parameter set and the enhanced script, using CosyVoice speech synthesis technology, large-scale feature matrix is block sampled and passed through The low-rank approximation method performs parallel computing optimization, accelerates speech feature extraction and synthesis, and outputs high-quality speech waveform data. Based on the high-quality speech waveform data, a prosody adjustment algorithm is applied to dynamically adjust the pauses, stress, and tone changes of the speech according to the speech synthesis markup language (SSML) markers, and the expressive speech data stream is output.
5. The method of claim 1, wherein, Based on the enhanced script and the speech data stream, key verbs, emphasized words, and emotional vocabulary are identified through semantic analysis, a mapping relationship between semantic content and appropriate body movements is established, the precise time points for action execution are calculated, and facial expression instructions consistent with the content emotion are generated, outputting a complete set of action expression instructions, including: Based on the enhanced script, key verbs, emphasized words, and emotional vocabulary are identified through semantic analysis technology, a mapping relationship between semantic content and appropriate body movements is established, and a semantic-action correlation mapping table is output. Based on the speech data stream and the semantic-action correlation mapping table, the time points for action execution are calculated by analyzing the rhythm, pauses, and stress characteristics of the speech, and a time-synchronized action sequence is output. Based on the emotional markers in the enhanced script and the tone changes in the speech data stream, facial expression instructions consistent with the content emotion are generated through an emotion-expression mapping model, and a time-synchronized expression sequence is output. Based on the time-synchronized action sequence and the time-synchronized expression sequence, the natural transition and consistent coordination between actions and expressions are ensured through an action-expression coordination algorithm, and a complete set of action-expression instructions is output.
6. The method of claim 5, wherein, The natural transition and consistent coordination between actions and expressions through the action-expression coordination algorithm include: Time alignment is performed on the time-synchronized action sequence and the time-synchronized expression sequence, the closeness of strong expressions in time is detected, and a time sequence conflict identifier is output. Based on the time sequence conflict identifier, a time staggering strategy is applied to ensure that important expressions have sufficient exclusive time windows by fine-tuning the time of actions or expressions, and a time sequence optimized action-expression sequence is output. Semantic consistency checks are performed on the time sequence optimized action-expression sequence to identify inconsistencies in the information conveyed by actions and expressions, and a semantic conflict list is output. Based on the semantic conflict list, the expression method with less conflict is adjusted according to the emotional tendency of the content to maintain the consistency of the overall expression, and a semantically coordinated action-expression sequence is output. Based on the semantically coordinated action-expression sequence, the strength of actions and expressions is dynamically adjusted according to the importance level of the content through a global strength balancing algorithm, and a strength-balanced complete action-expression instruction set is output.
7. The method of claim 1, wherein, The use of museTalk technology to drive digital human models includes: Based on a pre-defined digital human model library, a suitable lecturer image is selected according to the nature of the teaching content and the target audience, the skeletal structure and facial control points are initialized, and a ready-to-use digital human model is output. Based on the action expression instruction set, the action expression instruction set is parsed into a skeletal animation control flow, a facial animation control flow, and a mouth shape synchronization control flow, and three parallel control flows are output; Based on the skeletal animation control flow, the limb action of the digital person is controlled using key frame interpolation and programmatic animation methods, subtle body shaking, weight transfer, and momentum effects are added, and a natural skeletal animation sequence is output; Based on the facial animation control flow, the expression parameters are mapped to the digital person facial controller using parameterized facial expression control, natural blinking, micro-expression, and subtle head action are added, and a rich facial animation sequence is output; Based on the mouth shape synchronization control flow and the speech data flow, the phoneme sequence in the speech is converted into the corresponding mouth shape using phoneme-to-viseme mapping technology, and the mouth shape smoothing algorithm is applied to avoid the mouth shape changes being too mechanical, and a complete animation sequence synchronized with the speech is output.
8. The method of claim 7, wherein, The optimized parallel rendering algorithm integrates the digital person lecturer and teaching content into a unified visual scene, including: Based on the PPT content in the structured data, visual elements including slides, images, charts, and text points are extracted, content readability, lecturer visibility, content-lecturer relationship, and picture balance are analyzed, and the most suitable picture layout strategy is output; Based on the picture layout strategy and the complete animation sequence, the position, angle, and focal length of the virtual camera are automatically adjusted according to the content importance and lecturer action, dynamic visual emphasis effects including highlighting, zooming, and arrow pointing are added for important content, and a visual scene with shot language is output; Based on the visual scene with shot language, the interaction effect between the lecturer and the content is processed, transition effects including fade-in, fade-out, and zoom transition are applied at content switching points, and an interaction-enhanced scene sequence is output; Based on the interaction-enhanced scene sequence, color correction, lighting balance, and layer blending techniques are applied to ensure that the digital person lecturer is visually unified with the background and content, and a visually integrated complete scene sequence is output; Based on the visually integrated complete scene sequence, professional-level complete video scene sequences are output through high dynamic range rendering, temporal anti-aliasing, motion blur, and color grading processing.
9. The method of claim 8, wherein, The final lecturer teaching video is output, including: Based on the complete video scene sequence and the speech data flow, the entire video rendering task is decomposed into parallel processable time-sliced subtasks and space-sliced subtasks, and the decomposed rendering task set is output; Based on the decomposed rendering task set, the available CPU cores and GPU units are intelligently allocated for computing resources, the tasks on the critical path are prioritized to ensure balanced rendering progress, and resource-optimized rendering scheduling is output; Based on the resource-optimized rendering scheduling, GPU parallel computing capabilities are used to accelerate lighting calculation, material rendering, and post-processing, and a multi-resolution rendering strategy is adopted to first render at a lower resolution to obtain a preview, and a progressive quality rendering result is output; Based on the progressive quality rendering result, the quality is gradually improved to the target resolution, audio-visual synchronization verification, picture quality evaluation, and compatibility testing are performed, and a quality-verified high-definition video file is output; Based on the quality-verified high-definition video file, generate multiple format variants including MP4, WebM, low-bandwidth optimized, according to different usage scenarios, generate complete metadata including chapter markers, keyword index and content abstract, output the final lecturer teaching video.
Citation Information
Patent Citations
Text-to-video conversion method and device based on deep semantic analysis, equipment and medium
CN119399330A
Emotion synchronization 2D digital human model training method and device
CN119598346A
Intelligent dialogue method and system based on digital human
CN120216646A
Personalized digital human generation method based on single video
CN120298557A
System
JP2025054683A
Cited By
Voice service interaction method and device, storage medium and electronic equipment
CN121354564A
Digital human explanation video generation method and device
CN121691848A
Digital human commentary video generation method and device
CN121691848B
Network broadcasting method and system realized by adopting intelligent generation technology
CN121999759A
A webcasting method and system using intelligent generation technology
CN121999759B