Teaching-oriented multi-modal interactive digital human teaching assistant generation method

Through the multimodal interactive digital human teaching assistant generation method, the problem of semantic-move mismatch and single emotional expression in the digital human teaching assistant system is solved, and the teaching content is highly consistent with body language is achieved, and the knowledge transmission efficiency and interactive realism are improved.

CN120492587APending Publication Date: 2025-08-15CHINA UNIV OF MINING & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510639063.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing digital human teaching assistant system has semantic-action mismatch, single emotional expression and cross-modal timing mismatch in teaching scenarios, making it difficult to achieve efficient knowledge transmission and interaction realism.

Method used

The multimodal interactive digital human teaching assistant generation method is used to generate structured text answers through the SE-QA model, combined with emotional adaptive speech synthesis technology and ST-GCN action scheduling, video is generated using TCN, sound lip synchronization is optimized through DTW, and multimodal quality evaluation and mixed reinforcement learning are carried out to optimize the generation process.

Benefits of technology

It significantly improves the compatibility between teaching content and body language, improves the efficiency of knowledge transmission and interactive realism, and shows accurate behavior mapping ability in concept explanation and case analysis scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492587A_ABST
    Figure CN120492587A_ABST
Patent Text Reader

Abstract

The invention discloses a teaching-oriented multi-modal interactive digital human teaching-assisted generation method, and belongs to the technical field of artificial intelligence education, and the method comprises the steps: generating a structured answer through multi-modal input (voice, text and portrait graph) in combination with a semantic enhancement question and answer model (SE-QA); generating personalized voice by using an emotion adaptive voice synthesis technology; constructing a teaching action video library, extracting action features by using a space-time diagram convolutional network (ST-GCN), generating a video through a time sequence convolutional network (TCN), and optimizing audio-lip synchronization and micro expressions; and a'generation-evaluation-optimization 'closed loop is realized through a multi-modal evaluation and reinforcement learning optimization generation process. According to the teaching-oriented multi-mode interactive digital human teaching-assistant generation method, the technical bottlenecks of semantic-action mismatch, single emotion expression and the like of a traditional digital human system are broken through, the knowledge transmission efficiency and the interaction reality sense can be remarkably improved, and an innovative solution is provided for an intelligent education tool.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence education technology, and in particular to a method for generating a multimodal interactive digital human teaching assistant for teaching. Background Art

[0002] In recent years, the digital transformation of education has continued to deepen, driven by artificial intelligence and multimodal technologies. Intelligent digital human teaching assistants, due to their interactive immersion, precise knowledge transfer, and sustainable service, have become the core carrier of the smart education ecosystem. However, existing systems generally face technical challenges such as discretization of multimodal generation, emotional semantic disconnection, and cross-modal temporal mismatch. The traditional digital human construction process relies on manually choreographed movements, voice, and text mapping, resulting in rigid teaching expressiveness and weak adaptability. Especially in dynamic teaching scenarios, existing technologies struggle to achieve emotion-driven multimodal collaborative generation. Their semantic understanding limitations and rigid movement scheduling severely restrict the realism of teaching and the efficiency of knowledge transfer.

[0003] From the perspective of technical architecture, existing digital human generation systems can be divided into two categories: rule-driven modular frameworks and data-driven end-to-end models. The former ensures process controllability through predefined action libraries and voice templates, but is limited by the diversity of teaching scenarios and often suffers from problems such as mechanical emotional expression and semantic behavior mismatch. The latter, while leveraging deep learning to increase generation freedom, lacks fine-grained semantic constraints and cross-modal alignment mechanisms, and is prone to defects such as lip shape asynchrony and action logic discontinuity. The triple demands of teaching scenarios for knowledge accuracy, emotional adaptability, and natural interaction require the system to possess deep semantic analysis, multimodal feature coupling, and dynamic optimization capabilities. However, the current technical system has not yet formed a closed-loop association mechanism for teaching semantics, emotion, and action.

[0004] From a multimodal generation perspective, the high fidelity of digital human teaching assistants relies on the deep collaboration of speech synthesis, motion scheduling, and audio-visual synchronization technologies. Although prosody control based on models such as WaveNet and Tacotron has been achieved in the field of speech synthesis, the dynamic adaptation of its emotional characteristics to the user's input timbre still faces bottlenecks. Although motion generation technology captures spatiotemporal characteristics through models such as ST-GCN and LSTM, it lacks the support of a priori knowledge base of teaching scenarios, making it difficult to achieve semantic-driven refined scheduling. In addition, existing cross-modal alignment methods mostly use static temporal matching strategies, which cannot cope with the complex requirements of dynamic semantic focus switching in teaching interactions, resulting in insufficient consistency in audio-visual perception. Therefore, building a technical framework that integrates hierarchical semantic parsing, adaptive emotional generation, and closed-loop optimization mechanisms is of great significance for improving the interactive authenticity and teaching effectiveness of virtual teaching scenarios. Summary of the Invention

[0005] The purpose of this invention is to provide a multimodal interactive digital human teaching assistant generation method for teaching, which can break through the technical bottlenecks of traditional digital human systems such as semantic-action mismatch and single emotional expression, significantly improve the efficiency of knowledge transfer and interactive realism, and provide an innovative solution for intelligent educational tools.

[0006] To achieve the above-mentioned object, the present invention provides a method for generating a multimodal interactive digital human teaching assistant for teaching, comprising the following steps:

[0007] S1, the user inputs voice data, question text and full-body portrait image;

[0008] S2. Generate structured text teaching answers through the SE-QA model and teaching knowledge graph;

[0009] S3. Extract the emotional features of the user's input voice data, combine it with the text answer, and use emotion-adapted speech synthesis technology to generate personalized voice answers;

[0010] S4. Build a teaching action video database, use ST-GCN to extract spatiotemporal features, and schedule matching teaching videos;

[0011] S5. Use TCN to extract spatiotemporal features from the video, generate joint motion sequences to suppress sudden action changes, and generate preliminary action videos through the inverse mapping network.

[0012] S6. Using the cross-modal attention alignment mechanism, the dynamic time warping algorithm DTW is used to optimize the lip synchronization path and generate a refined video.

[0013] S7, face restoration and super-resolution reconstruction using the SRNTT model to enhance micro-expressions;

[0014] S8. Through a multimodal quality evaluation feedback system, visual quality, generation authenticity and teaching suitability indicators are integrated to comprehensively evaluate video quality and teaching suitability;

[0015] S9. Adopt the curriculum-driven hybrid reinforcement learning framework CD-HRL, combine the PPO-SAC dual-strategy mechanism with the learning progress-aware reward function, and realize the dynamic optimization of generation parameters.

[0016] Preferably, the semantically enhanced course question-answering model SE-QA in S2 adopts hierarchical semantic parsing technology to perform four-level parsing of syntax, semantics, intention, and action association on user questions to generate multi-granularity semantic vectors; associates the teaching knowledge base with the retrieval strategy enhanced by the teaching knowledge graph, and generates structured text teaching answers by the PEGASUS model; defines the teaching knowledge base as where d i =(c i ,fi ), c i ,f i Represent concept text and associated formula respectively, I represents the total number of knowledge units in the teaching knowledge base, and is jointly encoded into a semantic vector through Sentence-BERT and the formula syntax tree:

[0017] E i =[SBERT(c i ); GAT(f i )];

[0018] Among them, SBERT is the Sentence-BERT model, GAT is the graph attention network, and the parsing formula syntax tree structure;

[0019] User question Q is parsed through four levels to generate a multi-granularity semantic vector S Q =s syn ;s sem ;s intent ;s act ], where: Syntax parsing based on dependency syntax tree s syn =BiLSTM(Q), extract the main components; combine the teaching knowledge graph KG for semantic analysis sem =GNN(Q,KG), performs entity disambiguation and relationship reasoning; uses attention matching mechanism based on intent tag library for intent recognition intent =Attention(Q,D intent ), D intent is the intent tag library, Attention represents attention matching; action-related entities s are extracted from user questions through the CRF model act = CRF(Q);

[0020] Combined with the semantic vector similarity cos(S Q ,E i ) and keyword statistics matching TF-IDF(Q,d i ), dynamically select the most relevant knowledge items from the teaching knowledge base:

[0021] M(Q,E i )=cos(S Q ,E i )+γ·TF-IDF(Q,d i );

[0022] Among them, γ is the retrieval enhancement coefficient, which is adaptively adjusted according to the complexity of the question;

[0023] The PEGASUS model generates structured text answers, including explanations of core concepts. ans And the derivation process of formula fans :

[0024]

[0025] Among them, A ans Represents the text answer, L represents the matching score M(Q,E i )The number of Top-K most relevant knowledge items selected.

[0026] Preferably, the S3 adopts cross-modal emotion adaptation speech synthesis technology and extracts the user input speech data X based on the Wav2Vec2 emotion recognition model. voice Emotional characteristics E audio , extract text answer A based on RoBERTa text sentiment analysis model ans Semantic sentiment features E sem :

[0027] E audio =Wav2Vec2(X voice ),E sem =RoBERTa(A ans );

[0028] Calculate the speech-text sentiment fusion weight matrix W through cross-modal attention emo :

[0029]

[0030] Where d is the dimension of the sentiment feature vector, and T is the matrix transpose operation;

[0031] Extract user voiceprint features V through the adversarial voiceprint encoder AVE id =AVE(X voice ), and combined with the sentiment weight matrix W emo , using the StyleTTS2 model to achieve timbre-emotion-speech rate three-element adaptive migration, and obtain the voiceprint style vector V style =StyleTTS2(V id ,W emo );

[0032] The design of the module generation architecture is as follows: the fundamental frequency prediction module PitchNet is based on the audio emotion feature E audio Generate pitch curve F0; rhythm control module ProsoNet combined with text answer A ans Parsing sentence rhythm and stress pattern P; TimbreNet timbre generation module integrates voiceprint style vector V style Output personalized tone parameters T timbre ;;Finally, the neural vocoder is used to synthesize highly natural audio. Answer Aaudio :

[0033]

[0034] Preferably, the S4 is based on a self-constructed structured database of teaching action videos, defines classification standards through multi-dimensional teaching scenario surveys, systematically collects teaching action videos according to the classification standards, and performs multi-level semantic annotations; through a teaching action scheduling system with cross-modal spatiotemporal feature alignment, the spatiotemporal features of the teaching action videos are extracted using the spatiotemporal graph convolutional network ST-GCN, combined with text answers and sentiment weights, and the matching priority is dynamically optimized through the deep Q network DQN, and finally, based on the mixed similarity, teaching action videos with triple matching of emotion, semantics, and spatiotemporal are scheduled from the database.

[0035] Optimally, the statistical feature matrix X∈R based on the multi-dimensional teaching scenario survey N×M , where N is the number of teaching video samples, M is the number of statistical feature dimensions of each sample; define the classification standard set C = {c1, c2, ..., c n}, where c i (i=1,2,…,n) corresponds to different teaching scene types, including knowledge point explanation, experimental demonstration and case analysis; video dataset is constructed according to classification standard C where v i For teaching action video examples, L i =[f type ,f kp ,f grain ] is the annotation vector generated by the multi-level semantic annotation function, and the annotation level includes the action type f type , knowledge point relevance kp and semantic granularity f grain ; Based on the semantic annotation vector L i , generating a searchable structured database Where R is the semantic similarity matrix, satisfying R ij =sim(L i ,L j ), sim(·) is the preset semantic similarity calculation function;

[0036] The spatiotemporal graph convolutional network ST-GCN is used to model the teaching action video. The nodes of the graph structure correspond to the human joints, and the edges represent the spatiotemporal motion relationship between the joints. The local limb motion and global action pattern are captured through layered convolution, and the spatiotemporal feature vector F is output. st =ST_GCN(v i ), ST_GCN represents the spatiotemporal graph convolutional network; combined with the text answer A in S2 ans and the sentiment weight matrix W in S3 emoConstruct joint driving vector A drive :

[0037] A drive =FC([Flatten(W emo );A ans ]);

[0038] Among them, A drive The semantics, emotional intensity, and action relevance of teaching content are integrated to achieve multimodal collaborative driving. FC(·) represents the fully connected layer, and Flatten(·) represents flattening the multidimensional tensor into a one-dimensional vector to accommodate the input of the fully connected layer.

[0039] Define the state space and action space Optimizing scheduling strategies through deep Q-network DQN:

[0040]

[0041] Among them, s is a state sample, which is an element in the state space, a is an action sample, which is an element in the action space, and the reward function r = λ1·Accuracy+λ2·Delay -1 , balancing matching accuracy and real-time performance; λ1 and λ2 are weights in the reward function, used to balance the two objectives in scheduling optimization, Accuracy represents matching accuracy, and Delay represents time delay;

[0042] The hybrid similarity is calculated by combining spatiotemporal feature similarity with semantic relevance, and the optimal video is dynamically selected:

[0043] M(v i )=η1·cos(F st ,A drive )+η2·MLP(s act );

[0044] Among them, η1, η2 are weights in the hybrid similarity calculation, which respectively measure the influence of spatiotemporal features and semantic relevance on the final similarity calculation. η1+η2 satisfies η1+η2=1, s act The action-related entity vector obtained by parsing the user question in S2;

[0045] Final scheduling results: Among them v * Represents a video sequence of teaching actions.

[0046] Preferably, the S5 is based on a layered temporal convolutional network TCN from the teaching action video v * Extract multi-scale spatiotemporal features and generate joint motion sequence A seq :

[0047] A seq =TCN Enc (v * )∈R T×J×3 ;

[0048] in v t * The input teaching action video frame sequence is J, the number of human skeleton joints, 3 represents the three-dimensional space coordinate, T is the number of video frames, and t is the time step index in the video frame sequence, with a value range of t = 1, 2, ..., T. Enc (·) represents a layered temporal convolutional network encoder;

[0049] Introducing joint motion smoothness constraint loss function Suppress action mutations and obtain smoothed action sequence A refined :

[0050]

[0051] Where λ is the path smoothing coefficient;

[0052] Through the inverse mapping network G drive The smoothed action sequence A refined Converted into character portrait driving parameters to generate preliminary action video V pre :

[0053] V pre =G drive (A refined ,I in );

[0054] Among them I in A full-body portrait image input by the user.

[0055] Preferably, the S6 is based on a cross-modal attention alignment mechanism, using the audio answer A audio Mel spectrum feature M A With action video V pre Lip movement characteristics L t , calculate the time sequence alignment weight matrix W align ∈R T×T :

[0056]

[0057] Where T is the time step, d is the feature dimension; Q(·) and K(·) are learnable linear transformation matrices, and the Softmax(·) function converts similarity into a probability distribution;

[0058] Optimize the alignment path through the dynamic time warping algorithm DTW:

[0059]

[0060] Among them, π is the set of alignment paths, π * is the optimal alignment path, is the i-th audio Mel spectrum frame, is the jth lip action frame, i,j∈{1,2,...,T};

[0061] Fusion of alignment weights and action videos to generate refined lip-sync instructional videos V S :

[0062] V S =Conv1D(W align ·V pre );

[0063] Among them, Conv1D(·) is a one-dimensional temporal convolution layer for smoothing action transitions.

[0064] Preferably, the S7 uses the SRNTT model based on neural texture migration to train the teaching video V S Combined face restoration and super-resolution reconstruction are performed, and bidirectional LSTM is used to enhance dynamic textures based on the dynamic characteristics of micro-expressions in teaching scenes: high-resolution reference texture library T is used to ref , align the face area in the video frame through the adaptive attention mechanism, repair the occluded or blurred part, and generate the repaired frame F restored :

[0065]

[0066] Among them, W face is the attention weight matrix of the face area in the video frame, For structural repair modules;

[0067] Fusion repair frame F restored With T ref Multi-scale texture features and repaired frame features are used to reconstruct high-resolution intermediate frames through scale-adaptive convolution:

[0068]

[0069] in, is the high-resolution intermediate frame feature map after fusion at scale s, is the repair frame feature map at scale s, is the high-resolution reference texture library feature map at scale s, is the feature fusion operation corresponding to scale s;

[0070] Aiming at the dynamic characteristics of micro-expressions in teaching scenes, bidirectional LSTM is used to extract video temporal motion features E motion And perform gated fusion with multi-scale texture features for dynamic texture enhancement:

[0071]

[0072] Where σ is the Sigmoid function, [·‖·] represents feature concatenation, G( s ) is the s-th scale gate weight map, which controls the fusion ratio of high-resolution texture and temporal motion features. The final enhanced feature map that fuses texture details and motion features at the s-th scale;

[0073] Final output video

[0074] Preferably, the S8 integrates visual quality, generates authenticity and teaching adaptability indicators through a multimodal quality evaluation feedback system to achieve a comprehensive quality evaluation with dynamic weight adjustment:

[0075] The video quality evaluation index is defined as a multi-dimensional vector:

[0076]

[0077] Among them, PSNR measures video clarity, FID evaluates generation authenticity, and LPIPS quantifies perceptual similarity. is the adaptability of the teaching scene, where To quantify the normativeness of actions, For expression accuracy;

[0078] Construct a dynamic feedback controller to obtain the quality score Q total :

[0079] Q total =AdaWeight(Q metric ,ω);

[0080] Where AdaWeight(·) represents the adaptive weight assignment, ω is the adaptive weight of each quality dimension;

[0081] The gated recurrent unit GRU network is used to adaptively adjust the weights to meet

[0082]

[0083] ω i ∝UserFeedback(V F );

[0084] Among them, ω iis the weighted coefficient of the i-th quality indicator, and UserFeedback(·) represents user feedback.

[0085] Preferably, the S9 adopts the curriculum-driven hybrid reinforcement learning framework CD-HRL, combined with the PPO-SAC dual-strategy mechanism and the learning progress-aware reward function to achieve dynamic optimization of generation parameters:

[0086] Define the reinforcement learning quadruple (S, A, R, P):

[0087] State Space: Contains quality ratings, video feature encoding, and user historical interaction data; represents the feature encoding function, Represents user historical interaction data;

[0088] Action space: A = {Δθ}∈R d Indicates the adjustment amount of the generated model parameters; θ is the current model parameter, R d is the parameter dimension space;

[0089] Reward function: R = αQ total +βPPO clip (π)+γSAC entropy (Q), where α, β, and γ are the decay coefficients driven by learning progress;

[0090] Policy network: A dual-branch architecture is used to model PPO and SAC strategies separately:

[0091]

[0092] Optimization process iterative update: k is the current training iteration round, B k is the experience replay buffer for k iterations, π hybrid It is a mixed strategy network;

[0093] Output optimized video: V out =G(θ k ,V F ) ↑4K , where K is a hyperparameter representing the output video resolution.

[0094] Therefore, the present invention adopts the above-mentioned method for generating a multimodal interactive digital human teaching assistant for teaching, and builds a test environment by integrating subject teaching resources and real classroom data. The system has shown significant performance improvement:

[0095] (1) The knowledge question answering model based on dynamic retrieval enhancement can accurately analyze complex questions and generate professional answers that conform to teaching logic;

[0096] (2) The audio output by the speech synthesis module is significantly better than traditional speech engines in terms of naturalness of intonation and emotional expression;

[0097] (3) The semantic-driven action scheduling system achieves a high degree of consistency between teaching content and body language, especially demonstrating accurate behavior mapping capabilities in concept explanation and case analysis scenarios.

[0098] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] Figure 1 This is a flow chart of an embodiment of a method for generating a multimodal interactive digital human teaching assistant for teaching according to the present invention;

[0100] Figure 2 It is a schematic diagram of an embodiment of a method for generating a multimodal interactive digital human teaching assistant for teaching according to the present invention. DETAILED DESCRIPTION

[0101] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0102] Unless otherwise defined, technical or scientific terms used in the present invention shall have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.

[0103] Example 1

[0104] like Figure 1 As shown, the present invention provides a method for generating a multimodal interactive digital human teaching assistant for teaching. Figure 1 This is a flow chart of a method for generating a multimodal interactive digital human teaching assistant for teaching, according to one embodiment of the present invention. This method aims to improve the interactive authenticity and teaching effectiveness of virtual teaching scenarios through semantic-driven and closed-loop optimization technology. It can be performed by following these steps:

[0105] S1, the user inputs voice data, question text and full-body portrait image;

[0106] S2. The semantically enhanced course question-answering model (SE-QA) uses hierarchical semantic parsing technology to perform four-level parsing of user questions, namely syntax, semantics, intention, and action association, to generate multi-granularity semantic vectors; it associates the teaching knowledge base with the retrieval strategy enhanced by the teaching knowledge graph, and generates structured text teaching answers using the PEGASUS model.

[0107] Define the teaching knowledge base as where d i =(c i ,f i ), c i,f i Represent concept text and associated formula respectively, I represents the total number of knowledge units in the teaching knowledge base, and is jointly encoded into a semantic vector through Sentence-BERT and the formula syntax tree:

[0108] E i =[SBERT(c i ); GAT(f i )];

[0109] Among them, SBERT is the Sentence-BERT model, GAT is the graph attention network, and the parsing formula syntax tree structure;

[0110] User question Q is parsed through four levels to generate a multi-granularity semantic vector S Q =s syn ;s sem ;s intent ;s act ], where: Syntax parsing based on dependency syntax tree s syn =BiLSTM(Q), extract the main components; combine the teaching knowledge graph KG for semantic analysis sem =GNN(Q,KG), performs entity disambiguation and relationship reasoning; uses attention matching mechanism based on intent tag library for intent recognition intent =Attention(Q,D intent ), D intent is the intent tag library, Attention represents attention matching; action-related entities s are extracted from user questions through the CRF model act = CRF(Q);

[0111] Combined with the semantic vector similarity cos(S Q ,E i ) and keyword statistics matching TF-IDF(Q,d i ), dynamically select the most relevant knowledge items from the teaching knowledge base:

[0112] M(Q,E i )=cos(S Q ,E i )+γ·TF-IDF(Q,d i );

[0113] Among them, γ is the retrieval enhancement coefficient, which is adaptively adjusted according to the complexity of the question;

[0114] The PEGASUS model generates structured text answers, including explanations of core concepts. ans And the derivation process of formula f ans :

[0115]

[0116] Among them, A ans Represents the text answer, L represents the matching score M(Q,E i )The number of Top-K most relevant knowledge items selected.

[0117] S3, using cross-modal emotion adaptation speech synthesis technology, extracting user input speech data based on the Wav2Vec2 emotion recognition model X voice Emotional characteristics E audio , extract text answer A based on RoBERTa text sentiment analysis model ans Semantic sentiment features E sem :

[0118] E audio =Wav2Vec2(X voice ),E sem =RoBERTa(A ans );

[0119] Calculate the speech-text sentiment fusion weight matrix W through cross-modal attention emo :

[0120]

[0121] Where d is the dimension of the sentiment feature vector, and T is the matrix transpose operation;

[0122] Extract user voiceprint features V through the adversarial voiceprint encoder AVE id =AVE(X voice ), and combined with the sentiment weight matrix W emo , using the StyleTTS2 model to achieve timbre-emotion-speech rate three-element adaptive migration, and obtain the voiceprint style vector V style =StyleTTS2(V id ,W emo );

[0123] The design of the module generation architecture is as follows: the fundamental frequency prediction module (PitchNet) is based on the audio emotion feature E audio Generate pitch curve F0; ProsoNet combined with text answer A ans Parsing sentence rhythm and stress pattern P; the timbre generation module (TimbreNet) integrates the voiceprint style vector V style Output personalized tone parameters T timbre Finally, the neural vocoder is used to synthesize high-naturalness audio. Answer A audio :

[0124]

[0125] S4. Based on a self-built structured database of teaching action videos (classification standards are defined through multi-dimensional teaching scenario research, teaching action videos are systematically collected according to the classification standards, and multi-level semantic annotation is performed), a teaching action scheduling system with cross-modal spatiotemporal feature alignment is constructed. The spatiotemporal features of teaching action videos are extracted using the spatiotemporal graph convolutional network ST-GCN, and the matching priority is dynamically optimized through the deep Q network DQN, combining text answers with sentiment weights. Finally, teaching action videos with triple matching of sentiment, semantics, and spatiotemporal matching are scheduled from the database based on hybrid similarity. Specifically:

[0126] Statistical feature matrix X∈R based on multi-dimensional teaching scenario survey N×M , where N is the number of teaching video samples, M is the number of statistical feature dimensions of each sample; define the classification standard set C = {c1, c2, ..., c n}, where c i (i=1,2,…,n) corresponds to different teaching scene types, including knowledge point explanation, experimental demonstration and case analysis; video dataset is constructed according to classification standard C where v i For teaching action video examples, L i =[f type ,f kp ,f grain ] is the annotation vector generated by the multi-level semantic annotation function, and the annotation level includes the action type f type , knowledge point relevance kp and semantic granularity f grain ; Based on the semantic annotation vector L i , generating a searchable structured database Where R is the semantic similarity matrix, satisfying R ij =sim(L i ,L j ), sim(·) is the preset semantic similarity calculation function;

[0127] The spatiotemporal graph convolutional network ST-GCN is used to model the teaching action video. The nodes of the graph structure correspond to the human joints, and the edges represent the spatiotemporal motion relationship between the joints. The spatiotemporal feature vector F is output by capturing the local limb motion and the global action pattern through layered convolution. st =ST_GCN(v i ), ST_GCN represents the spatiotemporal graph convolutional network; combined with the text answer A in S2 ans and the sentiment weight matrix W in S3 emo Construct joint driving vector A drive :

[0128] A drive =FC([Flatten(W emo );A ans ]);

[0129] Among them, A drive The semantics, emotional intensity, and action relevance of teaching content are integrated to achieve multimodal collaborative driving. FC(·) represents the fully connected layer, and Flatten(·) represents flattening the multidimensional tensor into a one-dimensional vector to accommodate the input of the fully connected layer.

[0130] Define the state space and action space Optimizing scheduling strategies through deep Q-network DQN:

[0131]

[0132] Among them, s is a state sample, which is an element in the state space, a is an action sample, which is an element in the action space, and the reward function r = λ1·Accuracy+λ2·Delay -1 , balancing matching accuracy and real-time performance; λ1 and λ2 are weights in the reward function, used to balance the two objectives in scheduling optimization, Accuracy represents matching accuracy, and Delay represents time delay;

[0133] The hybrid similarity is calculated by combining spatiotemporal feature similarity with semantic relevance, and the optimal video is dynamically selected:

[0134] M(v i )=η1·cos(F st ,A drive )+η2·MLP(s act );

[0135] Among them, η1, η2 are weights in the hybrid similarity calculation, which respectively measure the influence of spatiotemporal features and semantic relevance on the final similarity calculation. η1+η2 satisfies η1+η2=1, s act The action-related entity vector obtained by parsing the user question in S2;

[0136] Final scheduling results: where v * Represents a video sequence of teaching actions.

[0137] S5. Based on the hierarchical temporal convolutional network (TCN), multi-scale spatiotemporal features are extracted from the teaching action video to generate joint motion sequences. The joint motion smoothness constraint loss is introduced to suppress motion mutations, and the preliminary action video is generated through the inverse mapping network.

[0138] Based on the layered temporal convolutional network TCN from the teaching action video v * Extract multi-scale spatiotemporal features and generate joint motion sequence A seq :

[0139] A seq =TCN Enc (v *) ∈R T×J×3 ;

[0140] in v t * The input teaching action video frame sequence is J, the number of human skeleton joints, 3 represents the three-dimensional space coordinate, T is the number of video frames, and t is the time step index in the video frame sequence, with a value range of t = 1, 2, ..., T. Enc (·) represents a layered temporal convolutional network encoder;

[0141] Introducing joint motion smoothness constraint loss function Suppress action mutations and obtain smoothed action sequence A refined :

[0142]

[0143] Where λ is the path smoothing coefficient;

[0144] Through the inverse mapping network G drive The smoothed action sequence A refined Converted into character portrait driving parameters to generate preliminary action video V pre :

[0145] V pre =G drive (A refined ,I in );

[0146] Among them I in A full-body portrait image input by the user.

[0147] S6. Based on the cross-modal attention alignment mechanism, combined with the audio answer A audio Mel spectrum feature M A With action video V pre Lip movement characteristics L t , calculate the time sequence alignment weight matrix W align ∈R T×T :

[0148]

[0149] Where T is the time step, d is the feature dimension; Q(·) and K(·) are learnable linear transformation matrices, and the Softmax(·) function converts similarity into a probability distribution;

[0150] Optimize the alignment path through the dynamic time warping algorithm DTW:

[0151]

[0152] Among them, π is the set of alignment paths, π * is the optimal alignment path, is the i-th audio Mel spectrum frame, is the jth lip action frame, i,j∈{1,2,...,T};

[0153] Fusion of alignment weights and action videos to generate refined lip-sync instructional videos V S :

[0154] V S =Conv1D(W align ·V pre );

[0155] Among them, Conv1D(·) is a one-dimensional temporal convolution layer for smoothing action transitions.

[0156] S7, using the SRNTT model based on neural texture transfer to train the teaching video V S Combined face restoration and super-resolution reconstruction are performed, and bidirectional LSTM is used to enhance dynamic textures based on the dynamic characteristics of micro-expressions in teaching scenes: high-resolution reference texture library T is used to ref , align the face area in the video frame through the adaptive attention mechanism, repair the occluded or blurred part, and generate the repaired frame F restored :

[0157]

[0158] Among them, W face is the attention weight matrix of the face area in the video frame, For structural repair modules;

[0159] Fusion repair frame F restored With T ref Multi-scale texture features and repaired frame features are used to reconstruct high-resolution intermediate frames through scale-adaptive convolution:

[0160]

[0161] in, is the high-resolution intermediate frame feature map after fusion at scale s, is the repair frame feature map at scale s, is the high-resolution reference texture library feature map at scale s, is the feature fusion operation corresponding to scale s;

[0162] Aiming at the dynamic characteristics of micro-expressions in teaching scenes, bidirectional LSTM is used to extract video temporal motion features E motion And perform gated fusion with multi-scale texture features for dynamic texture enhancement:

[0163]

[0164] Where σ is the Sigmoid function, [·‖·] represents feature concatenation, G( s ) is the s-th scale gate weight map, which controls the fusion ratio of high-resolution texture and temporal motion features. The final enhanced feature map that fuses texture details and motion features at the s-th scale;

[0165] Final output video

[0166] S8. Through a multimodal quality evaluation feedback system, visual quality, generation authenticity, and teaching adaptability indicators are integrated to achieve a comprehensive quality evaluation with dynamic weight adjustment:

[0167] The video quality evaluation index is defined as a multi-dimensional vector:

[0168]

[0169] Among them, PSNR measures video clarity, FID evaluates generation authenticity, and LPIPS quantifies perceptual similarity. is the adaptability of the teaching scene, where To quantify the normativeness of actions, For expression accuracy;

[0170] Construct a dynamic feedback controller to obtain the quality score Q total :

[0171] Q total =AdaWeight(Q metric ,ω);

[0172] Where AdaWeight(·) represents the adaptive weight assignment, ω is the adaptive weight of each quality dimension;

[0173] The gated recurrent unit GRU network is used to adaptively adjust the weights to meet

[0174]

[0175] ω i ∝UserFeedback(V F );

[0176] Among them, ω i is the weighted coefficient of the i-th quality indicator, and UserFeedback(·) represents user feedback.

[0177] S9. We use the curriculum-driven hybrid reinforcement learning framework CD-HRL, combined with the PPO-SAC dual-strategy mechanism and the learning progress-aware reward function, to achieve dynamic optimization of generation parameters:

[0178] Define the reinforcement learning quadruple (S, A, R, P):

[0179] State Space: Contains quality ratings, video feature encoding, and user historical interaction data; represents the feature encoding function, Represents user historical interaction data;

[0180] Action space: A = {Δθ}∈R d Indicates the adjustment amount of the generated model parameters; θ is the current model parameter, R d is the parameter dimension space;

[0181] Reward function: R = αQ total +βPPO clip (π)+γSAC entropy (Q), where α, β, and γ are the decay coefficients driven by learning progress;

[0182] Policy network: A dual-branch architecture is used to model PPO and SAC strategies separately:

[0183]

[0184] Optimization process iterative update: k is the current training iteration round, B k is the experience replay buffer for k iterations, π hybrid It is a mixed strategy network;

[0185] Output optimized video: V out =G(θ k ,V F ) ↑4K , where K is a hyperparameter representing the output video resolution.

[0186] The above-mentioned method for generating a multimodal interactive digital human teaching assistant for teaching has been verified in multiple dimensions in diverse teaching scenarios. In terms of user experience, most subjects reported that the digital human teaching assistant's lip synchronization accuracy and facial micro-expression presentation were close to that of a real teacher, and the audio-visual synchronization error was below the perceptible threshold. Experimental evaluation confirmed that compared with existing educational digital human systems, this solution significantly improved video fluency and interaction continuity by integrating hierarchical semantic parsing, adaptive emotion generation, and closed-loop optimization mechanisms, maintaining stable expressiveness and logical coherence in long-term teaching demonstrations.

[0187] Therefore, the present invention adopts the above-mentioned multimodal interactive digital human teaching assistant generation method for teaching, breaking through the technical bottlenecks of traditional digital human systems such as semantic-action mismatch and single emotional expression, which can significantly improve the efficiency of knowledge transfer and interactive realism, and provide an innovative solution for intelligent education tools.

[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for generating a multimodal interactive digital human teaching assistant for teaching, characterized by: The following steps are involved: S1, the user inputs voice data, question text and full-body portrait image; S2. Generate structured text teaching answers through the SE-QA model and teaching knowledge graph; S3. Extract the emotional features of the user's input voice data, combine it with the text answer, and use emotion-adapted speech synthesis technology to generate personalized voice answers; S4. Build a teaching action video database, use ST-GCN to extract spatiotemporal features, and schedule matching teaching videos; S5. Use TCN to extract spatiotemporal features from the video, generate joint motion sequences to suppress sudden action changes, and generate preliminary action videos through the inverse mapping network. S6. Using the cross-modal attention alignment mechanism, the dynamic time warping algorithm DTW is used to optimize the lip synchronization path and generate a refined video. S7, face restoration and super-resolution reconstruction using the SRNTT model to enhance micro-expressions; S8. Through a multimodal quality evaluation feedback system, visual quality, generation authenticity and teaching suitability indicators are integrated to comprehensively evaluate video quality and teaching suitability; S9. Adopt the curriculum-driven hybrid reinforcement learning framework CD-HRL, combine the PPO-SAC dual-strategy mechanism with the learning progress-aware reward function, and realize the dynamic optimization of generation parameters.

2. The method for generating a multimodal interactive digital human teaching assistant for teaching according to claim 1, characterized in that: The semantically enhanced course question-answering model SE-QA in S2 adopts hierarchical semantic parsing technology to parse user questions at four levels: syntax, semantics, intention, and action association, and generate multi-granularity semantic vectors; the teaching knowledge base is associated with the retrieval strategy enhanced by the teaching knowledge graph, and the structured text teaching answers are generated by the PEGASUS model; the teaching knowledge base is defined as where d i =c i ,f i ), c i ,f i Represent concept text and associated formula respectively, I represents the total number of knowledge units in the teaching knowledge base, and is jointly encoded into a semantic vector through Sentence-BERT and the formula syntax tree: IN i =[SBERT(c i );GAT(f i )]; Among them, SBERT is the Sentence-BERT model, GAT is the graph attention network, and the parsing formula syntax tree structure; User question Q is parsed through four levels to generate a multi-granularity semantic vector S Q =s syn ;s sem ;s intent ;s act ], where: Syntax parsing based on dependency syntax tree s syn =BiLSTM(Q), extract the main components; combine the teaching knowledge graph KG for semantic analysis sem =GNN(Q,KG), performs entity disambiguation and relationship reasoning; uses attention matching mechanism based on intent tag library for intent recognition intent =Attention(Q,D intent ), D intent is the intent tag library, Attention represents attention matching; action-related entities s are extracted from user questions through the CRF model act = CRF(Q); Combined with the semantic vector similarity cos(S Q ,E i ) and keyword statistics matching TF-IDF(Q,d i ), dynamically select the most relevant knowledge items from the teaching knowledge base: M(Q,E i )=cos(S Q ,E i )+γ·TF-IDF(Q,d i ); Among them, γ is the retrieval enhancement coefficient, which is adaptively adjusted according to the complexity of the question; The PEGASUS model generates structured text answers, including explanations of core concepts. ans And the derivation process of formula f ans : Among them, A ans Represents the text answer, L represents the matching score M(Q,E i )The number of Top-K most relevant knowledge items selected.

3. The method for generating a multimodal interactive digital human teaching assistant for teaching according to claim 2, characterized in that: The S3 adopts cross-modal emotion adaptation speech synthesis technology and extracts user input speech data X based on the Wav2Vec2 emotion recognition model. voice Emotional characteristics E audio , extract text answer A based on RoBERTa text sentiment analysis model ans Semantic sentiment features E sem : AND audio =Wav2Vec2(X voice ),AND sem =RoBERTa(A ans ); Calculate the speech-text sentiment fusion weight matrix W through cross-modal attention emo : Where d is the dimension of the sentiment feature vector, and T is the matrix transpose operation; Extract user voiceprint features V through the adversarial voiceprint encoder AVE id =AVE(X voice ), and combined with the sentiment weight matrix W emo , using the StyleTTS2 model to achieve timbre-emotion-speech rate three-element adaptive migration, and obtain the voiceprint style vector V style =StyleTTS2(V id ,W emo ); The design of the module generation architecture is as follows: the fundamental frequency prediction module PitchNet is based on the audio emotion feature E audio Generate pitch curve F0; rhythm control module ProsoNet combined with text answer A ans Parsing sentence rhythm and stress pattern P; TimbreNet timbre generation module integrates voiceprint style vector V style Output personalized tone parameters T timbre Finally, the neural vocoder is used to synthesize high-naturalness audio. Answer A audio :

4. The method for generating a multimodal interactive digital human teaching assistant for teaching according to claim 3, characterized in that: The S4 is based on a self-constructed structured database of teaching action videos. It defines classification standards through multi-dimensional teaching scenario surveys, systematically collects teaching action videos according to the classification standards, and performs multi-level semantic annotations. Through a teaching action scheduling system with cross-modal spatiotemporal feature alignment, the spatiotemporal features of the teaching action videos are extracted using the spatiotemporal graph convolutional network ST-GCN, and the matching priority is dynamically optimized through the deep Q network DQN, combining text answers with sentiment weights. Finally, teaching action videos with triple matching of sentiment, semantics, and spatiotemporal matching are scheduled from the database based on mixed similarity.

5. The method for generating a multimodal interactive digital human teaching assistant for teaching according to claim 4, characterized in that: Statistical feature matrix X∈R based on multi-dimensional teaching scenario survey N×M , where N is the number of teaching video samples, M is the number of statistical feature dimensions of each sample; define the classification standard set C = {c1, c2, ..., c n }, where c i (i=1,2,…,n) corresponds to different teaching scene types, including knowledge point explanation, experimental demonstration and case analysis; video dataset is constructed according to classification standard C where v i For teaching action video examples, L i =[f type ,f kp ,f grain ] is the annotation vector generated by the multi-level semantic annotation function, and the annotation level includes the action type f type , knowledge point relevance kp and semantic granularity f grain ; Based on the semantic annotation vector L i , generating a searchable structured database Where R is the semantic similarity matrix, satisfying R ij =sim(L i ,L j ), sim(·) is the preset semantic similarity calculation function; The spatiotemporal graph convolutional network ST-GCN is used to model the teaching action video. The nodes of the graph structure correspond to the human joints, and the edges represent the spatiotemporal motion relationship between the joints. The local limb motion and global action pattern are captured through layered convolution, and the spatiotemporal feature vector F is output. st =ST_GCN(v i ), ST_GCN represents the spatiotemporal graph convolutional network; combined with the text answer A in S2 ans and the sentiment weight matrix W in S3 emo Construct joint driving vector A drive : A drive =FC([Flatten(W emo );A ans ]); Among them, A drive The semantics, emotional intensity, and action relevance of teaching content are integrated to achieve multimodal collaborative driving. FC(·) represents the fully connected layer, and Flatten(·) represents flattening the multidimensional tensor into a one-dimensional vector to accommodate the input of the fully connected layer. Define the state space and action space Optimizing scheduling strategies through deep Q-network DQN: Among them, s is a state sample, which is an element in the state space, a is an action sample, which is an element in the action space, and the reward function r = λ1·Accuracy+λ2·Delay -1 , balancing matching accuracy and real-time performance; λ1 and λ2 are weights in the reward function, used to balance the two objectives in scheduling optimization, Accuracy represents matching accuracy, and Delay represents time delay; The hybrid similarity is calculated by combining spatiotemporal feature similarity with semantic relevance, and the optimal video is dynamically selected: M(v i )=η1·cos(F st ,A drive )+η2·MLP(s act ); Among them, η1, η2 are weights in the hybrid similarity calculation, which respectively measure the influence of spatiotemporal features and semantic relevance on the final similarity calculation. η1+η2 satisfies η1+η2=1, s act The action-related entity vector obtained by parsing the user question in S2; Final scheduling results: where v * Represents a video sequence of teaching actions.

6. The method for generating a multimodal interactive digital human teaching assistant for teaching according to claim 5, characterized in that: The S5 is based on the layered temporal convolutional network TCN from the teaching action video v * Extract multi-scale spatiotemporal features and generate joint motion sequence A seq : A seq =TCN Enc (v *) ∈R T×J×3 ; in v t * The input teaching action video frame sequence is J, the number of human skeleton joints, 3 represents the three-dimensional space coordinate, T is the number of video frames, and t is the time step index in the video frame sequence, with a value range of t = 1, 2, ..., T. Enc (·) represents a layered temporal convolutional network encoder; Introducing joint motion smoothness constraint loss function Suppress action mutations and obtain smoothed action sequence A refined : Where λ is the path smoothing coefficient; Through the inverse mapping network G drive The smoothed action sequence A refined Converted into character portrait driving parameters to generate preliminary action video V pre : V pre =G drive (A refined ,I in ); Among them I in A full-body portrait image input by the user.

7. The method for generating a multimodal interactive digital human teaching assistant for teaching according to claim 6, characterized in that: The S6 is based on the cross-modal attention alignment mechanism and uses the audio answer A audio Mel spectrum feature M A With action video V pre Lip movement characteristics L t , calculate the time sequence alignment weight matrix W align ∈R T×T : Where T is the time step, d is the feature dimension; Q(·) and K(·) are learnable linear transformation matrices, and the Softmax(·) function converts similarity into a probability distribution; Optimize the alignment path through the dynamic time warping algorithm DTW: Among them, π is the set of alignment paths, π * is the optimal alignment path, is the i-th audio Mel spectrum frame, is the jth lip action frame, i,j∈{1,2,...,T}; Fusion of alignment weights and action videos to generate refined lip-sync instructional videos V S : V S =Conv1D(W align ·V pre ); Among them, Conv1D(·) is a one-dimensional temporal convolution layer for smoothing action transitions.

8. The method for generating a multimodal interactive digital human teaching assistant for teaching according to claim 7, characterized in that: The S7 uses the SRNTT model based on neural texture transfer to train the teaching video V S Combined face restoration and super-resolution reconstruction are performed, and bidirectional LSTM is used to enhance dynamic textures based on the dynamic characteristics of micro-expressions in teaching scenes: high-resolution reference texture library T is used to ref , align the face area in the video frame through the adaptive attention mechanism, repair the occluded or blurred part, and generate the repaired frame F restored : Among them, W face is the attention weight matrix of the face area in the video frame, For structural repair modules; Fusion repair frame F restored With T ref Multi-scale texture features and repaired frame features are used to reconstruct high-resolution intermediate frames through scale-adaptive convolution: in, is the high-resolution intermediate frame feature map after fusion at scale s, is the repair frame feature map at scale s, is the high-resolution reference texture library feature map at scale s, is the feature fusion operation corresponding to scale s; Aiming at the dynamic characteristics of micro-expressions in teaching scenes, bidirectional LSTM is used to extract video temporal motion features E motion And perform gated fusion with multi-scale texture features for dynamic texture enhancement: Where σ is the Sigmoid function, [·‖·] represents feature concatenation, G( s ) is the s-th scale gate weight map, which controls the fusion ratio of high-resolution texture and temporal motion features. The final enhanced feature map that fuses texture details and motion features at the s-th scale; Final output video 9. The method for generating a multimodal interactive digital human teaching assistant for teaching according to claim 8, characterized in that: The S8 uses a multimodal quality evaluation feedback system to integrate visual quality, generation authenticity, and teaching adaptability indicators to achieve a comprehensive quality evaluation with dynamic weight adjustment: The video quality evaluation index is defined as a multi-dimensional vector: Among them, PSNR measures video clarity, FID evaluates generation authenticity, and LPIPS quantifies perceptual similarity. is the adaptability of the teaching scene, where To quantify the normativeness of actions, For expression accuracy; Construct a dynamic feedback controller to obtain the quality score Q total : Q total =AdaWeight(Q metric ,ω); Where AdaWeight(·) represents the adaptive weight assignment, ω is the adaptive weight of each quality dimension; The gated recurrent unit (GRU) network is used to adaptively adjust the weights to meet the following requirements: oh i ∝UserFeedback(V F ); Among them, ω i is the weighted coefficient of the i-th quality indicator, and UserFeedback(·) represents user feedback.

10. The method for generating a multimodal interactive digital human teaching assistant for teaching according to claim 9, characterized in that: The S9 adopts the curriculum-driven hybrid reinforcement learning framework CD-HRL, combined with the PPO-SAC dual-strategy mechanism and the learning progress-aware reward function to achieve dynamic optimization of generation parameters: Define the reinforcement learning quadruple (S, A, R, P): State Space: Contains quality ratings, video feature encoding, and user historical interaction data; represents the feature encoding function, Represents user historical interaction data; Action space: A = {Δθ}∈R d Indicates the adjustment amount of the generated model parameters; θ is the current model parameter, R d is the parameter dimension space; Reward function: R = αQ total +βPPO clip (π)+γSAC entropy (Q), where α, β, and γ are the decay coefficients driven by learning progress; Policy network: A dual-branch architecture is used to model PPO and SAC strategies separately: Optimization process iterative update: k is the current training iteration round, B k is the experience replay buffer for k iterations, π hybrid It is a mixed strategy network; Output optimized video: V out =G(θ k ,V F ) ↑4K , where K is a hyperparameter representing the output video resolution.

Citation Information

Cited By

  • Teaching plan content generation method based on generative adversarial network and hierarchical teaching target analysis

    CN121071193A