MOOC video method and device based on AIGC digital human
Through pre-training digital human model library and transfer learning technology, combined with VR/AR and tactile feedback, the problems of high labor costs and low automation in MOOC video production are solved, and efficient and personalized MOOC video generation is achieved, which improves teaching immersion and practical effect.
Patent Information
- Application Number
- CN202510428056.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-25
AI Technical Summary
The existing MOOC video production relies on manual recording and post-editing, and there are problems such as long production cycle, high labor cost, high model training complexity, insufficient synchrony between action and voice, limited personalization and interaction capabilities, and low end-to-end automation, which is difficult to meet the needs of large-scale course production.
Pre-trained digital human basic model library, transfer learning, timing model and reinforcement learning algorithm are adopted, combined with virtual engines and multimodal AI models, personalized MOOC videos are generated, and a collaborative teaching model is constructed through VR/AR and tactile feedback to achieve high synchronization of actions and speech and multimodal interaction.
It significantly reduces the deployment complexity and labor costs, and shortens the adaptation time from several weeks to several hours, improves the immersion and practical effect, and realizes "teaching according to aptitude", which is suitable for large-scale course production.
Smart Images

Figure CN120378707A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence-generated content (AIGC), and more particularly, to a method and device for MOOC videos based on AIGC digital humans. Background Art
[0002] With the rapid development of the online education industry, Massive Open Online Course (MOOC) videos have become an important part of the modern education system. The production of traditional MOOC videos highly relies on manual recording and post-editing by teachers, which has significant problems such as long production cycles and high labor costs. Especially in large-scale online education platforms, the frequent update of course content and personalized needs further magnify the efficiency bottleneck. In recent years, the rise of artificial intelligence-generated content (AIGC) technology has provided a new direction for video automation production. Among them, digital human technology gradually applies to the field of teaching video generation by simulating real human images and behaviors. However, the existing technologies still have the following key defects:
[0003] (1) High complexity of model training: Existing solutions (such as Chinese Patent Application CN115515002A) rely on 3D digital human modeling and massive data training, resulting in long deployment cycles and poor flexibility. For example, 3D models need to collect multi-angle human data and undergo long-term training, making it difficult to quickly adapt to diverse teaching scenarios;
[0004] (2) Insufficient synchronization of actions and voices: The matching accuracy of the limb movements, expressions, and voice rhythms of digital humans is low, and there are often lagging actions or mechanical repetitions, seriously affecting teaching expressiveness and the learning experience of learners;
[0005] (3) Limited personalization and interaction capabilities: Existing methods mostly adopt static generation modes, unable to dynamically adjust teaching content according to audience characteristics (such as knowledge level, learning preferences), and the interaction forms are limited to one-way output, lacking multi-modal (such as tactile feedback, VR / AR) collaboration;
[0006] (4) Low end-to-end automation level: Manual intervention is still required in links such as material matching and video synthesis, making it difficult to meet the efficiency requirements of large-scale course production. Summary of the Invention
[0007] The purpose of the present invention is to provide a method and device for MOOC videos based on AIGC digital humans to improve the immersion and practical effect, achieve "teaching students in accordance with their aptitude", reduce labor and time costs, and be applicable to large-scale course production, aiming at the deficiencies in the above-mentioned existing technologies.
[0008] To achieve the above purpose, the technical solutions adopted in the embodiments of the present application are as follows:
[0009] First aspect, an embodiment of the present application provides a method for generating a massive open online course (MOOC) video based on an AI-generated digital human, including: determining a target digital human base model that matches the target course theme from a pre-trained digital human base model library; the digital human models in the pre-trained digital human model library have functions of synthesizing actions, expressions, and voices; inputting feature data of a part of the target course into the target digital human base model, and adjusting the model parameters of the target digital human base model through transfer learning to obtain a target digital human model; inputting the target speech text into a temporal model, and generating a temporally aligned target action sequence based on the attention mechanism; the temporal model is trained by speech-action alignment samples; inputting the target action sequence into the target digital human model to drive the target digital human model to execute corresponding target actions according to the target action sequence to obtain a target digital human teaching video; constructing a target virtual teaching scene using a virtual engine, and embedding the target digital human model into the target virtual teaching scene to obtain a target teaching digital human model; using a multi-modal AI model to determine a video background image, decorations of the target digital human model, and speech parameters according to the target courseware; generating a MOOC video based on a preset template, video parameters, the target digital human teaching video, the target teaching digital human model, the video background image, the decorations of the target digital human model, and the speech parameters.
[0010] In one implementation, after driving the target digital human model to execute corresponding target actions according to the target action sequence, the method further includes: optimizing the matching degree between the target action and the target voice corresponding to the target speech text by using a reinforcement learning algorithm.
[0011] In one implementation, the reward function in the reinforcement learning algorithm is as follows:
[0012] R total =R sync +α·R diversity +β·R emotion
[0013] where R total is the total reward; R sync is the synchronization error reward, which is used for the time alignment error between the action and the voice; R diversity is the action diversity reward, which is used for calculating the entropy value of the action sequence; R emotion is the emotion consistency reward; α and β are hyperparameters used to control the weights of diversity and emotion consistency;
[0014] The policy output by the reinforcement learning algorithm is as follows:
[0015]
[0016] Among them, γ is the discount factor, which is used to balance the immediate reward and the long-term cumulative reward; R(s t ,a t ) is the immediate reward; a is the action; s is the learner's behavior data; E s,a is the expected value of the long-term cumulative reward under the policy π.
[0017] In one implementation, after obtaining the target digital human teaching video, the method further includes: obtaining the learner's behavior data from the learning platform; using a clustering algorithm to construct a learner profile according to the learner's behavior data; and adjusting the target digital human teaching video according to the learner profile.
[0018] In one implementation, the adjusting the target digital human teaching video according to the learner profile includes: adjusting the depth of knowledge points according to the learner profile; analyzing the learner's classroom performance through NLP technology, and using the CLIP model based on the knowledge graph to generate supplementary explanation segments according to the depth of knowledge points; and adding the supplementary explanation segments to the target digital human teaching video.
[0019] In one implementation, the CLIP model based on the knowledge graph is as follows:
[0020] sim enhanced (I,T) = sim(I,T) + λ·KG-Align(I,T)
[0021] Among them, sim enhanced (I,T) represents the similarity; sim(I,T) represents the CLIP similarity; KG-Align(I,T) represents the semantic alignment score of the CLIP model based on the knowledge graph; λ is a hyperparameter.
[0022] In one implementation, the loss function of the transfer learning is as follows:
[0023] L total = L CE + λ·L MMD + μ·L MT
[0024] Among them, L total is the total loss; L CE is the cross-entropy loss; L MMD is the maximum mean discrepancy loss, which is used to align the distribution difference between the pre-training data and the teaching scenario data, and L MT is the multi-task loss, which is used to jointly perform action and speech synthesis; λ and μ are hyperparameters, which are used to control the relative importance of adaptive and multi-task learning.
[0025] In one implementation, after embedding the target digital human model into the target virtual teaching scenario, the method further includes: combining with a wearable device to convert the target action of the target digital human model into physical feedback.
[0026] In one implementation, the attention mechanism is as follows:
[0027]
[0028] Where A is the speech feature; B is the action feature; Q i is the query vector of the speech feature; is the transposed matrix of the key vector of the action feature; V j is the value vector of the action feature.
[0029] In a second aspect, an AIGC digital human-based MOOC video generation device provided by an embodiment of the present application includes: a first determination module configured to determine a target digital human base model matching a target course theme from a pre-trained digital human base model library; the digital human models in the pre-trained digital human model library have the functions of synthesizing actions, expressions, and voices; an adjustment module configured to input feature data of a part of the target course into the target digital human base model, and adjust the model parameters of the target digital human base model through transfer learning to obtain a target digital human model; a timing alignment module configured to input a target speech text into a timing model, and generate a target action sequence with timing alignment based on the attention mechanism; the timing model is trained with speech-action alignment samples; a teaching video generation module configured to input the target action sequence into the target digital human model, drive the target digital human model to execute corresponding target actions according to the target action sequence, and obtain a target digital human teaching video; a teaching digital human generation module configured to use a virtual engine to construct a target virtual teaching scenario, and embed the target digital human model into the target virtual teaching scenario to obtain a target teaching digital human model; a second determination module configured to use a multimodal AI model to determine a video background image, decorations of the target digital human model, and voice parameters according to a target courseware; a MOOC video generation module configured to generate a MOOC video based on a preset template, video parameters, the target digital human teaching video, the target teaching digital human model, the video background image, the decorations of the target digital human model, and the voice parameters.
[0030] In one implementation, the timing alignment module is configured to: optimize the matching degree between the target action and the target voice corresponding to the target speech text by using a reinforcement learning algorithm.
[0031] In one implementation, the reward function in the reinforcement learning algorithm is as follows:
[0032] R total = R sync + α·R diversity + β·R emotion
[0033] Wherein, R total is the total reward; R sync is the synchronization error reward, which is used for the time alignment error between actions and voices; R diversity is the action diversity reward, which is used to calculate the entropy value of the action sequence; R emotion is the emotion consistency reward; α and β are hyperparameters used to control the weights of diversity and emotion consistency;
[0034] The policy output by the reinforcement learning algorithm is as follows:
[0035]
[0036] Wherein, γ is the discount factor, which is used to balance the immediate reward and the long-term cumulative reward; R(s t , a t ) is the immediate reward; a is the action; s is the learner's behavior data; E s,a is the expected value of the long-term cumulative reward under the policy π.
[0037] In one implementation, the teaching video generation module is configured to: obtain the learner's behavior data from the learning platform; use a clustering algorithm to construct a learner profile according to the learner's behavior data; and adjust the target digital human teaching video according to the learner profile.
[0038] In one implementation, the teaching video generation module is configured to: adjust the depth of knowledge points according to the learner profile; analyze the learner's classroom performance through NLP technology, and use a CLIP model based on a knowledge graph to generate supplementary explanation segments according to the depth of knowledge points; and add the supplementary explanation segments to the target digital human teaching video.
[0039] In one implementation, the CLIP model based on the knowledge graph is as follows:
[0040] sim enharnced (I, T) = sim(I, T) + λ·KG-Align(I, T)
[0041] Wherein, sim enhanced (I, T) represents the similarity; sim(I, T) represents the CLIP similarity; KG-Align(I, T) represents the semantic alignment score of the CLIP model based on the knowledge graph; λ is a hyperparameter.
[0042] In one embodiment, the loss function of the transfer learning is as follows:
[0043] L total = L CE + λ · L MMD + μ · L MT
[0044] where L total is the total loss; L CE is the cross-entropy loss; L MMD is the maximum mean discrepancy loss, which is used to align the distribution differences between the pre-trained data and the teaching scenario data, and L MT is the multi-task loss, which is used for joint action and speech synthesis; λ and μ are hyperparameters used to control the relative importance of adaptive and multi-task learning.
[0045] In one embodiment, the teaching digital human generation module is configured to: combine with a wearable device to convert the target action of the target digital human model into physical feedback.
[0046] In one embodiment, the attention mechanism is as follows:
[0047]
[0048] where A is the speech feature; B is the action feature; Q i is the query vector of the speech feature; is the transposed matrix of the key vector of the action feature; V j is the value vector of the action feature.
[0049] In a third aspect, an embodiment of the present application provides a computer device, including: a processor, a storage medium, and a bus. The storage medium stores program instructions executable by the processor. When the computer device runs, the processor communicates with the storage medium through the bus, and the processor executes the program instructions to perform the steps of any one of the above methods.
[0050] In a fourth aspect, an embodiment of the present application provides a non-volatile computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it performs the steps of any one of the above methods.
[0051] The beneficial effects of the present application are as follows:
[0052] (1) Through the pre-trained model library and transfer learning technology, the present application significantly reduces the deployment complexity, and the adaptation time is shortened from several weeks to several hours; one-click completion from material matching to video export greatly reduces the labor and time costs, and is applicable to large-scale course production;
[0053] (2) By integrating VR / AR and haptic feedback, a "visual-auditory-tactile" collaborative teaching mode is constructed to enhance the immersion and practical operation effect;
[0054] (3) Optimize the teaching content based on real-time data feedback, break through the limitations of static generation, and achieve "teaching students in accordance with their aptitude". Description of the Drawings
[0055] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0056] Figure 1 A schematic flowchart of a method for MOOC videos based on AIGC digital humans provided by an embodiment of the present application;
[0057] Figure 2 A schematic flowchart of a method for MOOC videos based on AIGC digital humans provided by an embodiment of the present application;
[0058] Figure 3 A schematic flowchart of a method for MOOC videos based on AIGC digital humans provided by an embodiment of the present application;
[0059] Figure 4 A schematic structural diagram of a device for MOOC videos based on AIGC digital humans provided by an embodiment of the present application;
[0060] Figure 5 A schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed Embodiments
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0062] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0063] In the description of the present application, it should be noted that if terms such as "upper", "lower", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the products of this application are usually placed during use, it is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation to the present application.
[0064] In addition, terms such as "first", "second", etc. in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0065] It should be noted that, without conflict, the features in the embodiments of the present application can be combined with each other.
[0066] Figure 1 It is a schematic flowchart of a MOOC video method based on an AIGC digital human provided for the embodiments of the present application; as Figure 1 shown, the method includes the following steps 110 to step 170:
[0067] Step 110, determine a target digital human base model that matches the target course theme from the pre-trained digital human base model library.
[0068] Among them, the digital human models in the pre-trained digital human model library have the functions of synthesizing actions, expressions and voices.
[0069] A pre-trained model is a model that has been pre-trained on a large-scale dataset. These models usually have learned the general features or knowledge of the data. The main idea of the pre-trained model is to use a large amount of data to learn some common features or patterns, and these features are transferable between different tasks.
[0070] The digital human models in the digital human model library can be obtained through MetaHuman Animator and the Dream AI platform. Among them, MetaHuman Animator is a digital human animation tool developed by Epic Games, which supports driving facial expressions and limb movements through audio and can be directly integrated into technical solutions. The Dream AI platform provides a pre-trained 2D digital human action generation model that supports API calls for quickly generating basic action sequences.
[0071] In actual operation, the target digital human basic model that matches the target course theme can be determined; for example, if the theme is a scientific experiment, the image of a laboratory white coat is determined as the target digital human basic model.
[0072] Step 120: Input the feature data of some target courses into the target digital human basic model, and adjust the model parameters of the target digital human basic model through transfer learning to obtain the target digital human model.
[0073] Among them, the feature data of some target courses can be understood as the specific data of the target course, such as: experimental operation video clips.
[0074] Transfer Learning is a machine learning method aimed at applying the knowledge learned from one task to another related task to accelerate the learning process of the new task. This method improves the effect of learning the new task by transferring knowledge from the learned related tasks. Transfer learning does not require training the model from scratch, but uses the existing knowledge to accelerate the learning of the new task.
[0075] For example, the loss function of transfer learning is shown in the following formula (1):
[0076] L total =L CE +λ·L MMD +μ·L MT (1)
[0077] Among them, L total is the total loss; L CE is the cross-entropy loss; L MMD is the maximum mean discrepancy loss, which is used to align the distribution differences between the pre-trained data and the teaching scenario data. L MT is the multi-task loss, which is used for joint action and speech synthesis; λ and μ are hyperparameters used to control the relative importance of adaptive and multi-task learning.
[0078] In this step, through transfer learning, the model parameters of the general model (i.e., the above-mentioned target digital human basic model) can be quickly adjusted, and then the target digital human model that can be adapted to the specific teaching scenario can be obtained.
[0079] Step 130: Input the target speech text into the temporal model, and generate a temporally aligned target action sequence based on the attention mechanism.
[0080] Among them, the temporal model is trained with speech-action alignment samples.
[0081] The target speech text is the teaching script generated according to the MOOC courseware, and the speech corresponding to this teaching script (i.e., the target speech corresponding to the target speech text) is the teaching audio.
[0082] The temporal model is used to analyze the time series features of the speech text (such as phonemes, intonation, rhythm), and synchronously match them with the action sequence of the digital human (gestures, expressions, body postures). Further, the temporal model can adopt the Transformer model as the basic architecture. The Transformer model is widely used in speech recognition, natural language processing (NLP), and action generation tasks due to its parallel computing ability and effective capture of long sequence dependencies. Its attention mechanism can accurately associate the key time nodes in the speech text with the action parameters.
[0083] Exemplarily, the attention mechanism is shown in the following formula (2):
[0084]
[0085] Among them, A is the speech feature; B is the action feature; Q i is the query vector of the speech feature; is the transposed matrix of the key vector of the action feature; V j is the value vector of the action feature.
[0086] Step 140: Input the target action sequence into the target digital human model, drive the target digital human model to execute the corresponding target actions according to the target action sequence, and obtain the target digital human teaching video.
[0087] Among them, the temporal model and the digital human basic model are two core modules. Through division of labor, cooperation, and data interaction, they jointly achieve the efficient generation of digital human MOOC videos.
[0088] The pre-trained digital human basic model can generate a high-fidelity digital human image and actions based on the input parameters (action sequence, speech text); in addition, the pre-trained digital human basic model has an action library (such as standing, walking, gesture library) and speech synthesis function, which can ensure the naturalness and diversity of digital human actions.
[0089] In actual operation, after driving the target digital human model to execute the corresponding target actions according to the target action sequence, the target actions of the target digital human model can be further optimized; specifically, the MOOC video method based on AIGC digital humans provided in the embodiments of the present application can further include the following steps:
[0090] Adopt a reinforcement learning algorithm to optimize the matching degree between the target action and the target voice corresponding to the target voice text.
[0091] Among them, the reinforcement learning algorithm can adopt the PPO algorithm, the DQN algorithm, and there is no limitation here.
[0092] Exemplarily, adopt a reinforcement learning algorithm to optimize the matching degree between the target action and the target voice corresponding to the target voice text. For example, when the voice emphasizes keywords, the reinforcement learning model will drive the digital human to make more obvious gesture actions to enhance the expressiveness.
[0093] From step 110 to step 140, by combining the pre-trained digital human model with transfer learning and dynamic parameter adjustment, the complex process of training a 3D model from scratch in the traditional method can be avoided, and the training time can be reduced by more than 90%; the synchronization of actions and voices can be improved, and the error rate can be reduced from 15% in the traditional method to within 5%.
[0094] Exemplarily, the reward function in the reinforcement learning algorithm is shown in the following formula (3):
[0095] R total =R sync +α·R diversity +β·R emotion (3)
[0096] Among them, R total is the total reward; R sync is the synchronization error reward, which is used based on the time alignment error between the action and the voice; R diversity is the action diversity reward, which is used to calculate the entropy value of the action sequence; R emotion is the emotion consistency reward; α and β are hyperparameters, which are used to control the weights of diversity and emotion consistency.
[0097] Exemplarily, the policy output by the reinforcement learning algorithm is shown in the following formula (4):
[0098]
[0099] Among them, γ is the discount factor, which is used to balance the immediate reward and the long-term cumulative reward; R(s t ,a t ) is the immediate reward; a is the action; s is the behavior data of the learner; E s,ais the expected value of the long-term cumulative reward under policy π.
[0100] In actual operation, after obtaining the target digital human teaching video, personalized teaching videos can also be output for different learners; specifically, as Figure 2 shown, the MOOC video method based on AIGC digital humans provided by the embodiments of the present application can further include the following steps 210 to 230:
[0101] Step 210: Obtain the behavior data of the learner from the learning platform.
[0102] Among them, the behavior data of the learner includes course completion rate, question answering accuracy rate, interaction frequency, stay duration, repeated viewing segments, etc.
[0103] Step 220: Use a clustering algorithm to construct a learner profile according to the behavior data of the learner.
[0104] Among them, the clustering algorithm can be the K-means algorithm. Through the clustering algorithm, learners can be classified, for example: divided into efficient learners, learners who need to be strengthened, etc.
[0105] Step 230: Adjust the target digital human teaching video according to the learner profile.
[0106] Among them, corresponding teaching video adjustments can be made for different learner types; specifically, as Figure 3 shown, the above step 230 can further include the following steps 310 to 330:
[0107] Step 310: Adjust the depth of knowledge points according to the learner profile.
[0108] Exemplarily, for learners with weak foundations, the detail of the knowledge point explanations is automatically increased.
[0109] Step 320: Analyze the learner's classroom performance through NLP technology, and use the CLIP model based on the knowledge graph to generate supplementary explanation segments according to the depth of knowledge points.
[0110] Among them, Natural Language Processing (NLP) is an important branch in the field of artificial intelligence, aiming to enable computers to understand, process, and generate human languages to achieve natural human-computer interaction.
[0111] Exemplarily, the CLIP model based on the knowledge graph is shown in the following formula (5):
[0112] sim enhanced (I,T) = sim(I,T) + λ·KG - Align(I,T) (5)
[0113] Among them, sim enhanced (I, T) represents the similarity; sim(I, T) represents the CLIP similarity; KG - Align(I, T) represents the semantic alignment score of the CLIP model based on the knowledge graph; λ is a hyperparameter.
[0114] Step 330: Add the supplementary explanation segment to the target digital human teaching video.
[0115] Among them, in this step, the supplementary explanation segment can be automatically added to the appropriate position of the video for explaining the corresponding knowledge point in the target digital human teaching video.
[0116] Steps 310 to 330 solve the "one - size - fits - all" teaching mode of traditional MOOCs and can personalized match the needs of students; through real - time feedback optimization, the response time for updating course content is shortened to within 10 minutes.
[0117] Step 150: Use a virtual engine to construct a target virtual teaching scenario, and embed the target digital human model into the target virtual teaching scenario to obtain a target teaching digital human model.
[0118] Among them, the virtual engine can be Unity, Unreal Engine, etc., and there is no limitation here.
[0119] Exemplarily, embedding the target digital human model into the target virtual teaching scenario, for example: in a chemistry experiment course, students can observe the experimental steps demonstrated by the digital human through a VR headset and interact with it (such as selecting a reagent bottle).
[0120] In actual operation, for teaching videos that require interaction, haptic feedback can also be added; specifically, after embedding the target digital human model into the target virtual teaching scenario, the MOOC video method based on the AIGC digital human provided by the embodiments of the present application can further include the following steps:
[0121] In combination with a wearable device, convert the target actions of the target digital human model into physical feedback.
[0122] Exemplarily, when the digital human demonstrates "stirring the solution", the haptic device simulates the stirring resistance to enhance the sense of operation reality.
[0123] Step 150 can break the limitation of the one - way output of traditional MOOCs and achieve two - way interaction between students and digital humans; through haptic feedback, the teaching effect of complex operation - type courses is improved, and the error operation rate is reduced by 30%.
[0124] Step 160: Use a multimodal AI model to determine the video background image, the decoration of the target digital human model, and the voice parameters according to the target courseware.
[0125] Among them, in this step, a multi-modal AI model (such as CLIP) is used to analyze the courseware content, and automatically match the background image, the decoration of the digital human model, and the voice parameters (such as intonation). For example, for ancient Chinese courses, an ancient style background and a steady voice style are automatically matched.
[0126] In actual operation, in this step, the courseware can be first converted into a standardized format, such as converting PPT to a PNG sequence; then keywords are extracted; and then the CLIP model is used to analyze the keywords and the material library to recommend the best combination (such as "quantum mechanics" matching the starry sky background).
[0127] Step 170: Generate a massive open online course (MOOC) video based on a preset template, video parameters, the target digital human teaching video, the target teaching digital human model, the video background image, the decoration of the target digital human model, and the voice parameters.
[0128] Among them, in this step, through the preset template and video parameters (such as video resolution and duration), the full process automation from material input to video export is realized, reducing the manual editing link.
[0129] In actual operation, in this step, the video synthesis API (such as FFmpeg) can be called to splice the materials according to the template and export the final video (i.e., the MOOC video).
[0130] Step 160 and Step 170 reduce the time-consuming of manually matching materials in the traditional process by 80%, and support multi-platform format output (such as MP4, VR video stream), improving compatibility.
[0131] The MOOC video method based on AIGC digital humans provided by the embodiments of the present application first determines a target digital human base model matching the target course theme from a pre-trained digital human base model library; the digital human models in the pre-trained digital human model library have functions of synthesizing actions, expressions, and voices. Secondly, input some feature data of the target course into the target digital human base model, and adjust the model parameters of the target digital human base model through transfer learning to obtain a target digital human model. Thirdly, input the target speech text into a temporal model, and generate a temporally aligned target action sequence based on the attention mechanism; the temporal model is trained with speech-action alignment samples. Then, input the target action sequence into the target digital human model to drive the target digital human model to perform corresponding target actions according to the target action sequence, and obtain a target digital human teaching video. Then, use a virtual engine to construct a target virtual teaching scene, and embed the target digital human model into the target virtual teaching scene to obtain a target teaching digital human model. Then, use a multimodal AI model to determine a video background image, decorations of the target digital human model, and speech parameters according to the target courseware. Finally, based on a preset template and video parameters, generate a MOOC video according to the target digital human teaching video, the target teaching digital human model, the video background image, the decorations of the target digital human model, and the speech parameters. In this way, through the pre-trained model library and transfer learning technology, the deployment complexity is significantly reduced, and the adaptation time is shortened from several weeks to several hours; VR / AR and tactile feedback are integrated to construct a "visual-auditory-tactile" collaborative teaching mode, improving the immersion and practical operation effect; the teaching content is optimized based on real-time data feedback, breaking through the limitations of static generation, and realizing "teaching students in accordance with their aptitude"; from material matching to video export is completed in one key, greatly reducing the human and time costs, and is applicable to large-scale course production.
[0132] After introducing the MOOC video method based on AIGC digital humans of the exemplary embodiments of the present disclosure, next, reference is made to Figure 4 to describe the MOOC video device 400 of the exemplary embodiments of the present disclosure.
[0133] Reference is made to Figure 4, the massive open online course (MOOC) video device 400 based on the AI-generated digital human includes: a first determination module 410 configured to determine a target digital human basic model matching the target course theme from a pre-trained digital human basic model library; the digital human models in the pre-trained digital human model library have functions of synthesizing actions, expressions, and voices; an adjustment module 420 configured to input feature data of part of the target course into the target digital human basic model, and adjust the model parameters of the target digital human basic model through transfer learning to obtain a target digital human model; a timing alignment module 430 configured to input the target speech text into a timing model, and generate a target action sequence with timing alignment based on the attention mechanism; the timing model is trained by speech-action alignment samples; a teaching video generation module 440 configured to input the target action sequence into the target digital human model, drive the target digital human model to execute corresponding target actions according to the target action sequence, and obtain a target digital human teaching video; a teaching digital human generation module 450 configured to construct a target virtual teaching scene using a virtual engine, and embed the target digital human model into the target virtual teaching scene to obtain a target teaching digital human model; a second determination module 460 configured to use a multimodal AI model to determine a video background image, decorations of the target digital human model, and voice parameters according to the target courseware; a MOOC video generation module 470 configured to generate a MOOC video based on a preset template and video parameters according to the target digital human teaching video, the target teaching digital human model, the video background image, the decorations of the target digital human model, and the voice parameters.
[0134] In one implementation, the timing alignment module 430 is configured to: optimize the matching degree between the target action and the target voice corresponding to the target speech text by using a reinforcement learning algorithm.
[0135] In one implementation, the reward function in the reinforcement learning algorithm is as follows:
[0136] R total = R sync + α·R diversity + β·R emotion
[0137] where R total is the total reward; R sync is the synchronization error reward, which is used for the time alignment error between the action and the voice; R diversity is the action diversity reward, which is used to calculate the entropy value of the action sequence; R emotion is the emotion consistency reward; α and β are hyperparameters used to control the weights of diversity and emotion consistency.
[0138] In one implementation, the policy output by the reinforcement learning algorithm is as follows:
[0139]
[0140] Among them, γ is the discount factor, which is used to balance the immediate reward and the long-term cumulative reward; R(s t , a t ) is the immediate reward; a is the action; s is the learner's behavior data; E s,a is the expected value of the long-term cumulative reward under the policy π.
[0141] In one implementation, the teaching video generation module 440 is configured to: obtain the learner's behavior data from the learning platform; use a clustering algorithm to construct a learner profile according to the learner's behavior data; and adjust the target digital human teaching video according to the learner profile.
[0142] In one implementation, the teaching video generation module 440 is configured to: adjust the depth of knowledge points according to the learner profile; analyze the learner's classroom performance through NLP technology, and use a CLIP model based on a knowledge graph to generate supplementary explanation segments according to the depth of knowledge points; and add the supplementary explanation segments to the target digital human teaching video.
[0143] In one implementation, the CLIP model based on the knowledge graph is as follows:
[0144] sim enhanced (I, T) = sim(I, T) + λ·KG-Align(I, T)
[0145] Among them, sim enhanced (I, T) represents the similarity; sim(I, T) represents the CLIP similarity; KG-Align(I, T) represents the semantic alignment score of the CLIP model based on the knowledge graph; λ is a hyperparameter.
[0146] In one implementation, the loss function of the transfer learning is as follows:
[0147] L total = L CE + λ·L MMD + μ·L MT
[0148] Among them, L total is the total loss; L CE is the cross-entropy loss; L MMD is the maximum mean discrepancy loss, which is used to align the distribution difference between the pre-trained data and the teaching scenario data, and L MT is the multi-task loss, which is used to jointly perform action and speech synthesis; λ and μ are hyperparameters, which are used to control the relative importance of adaptive and multi-task learning.
[0149] In one embodiment, the teaching digital human generation module 450 is configured to: in combination with a wearable device, convert the target action of the target digital human model into physical feedback.
[0150] In one embodiment, the attention mechanism is as follows:
[0151]
[0152] where A is the speech feature; B is the action feature; Q i is the query vector of the speech feature; is the transposed matrix of the key vector of the action feature; V j is the value vector of the action feature.
[0153] The above device is used to execute the method provided in the foregoing embodiment, and its implementation principle and technical effects are similar, and will not be elaborated here.
[0154] The above modules may be one or more integrated circuits configured to implement the above method, for example: one or more application specific integrated circuits (ASICs), or, one or more microprocessors, or, one or more field programmable gate arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. Again, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0155] Figure 5 FIG. is a schematic diagram of a computer device provided by an embodiment of the present application. This device may be integrated into a terminal device or a chip of a terminal device, and the terminal may be a computing device with data processing capabilities.
[0156] The device includes: a processor 501, a storage medium 502, and a bus 503.
[0157] The storage medium 502 stores program instructions executable by the processor 501. When the computer device 500 runs, the processor 501 communicates with the storage medium 502 through the bus 503, and the processor 501 executes the program instructions to execute the above method embodiment. The specific implementation manner and technical effects are similar, and will not be elaborated here.
[0158] Optionally, the present invention further provides a program product, such as a computer-readable storage medium, including a program that, when executed by a processor, is used to execute the above method embodiments.
[0159] In several embodiments provided by the present invention, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the apparatus or unit can be in an electrical, mechanical or other form.
[0160] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0161] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a hardware plus software functional unit.
[0162] The above integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above software functional unit stored in a storage medium includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drive, mobile hard disk, read-only memory (English: Read-Only Memory, abbreviated as: ROM), random access memory (English: Random Access Memory, abbreviated as: RAM), magnetic disk or optical disk and other various media that can store program codes.
[0163] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for generating MOOC videos based on AIGC digital humans, characterized in that, Including: Determine a target digital human base model that matches the target course theme from a pre-trained digital human base model library; the digital human models in the pre-trained digital human model library have the functions of synthesizing actions, expressions, and voices; Input feature data of part of the target course into the target digital human base model, and adjust the model parameters of the target digital human base model through transfer learning to obtain a target digital human model; Input the target speech text into a temporal model, and generate a temporally aligned target action sequence based on the attention mechanism; The temporal model is trained with speech-action alignment samples; Input the target action sequence into the target digital human model, and drive the target digital human model to execute corresponding target actions according to the target action sequence to obtain a target digital human teaching video; Construct a target virtual teaching scene using a virtual engine, and embed the target digital human model into the target virtual teaching scene to obtain a target teaching digital human model; Adopt a multi-modal AI model to determine a video background image, decorations of the target digital human model, and speech parameters according to the target courseware; Generate a massive open online course (MOOC) video based on a preset template, video parameters, the target digital human teaching video, the target teaching digital human model, the video background image, decorations of the target digital human model, and speech parameters.
2. The method according to claim 1, wherein After driving the target digital human model to execute corresponding target actions according to the target action sequence, the method further includes: Adopt a reinforcement learning algorithm to optimize the matching degree between the target action and the target voice corresponding to the target speech text.
3. The method according to claim 2, wherein The reward function in the reinforcement learning algorithm is as follows: R total = R sync + α·R diversity + β·R emotion where R total is the total reward; R sync is the synchronization error reward, which is used based on the time alignment error between the action and the speech; R diversity is the action diversity reward, which is used to calculate the entropy value of the action sequence; R emotion is the emotional consistency reward; α and β are hyperparameters used to control the weights of diversity and emotional consistency; The policy output by the reinforcement learning algorithm is as follows: Among them, γ is the discount factor, which is used to balance the immediate reward and the long-term cumulative reward; R(s t ,a t ) is the immediate reward; a is the action; s is the learner's behavior data; E s,a is the expected value of the long-term cumulative reward under the policy π.
4. The method according to claim 1, wherein After obtaining the target digital human teaching video, the method further includes: Obtain the behavior data of learners from a learning platform; Adopt a clustering algorithm to construct a learner portrait according to the behavior data of the learners; Adjust the target digital human teaching video according to the learner portrait.
5. The method according to claim 4, characterized in that, The adjusting the target digital human teaching video according to the learner portrait includes: Adjust the depth of knowledge points according to the learner portrait; Analyze the classroom performance of learners through natural language processing (NLP) technology, and adopt a CLIP model based on a knowledge graph to generate supplementary explanation segments according to the depth of knowledge points; Add the supplementary explanation segments to the target digital human teaching video.
6. The method according to claim 5, characterized in that, The CLIP model based on the knowledge graph is as follows: sim enhanced (I, T) = sim(I, T) + λ·KG - Align(I, T) Among them, sim enhanced (I, T) represents the similarity; sim(I, T) represents the CLIP similarity; KG-Align(I, T) represents the semantic alignment score of the CLIP model based on the knowledge graph; λ is a hyperparameter.
7. The method according to claim 1, wherein The loss function of the transfer learning is as follows: L total = L CE + λ·L MMD + μ·L MT Among them, L total is the total loss; L CE is the cross-entropy loss; L MMD is the maximum mean discrepancy loss, which is used to align the distribution differences between the pre-trained data and the teaching scenario data, and L MT is the multi-task loss, which is used for joint action and speech synthesis; λ and μ are hyperparameters used to control the relative importance of adaptive and multi-task learning.
8. The method according to claim 1, wherein After embedding the target digital human model into the target virtual teaching scene, the method further includes: Combine with a wearable device to convert the target actions of the target digital human model into physical feedback.
9. The method according to claim 1, wherein The attention mechanism is as follows: Among them, A is the speech feature; B is the action feature; Q i is the query vector of the speech feature; is the transposed matrix of the key vector of the action feature; V j is the value vector of the action feature.
10. A massive open online course (MOOC) video generation device based on an AI-generated digital human, characterized in that, Including: A first determination module configured to determine a target digital human base model that matches the target course theme from a pre-trained digital human base model library; the digital human models in the pre-trained digital human model library have the functions of synthesizing actions, expressions, and voices; An adjustment module, configured to input feature data of part of the target courses into the target digital human base model, and adjust the model parameters of the target digital human base model through transfer learning to obtain a target digital human model; A timing alignment module, configured to input the target speech text into a timing model, and generate a target action sequence with timing alignment based on the attention mechanism; The timing model is trained by speech-action alignment samples; A teaching video generation module, configured to input the target action sequence into the target digital human model, drive the target digital human model to execute corresponding target actions according to the target action sequence, and obtain a target digital human teaching video; A teaching digital human generation module, configured to construct a target virtual teaching scene using a virtual engine, and embed the target digital human model into the target virtual teaching scene to obtain a target teaching digital human model; A second determination module, configured to use a multimodal AI model to determine a video background image, decorations of the target digital human model, and voice parameters according to a target courseware; A MOOC video generation module, configured to generate a MOOC video based on a preset template, video parameters, the target digital human teaching video, the target teaching digital human model, the video background image, the decorations of the target digital human model, and voice parameters.
Citation Information
Patent Citations
Intelligent MOOC generation method and device based on virtual digital human, and storage medium
CN115515002A
Cited By
Dynamic video synthesis method and device based on instruction sequence and instruction editing system
CN120676223A
A method, apparatus, and instruction editing system for dynamic video synthesis based on instruction sequences.
CN120676223B
Performance description meeting video generation method and device based on artificial intelligence
CN121037640A