Teaching-friendly digital human online MOOC construction system, method, device and medium

By using large language models and 3D point cloud reconstruction technology, combined with UX tools, a teaching digital human system was constructed. This solved the problems of dependence on teachers and long recording cycles in the MOOC recording process, realized the adaptability of teaching content and the teaching expression of gestures, reduced production costs and improved teaching effectiveness.

CN119181283BActive Publication Date: 2025-11-04BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411389921.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-08
Publication Date
2025-11-04
Estimated Expiration
2044-10-08

AI Technical Summary

Technical Problem

Existing technologies rely heavily on teachers during MOOC recording, have long recording cycles, and fail to fully utilize the advantages of virtual agents, lacking adaptability to teaching content and pedagogical expression through gestures.

Method used

By employing large language model technology, text-to-speech technology, intelligent digital human full-body generation technology with accompanying speech, and 3D point cloud reconstruction technology, combined with UX tools, an intelligent generation system for teaching digital humans is constructed. Through teaching action design and scene design, action sequences with enhanced teaching semantics are generated and coordinated with teaching materials to achieve intelligent driving of teaching digital humans.

Benefits of technology

It reduces the time cost for teachers in MOOC production, improves the teaching effectiveness and scenario adaptability of digital human-based teaching, and enhances the semantic expression of gestures and the synergy of teaching materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119181283B_ABST
    Figure CN119181283B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of artificial intelligence, and relates to a teaching-friendly digital human online MOOC construction system, method, equipment and medium. The system comprises a teaching digital human intelligent construction module, a teaching digital human intelligent driving action construction module, a teaching digital human intelligent scene construction module and a teaching digital human intelligent driving module. The teaching digital human intelligent construction module constructs a teaching digital human. The teaching digital human intelligent driving action construction module forms a teaching digital human intelligent driving action. The teaching digital human intelligent scene construction module constructs an intelligent scene. The teaching digital human intelligent driving module drives the teaching digital human in the intelligent scene based on the teaching digital human intelligent driving action. The intelligent scene intelligent adjustment module realizes intelligent adjustment of the intelligent scene. The application can help improve the effectiveness of teaching and reduce the time cost of teachers in MOOC production.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence, and relates to an online MOOC construction system, method, device and medium, in particular to a teaching-friendly digital human online MOOC construction system, method, device and medium. BACKGROUND

[0002] In recent years, the rapid development and continuous technological innovation of Internet technology have further promoted the reform of the form of education. As one of the main forms of online education in China, MOOC (Massive Open Online Course) online education has the characteristics of large scale, openness and online. "Large scale" means the consideration of learners of different levels and needs, "openness" means the transparent peer review of high-quality course standards, and "online" means the change of the teaching medium for students, and students cannot achieve all-round observation and interaction of teaching aids. High-quality MOOC teaching videos are an important carrier for delivering high-quality MOOC teaching content, and the construction system thereof covers various elements such as teaching teachers, students, recording and broadcasting systems, recording personnel, editing personnel and environment. The recording types include traditional studio, studio automatic recording, live shooting and classroom recording. From the arrangement of the site to the scheduling of personnel, from the hardware system to the communication of the camera, the traditional MOOC recording method has a high dependence on the site and manpower, and the period of MOOC recording is three months to half a year. As the core of teaching design and teaching implementation, teachers not only need to arrange the MOOC recording script and plan the course content outline, but also need to make sufficient preparations for content shooting and avoid mistakes in shooting, which will greatly consume the energy of teachers. Therefore, some technologies for assisting in MOOC production have been invented:

[0003] 1. Intelligent driving of 3D digital human, which mainly includes face driving and body movement driving.

[0004] Talking-face Generation (TFG) synthesizes the corresponding facial features and expressions by giving the voice of the speaker. In the coding part of the face, there are three coding methods based on landmarks, vertices and coefficients. Among the three coding methods, landmark depends on the accurate key point feature position, vertices may have part of the three-dimensional coordinate missing, and the method based on coefficient needs a large amount of training data. In the face movement, the movement of the mouth and the content of the speaker have strong correlation, the expression of the landscape face, the posture of the head and the blinking have weak correlation with the content of the speaker, but affect the natural expression of the face generation. In addition, using a pre-trained model and multi-modal input or a hierarchical method to drive the face is the current development trend.

[0005] Digital human gesture generation, which generates natural and expressive body movements from the input of multi-modal information such as speech audio, text or speaker identity, has attracted extensive attention in recent years. Early gesture generation methods were rule-based, relying on pre-set correspondences between human dialogues or speeches and body movements. However, this approach has limitations, as the quality of the results is greatly influenced by the scale of pre-defined actions and the setting of rules, and requires a large amount of human labor. To overcome these problems, deep learning-based techniques have been widely used in digital human gesture generation in recent years. Recurrent neural networks (RNN), long short-term memory recurrent neural networks (LSTM), generative adversarial networks (GAN), vector quantization variational autoencoder (VQ-VAE), diffusion models, and other architectures have been used to generate body movements from speech. Although these architectures can achieve alignment between generated gestures and audio in rhythm, the generation of audio-visual gestures with good semantic expression is still an ongoing exploration. Some methods use text transcriptions of speech as input to improve the semantic nature of gestures. Some methods use specific structures to better learn the semantic information of audio. Some methods use contrastive learning to learn the mapping relationship between text and motion sequences in the latent space. However, the quality of the training results is often determined by the data quality and data distribution of the dataset. The semantic words in existing digital human gesture datasets are relatively sparse, making it difficult for deep learning methods to represent semantic words with low frequency or no frequency. The mapping relationship between semantic words and body movements is many-to-many. Based on the current situation and the ability of LLM to understand the complex relationships and metaphorical meanings between words, some researchers have used LLM to enhance the semantic expression in audio-visual gesture generation tasks by expanding the training dataset and matching external semantic action datasets.

[0006] 2. Design related to Pedagogical Agents

[0007] Researchers are interested in Visual Learning Environments (VLE). There are some related patents that apply digital human to the design of distance learning, a narrative fusion teaching aid automatic generation method aims to establish the mapping relationship between text and voice and teaching aid, focusing on the fusion between teaching aid objects and environment, a digital human teacher personalized teaching device focuses on personalized design, aiming to realize the personalized support and assistance of digital human in the teaching scene through data collection, interaction and personalized analysis and evaluation feedback. A teacher-end virtual human classroom recording and broadcasting video processing method mainly processes the recording and broadcasting video of the teacher, captures the facial features and body movements of the teacher, reconstructs the three-dimensional model of the teacher, and generates movements according to the recording and broadcasting video to drive the teacher model to generate movements, in order to solve the problem of poor recording effect. However, they do not generate content other than the original recording and broadcasting data.

[0008] Pedagogical Agents (PA) as an important part of virtual learning environment, convey teaching information through language and nonverbal behavior. Nonverbal communication is all elements in communication except language, including paralanguage in voice (intonation, pitch) and body language in non-voice (facial expression, eye contact, body movement, etc.). A number of studies have shown that teachers in real classrooms, and PAs in online virtual learning environments, will help enhance students' learning experience and performance through some indicative, metaphorical and symbolic gestures. A working method of virtual teaching system based on AI assistant designs teachers and students to be in a virtual classroom environment, through facial recognition and avatar tracking technology, by annotating the classroom environment and teaching activities, constructing teacher action dataset, classroom scene dataset, teaching activity dataset, constructing the image of intelligent AI assistant, and constructing the spatio-temporal association of narrative, object, and scene. Then through the ways of voice, gesture, and virtual teaching drawing, collaborative teaching content teaching, and teaching content display to organize teaching activities. Among them, the design of teachers' and teaching environment's actions is the movement and stay of the position in the teaching environment, and the design of gestures mainly uses 2D video gesture feature point recognition (there may be cases where it cannot be recognized - this is only the construction of the dataset) + Unity hand interaction library + fuzzy position smoothing. It can be seen that their data quality is not enough, and they have not done very detailed adaptive teaching gesture design and research. Regarding the fine-grained coding of classroom gestures, some scholars have collected classroom teaching videos and analyzed and sorted out the behavior representations in classroom teaching, including symbolic actions, explanatory actions, expressive actions, adaptive actions, regulatory actions, distance actions, and instrumental actions. Among them, instrumental actions reflect the cooperation between teachers and teaching materials in the classroom teaching scene. Some people consider teaching gestures from the perspectives of demonstration, guidance, attention, and emphasis, and finally determine six common classroom gestures, including intentional / unintentional, habitual, guidance, interaction, emphasis, construction, and visualization gestures. Some people focus on how to design reusable standard teaching gestures for PAs applied in holographic projection in intelligent teaching systems. They designed a two-person collaboration task and captured, analyzed, and clustered the gestures of the demonstrator, including seven gesture categories: indication, sign, metaphor, symbol, beat, transition, imitation, and interaction gestures, and designed the corresponding hand movement descriptions. These works explore the division of gestures that are effective for PA teaching in different media.

[0009] Compared with the prior art, although they realize the method of teaching digital person in teaching scene, individualized education support, virtual-real integrated teaching aid automatic generation, intelligent MOOC generation and other digital person applications in teaching scene, in the design of the scene, the scanning reconstruction of the real teaching scene is often considered, and the advantages of digital person as a virtual agent in a virtual environment are rarely considered, that is, the adaptive virtual scene switching in real teaching according to teaching content is considered; in the intelligent driving part of the teaching digital person, although the existing work has considered the language and non-language behavior (facial expression, eye contact, orientation, body movement), some work uses motion capture to obtain the behavior of real teachers to mirror drive digital person, some work extracts audio representation as input to realize the generation and driving of digital person's facial and body movements, and further optimization of action generation--with the help of external semantic-gesture dataset + LLM retrieval matching corresponding gestures to realize semantic enhancement. The former needs the participation and cooperation of teachers, and the latter generated driving gestures, although considering the semantic expression of gesture action, lack of further consideration of the teaching property and teaching material space property of gesture action.

[0010] Therefore, in view of the defects in the prior art, a novel teaching-friendly digital person online MOOC construction system, method, device and medium are needed. SUMMARY

[0011] In order to overcome the defects of the prior art, the present application proposes a teaching-friendly digital person online MOOC construction system, method, device and medium, which aims to build an intelligent generation system of teaching digital person by means of large language model technology, text-to-speech technology, intelligent digital person full-body generation technology with accompanying voice, 3D point cloud reconstruction technology and UE tool, design effective teaching action and teaching space scene, and render efficient, intelligent and accurate 3D teaching scene, so as to reduce the time cost of teachers in MOOC production while improving the effectiveness of teaching.

[0012] In order to achieve the above purpose, the present application provides the following technical solutions:

[0013] A teaching-friendly digital person online MOOC construction system, characterized in that it comprises:

[0014] A teaching digital person intelligent construction module for constructing a teaching digital person close to the target teacher in appearance;

[0015] The teaching digital human intelligent driving action construction module is configured to generate a basic action sequence based on the voice and the text, obtain an enhanced action sequence based on the text using a large language model after fine-tuning, and fuse the basic action sequence and the enhanced action sequence to obtain a teaching semantic enhanced action sequence of the teaching digital human. Meanwhile, the teaching semantic enhanced action sequence is coordinated with teaching materials in a teaching space to form an intelligent driving action of the teaching digital human.

[0016] The teaching digital human intelligent scene construction module is configured to construct an intelligent scene of the teaching digital human.

[0017] The teaching digital human intelligent driving module is configured to drive the teaching digital human in the intelligent scene based on the teaching digital human intelligent driving action.

[0018] The intelligent scene intelligent adjustment module is configured to obtain a time at which each intelligent scene is activated, a time at which the digital human changes orientation, and a change angle based on text and time information corresponding to the teaching digital human intelligent driving action and angle information corresponding to each intelligent scene, and implement intelligent adjustment of the intelligent scene based on the same.

[0019] Preferably, the teaching digital human intelligent driving action construction module comprises:

[0020] The accompaniment rhythm type action generation submodule is configured to extract low-dimensional and high-dimensional audio feature information in the voice using an audio analysis library and an automatic speech recognition model, extract corresponding word embedding in the text using a word vector model, and generate a basic action vector z q * through a decoder of a pre-trained Transformer model.

[0021] The teaching collaborative action enhancement submodule is configured to obtain an enhanced action index based on the text using a large language model after fine-tuning, retrieve a matched text action from a text action data set based on the enhanced action index, encode the matched text action using an encoder of a pre-trained vector quantization variational autoencoder to obtain a quantized label action vector z e , realize fusion of z q and z e by weighted fusion to obtain a semantic enhanced label action vector z e-argu *, and decode the semantic enhanced label action vector z e-argu * using a decoder of the pre-trained vector quantization variational autoencoder to obtain a teaching semantic enhanced action sequence of the teaching digital human.

[0022] The teaching material collaborative action enhancement sub-module is configured to set different types of collaborative actions according to the type of the teaching material, and to achieve the collaboration of the teaching semantic enhanced action sequence and the teaching material based on different types of collaborative actions.

[0023] Preferably, the fusion of z q *and z e is achieved by weighted fusion to obtain a semantic enhanced label action vector z e-argu Specifically includes:

[0024] According to the position of the enhanced action index, the time range of fusion is determined, the base action is segmented according to the time range to form a plurality of base action segments, the motion speed change of the base action is calculated, and the position with the largest speed change is determined as the splicing point of fusion;

[0025] The label action vector is used to replace the corresponding base action vector at the splicing point position, and weighted merging operation is performed on z q *and z e before and after the replacement point.

[0026] Preferably, the type of teaching material includes 2D plane class, 3D scene class, 3D fixed object class and 3D handheld object class, the collaborative action corresponding to the 2D plane class, 3D scene class and 3D fixed object class is a pointing action, the collaborative action corresponding to the 3D handheld object class is a presentation action, the pointing action adopts an IK inverse settlement method to achieve the collaboration of the teaching semantic enhanced action sequence and the teaching material, and the presentation action adopts a preset action library to achieve the collaboration of the teaching semantic enhanced action sequence and the teaching material.

[0027] Preferably, the teaching digital human close in appearance to the target teacher is specifically constructed by: shooting the target teacher to obtain a scanning video, extracting facial feature points after converting the scanning video into a TIFF format, reconstructing a head grid after obtaining a head point cloud, and performing mapping on the head grid according to the facial feature points to obtain head information of the target teacher, replacing head information of an existing model with the head information of the target teacher to obtain the teaching digital human close in appearance to the target teacher.

[0028] In addition, the present application also provides a teaching-friendly digital human online MOOC construction method, characterized by comprising:

[0029] Constructing a teaching digital human close in appearance to the target teacher;

[0030] The basic action sequence is generated based on voice and text, an enhanced action sequence is obtained based on the large language model after fine-tuning based on the text, and the basic action sequence and the enhanced action sequence are fused to obtain a teaching semantic enhanced action sequence of the teaching digital person, and the teaching semantic enhanced action sequence is coordinated with teaching materials in a teaching space to form intelligent driving actions of the teaching digital person.

[0031] An intelligent scene of the teaching digital person is constructed.

[0032] The teaching digital person is driven based on the intelligent driving actions of the teaching digital person in the intelligent scene.

[0033] Based on the text and time information corresponding to the intelligent driving actions of the teaching digital person and the angle information corresponding to each intelligent scene, the time of activation of each intelligent scene and the time and change angle of the orientation change of the digital person are obtained, and intelligent adjustment of the intelligent scene is realized based thereon.

[0034] Preferably, low-dimensional and high-dimensional audio feature information in the voice is extracted using an audio analysis library and an automatic speech recognition model, corresponding word embeddings in the text are extracted using a word vector model, and a basic action vector z q is generated by a decoder of a pre-trained Transformer model.

[0035] An enhanced action index is obtained based on the text using a large language model after fine-tuning, and a matched text action is retrieved from a text action data set based on the enhanced action index, the matched text action is encoded by an encoder of a pre-trained vector quantization variational autoencoder to obtain a quantized label action vector z e , the fusion of z q and z e is realized by weighted fusion to obtain a semantic enhanced label action vector z e-argu , the semantic enhanced label action vector z e-argu is decoded by a decoder of a pre-trained vector quantization variational autoencoder to obtain a teaching semantic enhanced action sequence of the teaching digital person.

[0036] Different types of collaborative actions are set according to the type of teaching materials, and the collaboration of the teaching semantic enhanced action sequence and the teaching materials is realized based on different types of collaborative actions using different methods.

[0037] Preferably, the types of teaching materials include 2D plane class, 3D scene class, 3D fixed object class and 3D handheld object class, the corresponding collaborative actions of the 2D plane class, 3D scene class and 3D fixed object class are pointing actions, the corresponding collaborative action of the 3D handheld object class is a presenting action, the pointing action realizes the collaboration between the teaching semantic enhanced action sequence and the teaching material by using an IK reverse settlement method, and the presenting action realizes the collaboration between the teaching semantic enhanced action sequence and the teaching material by using a preset action library.

[0038] Furthermore, the present application also provides a teaching-friendly digital human online MOOC construction device, characterized in that it comprises:

[0039] one or more processors;

[0040] a memory for storing one or more programs;

[0041] when the one or more programs are executed by the one or more processors, the one or more processors realize the teaching-friendly digital human online MOOC construction method as described above.

[0042] Finally, the present application provides a computer readable storage medium having a computer program stored thereon, characterized in that the program is executed by a processor to realize the steps of the teaching-friendly digital human online MOOC construction method as described above.

[0043] Compared with the prior art, the teaching-friendly digital human online MOOC construction system, method, device and medium of the present application have one or more of the following beneficial technical effects:

[0044] 1. In the intelligent driving part of the teaching digital human, the present application considers the teaching nature of the action and the spatial attribute of the teaching material, and further integrates common teaching audio actions and action parts that are coordinated with the teaching material in teaching.

[0045] 2. In the design of the scene, the present application considers the innate advantages of the digital human as a virtual agent in a virtual environment, i.e. considering switching the virtual scene according to the adaptability of the teaching content in real teaching. BRIEF DESCRIPTION OF DRAWINGS

[0046] Figure 1 is a constituent schematic diagram of the teaching-friendly digital human online MOOC construction system of the present application.

[0047] Figure 2 is a collaborative work flow of the teaching semantic enhanced action sequence and the teaching material of the present application.

[0048] Figure 3 is an exemplary intelligent scene of the present application.

[0049] Figure 4 is a flowchart of the teaching-friendly digital human online MOOC construction method of the present application. DETAILED DESCRIPTION

[0050] Before any embodiments of the application are explained in detail, it is to be understood that the application is not limited in its application to the details of construction and the arrangement of components set forth in the following description or illustrated in the following drawings. The application is capable of other embodiments and of being practiced or being carried out in various ways. Also, it is to be understood that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of "including," "comprising," or "having" and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms "mounted," "connected," "supported," and "coupled" and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Further, "connected" and "coupled" are not restricted to physical or mechanical connections or couplings.

[0051] Also, in the disclosure of the present application, the terms "longitudinal", "lateral", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", and the like indicate the orientation or positional relationship shown in the drawings, which are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore the above terms cannot be understood as limiting the present application; secondly, the term "one" should be understood as "at least one" or "one or more", that is, in one embodiment, the number of one element can be one, and in another embodiment, the number of the element can be multiple, and the term "one" cannot be understood as limiting the number.

[0052] The present application relates to a teaching-friendly digital human online MOOC construction system, method, device and medium, which builds a teaching digital human intelligent generation system by means of large language model technology, text-to-speech technology, intelligent digital human whole body generation technology, 3D point cloud reconstruction technology and UE tools, designs effective teaching actions, teaching space scenes, and renders high-efficiency, intelligent and accurate 3D teaching scenes, thereby reducing the time cost of teachers in MOOC production while improving the effectiveness of teaching.

[0053] Figure 1 The constitution schematic diagram of the teaching-friendly digital human online MOOC construction system of the present application is shown. Figure 1 As shown in the drawings, the teaching-friendly digital human online MOOC construction system of the present application comprises:

[0054] I. Teaching digital human intelligent construction module.

[0055] The teaching digital human intelligent construction module is used to construct a teaching digital human close in appearance to a target teacher.

[0056] In the present application, when constructing a teaching digital human close in appearance to a target teacher, the target teacher is first photographed to obtain a scanning video, and the scanning video is converted into a TIFF format for facial feature point extraction; then, after obtaining head point cloud, the head mesh is reconstructed and mapping is performed on the head mesh according to the facial feature points to obtain the head information of the target teacher; finally, the head information of the existing model is replaced with the head information of the target teacher, and a teaching digital human close in appearance to the target teacher is obtained.

[0057] Specifically, the teaching digital human intelligent construction module is mainly realized by means of UE5 Metahuman. In order to ensure the consistency of the appearance of the teaching digital human and the target teacher, at least the head information of the target teacher needs to be obtained. The present application obtains a scanning video by photographing the target teacher, then converts it into a TIFF format for facial feature point extraction using MetaShape, obtains the head point cloud of the target teacher, then reconstructs the head mesh and performs material mapping based on the facial feature points to obtain the head information of the target teacher, and then replaces the head information of the existing model by means of UE Metahuman plug-in to obtain a teaching digital human close in appearance to the target teacher.

[0058] II. Teaching digital human intelligent driving action construction module.

[0059] The teaching digital human intelligent driving action construction module is used to generate a basic action sequence based on voice and text, obtain an enhanced action sequence based on text using a fine-tuned large language model, and fuse the basic action sequence and the enhanced action sequence to obtain a teaching semantic enhanced action sequence of the teaching digital human, while coordinating the teaching semantic enhanced action sequence with the teaching materials in the teaching space to form a teaching digital human intelligent driving action.

[0060] In the present application, the teaching digital human intelligent driving action construction module mainly uses a deep learning network to predict and generate an action sequence, therefore, a data set needs to be used to pre-train the deep learning network, for which a data set needs to be prepared in advance.

[0061] Firstly, the existing open-source BEAT2 dataset with speech, text transcription alignment information, and Mo-Cap motion data is selected. However, the semantic of this dataset is not strong enough, and it is necessary to enhance the semantic expression of teaching actions and the design of teaching materials. Therefore, it is necessary to further construct a text action dataset. Since teaching actions are mainly gestures, only gesture actions are discussed here. According to the research of relevant scholars on gestures, the common gesture types include five types of pointing gestures, beat gestures, metaphor gestures, symbolic gestures, and transitional gestures (Table 1). According to the research and verification of relevant scholars on meaningful teaching gestures, they include four meaningful teaching gestures of guidance, emphasis, structure, and example (Table 2). Based on these teaching intentions, combined with the teaching behavior of teachers, the existing MOOC teaching video is analyzed, and the teaching action design is performed from five body parts of fingers, palms, arms, whole body, and others (Table 3). Then, the teaching action files are obtained by optical motion capture and motion capture gloves to construct the text action dataset.

[0062] In the text action dataset, each teaching action is stored in the form of index, main time, action description, action meaning, text example, and action file (Table 4). The complete gesture includes preparation, main body, and withdrawal three stages, and the main time represents the effective action time. The index consists of two parts: letter + number, where the letter A ~ E represents the body part to which the action belongs (i.e., fingers, palms, arms, whole body, and others), and the number part is the word embedding obtained by the T5 model according to the corresponding action description, to perform hierarchical clustering according to semantics and obtain the number of each position. Actions with similar semantic information have high similarity in numbers, such as hand clapping and hand waving, which have negative meanings, and their indexes may be B201 and B202.

[0063] Table 1 Gesture classification

[0064] Gesture type Definition Deictic gesture A gesture that points to objects and events in the specific world, or to abstract objects. Beat gesture A hand moves in rhythm with speech. Metaphoric gesture A gesture that depicts an abstract concept or a concrete metaphor of a visual and motoric image. Iconic gesture A gesture that is closely form-related to the semantic content of speech. Conjunctive gesture A gesture that serves to maintain consistency with previous discourse.

[0065] Table 2 Meaningful teaching gestures

[0066]

[0067] Table 3 Teaching action design (part)

[0068]

[0069]

[0070] Table 4 Data format stored for each teaching action

[0071]

[0072] With the above existing BEAT2 dataset and the constructed text action dataset, the deep learning network used can be pre-trained. Since the VQ-VAE (Vector Quantized Variational Autoencoder) and the Decoder of the Transformer model are mainly used in the present application, the VQ-VAE (Vector Quantized Variational Autoencoder) and the Decoder of the Transformer model are mainly pre-trained.

[0073] First, the VQ-VAE is pre-trained. The training data comes from the BEAT2 dataset and the constructed text action dataset. Three losses need to be calculated for training, the first one is the reconstruction loss used only to train the Encoder and Decoder of the VQ-VAE, the gradients of the two are consistent. The second one is used to train the Codebook of the VQ-VAE, so that the encoded Z e (Z e The gradient is fixed) is close. The third one is used to train the Encoder of the VQ-VAE while fixing the Codebook gradient, to ensure the stability of the output of the Encoder. Wherein, L represents the loss function, sg[·] represents the gradient stop, x represents the input, z e (x) represents the Encoder vector of the input x, z q (x) represents the generated quantization vector, which is passed to the Decoder of the VQ-VAE as input.

[0074]

[0075] Next, the Decoder of the Transformer model is trained. The training data comes from the BEAT2 dataset. The data input is audio and text, and the output is the z q (Notice that the parameters of the Encoder and Decoder of the VQ-VAE need to be frozen at this time to prevent further gradient descent). Among them, for audio, Librosa and wav2vec pre-training models are used to extract low-dimensional and high-dimensional audio feature information respectively; for text, the corresponding word embedding is obtained through the FastText pre-training model. Based on the low-dimensional and high-dimensional audio feature information and the corresponding word embedding, the Decoder of the Transformer model is used to generate z q *。z q * is restored by the Decoder of the pre-trained VQ-VAE* The training loss function includes z. q * and z q x * The reconstruction loss of both x and x.

[0076] In this invention, the teaching digital human intelligent driving action construction module includes an audio-accompanied rhythmic action generation submodule, a teaching collaborative action enhancement submodule, and a textbook collaborative action enhancement submodule.

[0077] The audio-accompanied rhythmic action generation submodule uses an audio analysis library (e.g., Librosa) and an automatic speech recognition model (e.g., a wav2vec pre-trained model) to extract low-dimensional and high-dimensional audio feature information from the speech, respectively. It uses a word vector model (e.g., a FastText pre-trained model) to extract the corresponding word vectors / word embeddings from the text, and generates the basic action vector z using the decoder of the pre-trained Transformer model. q *

[0078] The teaching collaborative action enhancement submodule is used to obtain an enhanced action index based on the text using a fine-tuned large language model. Based on the enhanced action index, matching text actions are retrieved from the text action dataset. The matching text actions are then encoded by a pre-trained vector quantization variational autoencoder to obtain a quantized label action vector z. e z is achieved through weighted fusion q * and z e The fusion of these elements yields a semantically enhanced tag action vector z. e-argu *, the semantically enhanced tag action vector z e-argu The decoder of the pre-trained vector quantization variational autoencoder decodes the action sequence with enhanced instructional semantics for the instructional digital human.

[0079] Since the fine-tuned large language model is to be used, an instruction data set for fine-tuning the large language model needs to be constructed first. In the present application, 20 video clips of 10-20 minutes are collected from online MOOC videos and TED speeches according to the text action data set that has been constructed, the video audio is transcribed by ASR (Automatic Speech Recognition), and after the corresponding text is obtained, the gesture expression of the video is compared, and the index is manually annotated and added to construct an instruction data set that can be used to fine-tune the large language model. After fine-tuning the large language model using the instruction data set, the text is input into the fine-tuned large language model, and an enhanced action index can be obtained. With the action index, the matching text action can be retrieved from the text action data set, including subject time, action description, action meaning, text example, etc. Then, it is input into the Decoder (encoder) of the pre-trained VQ-VAE to obtain the quantized label action vector z e .

[0080] In the fusion of z q and z e , first, the time range of the fusion is determined according to the position of the enhanced action index, and the basic action is divided into multiple basic action segments according to the time range, the motion speed change of the basic action is calculated, and the position with the largest speed change is determined as the splicing point of the fusion. Then, the label action vector is used to replace the corresponding basic action vector at the splicing point position, and weighted merging operation is performed on z q and z e before and after the replacement point.

[0081] Preferably, in the weighted merging operation, the pre- replacement point time to the post-replacement point time is multiplied by a half-cosine curve from 0.3 to 0.7 to 0.3 from the weight (ws).z e-argu The merging weight of z e-argu is set to 1-ws to ensure that the sum is 1. Then, the spliced z e-argu vector is decoded by the decoder of the quantized variational autoencoder to obtain the teaching semantic enhanced action sequence of the teaching digital person.

[0082] The teaching material collaborative action enhancement sub-module is used to set different types of collaborative actions according to the type of teaching material, and to realize the collaboration of the teaching semantic enhanced action sequence and the teaching material based on different types of collaborative actions.

[0083] According to the spatial attributes of the teaching materials, they are divided into 2D teaching materials and 3D teaching materials. The application emphasizes the automatic generation of collaborative gestures, but the interaction between the teacher and the teaching materials is various, and the strong interaction (grasping, playing, etc.) is not considered here. The application starts from the emphasis intention in the teaching scene, and focuses on the use of pointing and indicating gestures.

[0084] The application does not consider the automation of the time of collaborative gesture generation. Therefore, the teacher needs to manually provide when to perform the collaborative gesture. In the preparation stage of the teaching materials and the teaching script, the teacher performs the annotation work of the key trigger words of the teaching script and the corresponding spatial attribute information of the teaching materials. Through TTS (Text-to-Speech, text-to-speech) and MFA (Montreal Forced Alignment, audio text alignment), the corresponding timestamp of the key word can be restored, so as to obtain the time information of the generation of the collaborative gesture. And through the index of the scene to which the teaching material belongs and the spatial attribute information, the IK (Inverse Kinematics, inverse dynamics) is used to inversely solve the action animation required for the pointing action of the bound skeleton in the corresponding time range according to the spatial position information of the target.

[0085] Among them, the 2D teaching materials mainly include PPT or video information presented on the 2D screen, and the teaching collaborative gesture is to point to the position in the 2D plane to guide the students to focus their attention on the emphasized area. According to the spatial attributes of the 3D teaching materials, the scene direction to which they belong, the mutual relationship with the digital person, etc. According to the mutual relationship, the teaching materials can be divided into the following three categories: 1. Scene type: long-distance scene teaching materials, digital people are often in them, generally in groups or in large scale, including buildings, venues, historical sites, etc.; 2. Fixed object type: medium-distance single 3D element teaching material, generally and the digital person exist in space, with fixed spatial position; 3. Hand-held object type: close-range single 3D element teaching material, which can be held and displayed by the digital person. This category of teaching materials is not specific to a certain scene and does not exist fixedly in space.

[0086] The performance mode of the teaching gesture is designed differently according to the different mutual relationship with the spatial object.

[0087] The scene type and the fixed object type teaching materials are far away from the digital person, and generally exist fixedly in the scene. When uploading, the scene index, distance information, and the position information of the emphasized text and the emphasized area need to be annotated first. This part only uses the pointing emphasis action obtained by IK inverse calculation.

[0088] For the class of objects that can be held in hand, due to its strong correlation with the appearance and text content, and there is no fixed position. And regardless of the complexity and the interaction of this class of objects (including hand-held or gestures, spatial size attributes, etc.), only the automation of the generation of the presentation action of this class of objects is implemented. The data preparation stage selects the standard presentation action library according to the previous text action data set, and pre-acquires the position information of the presented object relative to the teaching digital human. When uploading this type of teaching material, the corresponding keyword text of the object appearance needs to be labeled, and the presentation action is selected. Then the action splicing is completed within the time segment range corresponding to the keyword. In addition to synthesizing the presentation action of the digital human, this part will also generate the animation of the appearance and disappearance of the object.

[0089] Figure 2 The specific workflow is shown. Among them, uploading and labeling are completed through the front and back end platforms built by django+vue.js+3D.js, and the labeling key information (text, keyword text, teaching material index, teaching material attribute, emphasized position coordinate / presentation action index) is transmitted to the back end in the form of json through uploading. The blender python api is used to call the blender to set the constraint information of the matching bone node, and the absolute spatial coordinates of the emphasized position of the teaching material in the space are calculated according to the preset spatial layout. The time range t (t1, t2) is obtained by using the TTS+MFA method through the keyword text. The IK tool in the blender is called to solve the directional gesture animation, and the semantic enhanced action sequence output by the algorithm and the splicing of the directional gesture action sequence output are realized by using the Kalman filter smoothing algorithm, so as to obtain the final limb action sequence.

[0090] Table 5 Quantitative information and collaborative gestures

[0091]

[0092] III. Teaching digital human intelligent scene construction module.

[0093] The teaching digital human intelligent scene construction module is used to construct the intelligent scene of the teaching digital human.

[0094] The teaching materials that can be placed in teaching mainly include 2D and 3D teaching materials. Unlike the fixed classroom scene mode of the real classroom, the teaching scene in the virtual world hopes to be diverse and changeable, but at the same time, the differences in the specifications of the teaching materials themselves (including dimensions, scene size, etc.) need to be balanced. The goal of the teaching scene designed by the present application is to hope that the digital human can complete the switching of the teaching environment displayed in the visual range with the help of the lens at the minimum moving step. Therefore, the present application designs a teaching scene layout centered on the digital human in the form of scene surrounding (such as Figure 3The intelligent scene includes a 2D screen mainly used for displaying 2D ppt and videos in teaching materials and other 3D scene areas (animation of long-distance large objects / medium-distance medium objects / near-distance small objects is directly added to the whole scene from the previous step). Each 3D scene corresponds to a separate camera and lighting.

[0095] IV. Teaching digital human intelligent driving module.

[0096] The teaching digital human intelligent driving module is used to drive the teaching digital human in the intelligent scene based on the teaching digital human intelligent driving action.

[0097] In the present application, the teaching digital human intelligent driving action is bound to the corresponding channel of the teaching digital human by establishing a port link through the UE, so that the teaching digital human intelligent driving action can be used to drive the teaching digital human. The driven teaching digital human, that is, the teaching digital human animation, can be imported into the intelligent scene.

[0098] V. Intelligent scene intelligent adjustment module.

[0099] The intelligent scene intelligent adjustment module is used to obtain the time of activating each intelligent scene and the time and angle of changing the orientation of the digital human based on the corresponding text and time information of the teaching digital human intelligent driving action and the angle information corresponding to each intelligent scene, and to realize intelligent adjustment of the intelligent scene based thereon.

[0100] Specifically, after the teaching digital human animation is imported into the intelligent scene, the time of activating the scene (lens activation, lighting activation) and the time and angle of changing the orientation of the digital human are obtained according to the original text and time information corresponding to the cooperative action and the fixed angle information corresponding to each intelligent scene, and are respectively transmitted to the lens and lighting control and the rotation angle of the digital human, and key frame information is added, so as to realize intelligent adjustment of the teaching scene.

[0101] As Figure 4 As shown in the figure, when the teaching-friendly digital human online MOOC construction system is used to construct a teaching-friendly digital human online MOOC, first, the teaching digital human intelligent construction module is used to construct a teaching digital human close to the target teacher in appearance.

[0102] Then, the teaching digital person intelligent driving action construction module is adopted to generate a basic action sequence based on voice and text, an enhanced action sequence is obtained based on text by using a large language model after fine-tuning, and the basic action sequence and the enhanced action sequence are fused to obtain a teaching semantic enhanced action sequence of the teaching digital person, and the teaching semantic enhanced action sequence is coordinated with teaching materials in a teaching space to form an intelligent driving action of the teaching digital person.

[0103] Then, the teaching digital person intelligent scene construction module is adopted to construct an intelligent scene of the teaching digital person.

[0104] After that, the teaching digital person intelligent driving module is adopted to drive the teaching digital person in the intelligent scene based on the intelligent driving action of the teaching digital person.

[0105] Finally, the intelligent scene intelligent adjustment module is adopted to obtain the time of each intelligent scene being activated and the time and angle of change of the orientation of the digital person based on the text and time information corresponding to the intelligent driving action of the teaching digital person and the angle information corresponding to each intelligent scene, and to realize intelligent adjustment of the intelligent scene based thereon.

[0106] In the present application, a teaching-friendly digital person online MOOC construction device is also involved, which comprises one or more processors, a memory for storing one or more programs, and the one or more programs are executed by the one or more processors to enable the one or more processors to implement the teaching-friendly digital person online MOOC construction method as described above.

[0107] Furthermore, the present application also relates to a computer readable storage medium having a computer program stored thereon, wherein the program is executed by a processor to implement the steps of the teaching-friendly digital person online MOOC construction method as described above.

[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit the protection scope of the present application. Based on the technical solutions of the present application, those skilled in the art can modify or equivalently replace the technical solutions of the present application without departing from the essence and scope of the technical solutions of the present application.

Claims

1. A pedagogically friendly digital human online MOOC building system, characterized by, Comprise: A teaching digital human intelligent construction module for constructing a teaching digital human that is visually close to a target teacher; A teaching digital human intelligent driving action construction module for generating a basic action sequence based on voice and text, obtaining an enhanced action sequence based on text using a fine-tuned large language model, and fusing the basic action sequence and the enhanced action sequence to obtain a teaching semantic enhanced action sequence of the teaching digital human, while coordinating the teaching semantic enhanced action sequence with teaching materials in a teaching space to form a teaching digital human intelligent driving action; A teaching digital human intelligent scene construction module for constructing an intelligent scene of a teaching digital human; A teaching digital human intelligent driving module for driving the teaching digital human in the intelligent scene based on the teaching digital human intelligent driving action; An intelligent scene intelligent adjustment module for obtaining the time of each intelligent scene activation and the time and angle of change of the digital human orientation based on the text and time information corresponding to the teaching digital human intelligent driving action and the angle information corresponding to each intelligent scene, and realizing intelligent adjustment of the intelligent scene based thereon; The teaching digital human intelligent driving action construction module comprises: The audio rhythm type action generation submodule is used for extracting low-dimensional and high-dimensional audio feature information in the voice respectively using an audio analysis library and an automatic speech recognition model, extracting corresponding word embedding in the text using a word vector model, and generating a basic action vector z through a decoder of a pre-trained Transformer model q *; a teaching collaborative action enhancement sub-module, which is configured to obtain an enhanced action index based on the fine-tuned large language model based on the text, retrieve a matched text action from a text action data set based on the enhanced action index, and encode the matched text action through an encoder of a pre-trained vector quantization variational autoencoder to obtain a quantized label action vector z e , implement fusion of z q and z e by weighted fusion, obtain a semantically enhanced label action vector z e-argu *, and decode the semantically enhanced label action vector z e-argu * through a decoder of the pre-trained vector quantization variational autoencoder to obtain a teaching semantically enhanced action sequence of the teaching digital person; A teaching material cooperative action enhancement sub-module for setting different types of cooperative actions according to the type of teaching materials, and implementing the coordination of the teaching semantic enhanced action sequence and the teaching materials based on different types of cooperative actions using different methods.

2. The pedagogically friendly digital human online MOOC building system of claim 1, wherein, The fusion of z q and z e is achieved by weighted fusion, resulting in a semantically enhanced label action vector z e-argu Specifically includes: Determine the time range of fusion according to the position of the enhanced action index, and cut the basic action according to the time range to form a plurality of basic action segments, calculate the motion speed change of the basic action, and determine the position with the maximum speed change as the splicing point of fusion; The tag action vector is used to replace the corresponding base action vector for the splice point location, and a weighted merge operation is performed on z q * and z e ​ 3. The pedagogically friendly digital human online MOOC building system of claim 2, wherein, The type of teaching materials includes 2D plane, 3D scene, 3D fixed object and 3D handheld object, the cooperative action corresponding to the 2D plane, 3D scene and 3D fixed object is a pointing action, and the cooperative action corresponding to the 3D handheld object is a presentation action, the pointing action uses an IK inverse settlement method to realize the coordination of the teaching semantic enhanced action sequence and the teaching materials, and the presentation action uses a preset action library to realize the coordination of the teaching semantic enhanced action sequence and the teaching materials.

4. The pedagogically friendly digital human online MOOC building system according to any one of claims 1-3, wherein, The teaching digital human that is visually close to the target teacher is specifically constructed by: shooting the target teacher to obtain a scanning video, converting the scanning video into TIFF format and extracting facial feature points, reconstructing a head grid after obtaining a head point cloud and pasting the facial feature points on the head grid to obtain head information of the target teacher, replacing the head information of an existing model with the head information of the target teacher, and obtaining a teaching digital human that is visually close to the target teacher.

5. A pedagogically friendly digital human online MOOC construction method, characterized in that, Comprise: Construct a teaching digital human that is visually close to a target teacher; Based on speech and text, a basic action sequence is generated. An enhanced action sequence is obtained from the text using a fine-tuned large language model. The basic and enhanced action sequences are then fused to obtain a semantically enhanced action sequence for the teaching digital human. Simultaneously, this semantically enhanced action sequence is coordinated with teaching materials in the teaching space to form intelligently driven actions for the teaching digital human. Specifically, forming these intelligently driven actions includes: extracting low-dimensional and high-dimensional audio features from the speech using an audio analysis library and an automatic speech recognition model; extracting corresponding word embeddings from the text using a word vector model; and generating a basic action vector z using a pre-trained Transformer model decoder. q *; An enhanced action index is obtained based on the text using a fine-tuned large language model. Matching text actions are retrieved from the text action dataset based on this enhanced action index. The matched text actions are then encoded by a pre-trained vector quantization variational autoencoder to obtain a quantized label action vector z. e z is achieved through weighted fusion q * and z e The fusion of these elements yields a semantically enhanced tag action vector z. e-argu *, the semantically enhanced tag action vector z e-argu The decoder of the pre-trained vector quantization variational autoencoder decodes the semantically enhanced action sequence of the teaching digital human; different types of collaborative actions are set according to the type of teaching materials, and different methods are used to achieve the collaboration between the semantically enhanced action sequence and the teaching materials based on the different types of collaborative actions; Construct an intelligent scene of a teaching digital human; Drive the teaching digital human in the intelligent scene based on the teaching digital human intelligent driving action; Based on the text and time information corresponding to the intelligent driving action of the teaching digital human and the angle information corresponding to each intelligent scene, the time of activation of each intelligent scene and the time and angle of change of the digital human orientation are obtained, and intelligent adjustment of the intelligent scene is realized based thereon.

6. The pedagogically friendly digital human online MOOC construction method of claim 5, wherein, The types of teaching materials include 2D plane class, 3D scene class, 3D fixed object class and 3D handheld object class, the corresponding collaborative actions of the 2D plane class, 3D scene class and 3D fixed object class are pointing actions, the corresponding collaborative action of the 3D handheld object class is a presentation action, the pointing action adopts an IK reverse settlement method to realize the collaboration of the teaching semantic enhanced action sequence and the teaching material, and the presentation action adopts a preset action library to realize the collaboration of the teaching semantic enhanced action sequence and the teaching material.

7. A pedagogically friendly digital human online MOOC construction device, characterized by, Comprise: One or more processors; Memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the teaching-friendly digital human online MOOC construction method according to any one of claims 5-6.

8. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the steps of the teaching-friendly digital human online MOOC construction method according to any one of claims 5-6.

Citation Information

Patent Citations

  • Intelligent MOOC generation method and device based on virtual digital human, and storage medium

    CN115515002A

  • Digital human motion intelligent generation method and digital human motion intelligent generation equipment

    CN117093669A