Multi-modal teaching method and device based on generative artificial intelligence, medium and product
Through the multimodal teaching method of generative artificial intelligence, using Wenshengwen, Wensheng, Zhisheng, Video Model, and AI digital human technology, the problem of insufficient dynamics and resource combination in multimodal teaching is solved, the teaching quality and learning experience are improved, and the efficient and fair provision of educational resources is achieved.
Patent Information
- Application Number
- CN202510424243.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-18
AI Technical Summary
The existing generative artificial intelligence has problems such as insufficient dynamics, difficult to grasp resource combinations, and automatic generation of resource abnormal pictures in multimodal teaching, which affects the teaching quality and immersion.
Multimodal teaching method based on generative artificial intelligence is adopted, and the text of educational resource is obtained for preprocessing, multimodal resources are generated using Wenshengwen, Wensheng diagram, and graphic video models, and assisted teaching is combined with AI digital human technology to monitor learners' physiological data in real time to optimize teaching methods.
It improves the quality and effectiveness of teaching resources, enhances the learning experience, reduces students' cognitive load, realizes 24-hour uninterrupted teaching services, reduces labor costs, and promotes educational equity.
Smart Images

Figure CN120339005A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and particularly to a multi-modal teaching method, device, medium and product based on generative artificial intelligence. Background Art
[0002] Artificial Intelligence Generated Content (AIGC) technology, that is, artificial intelligence generating content, relies on core technologies such as machine learning, deep learning, and natural language processing technologies, and trains a model that can generate high-quality content by analyzing a large amount of data. The development of AIGC benefits from breakthroughs in algorithms, technologies, computing power, and data. In terms of algorithms, taking Chat4.0 as an example, through human feedback reinforcement learning technology, after three training processes of supervised fine-tuning, constructing a reward model, and improving the language model, the quality of the output text is continuously improved. In terms of computing power, ChatGPT relies on the powerful infrastructure of the AzureAI supercomputing platform, including a large number of CPU cores and GPUs, as well as high-speed network transmission bandwidth, enabling a large amount of model training data and fast information dissemination speed. In terms of data, from GPT-1 to GPT-3, the training models and data materials have been continuously expanded. GPT-3 integrates data from multiple sources, has 175 billion parameters and 45TB of data volume, enabling ChatGPT to adjust answers according to user feedback and provide a smooth and realistic conversation experience. These technological advancements have enabled AIGC to demonstrate great potential and advantages in content generation. As an important direction in the field of artificial intelligence, multi-modal technology is in a stage of rapid development and evolution. With the expansion of data scale, the improvement of computing resources, and the innovation of algorithms, multi-modal technology presents new trends and opportunities in aspects such as model architecture, learning methods, and application scenarios. Large-scale multi-modal pre-training models are becoming a new research hotspot. These models pre-train on a large amount of multi-modal data, learn general, multi-modal feature representations, and can perform well in downstream tasks.
[0003] Combining generative artificial intelligence with education provides a new way to produce high-quality digital educational resources, which is conducive to the integration of multimodal educational resources, but there are still certain limitations. First, the dynamics of the videos synthesized using image synthesis are limited. Compared with video resources such as carefully crafted animated cartoons, the current synthesized videos are insufficient in simulating scenarios such as birds flying and people walking, thus reducing the immersion of learners during viewing. Second, the combination of existing high-quality digital educational resources and the educational resources generated by generative artificial intelligence can improve the understanding of landscape poems, but it is difficult to grasp the respective roles played by the two. Due to the uncertainty of the results of generative artificial intelligence in generating digital resources and the influence of the given text content, it is currently very difficult to determine the proportion of the resources generated by generative artificial intelligence in the integration of multimodal resources. Finally, abnormal images occasionally appear in the images of the automatically generated landscape poems, such as fishing on the grass. This situation is affected by various factors.
[0004] Based on the above problems, in order to make full use of generative artificial intelligence and improve the quality and effect of teaching resources, there is an urgent need to provide a multimodal teaching method based on generative artificial intelligence. Summary of the Invention
[0005] The purpose of this application is to provide a multimodal teaching method, device, medium and product based on generative artificial intelligence, which can improve the quality and effect of teaching resources.
[0006] To achieve the above purpose, this application provides the following solutions:
[0007] In the first aspect, this application provides a multimodal teaching method based on generative artificial intelligence, and the multimodal teaching method based on generative artificial intelligence includes:
[0008] Obtain educational resource text; the educational resource text includes: landscape poems;
[0009] Preprocess the educational resource text; the preprocessing includes: data cleaning, text standardization, word segmentation, stop word removal, and text vectorization;
[0010] For the preprocessed educational resource text, a multimodal resource model constructed based on generative artificial intelligence is used to obtain multimodal resources; the multimodal resources include: text, images, audio, and video corresponding to the preprocessed educational resource text; the multimodal resource model includes: a text-to-text model, a text-to-image model, and an image-to-video model; the text-to-text model is used to output the text corresponding to the preprocessed educational resource text based on natural language processing technology, text mining and machine learning technology, and deep learning models; the text-to-image model is used to output the image corresponding to the text based on deep learning models and image synthesis algorithms; the image-to-video model is used to generate the corresponding video according to the images, text, and audio data corresponding to the preprocessed educational resource text using video editing tools;
[0011] Combine the multimodal resources with the educational resource text and use AI digital human technology for assisted teaching.
[0012] Optionally, the text-to-text model includes: a feature extraction module, a text generation module, a discrimination module, and a first output module;
[0013] The feature extraction module is used to extract content class features and digital image class features from the preprocessed educational resource text using natural language processing technology, text mining, and machine learning technology;
[0014] The text generation module is used to obtain the text corresponding to the preprocessed educational resource text based on the content class features and digital image class features using a deep learning model;
[0015] The discrimination module is used to compare the text corresponding to the preprocessed educational resource text with the corresponding materials obtained from the database to obtain a comparison result; and correct the text corresponding to the preprocessed educational resource text according to the comparison result;
[0016] The first output module is used to output the corrected text.
[0017] Optionally, the text-to-image model includes: a text-to-image module, an image-to-text module, an evaluation module, a screening module, and a second output module;
[0018] The text-to-image module is used to extract key elements from the corrected text, and use word frequency statistics and TF-IDF methods to extract landscape keywords; according to the landscape keywords, match landscape elements in combination with the natural landscape vocabulary library; according to the text context, determine the color meaning by combining context analysis; according to the text context, use the sentiment analysis model to identify sentiment words and sentiment tendencies, determine the sentiment intensity, and then determine the emotional tone according to the sentiment intensity; according to the emotional tone and text context, use a deep learning model to generate multiple preliminary images;
[0019] The image-to-text module is used to generate the text corresponding to each preliminary image by using the cognitive vision language model CogVLM;
[0020] The evaluation module is used to determine the matching degree between the text corresponding to each preliminary image and the corrected text;
[0021] The screening module is used to screen the preliminary images according to the matching degree to obtain the screened images;
[0022] The second output module is used to obtain the image corresponding to the preprocessed educational resource text by using a style transfer algorithm based on the screened images and the corrected text.
[0023] Optionally, the image-to-video model includes: an audio data conversion module, an audio feature extraction module, a key frame extraction module, and a third output module;
[0024] The audio data generation module is used to determine the text speech according to the corrected text, and then convert the text speech into audio data by using a web crawler or a deep learning model;
[0025] The audio feature extraction module is used to extract features from the audio data to obtain audio features; the audio features include: frequency features, rhythm features, and emotional features;
[0026] The key frame extraction module is used to extract features from the image corresponding to the preprocessed educational resource text to obtain key frames;
[0027] The third output module is used to synchronize the audio features and the key frames and generate the corresponding video by using a video editing tool.
[0028] Optionally, the image-to-video model further includes: a speech recognition module and a commentary copywriting generation module;
[0029] The speech recognition module is used to convert the audio in the video into text through speech recognition technology;
[0030] The commentary copywriting generation module is used to synchronize and format the text with the video timeline, add subtitles or use a large language model to understand the video and automatically generate commentary copywriting.
[0031] Optionally, combine the multimodal resources with the educational resource text and use AI digital human technology for auxiliary teaching, and then further include:
[0032] Obtain the physiological data of the learner through wearable electroencephalogram sensors and galvanic skin sensors; the physiological data includes: brain wave frequency bands and corresponding psychological and physiological characteristics and skin conductance fluctuations;
[0033] Determine the cognitive load of the learner based on the learner's physiological data;
[0034] Continuously optimize the multimodal teaching method according to the cognitive load.
[0035] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the multimodal teaching method based on generative artificial intelligence.
[0036] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the multimodal teaching method based on generative artificial intelligence.
[0037] In a fourth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the multimodal teaching method based on generative artificial intelligence.
[0038] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0039] The present application provides a multimodal teaching method, device, medium, and product based on generative artificial intelligence, which utilizes multimodal digital resources that combine existing excellent educational resources and resources generated by generative artificial intelligence, emphasizing giving full play to the advantages of both to meet the diverse needs of multimedia teaching resources; through digital humans, teaching services can be provided continuously for 24 hours, without being restricted by time and place, while reducing labor costs. Secondly, in the practice of multimodal intelligent classrooms and rich learning experiences, the present application can, according to the teaching content, generate text from text, generate images from text, and generate videos from images, reducing the cognitive load of students and improving students' understanding and memory of learning content, thereby enhancing learning effects. In addition, AI digital human technology saves costs and improves efficiency. After using AI digital human technology, the teaching video content that originally required a team to collaborate and produce one by one can now be produced on a large scale with just an operator uploading the courseware content and performing simple operations in the background. Finally, this method helps to promote educational equity by breaking geographical and resource limitations, enabling more students to enjoy high-quality educational resources. In summary, the present application not only improves the quality of education but also brings a richer and more efficient learning experience to learners, reduces the cognitive load of students' learning, and improves teaching efficiency.
[0040] The multimodal digital resources generated by the present application can greatly reduce the cognitive load in the learning process and enhance students' in-depth understanding of landscape poetry. Furthermore, it explores the potential of the combination of generative artificial intelligence and education, especially in the construction of educational digital resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0042] Figure 1 It is a schematic flowchart of a multi-modal teaching method based on generative artificial intelligence in an embodiment of the present application. Specific embodiments
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0044] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0045] In an exemplary embodiment, as Figure 1 shown, a multi-modal teaching method based on generative artificial intelligence is provided, and this method includes the following S101 to S104. Among them:
[0046] S101, obtaining educational resource texts; the educational resource texts include but are not limited to: landscape poems;
[0047] S102, preprocessing the educational resource texts; the preprocessing includes: data cleaning, text standardization, word segmentation, stop word removal, and text vectorization;
[0048] The process of preprocessing is:
[0049] (1) Data cleaning, removing irrelevant information in educational resource texts, such as HTML tags, URL links, special symbols, etc. Delete redundant spaces and line breaks. Unify the encoding format to ensure data consistency. (2) Text standardization, reducing lexical diversity. Remove spelling or grammar mistakes. Conduct stemming or lemmatization to simplify words to their basic forms. (3) Word segmentation. For Chinese texts, use word segmentation tools (such as jieba) to split continuous character sequences into words. For English or other languages with space separation, further process abbreviations or compound words. Among them, word segmentation includes rule-based word segmentation, statistics-based word segmentation, and deep learning-based word segmentation; statistics-based word segmentation includes: Hidden Markov Model and Conditional Random Field; deep learning-based word segmentation includes: Recurrent Neural Network and WordPiece segmentation of BERT. (4) Remove stop words, deleting high-frequency and less semantically contributing words in educational resource texts, such as "de", "le", "he", etc. (5) Text vectorization (involving algorithms Word2Vec, GloVe, Bert), converting text into numerical forms, such as using one-hot encoding, word2vec, or word embedding, cleaning irrelevant characters, and performing word segmentation.
[0050] S103. For the preprocessed educational resource text, use a multi-modal resource model constructed based on generative artificial intelligence to obtain multi-modal resources; the multi-modal resources include: text, images, audio, and video corresponding to the preprocessed educational resource text; the multi-modal resource model includes: a text-to-text model, a text-to-image model, and an image-to-video model; the text-to-text model is used to output the text corresponding to the preprocessed educational resource text based on natural language processing technology (Natural Language Processing, NLP), text mining, machine learning technology, and deep learning models; the text-to-image model is used to output the image corresponding to the text based on deep learning models and image synthesis algorithms; the image-to-video model is used to generate the corresponding video according to the image, text, and audio data corresponding to the preprocessed educational resource text, using video editing tools.
[0051] In the text-to-text model, the feature extraction module is used to extract content class features and digital image class features from the preprocessed educational resource text using natural language processing technology, text mining, and machine learning technology.
[0052] The text generation module is used to obtain the text corresponding to the preprocessed educational resource text based on the content class features and digital image class features using a deep learning model.
[0053] The discrimination module is used to compare the text corresponding to the preprocessed educational resource text with the corresponding materials obtained from the database to obtain a comparison result; and correct the text corresponding to the preprocessed educational resource text according to the comparison result;
[0054] The first output module is used to output the corrected text.
[0055] Specifically, the feature extraction module specifically includes: sentiment analysis technology, theme extraction, image feature extraction, and knowledge graph construction;
[0056] Among them, sentiment analysis technology (Pathways Language Model, PaLM) is used for sentiment analysis to extract sentiment features, and Transformers-based Topic Modeling is used for theme extraction; U-Net is used for image feature extraction; Knowledge GraphEmbedding is used for knowledge graph construction.
[0057] The features of the content category include but are not limited to keywords and phrases; the extraction algorithms for keywords and phrases include but are not limited to: Bag of Words (BOW), TF-IDF algorithm (TermFrequency-InverseDocumentFrequency);
[0058] Natural language processing techniques are used to identify the emotional tendencies expressed in poems, such as joy, sadness, anger, etc. Modern sentiment analysis is usually based on deep learning models, such as BERT or GPT series, which can understand the metaphors and symbols in poems and thus accurately capture the emotional color.
[0059] As a specific embodiment, according to keywords and phrases, a topic model, context understanding, and semantic association are used for topic recognition; the algorithms involved in semantic association include Latent Dirichlet Allocation (LDA), Non-negative Matrix Factorization (NMF); the algorithms involved in context modeling include Transformer architecture, Recurrent Neural Network (RNN); the algorithms involved in context understanding include but are not limited to: Transformer architecture, BERT model, GPT, and natural language processing techniques;
[0060] Visualization is performed according to the recognized theme. Among them, visualization tools: such as Matplotlib, Seaborn, Plotly, etc., are used to draw word clouds, bar charts, line charts, etc.; sentiment analysis visualization shows the theme structure of the text through a theme distribution map;
[0061] Extract image features based on keywords and phrases using word vector models and dependency syntactic analysis techniques;
[0062] Construct a knowledge graph based on keywords and phrases using a knowledge graph and a deep learning model to obtain the semantic network of the preprocessed educational resource text; the semantic network is used to analyze the cultural background and historical information of the preprocessed educational resource text;
[0063] The text context identifies the emotional color and atmosphere image of the poem through sentiment analysis, thereby comprehensively and accurately identifying the core theme and rich images of the landscape poem. Such as natural landscapes, emotional expressions, and seasonal characteristics.
[0064] For example, the features extracted from the poem "All birds have flown away from thousands of mountains, and all human tracks have disappeared on ten thousand paths" are:
[0065] Theme: Mountains and paths;
[0066] Action: Static picture;
[0067] Context: Cold winter snow day;
[0068] Environment: Outdoor, mountains and big rivers;
[0069] Lighting: Soft;
[0070] Artist: Wu Zhen;
[0071] Style: Chinese ink painting on rice paper texture;
[0072] Digital painting: An art form of painting using digital devices;
[0073] Type: Landscape;
[0074] Color scheme: Soft;
[0075] Computer graphics: 3D Octane Render;
[0076] Season: Winter;
[0077] Weather: Heavy snowfall and extreme cold;
[0078] Scene: Countless mountains and paths, all covered by heavy snow;
[0079] Atmosphere: Silent, desolate, without any human tracks, as if the whole world has been frozen by ice and snow;
[0080] As a specific embodiment, the discrimination module specifically includes: constructing a prompt template from the corresponding materials obtained from the database, generating positive prompts in combination with the specific content of the poem, describing scenes that match the artistic conception of the poem, and generating negative prompts for comparison; then, through semantic similarity calculation ( where A and B respectively represent two words or sentences) and context consistency verification, ensure that the generated prompts highly match the poem; finally, optimize the model according to the effects and feedback to improve the quality and accuracy of the generated prompts, so as to generate prompts that match the poem context, help students better understand the artistic conception and emotions of the poem, and thus enhance their learning experience and understanding of landscape poems.
[0081] The text-to-image module in the text-to-image model specifically includes:
[0082] 1. Extract key elements from the preprocessed educational resource text; for example, perform word segmentation and part-of-speech tagging on the poem, understand the sentence structure through dependency parsing, and identify natural landscapes;
[0083] 2. Extract landscape keywords according to the key elements using word frequency statistics and TF-IDF method; TF-IDF(t, d) = TF(t, d) × IDF(t), where TF(t, d) is the frequency of word t in document d, and IDF(t) is the inverse document frequency of word t;
[0084] 3. Match landscape elements according to the landscape keywords in combination with the natural landscape vocabulary library; identify words directly describing colors and metaphorical color expressions;
[0085] 4. Determine the color meaning according to the text context in combination with context analysis;
[0086] 5. Use the sentiment analysis model to identify sentiment words and sentiment tendencies according to the text context, determine the sentiment intensity, and then determine the sentiment tone according to the sentiment intensity;
[0087] Specifically, use the sentiment analysis model to identify sentiment words and sentiment tendencies, extract topic words and analyze their associations with the atmosphere, comprehensively judge the poem atmosphere and finally identify words directly expressing emotions, calculate the sentiment intensity, and fuse the context sentiment information to determine the sentiment tone of the poem. Among them, the sentiment intensity S can be expressed as where S i is the intensity value of sentiment word i, and W i is the weight of sentiment word i;
[0088] 6. Generate a preliminary image according to the sentiment tone and text context using deep learning models such as generative adversarial networks or variational autoencoders;
[0089] To prevent errors in the images corresponding to text-to-image generation, the image-to-text module in this application uses the cognitive vision language model CogVLM to generate the text corresponding to each preliminary image based on the preliminary image; and uses an evaluation module to determine the matching degree between the text corresponding to each preliminary image and the corrected text; furthermore, the preliminary images are screened according to the matching degree to obtain the screened images until the images accurately reflect the artistic conception of the poem;
[0090] Finally, according to the screened images and the corrected text, a style transfer algorithm is used to obtain the image corresponding to the preprocessed educational resource text.
[0091] By using a style transfer algorithm to combine the artistic conception of the poem with a specific artistic style, an image with artistic beauty is generated, effectively conveying the deep meaning and artistic essence of the poem, and creating an image that matches the artistic conception of the poem. The image not only captures the intuitive picture of the poem but also conveys its inner emotions and atmosphere, such as peaceful mountains and waters, hazy fog, or the colors of seasonal changes. In this way, the generated image becomes an intuitive display of the artistic conception of the poem, helping students to understand the meaning of the poem more deeply.
[0092] The image-to-video model generates the corresponding video according to the image, text, and audio data corresponding to the preprocessed educational resource text by using video editing tools; the static image serves as the visual basis of the video, and then the corresponding background music and environmental sound effects are selected or generated according to the theme and emotion of the poem; for example, if the poem depicts peaceful mountains and waters, soft natural sounds such as the sound of a stream or bird calls may be added;
[0093] Specifically, technologies such as generative adversarial networks (GANs), diffusion models, or Transformer architectures are used to generate high-quality images and video frames, and the text-to-speech is determined according to the corrected text, and then the text-to-speech is converted into audio data by using a web crawler or a deep learning model. Finally, with the help of video editing tools (such as FFmpeg), the image sequence is combined with the audio to achieve multimodal fusion and generate complete video content.
[0094] Specifically, the audio data generation module is used to determine the text-to-speech according to the corrected text, and then the text-to-speech is converted into audio data by using a web crawler or a deep learning model;
[0095] The audio feature extraction module is used to extract features from the audio data to obtain audio features; the audio features include: frequency features, rhythm features, and emotional features;
[0096] Use the formula for feature extraction; where x(t) is the audio signal and x(f) is the spectrum;
[0097] The key frame extraction module is used to extract features from the image corresponding to the preprocessed educational resource text to obtain key frames;
[0098] The third output module is used to synchronize the audio features and key frames and generate the corresponding video using video editing tools.
[0099] Specifically, the Loopy model can be used to generate facial expressions and head movements; align the audio and images on the timeline and generate intermediate frames through interpolation algorithms to make the video actions smooth and natural; finally, optimize the model through multi-modal joint training to improve the understanding and generation ability of different modal data, so as to achieve precise synchronization of audio and images and create dynamic video content. Make the images change with the rhythm and emotion of the music, enhancing the coordination of vision and hearing;
[0100] The interpolation algorithm is as follows: Among them, f(x0) and f(x1) are the image features of the known key frames;
[0101] The image-to-video model also includes: a speech recognition module and an explanatory text generation module;
[0102] The speech recognition module is used to convert the audio in the video into text through speech recognition technology;
[0103] The explanatory text generation module is used to synchronize the text with the video timeline and format it, add subtitles or use large language models to understand the video and automatically generate explanatory text. Convert the text into explanatory speech through text-to-speech technology and synchronize it with the video, and at the same time automatically adjust the video clip according to the explanatory content, so as to explain the video and further explain the content and background of the poem, making the video a complete multi-sensory learning tool to help students understand landscape poems more comprehensively.
[0104] S104, combine the multi-modal resources with the educational resource text and use AI digital human technology for auxiliary teaching.
[0105] AI digital human technology provides an interactive learning partner for learners by simulating the image and voice of real people. In the study of landscape poetry, AI digital human technology is used to read poems aloud. Its speech synthesis technology can imitate the intonation and rhythm of real people. The speech synthesis technology generates natural and expressive speech by extracting the acoustic features of speech signals (such as fundamental frequency, energy, duration) and using deep learning models (such as WaveNet, Tacotron, FastSpeech, etc.). The latest technological advancements include few-shot speech cloning, enhanced expressiveness, multilingual support, and fine control of synthesized speech through natural language instructions, enabling the speech synthesis to imitate the intonation and rhythm of real people and vividly express the rhyme and emotion of poems. Speech synthesis technologies include: (1) Concatenative synthesis: Generating speech by splicing pre-recorded speech segments, with relatively high naturalness but poor fluency. (2) Parametric synthesis: Describing the spectral features of speech through an acoustic model and then driving a vocoder to synthesize speech. (3) Deep learning-driven synthesis: Using a neural network to directly map from text features to acoustic features to generate natural and expressive speech. And the latest speech synthesis technologies are: (1) Few-shot / zero-shot speech synthesis: Quickly cloning the voice of a speaker through a small number of speech samples. (2) Expressive speech synthesis: Enhancing the emotional expressiveness of synthesized speech and being able to automatically adjust intonation, rhythm, and stress according to the text content. (3) Multilingual speech synthesis: Supporting multiple languages and dialects and even achieving cross-lingual voice cloning. (4) Enhanced controllability: Achieving fine control of aspects such as the rhyme and emotion of synthesized speech through natural language descriptions or instructions.
[0106] AI digital human technology can also answer students' questions about poem content, author background, cultural significance, etc. according to a preset script or through machine learning algorithms. The preset script details questions and answers covering aspects such as poem content, author background, cultural significance, etc. When students ask questions, natural language processing technology is used to match the questions in the script and retrieve the answers, and at the same time, multi-modal resources are presented; the machine learning algorithm is trained by collecting a large amount of poem-related data. The model can understand the intention and keywords of students' questions, extract information from the data to generate accurate answers, and continuously learn and optimize with the interaction to provide personalized and accurate instant feedback and in-depth analysis, meeting the different learning needs of students and providing instant feedback and in-depth analysis. This interactive learning method not only enhances students' sense of participation but also improves their efficiency in understanding and memorizing poems, making the learning process more personalized and interesting.
[0107] After S104, it also includes:
[0108] S201, obtain the physiological data of the learner through a wearable electroencephalogram (EEG) sensor and a galvanic skin response / electrodermal activity (GSR / EDA) sensor; the physiological data includes: brain wave frequency bands and corresponding psychological and physiological characteristics and skin conductance fluctuations;
[0109] EEG is a medical test used to record the electrical activity of the brain. Several major brain wave frequency bands in EEG and their corresponding psychological and physiological characteristics include: Delta (1 - 3Hz), Theta (4 - 7Hz), Alpha (8 - 13Hz), Beta (14 - 30Hz), Gamma (31 - 49Hz); EEG captures the synchronous electrical signals of brain neurons through electrodes, and these signals are closely related to human thinking, emotional states, as well as attention, memory, and cognitive effort. Since the voltage fluctuations measured at the electrodes are very weak, the collected data needs to be digitized and transmitted to an amplifier for amplification, and finally displayed as a series of voltage changes. To facilitate efficient data collection, EEG electrodes are usually installed in elastic caps, meshes, or rigid grids, which can ensure consistent data collection from the same scalp positions between different sessions or subjects. Through EEG technology, information about the brain's electrical activity can be directly obtained to help analyze brain functions such as the learner's attention concentration, memory processing, and cognitive effort.
[0110] EEG signal processing is the core step in the development of a brain - computer interface (BCI) system, involving multiple links such as signal pre - processing, feature extraction and selection, and signal classification. Among them, the specific processing process is as follows:
[0111] (1) Signal pre - processing: The original data of EEG signals is usually accompanied by a large amount of noise, so it must undergo strict pre - processing to improve clarity. Common pre - processing steps include band - pass filtering, artifact removal, and signal normalization. For example, band - pass filtering can effectively remove low - frequency drift and high - frequency noise, and algorithms such as independent component analysis (ICA) can eliminate interference signals such as electrooculogram artifacts and muscle artifacts. In recent years, some advanced algorithms such as empirical mode decomposition (EMD) and noise - assisted signal decomposition (NSSD) have also been introduced to more efficiently separate the target components and artifacts in EEG signals.
[0112] (2) Feature extraction and selection: The complexity of EEG signals requires the extraction of features that can reflect the core information of brain activities. Time-frequency feature analysis methods, such as wavelet transform and short-time Fourier transform, can reveal the spectral characteristics of signals. In addition, spatio-temporal features extracted by methods such as common spatial patterns (CSP), by combining the temporal and spatial information of EEG signals, help to capture the electroencephalogram activity patterns more accurately. After feature extraction is completed, feature selection is an important step to reduce the data dimension and improve the performance of classifiers. Emerging feature selection methods, such as mutual information-based feature selection (MIFS) and Laplacian score method, are being increasingly applied to EEG signal analysis.
[0113] (3) Signal classification: Classifier design is the last link in the signal processing of brain-computer interfaces. Traditional machine learning algorithms such as support vector machine (SVM), random forest (RF), etc. have been widely applied to electroencephalogram signal classification. In recent years, the introduction of deep learning algorithms, especially convolutional neural network (CNN) and recurrent neural network (RNN), has brought revolutionary breakthroughs to EEG signal classification;
[0114] Galvanic skin response (GSR / EDA) refers to the fluctuations in skin conductance caused by changes in sweat gland activities, and it is a physiological indicator related to emotional arousal and stress levels. GSR technology focuses on the dynamic changes in skin conductance, which is closely related to an individual's sweating condition, and sweating is an intuitive physiological sign of emotional activation and stress response. The measurement mechanism of GSR is based on the sensitive changes in the skin's ability to conduct minute electric currents. When the human body encounters emotional activation, whether it is excitement, stress or tension, the sympathetic nervous system will operate more rapidly, prompting a significant increase in sweat secretion. The water and electrolyte components (such as salts) in sweat act together on the skin, enhancing its electrical conductivity and enabling the current to pass through the skin more smoothly. The number of sweat glands varies in different parts of the human body, but the hands and feet have the largest number of sweat glands (200 - 600 sweat glands per square centimeter), and GSR signals are usually collected from these parts. The calculation of GSR indicators usually focuses on its peak value, because the peak value can reflect an individual's immediate emotional response and physiological changes under specific stimuli or situations. Among them, the main indicators for GSR peak detection include: 1. Peak count: The number of peaks detected by the respondents in a specific stimulus or scenario. 2. Peaks / minute: The total number of peaks that appear in this stimulus or scenario divided by the duration of the stimulus or scenario, with the unit of peaks / minute. 3. Average peak amplitude: The average amplitude of all detected peaks in this stimulus or scenario.
[0115] The signal processing processes involved in the GSR peak detection include:
[0116] (1) Data acquisition: Use GSR / EDA sensor devices such as the PPB-Bio multi-channel physiological signal acquisition and analysis system, EmbracePlus, Shimmer3 GSR+Unit, etc. Ensure good contact between the electrodes and the skin during data acquisition, and use conductive gel to improve the signal quality.
[0117] (2) Preprocessing:
[0118] 1. Filtering: Use a low-pass filter to remove high-frequency noise. For example, use a Blackman filter with a cut-off frequency of 1 Hz.
[0119] 2. Downsampling: Reduce the sampling rate from a higher frequency (such as 1000 Hz) to a lower frequency (such as 250 Hz) to reduce the amount of data.
[0120] 3. Artifact removal: Remove motion artifacts through wavelet filters or Butterworth filters.
[0121] (3) Feature extraction:
[0122] 1. Time-domain features: Extract the mean, variance, range, signal magnitude area (SMA), etc. of the signal.
[0123] 2. Frequency-domain features: Use discrete wavelet transform (DWT) or stationary wavelet transform (SWT) to extract the spectral power of the signal.
[0124] 3. Mel-frequency cepstral coefficients (MFCC): Map the spectrum to the Mel scale through Mel filters and extract the cepstral coefficients
[0125] In this application, the data obtained through electroencephalogram (EEG) sensors and galvanic skin response (GSR) sensors can effectively evaluate cognitive load. EEG collects brain electrical activities through electrodes and extracts features such as band power, reflecting the activity state of the brain; GSR measures the skin conductance level (SCL) and skin conductance response (SCR) through electrodes, reflecting the activation degree of the autonomic nervous system.
[0126] S202, Determine the cognitive load of the learner according to the physiological data of the learner;
[0127] Using multimodal fusion and machine learning models, the cognitive load changes of learners when exposed to different teaching resources can be monitored more accurately in real time. Researchers can monitor the cognitive load changes of learners when exposed to different teaching resources such as traditional text materials, images, videos, or digital human-assisted learning in real time.
[0128] S203, Continuously optimize the multimodal teaching method according to the cognitive load.
[0129] Compare the differences in teaching effectiveness between traditional teaching materials and multimodal digital resources, especially their effectiveness in helping students understand and memorize landscape poetry. In the experiment, the students were divided into two groups, one using traditional materials and the other using multimodal digital resources. The researchers collected a variety of data including test scores, questionnaire feedback, engagement, and physiological indicators to evaluate the impact of the two methods on learning outcomes and explore the potential of multimodal resources in improving teaching effectiveness. Comprehensively evaluate the effects of multimodal digital resources and digital human-assisted learning, as well as the relationship between cognitive load and learning outcomes. Through experimental comparison, questionnaire surveys, physiological data measurements and other methods, analyze the application effects of AI-generated images and videos in teaching, the impact of digital human interaction on learning experience, and how different cognitive load levels affect learning efficiency and memory retention, thereby providing data support and optimization guidance for educational practice.
[0130] This application designs a multimodal digital resource framework that combines traditional education and generative AI technology to improve primary school students' understanding of landscape poetry and reduce cognitive load. The framework uses text generation, multimodal data fusion, physiological signal measurement, and digital human technology to generate images and videos, integrate different resources, monitor cognitive load, and provide interactive learning partners to improve the quality and effectiveness of teaching resources, providing a new method for AI applications in the field of education.
[0131] This application uses generative artificial intelligence to automatically extract and understand intent information from instructions, and generate content based on its knowledge and intent information. The content creation process is made more efficient and accessible, so that high-quality content can be generated at a faster rate, especially in the generation of text, images, audio, video and other content. This application converts abstract text information into intuitive images and dynamic videos through text-to-text, text-to-image, and image-to-video technology, which increases the interactivity and fun of learning. The multimodal learning method in this application can stimulate students' multiple senses, improve the efficiency of information absorption and memory, and make the learning content more vivid and easy to understand, thereby enhancing learning motivation and participation.
[0132] In an exemplary embodiment, a computer device is provided, which may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a multi-modal teaching method based on generative artificial intelligence.
[0133] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0134] In an exemplary embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0135] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0136] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0137] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0138] In the present application, all actions of obtaining signals, information, or data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining authorization from the owner of the corresponding device.
[0139] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0140] In this article, specific examples are used to illustrate the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A multi-modal teaching method based on generative artificial intelligence, characterized in that, The multi-modal teaching method based on generative artificial intelligence includes: Obtain educational resource texts; the educational resource texts include: landscape poems; Preprocess the educational resource texts; the preprocessing includes: data cleaning, text standardization, word segmentation, stop word removal, and text vectorization; For the preprocessed educational resource texts, use a multi-modal resource model constructed based on generative artificial intelligence to obtain multi-modal resources; the multi-modal resources include: texts, images, audio, and videos corresponding to the preprocessed educational resource texts; the multi-modal resource model includes: a text-to-text model, a text-to-image model, and an image-to-video model; the text-to-text model is used to output the text corresponding to the preprocessed educational resource texts based on natural language processing technology, text mining and machine learning technology, and deep learning models; the text-to-image model is used to output the image corresponding to the text based on deep learning models and image synthesis algorithms; the image-to-video model is used to generate the corresponding video by using video editing tools according to the images, texts, and audio data corresponding to the preprocessed educational resource texts; Combine the multi-modal resources with the educational resource texts and use AI digital human technology for auxiliary teaching.
2. The multimodal teaching method based on generative artificial intelligence according to claim 1, wherein, The text-to-text model includes: a feature extraction module, a text generation module, a discrimination module, and a first output module; The feature extraction module is used to extract content-based features and digital image-based features from the preprocessed educational resource texts by using natural language processing technology, text mining and machine learning technology; The text generation module is used to obtain the text corresponding to the preprocessed educational resource texts based on the content-based features and digital image-based features by using a deep learning model; The discrimination module is used to compare the text corresponding to the preprocessed educational resource texts with the corresponding materials obtained from the database to obtain a comparison result; and correct the text corresponding to the preprocessed educational resource texts according to the comparison result; The first output module is used to output the corrected text.
3. The multimodal teaching method based on generative artificial intelligence according to claim 2, wherein The text-to-image model includes: a text-to-image module, an image-to-text module, an evaluation module, a screening module, and a second output module; The text-to-image module is used to extract key elements from the corrected text, and use word frequency statistics and TF-IDF methods to extract landscape keywords; according to the landscape keywords, match landscape elements by combining with the natural landscape vocabulary library; according to the text context, determine the color meaning by combining with the context analysis; according to the text context, use the sentiment analysis model to identify sentiment words and sentiment tendencies, determine the sentiment intensity, and then determine the emotional tone according to the sentiment intensity; according to the emotional tone and text context, use a deep learning model to generate multiple preliminary images; The image-to-text module is used to generate the text corresponding to each preliminary image by using the cognitive vision language model CogVLM; The evaluation module is used to determine the matching degree between the text corresponding to each preliminary image and the corrected text; The screening module is used to screen the preliminary images according to the matching degree to obtain the screened images; The second output module is used to obtain the image corresponding to the preprocessed educational resource text by using a style transfer algorithm based on the filtered image and the corrected text.
4. The multimodal teaching method based on generative artificial intelligence according to claim 3, wherein The image-to-video model includes: an audio data conversion module, an audio feature extraction module, a key frame extraction module, and a third output module; The audio data generation module is used to determine the text speech according to the corrected text, and then convert the text speech into audio data by using a web crawler or a deep learning model; The audio feature extraction module is used to extract features from the audio data to obtain audio features; the audio features include: frequency features, rhythm features, and emotional features; The key frame extraction module is used to extract features from the image corresponding to the preprocessed educational resource text to obtain key frames; The third output module is used to synchronize the audio features and the key frames, and generate the corresponding video by using a video editing tool.
5. The multimodal teaching method based on generative artificial intelligence according to claim 4, wherein The image-to-video model further includes: a speech recognition module and an explanatory caption generation module; The speech recognition module is used to convert the audio in the video into text through speech recognition technology; The explanatory caption generation module is used to synchronize the text with the video timeline and format it, add subtitles, or use a large language model to understand the video and automatically generate explanatory captions.
6. The multimodal teaching method based on generative artificial intelligence according to claim 1, wherein Combining multi-modal resources with the educational resource text and using AI digital human technology for assisted teaching, and then further including: Obtaining the physiological data of the learner through a wearable electroencephalogram sensor and a galvanic skin sensor; the physiological data includes: brain wave frequency bands and corresponding psychological and physiological characteristics and galvanic skin conductance fluctuations; Determining the cognitive load of the learner according to the physiological data of the learner; Continuously optimizing the multi-modal teaching method according to the cognitive load.
7. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the multi-modal teaching method based on generative artificial intelligence according to any one of claims 1-6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multi-modal teaching method based on generative artificial intelligence according to any one of claims 1-6.
9. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the multi-modal teaching method based on generative artificial intelligence according to any one of claims 1-6.