Gamification optimization method and device for online learning, equipment and storage medium

By collecting and processing learners' multimodal data in real time and using time transformers and fusion network models to select appropriate educational games, the problem of lagging teaching rhythm adjustment in online learning is solved, and personalized learning experience is achieved and learning effects are improved.

CN120673636APending Publication Date: 2025-09-19HUBEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511007884.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing online learning methods find it difficult to capture students' non-verbal interactive feedback such as expressions and movements in real time, and are unable to adopt appropriate interactive methods to stimulate learners' learning interest based on their learning status, resulting in delayed adjustment of teaching rhythm and inability to provide a personalized learning experience.

Method used

By collecting learners' multimodal learning state data in real time, using time transformer and fusion network model to process learners' dynamic and static features, calculating the cosine similarity between learning state feature vector and educational game feature vector, and selecting appropriate educational games for optimization.

Benefits of technology

It achieves dynamic adjustment of teaching strategies based on learners’ real-time feedback, improves learners’ online learning effects and learning interests, and provides personalized learning experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673636A_ABST
    Figure CN120673636A_ABST
Patent Text Reader

Abstract

The invention discloses a gamification optimization method and device for online learning, equipment and a storage medium, and relates to the technical field of online learning, and the method comprises the steps: extracting learner dynamic features and learner static features from learner multi-modal learning state data collected in real time; performing time information processing on the learner dynamic characteristics and the learner static characteristics by using a time converter, and inputting the processed information to a fusion network model to determine a learning state characteristic vector of the learner dynamically changing along with time; processing the text information and the image information of the educational game through a language image pre-training model, a text encoder and an image encoder to obtain an educational game feature vector; and calculating the cosine similarity between the learning state feature vector and the educational game feature vector, and selecting an educational game based on the cosine similarity to optimize online learning. The state of the learner is analyzed, learning state features are determined, the educational games are individually selected according to the educational game information, and the teaching strategy is dynamically adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of online learning technology, and more specifically, to a gamification optimization method, apparatus, device, and storage medium for online learning. Background Art

[0002] With the rapid development of information technology, traditional teaching models are gradually failing to meet the personalized and interactive demands of modern education. As a product of the integration of information technology and education, online learning has become a key form of education worldwide, driven by technological advancements and growing demand. Mathematics, a subject with a strong logic component, often feels dull and tedious to learners, resulting in low classroom participation and poor learning outcomes.

[0003] Currently, online learning technologies have formed a relatively complete ecosystem, covering multiple aspects such as teaching resource delivery, real-time interaction, learning process management, and effectiveness evaluation. Large language models have further advanced online learning. With their powerful natural language understanding and generation capabilities, large language models can be used to analyze learners' language input and identify their emotions and understanding.

[0004] However, existing online learning methods have difficulty capturing students' non-verbal interactive feedback such as expressions and movements in real time, cannot capture and analyze learners' emotional state and participation in real time, and cannot adopt appropriate interactive methods to stimulate learners' learning interest based on the learning status of different learners. It is impossible to dynamically adjust teaching strategies and provide personalized learning experiences, resulting in a lag in the adjustment of teaching rhythm. Summary of the Invention

[0005] In response to at least one defect or improvement need in the prior art, the present invention provides a gamification optimization method, device, equipment and storage medium for online learning, which is used to solve the problems that the online learning methods in the prior art cannot capture and analyze the learners' emotional state and participation in real time, it is difficult to adopt appropriate interactive methods according to the learning status of different learners to stimulate learners' learning interest, and it is impossible to dynamically adjust teaching strategies and provide a personalized learning experience, resulting in a lag in the adjustment of teaching rhythm.

[0006] To achieve the above objectives, according to a first aspect of the present invention, a gamification optimization method for online learning is provided, comprising: Extract learner dynamic features and learner static features from learner multimodal learning state data collected in real time; The learner's dynamic features and static features are processed by the time transformer and input into the fusion network model to determine the learner's learning state feature vector that changes dynamically over time. The text information and image information of the educational game are processed through the language image pre-training model, text encoder and image encoder to obtain the educational game feature vector; The cosine similarity between the learning state feature vector and the educational game feature vector is calculated, and educational games are selected based on the cosine similarity to optimize online learning.

[0007] In a possible implementation, extracting learner dynamic features and learner static features from learner multimodal learning state data collected in real time also includes: Real-time collection of learners' facial expression information, voice information and body movement information in the classroom environment; The residual network and global convolutional attention module are used to process facial expression information, voice information and body movement information to extract learner dynamic features and learner static features.

[0008] In a possible implementation, a time transformer is used to process the learner's dynamic features and the learner's static features for time information processing, and the processed features are input into a fusion network model to determine a learning state feature vector of the learner that changes dynamically over time. The method also includes: According to classroom teaching, the time position embedding function of the time transformer is set to process the time information of the learner's dynamic characteristics and the learner's static characteristics; The fusion network model is used to classify the learner's expression according to the learner's static features, and the classification results are adjusted by the intensity-aware loss function to obtain the learner's expression state; Analyze learner dynamic characteristics based on learner historical behavior data, learner dynamic characteristics and fusion network model to determine learner behavior status; The learner's expression state and behavior state are fused to obtain the learner's learning state feature vector that changes dynamically over time.

[0009] In one possible implementation, a fusion network model is used to classify the learner's expression based on the learner's static features, and the classification results are adjusted using an intensity-aware loss function to obtain the learner's expression state, which also includes: According to the preset judgment rules, the learner's static features are preliminarily classified to obtain the initial expression information; The mapping relationship between expression intensity and learner's cognitive state is established through learner's historical expression data, and the initial expression information is calibrated; The learner's expression state is obtained by assigning weights through the intensity-aware loss function and adjusting the calibrated initial expression information.

[0010] In one possible implementation, analyzing the learner's dynamic characteristics based on the learner's historical behavior data, the learner's dynamic characteristics, and the fusion network model to determine the learner's behavior state further includes: Utilize the fusion network model to analyze the learner's historical behavior data and learner's dynamic characteristics to determine the learner's historical behavior vector and current behavior vector; Calculate the learner's behavioral deviation degree based on the historical behavior vector and the current behavior vector; Determine the learner's behavioral status based on the degree of behavioral deviation and preset behavioral rules.

[0011] In one possible implementation, the text information and image information of the educational game are processed by the language image pre-training model, the text encoder, and the image encoder to obtain the educational game feature vector, further comprising: Obtaining a text encoding feature vector, a text model feature vector, an image encoding feature vector, and an image model feature vector from the text information and image information of the educational game respectively through a language image pre-training model, a text encoder, and an image encoder; The text encoding feature vector and the text model feature vector are fused to obtain a text fusion feature vector, and the image encoding feature vector and the image model feature vector are fused to obtain an image fusion feature vector; The text fusion feature vector and the image fusion feature vector are spliced ​​and reconstructed to obtain the educational game fusion feature vector.

[0012] In a possible implementation, the text fusion feature vector and the image fusion feature vector are concatenated and reconstructed to obtain the educational game fusion feature vector, further comprising: The text fusion feature vector and the image fusion feature vector are concatenated and passed through a fully connected layer to obtain the initial fusion feature; A deep autoencoder is used to compress and reconstruct the initial fusion features to obtain the educational game fusion feature vector.

[0013] According to a second aspect of the present invention, there is also provided a gamification optimization device for online learning, characterized by comprising: A learning feature extraction module configured to extract learner dynamic features and learner static features from learner multimodal learning state data collected in real time; A learning feature fusion module is configured to process the learner's dynamic features and the learner's static features using a time transformer, and input the processed information into a fusion network model to determine a learning state feature vector of the learner that changes dynamically over time; a game feature extraction module configured to process text information and image information of the educational game using a language image pre-training model, a text encoder, and an image encoder to obtain an educational game feature vector; The game optimization module is configured to calculate the cosine similarity between the learning state feature vector and the educational game feature vector, and select an educational game based on the cosine similarity to optimize the online learning.

[0014] According to a third aspect of the present invention, a gamification optimization device for online learning is also provided, which includes at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the computer program is executed by the processing unit, the processing unit performs the steps of any one of the above-mentioned gamification optimization methods for online learning.

[0015] According to a fourth aspect of the present invention, a storage medium is also provided, which stores a computer program that can be executed by a gamification optimization device for online learning. When the computer program runs on the gamification optimization device for online learning, the gamification optimization device for online learning executes the steps of any one of the above-mentioned gamification optimization methods for online learning.

[0016] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art: The present invention provides a gamified optimization method for online learning. This method collects learners' multimodal learning state data in real time and extracts dynamic and static features from it to capture and analyze learners' emotional states and engagement in real time. A time transformer is used to process learner features for temporal information, and this information is input into a fusion network model to determine the learner's learning state feature vector, which changes dynamically over time. This helps to more accurately characterize the learner's learning state. A language image pre-training model, a text encoder, and an image encoder are used to process the text and image information of educational games to obtain educational game feature vectors. This method can fully exploit the effective information in educational games and understand the characteristics of different educational games. The cosine similarity between the learning state feature vector and the educational game feature vector is calculated, and educational games are selected based on this cosine similarity to optimize online learning. This method can select educational games that match the learner's learning state, thereby adopting appropriate interactive methods to stimulate learners' learning interest and provide a personalized learning experience. Teaching strategies can also be dynamically adjusted based on learners' real-time feedback, improving learners' online learning outcomes. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 A schematic diagram of a flow chart of an embodiment of the gamification optimization method for online learning provided by the present invention; Figure 2 A schematic diagram of a flow chart of an embodiment of obtaining online learning features provided by the present invention; Figure 3 The present invention provides Figure 1 A flow chart of an embodiment of step S102; Figure 4 The present invention provides Figure 3 A flow chart of an embodiment of step S302; Figure 5 The present invention provides Figure 3 A flow chart of an embodiment of step S303; Figure 6 The present invention provides Figure 1 A flow chart of an embodiment of step S103; Figure 7 A schematic structural diagram of an embodiment of the gamification optimization device for online learning provided by the present invention; Figure 8 A schematic diagram of the structure of a gamification optimization device for online learning provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0019] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0020] The terms "first," "second," "third," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0021] The present invention provides a gamification optimization method, device, equipment and storage medium for online learning, which are described below.

[0022] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of a gamification optimization method for online learning provided by the present invention. In a specific embodiment of the present invention, a gamification optimization method for online learning is disclosed, including: S101, extracting learner dynamic features and learner static features from learner multimodal learning state data collected in real time; S102, using a time transformer to process the learner's dynamic features and the learner's static features, and inputting the information into a fusion network model to determine a learning state feature vector of the learner that changes dynamically over time; S103, processing the text information and image information of the educational game through the language image pre-training model, the text encoder, and the image encoder to obtain an educational game feature vector; S104: Calculate the cosine similarity between the learning state feature vector and the educational game feature vector, and select an educational game based on the cosine similarity to optimize the online learning.

[0023] In the above embodiment, multimodal learning status data of learners during the learning process is collected in real time, covering facial expressions, voice information, and body movements. The collection of multimodal data can comprehensively and meticulously reflect the learner's learning status, providing a comprehensive data foundation for subsequent analysis. Separating and extracting dynamic and static features from the collected multimodal data can provide a deeper understanding of the learner's learning status and provide a basis for subsequent personalized optimization. Dynamic features mainly reflect the real-time changes of learners during the learning process, while static features focus on describing the learner's basic attributes and long-term stable learning habits.

[0024] The time transformer is used to process the extracted dynamic features and static features of the learner, which can capture the complex dependencies and changing patterns of feature data in the time dimension. Through the self-attention mechanism, the features at different time points are weighted and fused, so that the model can pay attention to the historical feature information that has a greater impact on the current learning state.

[0025] The fusion network model utilizes a multi-layered neural network structure, deeply integrating feature information from different sources and types through nonlinear transformations and feature combinations. During this fusion process, the model automatically learns the connections and importance between different features, generating a feature vector that comprehensively reflects the learner's dynamic learning state over time. This vector not only encompasses the learner's current state but also encompasses the historical evolution of their learning state, providing a precise basis for subsequent educational game selection.

[0026] Using language-image pre-training models (such as CLIP) pre-trained on large-scale language and image datasets, the model is fed with the text and image information from educational games. It automatically learns the semantic associations between text and images and maps them into a unified feature space. This approach fully exploits the rich semantic information in educational games.

[0027] Based on the pre-trained model, specialized text encoders and image encoders are used to process the text and image information of the educational game. The text encoder encodes the text, extracting information such as keywords, semantic structure, and sentiment, and converting it into a high-dimensional text feature vector. The image encoder extracts features from the image, capturing visual features such as color, shape, and texture, and generating an image feature vector. The collaborative work of the text encoder and image encoder allows for comprehensive and accurate extraction of feature information for the educational game, ultimately yielding a feature vector that comprehensively reflects the game's content and style.

[0028] Calculate the cosine similarity between the learning state feature vector and the educational game feature vector. Cosine similarity is a metric that measures the degree of directional similarity between two vectors in vector space, with a value range of -1 to 1. By calculating cosine similarity, we can quantitatively assess the degree of fit between the learner's current learning state and each educational game.

[0029] Based on the calculated cosine similarity, the educational game with the highest cosine similarity to the learner's learning state feature vector is selected to optimize online learning. This ensures that the selected educational game is highly compatible with the learner's current learning state, thereby better stimulating learner interest and engagement. By dynamically selecting appropriate educational games, personalized online learning optimization is achieved, improving learning outcomes and experience.

[0030] Compared to existing technologies, this embodiment provides a gamified optimization method for online learning. By collecting learners' multimodal learning state data in real time and extracting dynamic and static features from it, this method captures and analyzes learners' emotional states and engagement in real time. A time transformer is used to process learner features for temporal information, and this information is fed into a fusion network model to determine the learner's learning state feature vector, which changes dynamically over time. This helps more accurately characterize the learner's learning state. A language image pre-training model, a text encoder, and an image encoder are used to process the text and image information of educational games to generate educational game feature vectors. This fully exploits the effective information in educational games and understands the characteristics of different educational games. The cosine similarity between the learning state feature vector and the educational game feature vector is calculated, and educational games are selected based on this cosine similarity to optimize online learning. This method can select educational games that match the learner's learning state, thereby adopting appropriate interactive methods to stimulate learners' learning interest and provide a personalized learning experience. Teaching strategies can also be dynamically adjusted based on learners' real-time feedback, improving learners' online learning outcomes.

[0031] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of obtaining online learning features provided by the present invention. In some embodiments of the present invention, extracting learner dynamic features and learner static features from learner multimodal learning state data collected in real time also includes: Real-time collection of learners' facial expression information, voice information and body movement information in the classroom environment; The residual network and global convolutional attention module are used to process facial expression information, voice information and body movement information to extract learner dynamic features and learner static features.

[0032] In the above embodiment, in a mathematics classroom environment, devices such as cameras, microphones, and small sensors are used to collect learners' multimodal data, including facial expressions, voice information, and body movements, to provide a comprehensive data basis for subsequent analysis.

[0033] Use high-resolution (1080p), high-frame-rate (60fps) cameras installed at different locations in the classroom (such as the four corners and the middle of the front and back walls) to ensure that the facial expressions and body movements of each learner can be captured clearly and in all directions. The image data collected by the camera is ,in is the pixel coordinate of the image. During the acquisition process, the image is preprocessed in real time, including adjusting brightness, contrast, color balance, etc., to improve image quality and facilitate subsequent feature extraction.

[0034] A key frame extraction algorithm based on motion detection is used. When obvious movements of the learner (such as raising hands, standing, etc.) are detected, the key frame acquisition frequency is increased to reduce data redundancy, and finally a set of images with obvious features is obtained after processing.

[0035] Multiple high-sensitivity (e.g., sensitivity of 38dBV / Pa or higher), omnidirectional microphones should be reasonably arranged in the center of the classroom ceiling and on the front and back walls. The sampling rate should be set to 48kHz or higher and the bit depth should be 24 bits or higher to ensure that the voice information of each learner can be accurately captured.

[0036] Advanced adaptive beamforming, noise suppression, and echo cancellation algorithms are used to process the collected audio signals, effectively removing environmental noise (such as traffic noise outside the classroom, air conditioning operation noise, etc.) and echo interference.

[0037] According to the frequency characteristics of the speech, dynamic equalization technology is used to adjust the gain of the frequency band to highlight the key information in the learner's speech (such as speech content, intonation changes, etc.), thereby obtaining a clear learner interactive speech data set .

[0038] Equip learners with small inertial measurement unit (IMU) sensors that can be worn or mounted on tables and chairs. These sensors can accurately measure parameters such as the angle, velocity, and acceleration of limb joints to obtain limb movement data. .

[0039] After data acquisition, Kalman filtering and sliding average filtering are used to smooth the data and remove abnormal data caused by sensor jitter, signal interference, etc. A time synchronization algorithm is used to accurately align the data collected by different sensors on the time axis to ensure the temporal consistency of body movement data, facial expressions, and voice information.

[0040] Use residual network (such as ResNet18) and global convolutional attention module (GFCA) to collect 、 and Data is processed to extract important features and avoid overfitting.

[0041] When using the residual network (ResNet18), the network structure is optimized and a batch normalization layer is added to each residual block to accelerate network convergence and improve the generalization ability of the model.

[0042] The residual network formula is: ,in is the convolution operation, For input, is the output. Through this formula, the network can learn the input data The difference after the convolution operation , effectively extracting deep features and avoiding gradient vanishing or gradient exploding problems. For convolution operations , convolution kernels of different sizes are used for parallel convolution, and then the results are fused, so that features of different scales can be extracted.

[0043] In the Global Convolutional Attention (GFCA) module, an adaptive weight calculation method is designed to dynamically adjust attention weights based on the feature distribution of the input data. In addition to attention calculations in the spatial dimension, an attention mechanism is also introduced in the temporal dimension. For consecutive frames of data, attention weights are dynamically assigned based on temporal changes, thereby better capturing the evolution of features over time.

[0044] Attention mechanism formula: ,in This formula allows the model to focus on key information in the data and improve the effectiveness of feature extraction.

[0045] In online learning scenarios, a residual network performs basic processing on facial expressions, speech, and body movements, extracting the fundamental, deep, and sequential features of each modality. This processing is then enhanced using a global convolutional attention module, focusing on global information, focusing on key features, and understanding dynamic trends and semantic associations. Multimodal dynamic features are then fused and analyzed along the temporal dimension to extract dynamic features. Simultaneously, multimodal static features are integrated and filtered and optimized, ultimately accurately capturing both dynamic and static features of the learner, providing support for personalized learning optimization.

[0046] See also Figure 3 , Figure 3 The present invention provides Figure 1 FIG. 1 is a flow chart of an embodiment of step S102 in FIG. 1 . In some embodiments of the present invention, a time transformer is used to process the learner's dynamic features and the learner's static features for time information, and the processing is input into a fusion network model to determine the learner's learning state feature vector that changes dynamically over time. The process also includes: S301, setting a time position embedding function of a time transformer according to classroom teaching, and performing time information processing on learner dynamic features and learner static features; S302, using the fusion network model to classify the learner's expression according to the learner's static features, and adjusting the classification results through the intensity perception loss function to obtain the learner's expression state; S303, analyzing the learner's dynamic characteristics based on the learner's historical behavior data, the learner's dynamic characteristics, and the fusion network model to determine the learner's behavior state; S304: Fusing the learner's facial expression state and the learner's behavioral state to obtain a learning state feature vector that dynamically changes over time.

[0047] In the above embodiment, different classroom teaching methods have unique time rhythms and content arrangements. For example, a math class may include multiple knowledge point explanations and practice sessions, each with varying durations and importance. Meanwhile, a language class may prioritize interactive communication, with relatively flexible time allocation. Therefore, the time position embedding function of the time transformer is designed based on the specific classroom setting, fully considering the temporal structure of the classroom teaching and closely linking the learner's dynamic and static characteristics with the temporal dimension of the classroom teaching.

[0048] The temporal position embedding function encodes the temporal information of the learner's features, mapping each feature to a specific time point or time period. This allows the Time Transformer to understand the temporal sequence and relative position of the features. In this way, the Time Transformer can accurately process the temporal information of the learner's features and capture how the features change over time.

[0049] A fusion network model is used to classify learners' expressions based on their static features. These features influence their facial expressions during learning. By learning the correlation between these static features and facial expressions, the fusion network model can more accurately classify learners' expressions into different categories, such as happiness, sadness, anger, and confusion.

[0050] To more accurately describe the learner's facial expressions, we introduce an intensity-aware loss function to adjust the classification results. This loss function not only considers the category of the expression but also its intensity. By fine-tuning the classification results, the intensity-aware loss function enables the model to more accurately quantify the intensity of the expression, thereby obtaining a more detailed representation of the learner's facial expressions and more realistically reflecting the learner's emotional experience during the learning process.

[0051] Historical behavior data can reflect a learner's long-term learning patterns and behavioral trends, providing context for understanding their current learning status. For example, if a learner has frequently struggled with a particular knowledge point in their past studies, their dynamic characteristics when currently learning that knowledge point may be more easily identified as a behavioral state of difficulty.

[0052] The fusion network model comprehensively considers multiple dynamic features, such as body movements and learning operations, and combines them with historical behavioral data to determine the learner's behavioral state. This accurately captures the learner's behavioral performance during the learning process, providing an important basis for learning state analysis. By using an appropriate fusion strategy to integrate information from the two states and assigning different weights based on the degree of influence of facial expression and behavioral state on the learning state, the fused feature vector can more comprehensively reflect the learner's learning state.

[0053] The learning state feature vector, which dynamically changes over time, not only captures the learner's facial expressions and behaviors at different points in time but also reflects the changing trends of this information over time. For example, at different stages of a lesson, a learner's facial expressions may shift from initial anticipation to confusion and finally to understanding. Their behaviors may also shift from active participation to inattention and finally back to focus. The learning state feature vector accurately captures these changes, providing dynamic and comprehensive information support for personalized optimization of online learning, helping teachers or learning systems to timely adjust teaching strategies and improve learning outcomes.

[0054] See also Figure 4 , Figure 4 The present invention provides Figure 3 In some embodiments of the present invention, the process of classifying the learner's expression based on the learner's static features using a fusion network model and adjusting the classification results using an intensity perception loss function to obtain the learner's expression state further includes: S401, preliminarily classifying the learner's static features according to preset judgment rules to obtain initial expression information; S402: Establishing a mapping relationship between expression intensity and learner cognitive state through the learner's historical expression data, and calibrating the initial expression information; S403 , allocating weights through the intensity perception loss function, and adjusting the calibrated initial expression information to obtain the learner's expression state.

[0055] In the above embodiment, based on relevant theories of psychology and education, the potential correlations between different static features and facial expressions were analyzed to develop preset judgment rules. For example, introverted learners may be more likely to express nervousness and anxiety when faced with new knowledge or difficult problems; while extroverted learners may be relatively good at concealing their emotions, but their expressions may be more obvious when they are excited or interested. At the same time, combined with a large number of actual teaching cases and observation data, the common types of expressions learners experience under different combinations of static features were summarized, thus developing targeted preset judgment rules.

[0056] A feature matching algorithm compares a learner's personality, academic performance, learning style, and other characteristics with the feature thresholds or patterns specified in the rules to determine the learner's category. Based on the matching results, an initial expression is generated for the learner, providing a rough estimate of the learner's likely expression, such as anxiety, curiosity, or confidence. Because this initial expression is based solely on static features, it may contain some errors and requires further optimization.

[0057] The collected historical expression data are annotated and divided into different expression categories, such as happiness, sadness, anger, confusion, etc. At the same time, combined with the learner's cognitive state assessment results, the expression data is associated with the cognitive state.

[0058] The intensity of facial expressions is quantified, converting facial muscle movements into specific numerical indicators. Furthermore, a quantitative assessment of the learner's cognitive state is performed, such as scores on knowledge tests and the efficiency of learning task completion. Using machine learning algorithms such as regression analysis and neural networks, a mapping model between expression intensity and cognitive state is constructed, and the learner's cognitive state level is predicted based on the numerical expression intensity. The expression type in the initial expression information is mapped to the expression category in the historical expression data. Then, based on the intensity distribution of the expression type under different cognitive states, the initial expression information is assigned a corresponding intensity value.

[0059] Unlike traditional loss functions, the intensity-aware loss function focuses not only on the correctness of the expression classification but also on the accuracy of the predicted expression intensity. Based on the correlation between expression intensity and cognitive state, different weights are assigned to expressions of different intensities. Expression intensities that are closely related to cognitive state are given higher weights, placing greater emphasis on the accuracy of the predictions of these intensities during the optimization process.

[0060] The fusion network model is trained and optimized using an intensity-aware loss function. During training, the model continuously adjusts parameters based on feedback from the loss function, resulting in more accurate predictions of expression intensity and category. By combining the learner's current learning context and dynamic characteristics, and taking into account the category and intensity of the expression, the final learner's expression state is determined. This adjusted expression state more accurately reflects the learner's true emotional and cognitive state during the learning process, providing strong support for personalized instruction.

[0061] See also Figure 5 , Figure 5 The present invention provides Figure 3 FIG. 1 is a flow chart of an embodiment of step S303 in FIG. 1 . In some embodiments of the present invention, analyzing the learner's dynamic characteristics based on the learner's historical behavior data, the learner's dynamic characteristics, and the fusion network model to determine the learner's behavior state further includes: S501, using a fusion network model to analyze the learner's historical behavior data and the learner's dynamic characteristics to determine the learner's historical behavior vector and current behavior vector; S502, calculating the learner's behavior deviation degree based on the historical behavior vector and the current behavior vector; S503. Determine the learner's behavior state based on the degree of behavior deviation and preset behavior rules.

[0062] In the above embodiment, when analyzing learner behavior patterns, in addition to multimodal data and time series analysis, the learner's historical behavior data is also introduced. A long-term memory network (LSTM) is established to store and analyze the learner's behavior change trends in multiple classes. Then, the extracted features are encoded and converted into numerical vector form to obtain the learner's historical behavior vector and current behavior vector.

[0063] Compare the historical behavior vector with the current behavior vector. For example, for a learner, let the behavior vector of the current lesson be , the historical behavior pattern vector obtained through LSTM analysis is , degree of behavioral deviation ; like , it is judged as mild abnormality (such as occasional distraction but able to return to attention quickly), is the mild abnormality threshold; if , it is judged as moderate abnormality (such as being in a daze for a long time but still participating in some learning), where is the moderate abnormality threshold; if , it is considered a severe abnormality (e.g., almost no effective learning behavior throughout the entire class) and a warning is issued. By quantifying and evaluating the degree of behavioral deviation, abnormal behavior of learners can be detected in a timely manner, providing a basis for subsequent behavioral status judgment.

[0064] Once the learner's behavior status is determined, timely feedback is provided to the learner, teacher, or learning system. For normal behavior, affirmation and encouragement can be given; for mildly abnormal behavior, the learner can be reminded to adjust their learning state; for severe abnormal behavior, timely intervention measures can be taken, such as teacher intervention, adjustment of learning task difficulty, or personalized learning suggestions.

[0065] See also Figure 6 , Figure 6 The present invention provides Figure 1 FIG. 1 is a flow chart of an embodiment of step S103 in FIG. 1 . In some embodiments of the present invention, processing text information and image information of an educational game using a language image pre-training model, a text encoder, and an image encoder to obtain an educational game feature vector further includes: S601, obtaining a text encoding feature vector, a text model feature vector, an image encoding feature vector, and an image model feature vector from the text information and image information of the educational game respectively through a language image pre-training model, a text encoder, and an image encoder; S602: Fusing the text encoding feature vector and the text model feature vector to obtain a text fusion feature vector, and fusing the image encoding feature vector and the image model feature vector to obtain an image fusion feature vector; S603: Concatenate and reconstruct the text fusion feature vector and the image fusion feature vector to obtain an educational game fusion feature vector.

[0066] In the above embodiment, in educational game scenarios, a language-image pre-training model (such as the CLIP model) can simultaneously process text and image information, capturing potential connections between them. Through methods such as contrastive learning, the language-image pre-training model learns the semantic correspondence between text and image, thereby extracting text model feature vectors related to the image from the text information and image model feature vectors related to the text from the image information.

[0067] Text encoders can perform operations such as word segmentation, part-of-speech tagging, and syntactic analysis on text to understand the semantic and grammatical structure of the text. In educational games, text information can include, but is not limited to, game mission instructions, plot dialogues, and knowledge explanations. Through multi-layer neural network training, text encoders convert this text information into high-dimensional text encoding feature vectors that accurately represent the text's semantic content, sentiment, and thematic information.

[0068] The image encoder is responsible for feature extraction from educational game images, capturing visual features such as color, texture, shape, and object position. Educational game images can include, but are not limited to, game characters, props, and scenes. Through convolutional neural network operations such as convolution and pooling, the image encoder converts image information into an image encoding feature vector, effectively representing the image's visual information.

[0069] When processing the text and image information in educational games, the language-image pre-trained model, the text encoder, and the image encoder work in parallel. For each piece of text, the text encoder generates a text encoding feature vector, and the language-image pre-trained model generates a text model feature vector. For each image, the image encoder generates an image encoding feature vector, and the language-image pre-trained model generates an image model feature vector. These are represented as numerical matrices that contain rich semantic and visual information.

[0070] When fusing the text encoding feature vector and the text model feature vector, appropriate preprocessing is required, considering that the text encoding feature vector and the text model feature vector may have different dimensions and semantic emphases. To optimize the fusion effect, a cross-validation method can be used to adjust the parameters of the fusion process, such as weight values ​​and attention mechanism parameters, so that the fused text fusion feature vector can more accurately represent the semantic information of the educational game text.

[0071] For the image coding feature vector and the image model feature vector, the same fusion strategy is used for fusion. Since image information has multi-level characteristics, such as low-level texture and color features and high-level semantic features, a multi-level fusion method can be used.

[0072] During the fusion process, evaluation metrics are used to measure the fusion effect. Metrics such as image classification accuracy and object detection recall can be used to assess the performance of the fused image feature vector in image-related tasks. Based on the evaluation results, the fusion strategy and parameters are adjusted to improve the quality and expressiveness of the fused image feature vector.

[0073] In some embodiments of the present invention, the text fusion feature vector and the image fusion feature vector are concatenated and reconstructed to obtain the educational game fusion feature vector, further comprising: The text fusion feature vector and the image fusion feature vector are concatenated and passed through a fully connected layer to obtain the initial fusion feature; A deep autoencoder is used to compress and reconstruct the initial fusion features to obtain the educational game fusion feature vector.

[0074] In the above embodiment, since different model architectures may be used in the text and image feature extraction process, resulting in differences in the feature vector dimensions, before splicing the text fusion feature vector and the image fusion feature vector, it is necessary to ensure that the dimensions of the two are adaptable, and scale the values ​​of each feature dimension to a uniform range (such as [0, 1] or [-1, 1]) to eliminate the influence of different feature dimensions and value ranges, so that the spliced ​​feature vectors are more comparable and stable.

[0075] The order of concatenation affects the emphasis placed on text and image information in the initial fused features. This order can be determined based on the specific application scenario and task requirements of the educational game. A simple sequential concatenation method can be used to concatenate two feature vectors sequentially by dimension to form a longer feature vector. Alternatively, an attention mechanism can be introduced to dynamically adjust the concatenation weights based on the correlation between text and image features, ensuring that important features occupy a larger proportion of the concatenated vector.

[0076] The concatenated feature vectors are input into the fully connected layer, where the input features are transformed and integrated using linear transformations and nonlinear activation functions (such as ReLU and Sigmoid). During training, the weights and biases of the fully connected layer are optimized using backpropagation algorithms and optimizers (such as Adam and SGD). Parameters are adjusted by minimizing loss functions (such as cross-entropy and mean squared error) to enable the fully connected layer to learn effective fusion patterns between text and image features and generate representative initial fusion features.

[0077] A deep autoencoder consists of two parts: an encoder and a decoder. The encoder compresses the initial fused features and extracts their core feature representations; the decoder reconstructs the original features based on the compressed features output by the encoder. For one-dimensional vector data such as the initial fused features, an autoencoder with an MLP structure can be used. The encoder consists of multiple fully connected layers, with each layer gradually reducing the number of neurons to achieve dimensionality reduction and compression of the input features. The decoder, symmetrical to the encoder, consists of multiple fully connected layers, with each layer gradually increasing the number of neurons to gradually restore the dimensionality of the original features.

[0078] When the initial fused features are input into the encoder, the neurons in each layer process the input features through a linear transformation and an activation function. The linear transformation multiplies the input features by a weight matrix and adds a bias term, while the activation function performs a nonlinear mapping on the result of the linear transformation. In this way, the encoder extracts important information from the input features layer by layer and gradually reduces the feature dimension. Ultimately, the encoder outputs a low-dimensional compressed feature vector that contains the most critical information from the initial fused features.

[0079] Increasing the number of encoder layers or reducing the number of neurons per layer can improve compression, but this may result in excessive information loss, affecting reconstruction quality. Conversely, reducing the number of encoder layers or increasing the number of neurons per layer reduces compression, preserving more original information, but may increase model complexity and computational cost. Therefore, it is necessary to determine the appropriate compression level through experimentation and adjustment based on the specific educational game application scenario and task requirements.

[0080] The compressed feature vector output by the encoder is fed into the decoder, which gradually restores the dimensions and information of the original fused features by increasing the number of neurons layer by layer and performing linear transformations and activation function processing. In each decoder layer, neurons perform calculations based on the output of the previous layer and their own weights and biases, gradually expanding and refining the information in the compressed features.

[0081] During training, the error between the reconstructed features and the original initial fused features is calculated. Common error metrics include mean squared error (MSE) and mean absolute error (MAE). Through the backpropagation algorithm, the reconstruction error is propagated layer by layer from the output layer back to each layer of the encoder and decoder. The network weights and biases are updated based on the error gradient, continuously optimizing the model, reducing the reconstruction error, and improving the reconstruction quality.

[0082] The low-dimensional feature vector obtained after compression by the encoder is the educational game fusion feature vector, which integrates the text and image information in the initial fusion feature. Compared with the initial fusion feature, the fusion feature vector has a lower dimension, reduces the computational complexity and storage space, while retaining the important information in the original feature, which can better reflect the comprehensive characteristics of the educational game.

[0083] In order to better implement the gamification optimization method for online learning in the embodiment of the present invention, based on the gamification optimization method for online learning, please refer to Figure 7 , Figure 7 This is a schematic diagram of the structure of an embodiment of a gamification optimization device for online learning provided by the present invention. This embodiment of the present invention provides a gamification optimization device 700 for online learning, including: A learning feature extraction module 710 is configured to extract learner dynamic features and learner static features from learner multimodal learning state data collected in real time; The learning feature fusion module 720 is configured to use a time transformer to process the learner's dynamic features and the learner's static features, and input the processed information into a fusion network model to determine the learner's learning state feature vector that changes dynamically over time; A game feature extraction module 730 is configured to process text information and image information of the educational game using a language image pre-training model, a text encoder, and an image encoder to obtain an educational game feature vector; The game optimization module 740 is configured to calculate the cosine similarity between the learning state feature vector and the educational game feature vector, and select an educational game to optimize the online learning based on the cosine similarity.

[0084] It should be noted here that the device 700 provided in the above embodiment can implement the technical solutions described in the above method embodiments. The specific implementation principles of the above modules or units can be found in the corresponding contents in the above method embodiments, which will not be repeated here.

[0085] See also Figure 8 , Figure 8 This is a schematic diagram of the structure of a gamified optimization device for online learning provided by an embodiment of the present invention. Based on the aforementioned gamified optimization method for online learning, the present invention also provides a corresponding gamified optimization device for online learning. The gamified optimization device for online learning can be a computing device such as a mobile terminal, desktop computer, notebook, PDA, or server. The gamified optimization device 800 for online learning includes a processor 810, a memory 820, and a display 830. Figure 8 Only some components of the gamification optimization device for online learning are shown, but it should be understood that it is not required to implement all of the shown components, and more or fewer components may be implemented instead.

[0086] In some embodiments, the memory 820 can be an internal storage unit of the online learning gamification optimization device 800, such as the hard drive or memory of the online learning gamification optimization device 800. In other embodiments, the memory 820 can also be an external storage device of the online learning gamification optimization device 800, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped with the online learning gamification optimization device 800. Furthermore, the memory 820 can include both the internal storage unit of the online learning gamification optimization device 800 and an external storage device. The memory 820 is used to store application software installed in the online learning gamification optimization device 800 and various data, such as the program code for installing the online learning gamification optimization device 800. The memory 820 can also be used to temporarily store data that has been output or is about to be output. In one embodiment, the memory 820 stores an online learning gamification optimization program 840 , which can be executed by the processor 810 , thereby implementing the online learning gamification optimization method of each embodiment of the present application.

[0087] In some embodiments, the processor 810 may be a central processing unit (CPU), a microprocessor, or other data processing chip, configured to execute program codes stored in the memory 820 or process data, such as executing a gamification optimization method for online learning.

[0088] In some embodiments, display 830 can be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. Display 830 is used to display information on the gamified optimization device 800 for online learning and to display a visual user interface. Components 810-830 of the gamified optimization device 800 for online learning communicate with each other via a system bus.

[0089] In one embodiment, when the processor 810 executes the gamification optimization program 840 for online learning in the memory 820 , the steps in the above-mentioned gamification optimization method for online learning are implemented.

[0090] This embodiment further provides a computer-readable storage medium storing a gamification optimization program for online learning. When the gamification optimization program for online learning is executed by a processor, the following steps are implemented: Extract learner dynamic features and learner static features from learner multimodal learning state data collected in real time; The learner's dynamic features and static features are processed by the time transformer and input into the fusion network model to determine the learner's learning state feature vector that changes dynamically over time. The text information and image information of the educational game are processed through the language image pre-training model, text encoder and image encoder to obtain the educational game feature vector; The cosine similarity between the learning state feature vector and the educational game feature vector is calculated, and educational games are selected based on the cosine similarity to optimize online learning.

[0091] In summary, the present invention provides a gamified optimization method for online learning. This method captures and analyzes learners' emotional states and engagement levels in real time by collecting multimodal learning state data from learners in real time and extracting their dynamic and static features. A time transformer is used to process learner features for temporal information, which is then fed into a fusion network model to determine the learner's learning state feature vector, which dynamically changes over time. This helps more accurately characterize the learner's learning state. A language image pre-training model, a text encoder, and an image encoder are used to process the text and image information of educational games to generate educational game feature vectors. This fully exploits the effective information in educational games and understands the characteristics of different educational games. The cosine similarity between the learning state feature vector and the educational game feature vector is calculated, and educational games are selected based on this cosine similarity to optimize online learning. This method can select educational games that match the learner's learning state, thereby stimulating learners' learning interest through appropriate interactive methods and providing a personalized learning experience. Teaching strategies can also be dynamically adjusted based on learners' real-time feedback, improving online learning outcomes.

[0092] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above method. The computer-readable storage medium may include, but is not limited to, any type of disk, including a floppy disk, an optical disk, a DVD, a CD-ROM, a microdrive, a magneto-optical disk, a ROM, a RAM, an EPROM, an EEPROM, a DRAM, a VRAM, a flash memory device, a magnetic card or an optical card, a nanosystem (including a molecular memory IC), or any type of medium or device suitable for storing instructions and / or data.

[0093] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0094] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0095] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable memory, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0096] The above is only an exemplary embodiment of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the specification and practicing the disclosure herein, those skilled in the art will easily think of the implementation scheme of the present disclosure. This application is intended to cover any variation, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the art that are not recorded in the present disclosure. The description and examples are to be regarded as exemplary only, and the scope and spirit of the present disclosure are defined by the claims.

[0097] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0098] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A gamification optimization method for online learning, characterized in that: include: Extract learner dynamic features and learner static features from learner multimodal learning state data collected in real time; Using a time transformer to process the learner's dynamic features and the learner's static features, and inputting the processed features into a fusion network model to determine a learning state feature vector of the learner that changes dynamically over time; The text information and image information of the educational game are processed through the language image pre-training model, text encoder and image encoder to obtain the educational game feature vector; The cosine similarity between the learning state feature vector and the educational game feature vector is calculated, and an educational game is selected based on the cosine similarity to optimize online learning.

2. The gamification optimization method for online learning according to claim 1, characterized in that: The method of extracting learner dynamic features and learner static features from the learner multimodal learning state data collected in real time further includes: Real-time collection of learners' facial expression information, voice information and body movement information in the classroom environment; The facial expression information, the voice information and the body movement information are processed using a residual network and a global convolutional attention module to extract learner dynamic features and learner static features.

3. The gamification optimization method for online learning according to claim 1, characterized in that: The method further includes: processing the learner's dynamic features and the learner's static features using a time transformer, and inputting the processing information into a fusion network model to determine the learner's learning state feature vector that changes dynamically over time; Setting a time position embedding function of a time transformer according to classroom teaching, and performing time information processing on the learner's dynamic features and the learner's static features; Using the fusion network model to classify the learner's expression according to the learner's static features, and adjusting the classification results through the intensity perception loss function to obtain the learner's expression state; Analyzing the learner's dynamic characteristics based on the learner's historical behavior data, the learner's dynamic characteristics, and the fusion network model to determine the learner's behavior state; The learner's expression state and the learner's behavior state are fused to obtain a learning state feature vector that dynamically changes with time.

4. The gamification optimization method for online learning according to claim 3, characterized in that: The method further comprises: utilizing the fusion network model to classify the learner's expression according to the learner's static features, and adjusting the classification results by using an intensity perception loss function to obtain the learner's expression state. Preliminarily classifying the learner's static features according to preset judgment rules to obtain initial expression information; Establishing a mapping relationship between expression intensity and learner cognitive state through learner's historical expression data, and calibrating the initial expression information; The learner's expression state is obtained by assigning weights through the intensity-aware loss function and adjusting the calibrated initial expression information.

5. The gamification optimization method for online learning according to claim 3, characterized in that: The step of analyzing the learner's dynamic characteristics based on the learner's historical behavior data, the learner's dynamic characteristics, and the fusion network model to determine the learner's behavior state further includes: Analyzing the learner's historical behavior data and the learner's dynamic characteristics using the fusion network model to determine the learner's historical behavior vector and current behavior vector; Calculating the learner's behavior deviation degree based on the historical behavior vector and the current behavior vector; The learner's behavior state is determined based on the degree of behavior deviation and preset behavior rules.

6. The gamification optimization method for online learning according to claim 1, characterized in that: The method of processing the text information and image information of the educational game by using the language image pre-training model, the text encoder, and the image encoder to obtain the educational game feature vector also includes: Obtaining a text encoding feature vector, a text model feature vector, an image encoding feature vector, and an image model feature vector from the text information and image information of the educational game respectively through a language image pre-training model, a text encoder, and an image encoder; The text encoding feature vector and the text model feature vector are fused to obtain a text fusion feature vector, and the image encoding feature vector and the image model feature vector are fused to obtain an image fusion feature vector; The text fusion feature vector and the image fusion feature vector are concatenated and reconstructed to obtain the educational game fusion feature vector.

7. The gamification optimization method for online learning according to claim 6, characterized in that: The step of concatenating and reconstructing the text fusion feature vector and the image fusion feature vector to obtain the educational game fusion feature vector also includes: The text fusion feature vector and the image fusion feature vector are concatenated and passed through a fully connected layer to obtain the initial fusion feature; A deep autoencoder is used to compress and reconstruct the initial fusion features to obtain an educational game fusion feature vector.

8. A gamification optimization device for online learning, characterized in that: include: A learning feature extraction module configured to extract learner dynamic features and learner static features from learner multimodal learning state data collected in real time; a learning feature fusion module configured to process the learner's dynamic features and the learner's static features using a time transformer, and input the processed information into a fusion network model to determine a learning state feature vector of the learner that changes dynamically over time; a game feature extraction module configured to process text information and image information of the educational game using a language image pre-training model, a text encoder, and an image encoder to obtain an educational game feature vector; The game optimization module is configured to calculate the cosine similarity between the learning state feature vector and the educational game feature vector, and select an educational game based on the cosine similarity to optimize online learning.

9. A gamification optimization device for online learning, characterized in that: The method comprises at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the computer program is executed by the processing unit, the processing unit executes the steps of the gamification optimization method for online learning according to any one of claims 1 to 7.

10. A storage medium, characterized in that: It stores a computer program that can be executed by a gamification optimization device for online learning. When the computer program runs on the gamification optimization device for online learning, the gamification optimization device for online learning executes the steps of the gamification optimization method for online learning described in any one of claims 1 to 7.