Personalized English education system and method based on multi-modal sentiment analysis

Through multimodal interaction and emotion analysis technology, combined with knowledge graph constraints, the shortcomings of the existing education system in emotion analysis and personalized content generation are solved, and higher emotional recognition accuracy and the generation of personalized learning content are achieved, which improves learning effect and students' learning enthusiasm.

CN120086805AActive Publication Date: 2025-06-03XIAMEN UNIV

Patent Information

Application Number
CN202510567885.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The existing personalized education system has shortcomings in emotion analysis, personalized generation of learning content and real-time feedback. It cannot fully identify students' emotional states, ignore students' individual differences and real-time learning progress, resulting in the learning content not meeting students' actual needs.

Method used

The multimodal interaction module is used to receive and process students' multimodal data, and the sentiment analysis of multimodal data such as voice and images is carried out through the emotion recognition module. The real-time learning status of students is evaluated in combination with the personalized module, and a personalized learning feedback report is generated. The core brain module builds a knowledge graph to avoid knowledge beyond the syllabus and ensures that the generated content conforms to the current knowledge level of students.

Benefits of technology

It improves the accuracy of emotional recognition, and the generated personalized learning content meets students' needs more accurately, avoids the problem of over-the-syllabus and over-simplification of knowledge, and improves learning effect and students' learning enthusiasm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086805A_ABST
    Figure CN120086805A_ABST
Patent Text Reader

Abstract

The invention provides a personalized English education system and method based on multi-modal sentiment analysis. The system comprises a multi-modal interaction module for receiving and processing multi-modal data of students, an emotion recognition module for performing emotion analysis on the input multi-modal data, and a personalized module for evaluating real-time learning states of the students based on historical learning data of the students, and the core brain module is used for dynamically adjusting interactive feedback according to the emotional state and the real-time learning state. The emotion recognition module comprises emotion information fusion, the emotion information fusion adopts a weighting strategy, and the final output emotion state is adjusted through emotion consistency constraint and a conflict correction mechanism. According to the invention, through an emotion consistency loss function, a conflict correction mechanism and a knowledge graph-based super-outline control mechanism, the emotion and cognitive states of the students are accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of new generation information technology, and particularly to a virtual digital human education system based on technologies such as multimodal interaction, sentiment analysis, dynamic knowledge graph constraint, and large models, which can be applied to personalized education and intelligent learning tutoring, and is particularly suitable for the field of English learning. Background Art

[0002] With the progress of artificial intelligence technology, personalized education systems have made certain progress in improving learning effects, enhancing students' participation and initiative. However, the existing technologies still face various challenges, especially in aspects such as sentiment analysis, personalized generation of learning content, and real-time feedback. Existing personalized learning systems mainly rely on students' historical data (such as grades, homework completion status, etc.) to generate learning feedback and recommended content, but these systems have the following deficiencies: 1. Currently, most education systems rely on single-modal sentiment analysis (such as speech, facial expressions, or text input). These systems often fail to comprehensively identify students' emotional states. Even through multimodal fusion, there are still problems of emotional information conflicts. For example, a student's speech may convey confusion or frustration, while the facial expression may seem relatively calm. Traditional sentiment analysis methods do not have sufficient mechanisms to solve these conflicts, resulting in low accuracy in identifying emotional states, which in turn affects the effectiveness of personalized learning content and the authenticity of emotional feedback.

[0003] 2. Many existing systems use fixed syllabuses and preset knowledge points for teaching, ignoring students' individual differences and real-time learning progress. This leads to two situations for students in the learning process: one is that the content is too simple to stimulate students' learning interest; the other is that the content is too complex and beyond the students' current cognitive level, resulting in learning frustration.

[0004] 3. Although some systems attempt to provide personalized content through historical data analysis, in actual applications, many systems fail to fully consider students' cognitive fluctuations and emotional states. Therefore, the generated content often cannot fully meet students' needs. Although many education systems attempt to provide personalized recommendations based on data such as students' grades and learning progress, these systems often ignore the impact of emotional fluctuations on learning effects. Students' emotional states and learning motivations can greatly affect their learning effects. Therefore, personalized learning systems need to consider students' emotional fluctuations when generating feedback and suggestions to ensure that the feedback content not only matches students' learning progress but also can motivate students at the emotional level and enhance learning enthusiasm.

[0005] For this reason, in the prior art, a patent for invention with a publication number of CN117573904A discloses a method and system for generating a knowledge graph of multimedia teaching resources based on recognition and analysis, and a patent for invention with a publication number of CN117313852A discloses a method and system for updating a personalized teaching knowledge graph based on multimodal data. Both mention collecting multimodal information for feature processing, mining, and prediction to generate a personalized teaching knowledge graph.

[0006] For another example, a paper published by Feng Lei: Design and Implementation of an Intelligent Teaching Assistant Platform Based on a Multimodal Knowledge Graph, Hubei Normal University, 2023, mentions using YOLOv3-tiny to complete real-time supervision of classroom learning conditions, and using Paddle OCR and Deep KE to construct a fine-grained knowledge point graph. By constructing a logit regression model, the teaching video is automatically segmented, and the link and mapping of multimodal data such as fine-grained knowledge points, video teaching, and classroom learning conditions in the graph are realized, so as to complete multimodal comprehensive applications such as "intelligent learning condition analysis", "multi-dimensional quantitative assessment", "knowledge graph heat map", "personalized question bank generation", and "delimiting and positioning of learning condition deterioration".

[0007] For another example, an article: Intelligent Learning Analysis: The Innovative Fulcrum of Student Process Evaluation, Page 3, June 26, 2021, China Education News (https: / / www.sohu.com / a / 474156701_243614), mentions using learning analysis technologies such as knowledge graphs to diagnose students' dynamic cognitive levels in a timely manner, discover cognitive problems existing in students, and intelligently push appropriate learning resources; using learning analysis technologies such as computer vision and machine learning to dynamically monitor students' non-cognitive states such as emotions, and timely adjust and intervene in students' social emotions to achieve the purpose of optimizing learning effects; using learning analysis technologies such as dashboards to visually present students' diagnostic results, helping students understand themselves more intuitively, and enabling teachers and parents to better understand students.

[0008] For another example, an article published by Jiang Jie, Yu Wenting, Wang Haiyan, etc.: Research on Students' Learning Behaviors in Smart Classrooms Based on Multimodal Data, China Education Informatization, mentions analyzing multimodal data of students' learning behaviors in order to help students understand their own learning behaviors and states, and enable teachers to conduct personalized teaching.

[0009] In summary, even though the prior art integrates interactions from multiple modalities such as text, audio, image, and video, the data expressions of multiple modalities are inconsistent, and there are problems of multimodal data emotional conflicts in complex environments, making it impossible to accurately identify students' emotional fluctuations. Summary of the Invention

[0010] A brief overview of the embodiments of the present invention is given below to provide a basic understanding of certain aspects of the present invention. It should be understood that the following overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is only to present certain concepts in a simplified form as a prelude to the more detailed description to be discussed later.

[0011] The object of the present invention is to provide a personalized English education system based on multimodal interaction, emotion analysis and personalized recommendation, which can, through technologies such as multimodal data acquisition, emotion recognition, and knowledge graph constraint, obtain the learning status, emotion status and cognitive level of students in real time, so as to generate personalized learning content and feedback that conforms to the actual situation of students. Compared with the prior art, this system has higher emotion recognition accuracy, more accurate personalized content generation ability, and an innovative mechanism to effectively avoid knowledge beyond the syllabus through the knowledge graph.

[0012] According to one aspect of the present application, there is provided a personalized English education system based on multimodal emotion analysis, including: A multimodal interaction module: used to receive and process the multimodal data of students, where the multimodal includes text modal data, audio modal data, image modal data and video modal data. Through multimodal perception, the system can comprehensively understand the learning content, emotion status and attention distribution of students; An emotion recognition module: used to perform emotion analysis on the input multimodal data, including speech emotion recognition, video emotion recognition and emotion information fusion, and fuse the recognized emotion information (audio emotion information and image emotion recognition information), and finally output the emotion status; The emotion information fusion adopts a weighted strategy, and in the case of conflicts, the final output emotion status is adjusted through an emotion consistency constraint and a conflict correction mechanism; A personalized module: evaluates the real-time learning status of the student based on the student's historical learning data and generates a personalized learning feedback report. The learning feedback report includes learning progress, performance analysis, emotion fluctuation evaluation, and learning and interaction suggestions.

[0013] A core brain module: used to dynamically adjust the interaction feedback according to the emotion status of the emotion recognition module and the real-time learning status output by the personalized module. By constructing a knowledge graph to avoid the situation of knowledge beyond the syllabus, it is ensured that the generated content conforms to the current knowledge level of the student.

[0014] Among them, in the emotion recognition module, the specific process of the emotion consistency constraint includes: Input the voice signal and image data, and extract the voice emotion feature and the image emotion feature respectively; Align the speech emotion features and the image emotion features to obtain feature representations in the same emotion space; After the feature alignment, calculate the Euclidean distance between the speech emotion features and the image emotion features; establish an emotion consistency loss function based on the Euclidean distance and the cross-entropy loss function; The comprehensive loss obtained by jointly optimizing the emotion consistency loss function and the basic classification loss As a specific example, in the emotion recognition module, the specific process of the emotion consistency constraint includes: Input data preparation: Input speech signals and image data , and respectively extract the speech emotion features and the image emotion features ; Feature alignment: Use a shared fully connected layer to align the speech emotion features and the image emotion features to obtain feature representations in the same emotion space and : ; Among them, represents the alignment network; Consistency loss design: After the feature alignment, calculate the Euclidean distance between the speech emotion features and the image emotion features to constrain the consistency of the modal features; the emotion consistency loss function (cross-entropy loss) is as follows: ; Among them, is the actual emotion category; is the weight hyperparameter; is the batch size; and are respectively the feature representations after aligning the speech emotion features and the image emotion features by the alignment network , where represents the sample index in the i-th batch; the softmax function converts the output vector in the feature space into a probability distribution, that is, and are respectively the emotion prediction probabilities of the model in the speech and image modalities; is the cross-entropy loss function, and are respectively the differences between the prediction results and the true emotion labels, used to evaluate the prediction accuracy of the speech and image modalities; Joint optimization: The sentiment consistency loss and the base classification loss are jointly optimized; the base classification loss is as follows: ; wherein, is the batch size; is the number of sentiment categories; is the sample 's true label, which is 1 when belonging to category and 0 otherwise; is the predicted probability of the model for the category of the sample ; The total loss is the comprehensive loss obtained by jointly optimizing the base classification loss and the sentiment consistency loss. It is as follows: ; wherein, is the hyperparameter that adjusts the importance of consistency.

[0015] Existing multi-modal sentiment recognition methods mainly rely on the cross-entropy loss function to optimize single modalities and cannot effectively constrain the sentiment consistency between different modalities. This application proposes a sentiment consistency loss function, which ensures the consistent sentiment output of speech and images through the joint optimization of multi-modal feature alignment and classification tasks.

[0016] Among them, in the sentiment recognition module, the specific process of the conflict correction mechanism includes: Confidence calculation: Calculate the confidence of each modality using the aligned feature space, including calculating the speech sentiment classification confidence and the image sentiment classification confidence : ; Conflict detection: First, regard the sentiment with the highest confidence in each modality as the recognition result. The highest classification confidence of a certain modality in all sentiment categories is . Among them, or ); is the sentiment classification confidence of modality . Then, set the confidence threshold , if , no conflict judgment is performed, and it is considered that the model is uncertain in this modality and is directly marked as "low confidence" for output; if the highest confidence categories of each modality are different, the semantic similarity of the categories is calculated. If the two are semantically close, that is, less than the set category similarity threshold, no conflict is determined; calculate the confidence difference between the highest categories of the two modalities. If it is higher than the set confidence difference threshold, the prediction result of the modality with higher confidence is directly adopted without determining it as a conflict. In short, a conflict is truly determined only when all of the following conditions are met simultaneously: different predicted categories, low category semantic similarity, high confidence in both modalities, and a small confidence difference.

[0017] Priority setting: After a conflict is detected, the priority is set according to the directness of the modality expression and the weight is adjusted dynamically. The calculation formula for the priority is as follows: ; where is a measure of the directness of the emotion expressed by this modality, with a value range of [0,1]; and are hyperparameters that adjust the influence of confidence and expression directness on the weight respectively.

[0018] The final weight of each modality is calculated through normalization , and the formula is as follows: ; where is all modalities, including speech and image modalities, is the normalized weight of the corresponding modality ; therefore, according to the above formula, and can be calculated, that is, the final weights of the speech modality and the image modality after normalization.

[0019] Multimodal fusion: The speech emotion features in the unified space mapped by the alignment network are fused with the image emotion features to obtain the fused feature .

[0020] Classification output: The fused feature is input into the classifier to generate the final emotion classification result When fusing multimodal emotion information, due to the heterogeneity between modalities and the differences in expression methods, emotional conflicts may occur (such as the speech conveying doubt but the facial expression being calm). This application overcomes this problem through the above conflict correction mechanism based on priority and dynamic weight adjustment.

[0021] As a specific implementation solution, the multi-modal interaction module includes an input end and an output end; the input end includes text, audio, images, and videos, where the input end includes text, audio, images, and videos, and the text is the text information for students to ask questions or give feedback; the audio is the voice signal input by students during learning interactions; the images are the picture content uploaded by students; the videos are the learning videos of students captured by the camera; The output end includes text output, audio output, and image output; the text output is the text reply generated for students to read; the audio output is to convert the text reply into audio through voice synthesis technology; the image output is to generate pictures for auxiliary teaching using deep learning models such as GAN or Diffusion Models.

[0022] As a specific implementation solution, in the emotion recognition module, the speech emotion recognition includes: Speech preprocessing and feature extraction: Preprocess the input student speech and obtain the speech signal features; Emotion classification and modeling: Based on the extracted speech signal features, identify the emotion category and intensity; The image emotion recognition includes: Extract frames from the obtained video at intervals to get images; Facial expression recognition: Identify the movement of facial key points in the image to analyze the expression changes; extract the facial features in the image and classify the expressions into basic emotions; Deep pose estimation: Extract the pose and movement trajectory of the human body through a deep learning-based pose estimation model; Eye movement tracking: Use a convolutional neural network to analyze the eye movement trajectory and fixation points of students to infer their attention concentration.

[0023] The personalization module includes: Data collection: Collect real-time updated student learning data, including historical learning records, homework and exam scores, and interaction logs; Data analysis and feedback report generation: Generate a comprehensive learning feedback report based on the collected student learning data.

[0024] The learning feedback report includes learning progress, performance analysis, emotional state, and learning and interaction suggestions; Among them, the learning progress includes the tasks completed by students within a specific period, the knowledge point mastery, and the learning speed; The performance analysis includes the scoring situation, progress trend, and weak links of students in each subject and each chapter; The emotional state evaluates the emotional fluctuations of students during the recent learning process according to the emotional information provided by the emotion analysis module; Learning and interaction suggestions generate personalized learning and interaction suggestions based on students' learning progress, performance analysis, and emotional state.

[0025] The core brain module includes data collection and processing, fine-tuning of the large model, and knowledge overstepping control: Among them, the data collection and processing include data collection, data annotation, and knowledge graph construction; data collection is to collect multi-disciplinary standardized textbooks, syllabuses, tutoring materials, and courseware; data annotation is to organize the collected data by knowledge points and clearly label the difficulty, learning objectives, and relevance information of each knowledge point; knowledge graph construction is to construct a knowledge graph of the field, forming a structured graph of concepts, knowledge points, and their relationships in different disciplines; The fine-tuning of the large model includes model selection, construction of the training dataset, and model fine-tuning; among them, model selection: select a basic large language model as the basic model for generating and regulating conversations; model fine-tuning: input the collected dataset into the basic large language model, set specific fine-tuning parameters; conduct multiple rounds of training and fine-tuning processes to enable the model to produce accurate and emotional feedback on educational content; The knowledge overstepping control includes construction of the domain knowledge graph, establishment of the student knowledge graph, knowledge overstepping control mechanism based on the knowledge graph, and post-processing and correction; Among them, constructing the domain knowledge graph is to establish a knowledge graph covering the required subject fields; Establishing the student knowledge graph is specifically to update the student knowledge graph through the student's learning history data; The knowledge overstepping control mechanism based on the knowledge graph specifically includes: Step 1: Define knowledge constraint conditions: Suppose in the knowledge graph, a certain knowledge point depends on another knowledge point , and use a knowledge dependency matrix to represent the dependency relationship between knowledge points : , indicating that knowledge point depends on knowledge point ; , indicating that knowledge point is not directly related to knowledge point ; Define the student's knowledge state vector , where: , indicating that the student has mastered knowledge point ; , indicating that the student has not mastered knowledge point ; Step 2: Constraint Mechanism during Generation: When generating text, the model will refer to the knowledge graph to restrict the generated content. For example, assume the set of knowledge points currently mastered by the student is , and the input to the generation model includes the student's current learning task . When generating content, the following constraint conditions need to be met: For any knowledge point , if it is required to be generated in the current learning task , but the student has not mastered the relevant prerequisite knowledge points , then the generator should avoid generating content related to that knowledge point ; That is, the model is subject to the following constraints during generation: ; where: is the probability distribution of the model without constraints, representing the generated content given the learning task ; is an indicator function, indicating that if the student masters the knowledge point , that is, , then that knowledge point can appear in the generated content; otherwise , that knowledge point is excluded from the generated content.

[0026] The post - processing and correction include knowledge consistency checking and a fallback mechanism; Knowledge consistency checking: Perform knowledge consistency checking on the content generated by the model to ensure that the output knowledge points are consistent with the knowledge currently mastered by the student. For example, if the generated content contains an advanced knowledge point beyond the student's current knowledge scope, then that content will be marked as out of syllabus by the system; Fallback mechanism: If the generated content violates the knowledge graph constraints, the system regenerates content that meets the constraints.

[0027] According to another aspect of the present application, there is also provided a personalized English education method based on multi - modal sentiment analysis, including: Receiving and processing multi - modal data of students, where the multi - modal includes text, audio, image, and video input signals; Performing sentiment analysis on the input multi - modal data, including text sentiment recognition, speech sentiment recognition, video sentiment recognition, and sentiment information fusion, and fusing the recognized sentiment information, and finally outputting the sentiment state; The sentiment information fusion adopts a weighted strategy and adjusts the finally output sentiment state through sentiment consistency constraints and a conflict correction mechanism; Evaluate the real-time learning status of a student based on the student's historical learning data; Dynamically adjust the interactive feedback according to the emotional state and the real-time learning status.

[0028] This system involves cross-applications in multiple complex fields, solves problems such as multi-modal emotional data conflict handling and knowledge overstepping control in the prior art, and deeply combines multiple technical fields (such as sentiment analysis, knowledge graph construction, large model fine-tuning, etc.) to form an integrated intelligent education system. In particular, the proposed sentiment consistency loss function, conflict correction mechanism, and knowledge graph-based overstepping control mechanism make the present invention have significant advantages over the prior art in accurately identifying students' emotional and cognitive states and generating personalized learning content. Brief Description of the Drawings

[0029] The present invention can be better understood by referring to the description given below in conjunction with the accompanying drawings, in which the same or similar reference numerals are used throughout the drawings to denote the same or similar components. The accompanying drawings, together with the following detailed description, are included in this specification and form a part of this specification, and are further used to illustrate the preferred embodiments of the present invention and to explain the principles and advantages of the present invention. In the drawings: Figure 1 It is a module diagram of the personalized English education system according to an embodiment of the present invention; Figure 2 It is a flowchart of the personalized English education system according to an embodiment of the present invention. Detailed Embodiments

[0030] The embodiments of the present invention will be described below with reference to the accompanying drawings. The elements and features described in one drawing or one embodiment of the present invention can be combined with the elements and features shown in one or more other drawings or embodiments. It should be noted that, for the sake of clarity, the representation and description of components and processes irrelevant to the present invention and known to those of ordinary skill in the art are omitted in the drawings and the description.

[0031] The personalized English education system of the present application involves cross-applications in multiple complex fields, solves problems such as multi-modal emotional data conflict handling and knowledge overstepping control in the prior art, and deeply combines multiple technical fields (such as sentiment analysis, knowledge graph construction, large model fine-tuning, etc.) to form an integrated intelligent education system. In particular, the proposed sentiment consistency loss function, conflict correction mechanism, and knowledge graph-based overstepping control mechanism make the present invention have significant advantages in accurately identifying students' emotional and cognitive states and generating personalized learning content.

[0032] As a specific embodiment, refer to Figure 1, the personalized English education system based on multimodal sentiment analysis mainly includes the following modules: (1) Multimodal interaction module: This module is used to receive and process the multimodal data of students. The multimodal data is a variety of input signals, including text, audio, images, and videos. Through multimodal perception, the system can comprehensively understand the learning content, emotional state, and attention distribution of students.

[0033] (2) Emotion recognition module: This module is responsible for performing sentiment analysis on the input multimodal data, including speech emotion recognition, image emotion recognition, etc. The fusion of emotion information adopts a weighted strategy, combining multimodal information such as speech and images, and adjusts the final emotion output through emotion consistency constraints and conflict correction mechanisms in case of conflicts.

[0034] (3) Personalization module: This module is based on the historical learning data of students, conducts real-time learning status evaluation, and generates personalized learning feedback reports. The reports include learning progress, performance analysis, emotional fluctuation evaluation, and learning and interaction suggestions.

[0035] (4) Core brain module: This module is responsible for dynamically adjusting interaction feedback according to the learning progress and emotional state of students. By constructing a knowledge graph to avoid the situation of knowledge exceeding the syllabus, it ensures that the generated content meets the current knowledge level of students.

[0036] The following specifically introduces the detailed structure of each module: 1 Multimodal interaction module 1.1 Input end (1) Text: Students ask questions or give feedback through text input.

[0037] (2) Audio: Students conduct learning interactions through voice input. The audio signal is converted into text through speech recognition technology and transmitted to the emotion recognition module.

[0038] (3) Image: Through algorithms such as OCR (Optical Character Recognition) and object detection, understand the content of the pictures uploaded by students.

[0039] (4) Video: Capture the learning videos of students through the camera and provide them to the emotion recognition module to analyze the emotional information such as facial expressions and postures of students.

[0040] 1.2 Output end (1) Text: Generate text responses for students to read.

[0041] (2) Audio: Convert text into audio through speech synthesis technology.

[0042] (3) Image: Use deep learning models, such as GAN or Diffusion Models, to generate pictures for auxiliary teaching.

[0043] 2 Emotion Recognition Module Perform emotion analysis on the data obtained by the interaction module.

[0044] 2.1 Speech Emotion Recognition (1) Speech preprocessing and feature extraction: Preprocess the student's speech input, including noise removal, speech enhancement, etc. Use Perceptual Linear Prediction (PLP) to obtain the speech signal features; (2) Emotion classification and modeling: Based on the extracted speech features, use Convolutional Neural Network (CNN) to identify the emotion category and intensity.

[0045] 2.2 Image Emotion Recognition (1) Extract frames from the obtained video at intervals to get images; (2) Facial expression recognition: Analyze the movement of facial key points such as eyes, mouth, and eyebrows through MediaPipe to analyze the expression changes. Classify the expressions into basic emotions according to the extracted facial features; (3) Deep pose estimation: Extract the human pose and movement trajectory through a deep learning-based pose estimation model; (4) Eye movement tracking: Use Convolutional Neural Network to analyze the student's eye movement trajectory and fixation points to infer the degree of attention concentration; 2.3 Emotion information fusion: Combine the emotion information from speech and images through weighted fusion.

[0046] (1) Emotion consistency constraint: Existing multi-modal emotion recognition methods mainly rely on the cross-entropy loss function to optimize single-modal, and cannot effectively constrain the emotion consistency between different modalities. The present invention proposes an emotion consistency loss function, aiming to ensure the consistency of emotion outputs of speech and images through the joint optimization of multi-modal feature alignment and classification tasks. The algorithm steps are as follows: Input data preparation: Input speech signal and image data , and extract speech emotion features and image emotion features respectively; Feature alignment: Use a shared fully connected layer to and for alignment to obtain feature representations in the same emotion space: ; wherein, represents the alignment network.

[0047] Consistency loss design: After feature alignment, calculate the Euclidean distance between speech and image features , used to constrain the consistency of modal features; further, combined with the sentiment classification loss (cross-entropy loss) for each modality: ; Among them, is the actual sentiment category; is the weight hyperparameter; is the batch size. and are the feature representations after aligning the speech emotion feature and the image emotion feature through the alignment network , where represents the sample index in the i-th batch; the softmax function converts the output vector in the feature space into a probability distribution, that is, and are the sentiment prediction probabilities of the model in the speech / image modality; is the cross-entropy loss function, and are the differences between the prediction results and the true sentiment labels, used to evaluate the prediction accuracy of each modality; Joint optimization: jointly optimize the sentiment consistency loss and the basic classification loss ; the basic classification loss is as follows: ; Among them, is the batch size; is the number of sentiment categories; is the true label of the sample , which is 1 when it belongs to the category , otherwise 0; is the prediction probability of the model for the category of the sample ; The total loss is the comprehensive loss obtained by jointly optimizing the basic classification loss and the sentiment consistency loss. It is expressed as follows: ; Among them, is the hyperparameter that adjusts the importance of consistency.

[0048] (2) Conflict correction mechanism: When fusing multi-modal sentiment information, due to the heterogeneity between modalities and the differences in expression methods, sentiment conflicts may occur (such as the speech conveys doubt but the facial expression is calm). The present invention proposes a conflict correction mechanism based on priority and dynamic weight adjustment. The algorithm steps are as follows: Confidence calculation: Calculate the confidence of each modality using the aligned feature space, including calculating the confidence of speech emotion classification and the confidence of image emotion classification : ; Conflict detection: First, take the emotion with the highest confidence of each modality as the recognition result. The highest classification confidence of a certain modality in all emotion categories is . Among them, represents any modality (such as or ); is the emotion classification confidence of modality . Then, set the confidence threshold . If , then no conflict judgment is made, and it is considered that the model is uncertain in this modality and is directly marked as "low confidence" for output; if the highest confidence categories of each modality are different, calculate the semantic similarity of the categories. If the two are semantically close, that is, less than the set category similarity threshold, then it is not determined as a conflict; calculate the confidence difference between the highest categories of the two modalities. If it is higher than the set confidence difference threshold, then directly adopt the prediction result of the modality with higher confidence and do not determine it as a conflict. In short, a conflict is truly determined only when the following conditions are all met: different prediction categories, low category semantic similarity, high confidence in both modalities, and small confidence difference.

[0049] Priority setting: After a conflict is detected, set the priority according to the directness of the modality expression and dynamically adjust the weights. The calculation formula for the priority is as follows: ; Among them, is the measure of the directness of the emotion expressed by this modality, with a value range of [0,1]; and are the hyperparameters that adjust the influence of confidence and directness of expression on the weights respectively.

[0050] Calculate the final weight of each modality through normalization processing , and the formula is as follows: ; Among them, is all modalities, including speech and image modalities, is the weight of the corresponding modality ; therefore, according to the above formula, and can be calculated, that is, the final weights of the speech modality and the image modality after normalization.

[0051] Multimodal Fusion: The speech emotion features in the unified space mapped by the alignment network and the image emotion features are fused to obtain the fused features , and the specific fusion formula is as follows: ; Classification Output: The fused features are input into the classifier to generate the final emotion classification result 3. Personalized Module 3.1 Data Collection: Collect real-time updated student learning data, including historical learning records, homework and exam scores, interaction logs, etc.

[0052] 3.2 Data Analysis and Feedback Report Generation Use data mining and large model technologies to generate a comprehensive learning feedback report. The report content includes but is not limited to: (1) Learning Progress: The tasks completed by the student within a specific time period, the knowledge point mastery situation, and the learning speed; (2) Score Analysis: The scoring situation, progress trend, and weak links of the student in each subject and each chapter; (3) Emotional State: Evaluate the emotional fluctuations of the student during the recent learning process according to the emotional information provided by the emotion analysis module; (4) Learning and Interaction Suggestions: Generate personalized learning and interaction suggestions based on the student's situation, such as recommending areas for intensive practice or the need to provide appropriate psychological counseling.

[0053] 4. Core Brain Module 4.1 Data Collection and Processing (1) Data Collection: Collect multi-disciplinary standardized textbooks, teaching syllabuses, tutoring materials, courseware, etc.

[0054] (2) Data Annotation: Organize the collected data by knowledge points and clearly label information such as the difficulty, learning objectives, and relevance of each knowledge point.

[0055] (3) Knowledge Graph Construction: Construct a knowledge graph of the field, forming a structured graph of concepts, knowledge points, and their relationships (such as sequence, dependency, synonyms, etc.) in different disciplines. See 4.3 for specific steps.

[0056] 4.2 Fine-tuning the Large Model (1) Model Selection: Select a basic large language model as the basic model for generating and regulating conversations.

[0057] (2) Construct the training dataset (3)Fine-tuning the model: Input the collected dataset into the base large language model and set specific fine-tuning parameters (such as learning rate, batch size, etc.). Conduct multiple rounds of training and fine-tuning on the base large language model so that the fine-tuned model can generate accurate and emotional feedback on educational content.

[0058] 4.3 Knowledge Out-of-Scope Control To avoid generating inaccurate content, the present invention proposes a generation mechanism based on knowledge graph constraints: (1)Constructing a domain knowledge graph Before implementing knowledge out-of-scope control based on a knowledge graph, it is first necessary to establish a knowledge graph covering the required subject area. A knowledge graph is a directed graph where each node represents a knowledge point (such as "algebraic equation" or "Newton's first law"), and the edges represent the relationships between knowledge points (such as "dependency" or "prerequisite relationship").

[0059] (2)Establishing a student's knowledge graph Each student's knowledge graph is dynamically changing, representing the knowledge points that the student has mastered and the knowledge points to be learned. The student's knowledge graph can be updated through the student's learning history data (such as homework grades, test results, interaction records, etc.).

[0060] (3)Knowledge out-of-scope control mechanism based on knowledge graph When the large model generates teaching content, it needs to be constrained based on the student's knowledge graph and the current learning progress. The core of this process is to prevent the content generated by the large model from exceeding the scope of knowledge points mastered by the student and ensure that the generated teaching content is consistent with the student's current knowledge state.

[0061] Step 1: Define knowledge constraint conditions Suppose in the knowledge graph, a certain knowledge point depends on another knowledge point , and use a knowledge dependency matrix to represent the dependency relationship between knowledge points: , indicating that knowledge point depends on knowledge point ; , indicating that knowledge point is not directly related to knowledge point .

[0062] Define the student's knowledge state vector , where: , indicating that the student has mastered knowledge point ; , indicating that the student has not mastered the knowledge point ; Step 2: Constraint mechanism during generation When generating text, the model will refer to the knowledge graph to restrict the generated content. For example, assume that the set of knowledge points currently mastered by the student is , the input of the generation model includes the current learning task of the student , and the following constraint conditions need to be met when generating content: For any knowledge point , if is required to be generated in the current learning task , but the student has not mastered the relevant prerequisite knowledge points , then the generator should avoid generating content related to that knowledge point .

[0063] That is, the model will be subject to the following constraints when generating: ; Where: is the probability distribution of the generation model without constraints, indicating the generated content under the given learning task .

[0064] is the indicator function, indicating that if the student masters the knowledge point i.e. , then that knowledge point can appear in the generated content; otherwise , that knowledge point is excluded from the generated content.

[0065] 4.4 Post - processing and correction After generating the preliminary feedback content, further post - processing can ensure that the model does not break the constraints of the knowledge graph and avoid over - generation.

[0066] (1) Knowledge consistency check: Perform a knowledge consistency check on the generated content to ensure that the output knowledge points are consistent with the knowledge currently mastered by the student. For example, if the generated content contains an advanced knowledge point beyond the current knowledge scope of the student , then the content will be marked as out of syllabus by the system.

[0067] (2) Fallback mechanism: If the generated content violates the constraints of the knowledge graph (out of syllabus), the system regenerates content that meets the constraints.

[0068] As another specific embodiment, the present application also provides a virtual digital human education method based on multimodal interaction and sentiment analysis, including: Receive and process the multimodal data of students, where the multimodal includes text, audio, image, and video input signals; Conduct sentiment analysis on the input multimodal data, including speech sentiment recognition, video sentiment recognition, and sentiment information fusion, and fuse the recognized sentiment information, and finally output the emotional state; the sentiment information fusion adopts a weighted strategy, and adjusts the finally output emotional state through sentiment consistency constraints and conflict correction mechanisms; Evaluate the real-time learning state of the student based on the student's historical learning data; Dynamically adjust the interactive feedback according to the emotional state and the real-time learning state.

[0069] The personalized English education method based on multimodal sentiment analysis specifically includes the following steps: Step 1: Collect the text information of the questions or feedback input by the student as text modal data, collect the voice signals input by the student during the learning interaction as audio modal data, collect the picture content uploaded by the student as image modal data, and collect the learning video of the student captured by the camera as video modal data; Step 2: Process the input multimodal data, use the text reply generated for the student to read as the text output, and convert the text into audio through text-to-speech technology as the audio output; use deep learning models, such as GAN or Diffusion Models, to generate pictures for auxiliary teaching as the image output.

[0070] Step 3: Analyze the processed multimodal data, including: Speech preprocessing and feature extraction: Preprocess the input student speech and obtain the speech signal features; Sentiment classification and modeling: Based on the extracted speech signal features, identify the sentiment category and intensity; Image sentiment recognition includes: Obtain images by interval frame extraction of the acquired video; Facial expression recognition: Identify the movement of the facial key points in the image to analyze the expression changes; extract the facial features in the image and classify the expressions into basic emotions; Deep pose estimation: Extract the human pose and movement trajectory through a deep learning-based pose estimation model; Eye movement tracking: Use a convolutional neural network to analyze the student's eye movement trajectory and fixation points to infer their attention concentration.

[0071] Step 4: Collect the real-time updated student learning data, including historical learning records, homework and exam scores, and interaction logs; generate a comprehensive learning feedback report based on the collected student learning data.

[0072] The learning feedback report includes learning progress, performance analysis, emotional state, and suggestions for learning and interaction; Among them, the learning progress includes the tasks completed by the student within a specific time period, the knowledge point mastery situation, and the learning speed; The performance analysis includes the scoring situation, progress trend, and weak links of the student in each subject and each chapter; The emotional state evaluates the emotional fluctuations of the student during the recent learning process based on the emotional information provided by the emotion analysis module; The suggestions for learning and interaction generate personalized suggestions for learning and interaction based on the student's learning progress, performance analysis, and emotional state.

[0073] Step 5: Collect data such as multi-disciplinary standardized teaching materials, syllabuses, tutoring materials, and courseware; organize the collected data by knowledge points, and clearly label the difficulty, learning objectives, and relevance information of each knowledge point; construct knowledge graphs for each field, and form a structured graph of concepts, knowledge points, and their relationships in different disciplines; Step 6: Select a basic large language model as the basic model for generating and regulating conversations; input the collected dataset into the basic large language model, and set specific fine-tuning parameters; conduct multiple rounds of training and fine-tuning processes to enable the model to generate accurate and emotional feedback on educational content; Step 7: Knowledge overstepping control, specifically including the process of constructing a domain knowledge graph, the process of establishing a student knowledge graph, the process of an overstepping control mechanism based on the knowledge graph, and the post-processing and correction process; The process of constructing a domain knowledge graph includes establishing a knowledge graph covering the required subject fields; The process of establishing a student knowledge graph specifically updates the student's knowledge graph through the student's learning history data; The process of an overstepping control mechanism based on the knowledge graph specifically defines knowledge constraints and limits the constraint mechanism during the generation process; The post-processing and correction process includes knowledge consistency checking and a fallback mechanism; Knowledge consistency checking: Conduct knowledge consistency checking on the content generated by the model to ensure that the output knowledge points are consistent with the knowledge currently mastered by the student. For example, if the generated content contains an advanced knowledge point beyond the student's current knowledge range , then this content will be marked as overstepping by the system; Fallback mechanism: If the generated content violates the knowledge graph constraints, the system regenerates content that meets the constraints.

[0074] By introducing a multi-modal interaction module, the present invention can process multiple input signals from text, audio, images, and videos simultaneously, thereby comprehensively understanding the learning content, emotional state, and attention distribution of students. In particular, the proposed emotional consistency loss function and conflict correction mechanism solve the problem of emotional conflicts in multi-modal data and ensure the consistency of emotional expression. This innovative mechanism improves the accuracy of emotion recognition. Especially in complex situations, it can more precisely identify the emotional fluctuations of students, providing a more accurate emotional basis for personalized learning feedback.

[0075] By constructing a knowledge graph of the subject field and combining it with the student's personal knowledge graph, the system can dynamically evaluate the student's knowledge mastery level and avoid generating teaching content beyond the student's cognitive ability. Compared with traditional static teaching syllabuses or fixed knowledge recommendation systems, the present invention can real-time evaluate the student's learning status, generate personalized content through the constraint of the knowledge graph, and effectively avoid the situation of knowledge being beyond the syllabus or too simple.

[0076] Most of the existing technologies rely on historical data (such as students' grades and homework situations) to generate learning feedback, but these methods cannot timely reflect the current learning status and emotional fluctuations of students. In contrast, the core brain module of the present invention can real-time integrate information from multiple data sources. For example, the real-time learning data, emotional state, grade analysis, and constraint information of the knowledge graph of students, etc., to generate feedback that conforms to the student's cognitive state, emotional needs, and learning goals.

[0077] Compared with the existing technologies, the present invention mainly has the following innovative points: (1) Multi-modal fusion and conflict correction of emotion recognition: Different from the single-modal emotion recognition methods in the existing technologies, by introducing a multi-modal interaction module, the present invention can process multiple input signals from text, audio, images, and videos simultaneously, thereby comprehensively understanding the learning content, emotional state, and attention distribution of students. In particular, the proposed emotional consistency loss function and conflict correction mechanism solve the problem of emotional conflicts in multi-modal data and ensure the consistency of emotional expression. This innovative mechanism improves the accuracy of emotion recognition. Especially in complex situations, it can more precisely identify the emotional fluctuations of students, providing a more accurate emotional basis for personalized learning feedback. This is one of the core innovations of the present invention because its application in multi-modal emotion analysis requires in-depth integration across fields, not only involving speech processing, image recognition, and deep learning algorithms, but also solving the problems of emotional consistency and conflict handling.

[0078] (2) Out-of-syllabus control and personalized recommendation based on knowledge graph: Another innovation of the present invention is the out-of-syllabus control mechanism based on knowledge graph. By constructing a knowledge graph in the subject area and combining it with the student's personal knowledge graph, the system can dynamically evaluate the student's knowledge mastery level and avoid generating teaching content that exceeds the student's cognitive ability. Compared with traditional static teaching outlines or fixed knowledge recommendation systems, the present invention can evaluate the student's learning status in real time, generate personalized content through knowledge graph constraints, and effectively avoid situations where the knowledge is out of syllabus or too simple.

[0079] (3) Information integration and response based on fine-tuning the large model: Most existing technologies rely on historical data (such as students’ grades and homework) to generate learning feedback, but these methods cannot timely reflect students’ current learning status and emotional fluctuations. In contrast, the core brain module of the present invention can integrate information from multiple data sources in real time, such as students’ real-time learning data, emotional status, performance analysis, and constraint information of knowledge graphs, etc., to generate feedback that meets students’ cognitive status, emotional needs, and learning goals.

[0080] In the above description of specific embodiments of the present invention, features described and / or shown for one embodiment may be used in the same or similar manner in one or more other embodiments, combined with features in other embodiments, or replace features in other embodiments.

[0081] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.

[0082] In addition, the method of the present invention is not limited to being executed in the time sequence described in the specification, and may also be executed in other time sequences, in parallel or independently. Therefore, the execution order of the method described in this specification does not limit the technical scope of the present invention.

[0083] Although the present invention has been disclosed above by describing specific embodiments of the present invention, it should be understood that all the above embodiments and examples are exemplary rather than restrictive. Those skilled in the art may design various modifications, improvements or equivalents of the present invention within the spirit and scope of the appended claims. These modifications, improvements or equivalents should also be considered to be included in the protection scope of the present invention.

Claims

1. A personalized English education system based on multimodal sentiment analysis, characterized by: include: Multimodal interaction module: used to receive and process students' multimodal data, wherein the multimodal data includes text modal data, audio modal data, image modal data and video modal data; Emotion recognition module: used to perform emotion analysis on the input multimodal data, including speech emotion recognition, video emotion recognition and emotion information fusion; wherein the emotion information fusion adopts a weighted strategy and adjusts the final output emotion state through emotion consistency constraints and conflict correction mechanisms; Personalization module: evaluates the student's real-time learning status based on the student's historical learning data and generates a personalized learning feedback report; Core Brain Module: Used to dynamically adjust interactive feedback based on the emotional state of the emotion recognition module and the real-time learning state output by the personalization module.

2. The personalized English education system according to claim 1, characterized in that: In the emotion recognition module, the specific process of emotion consistency constraint includes: Input speech signals and image data, and extract speech emotion features and image emotion features respectively; Align the speech emotion features and image emotion features to obtain feature representation in the same emotion space; After feature alignment, the Euclidean distance between speech emotion features and image emotion features is calculated; the emotion consistency loss function is established based on the Euclidean distance and the cross entropy loss function; The comprehensive loss is obtained by jointly optimizing the sentiment consistency loss function and the basic classification loss.

3. The personalized English education system according to claim 1 or 2, characterized in that: In the emotion recognition module, the specific process of the conflict correction mechanism includes: The aligned feature space is used to calculate the confidence of each modality, including the confidence of speech emotion classification and image emotion classification: Conflict detection: First, the emotion with the highest confidence in each modality is used as the recognition result; if the highest classification confidence of a certain modality in all emotion categories is less than the preset confidence threshold, no conflict judgment is made, and the model is considered uncertain on this modality, and it is directly marked as "low confidence" output; if the highest confidence categories of each modality are different, the semantic similarity of the categories is calculated. If the semantics of the two are close, that is, less than the set category similarity threshold, it is not determined as a conflict; the confidence difference of the highest categories of the two modalities is calculated. If it is higher than the set confidence difference threshold, the prediction result of the modality with higher confidence is directly used, and it is not determined as a conflict; in short, only when the following conditions are met at the same time, the conflict is truly determined: the predicted categories are different, the category semantic similarity is low, and the confidence of the two modalities is high and the confidence difference is small; Priority setting: After a conflict is detected, the priority is set according to the directness of the modal expression and the weight is adjusted dynamically; Multimodal fusion: The speech emotion features and image emotion features in the unified space mapped by the alignment network are fused to obtain the fusion features.

4. The personalized English education system according to claim 1, characterized in that: The multimodal interaction module includes an input end and an output end; The input end includes text, audio, image and video, where text is text information asked by students or feedback; audio is voice signal input by students in learning interaction; image is picture content uploaded by students; video is learning video captured by camera; The output end includes text output, audio output and image output; the text output is a text reply generated for students to read; the audio output is the conversion of text into audio through speech synthesis technology; the image output is a picture generated by a deep learning model to assist teaching.

5. The personalized English education system according to claim 1, characterized in that: In the emotion recognition module, speech emotion recognition includes: Speech preprocessing and feature extraction: preprocess the input student speech and obtain speech signal features; Emotion classification and modeling: Identify emotion categories and intensities based on extracted speech signal features; Image emotion recognition includes: Extract frames from the acquired video at intervals to obtain images; Facial expression recognition: Identify the movement of facial key points in an image to analyze expression changes; extract facial features in an image and classify expressions into basic emotions; Deep posture estimation: extracting the posture and motion trajectory of the human body through a posture estimation model based on deep learning; Eye tracking: Use convolutional neural networks to analyze students’ eye movements and gaze points to infer their level of concentration.

6. The personalized English education system according to claim 1, characterized in that: The personalization module includes: Data collection: Collect student learning data that is updated in real time, including historical learning records, homework and test scores, and interaction logs; Data analysis and feedback report generation: Generate comprehensive learning feedback reports based on the collected student learning data.

7. The personalized English education system according to claim 6, characterized in that: The learning feedback report includes learning progress, performance analysis, emotional state, and learning and interaction suggestions; Among them, learning progress includes the tasks completed by students within a specific time period, the mastery of knowledge points and the learning speed; The performance analysis includes the student's scores in each subject and chapter, their progress trends, and their weak links; Emotional state evaluates students’ emotional fluctuations during the recent learning process based on the emotional information provided by the sentiment analysis module; Learning and interaction suggestions: Generate personalized learning and interaction suggestions based on students' learning progress, performance analysis and emotional state.

8. The personalized English education system according to claim 1, characterized in that: The core brain modules include data collection and processing, fine-tuning of large models, and knowledge control: The data collection and processing includes data collection, data annotation and knowledge graph construction; data collection is to collect standardized textbooks, syllabi, tutoring materials and courseware for multiple disciplines; data annotation is to organize the collected data by knowledge points, and clearly mark the difficulty, learning objectives and relevance information of each knowledge point; knowledge graph construction is to construct a knowledge graph of the field, forming a structured graph of concepts, knowledge points and the relationships between different disciplines; The fine-tuning of the large model includes selecting a model, building a training data set and fine-tuning the model; wherein, selecting a model: selecting a basic large language model as a basic model for generating and regulating dialogues; fine-tuning the model: inputting the collected data set into the basic large language model and setting specific fine-tuning parameters; performing multiple rounds of training and fine-tuning on the basic large language model to obtain a fine-tuned model, so that the model can generate accurate and emotional feedback on educational content; The control of knowledge beyond the syllabus includes constructing a domain knowledge graph, establishing a student knowledge graph, a knowledge graph-based control mechanism for knowledge beyond the syllabus, and post-processing and correction; Among them, constructing a domain knowledge graph is to establish a knowledge graph covering the required subject areas; Establishing a student knowledge graph specifically involves updating the student knowledge graph through the student's learning history data.

9. A personalized English teaching method based on multimodal sentiment analysis, characterized by: include: Receive and process multimodal data of students, including text, audio, image and video input signals; Perform sentiment analysis on the input multimodal data, including speech sentiment recognition, video sentiment recognition and sentiment information fusion; sentiment information fusion adopts a weighted strategy and adjusts the final output sentiment state through sentiment consistency constraints and conflict correction mechanisms; Evaluate the student's real-time learning status based on the student's historical learning data; Dynamically adjust interactive feedback based on emotional state and real-time learning status.

Citation Information

Patent Citations

  • Personalized teaching knowledge graph updating method and system based on multi-modal data

    CN117313852A

  • Multimedia teaching resource knowledge graph generation method and system based on recognition analysis

    CN117573904A

  • Multi-modal speech emotion recognition method based on relative entropy alignment fusion

    CN117672268A

  • Microteaching actual effect evaluation system based on artificial intelligence

    CN119740916A

  • Real-time processing method of automobile data based on artificial intelligence

    CN119783051A

Cited By

  • Teaching platform management system and method based on multi-source data analysis

    CN120387739A

  • Intelligent customer service dialogue generation method and system based on large model

    CN120873126A

  • Intelligent customer service dialogue generation method and system based on large model

    CN120873126B

  • Cognitive diagnosis method and model based on emotional state

    CN121148609A

  • A cognitive diagnosis method and model based on emotional state

    CN121148609B