Multimodal large model feedback emotion enhancement method and device, equipment and storage medium

By analyzing student feedback and multimodal interaction data and recognizing emotions, personalized feedback content is generated, which solves the problem of lack of emotional care in artificial intelligence education and improves the adaptive enhancement efficiency of feedback literacy.

CN119598128BActive Publication Date: 2025-11-04SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411681553.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-11-04
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

In existing technologies, AI-enabled education feedback interactions lack personalized emotional care for students during the feedback process, resulting in low student motivation for feedback and low efficiency in improving feedback literacy.

Method used

By acquiring student feedback, multimodal interaction data, and feedback history data, we analyze and perform sentiment recognition to generate personalized feedback content. Based on the human-computer symbiotic paradigm, we optimize the feedback literacy adaptive enhancement.

Benefits of technology

It effectively improves the efficiency of adaptive enhancement of feedback literacy in the learning process empowered by artificial intelligence, and realizes the personalization and efficient generation of emotional feedback content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119598128B_ABST
    Figure CN119598128B_ABST
Patent Text Reader

Abstract

The application discloses a multimodal large model feedback emotion enhancement method and device, equipment and a storage medium, which can be applied to the technical field of emotional data processing. The application obtains feedback literacy scores and feedback literacy levels by analyzing and quantifying a plurality of feedback literacy elements, simultaneously performs emotional recognition on multimodal interaction data of the current student to obtain interaction emotional tags, and analyzes feedback history data of the current student to obtain feedback history features. The feedback emotion preference personalized model is inputted with the feedback literacy scores and the interaction emotional tags to obtain feedback emotion preference tags. The multimodal large model is inputted with a feedback content theme to be generated, the feedback literacy level and the interaction emotional tags to perform emotional enhancement and obtain emotional enhancement feedback content. The emotional enhancement feedback content is optimized based on the human-machine symbiosis paradigm, so that the efficiency of the artificial intelligence empowerment student learning process feedback literacy self-adaptive enhancement can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of emotion data processing technology, and in particular to a multimodal large model feedback emotion enhancement method, apparatus, device and storage medium. Background Technology

[0002] In related technologies, feedback literacy refers to students' ability and attitude to actively participate in feedback and utilize it to promote learning; it is an important component of learner literacy. Emotional feedback can enhance students' learning motivation, improve learning satisfaction, and promote learning outcomes. However, in current AI-enabled education, the interaction between large models and students focuses only on knowledge-based question-and-answer sessions, lacking personalized emotional care during the feedback process, resulting in low student feedback motivation. To address the problem of low feedback motivation, existing technologies increase emotional expression in feedback interactions through methods such as fine-tuning or full-scale tuning of large models or simple prompt word engineering. However, these methods also have low corresponding model continuous optimization and learning efficiency, thus resulting in low efficiency in improving students' feedback literacy in the AI-enabled learning process.

[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention

[0004] The main objective of this application is to propose a multimodal large-model feedback emotion enhancement method, device, equipment, and storage medium, which can effectively improve the efficiency of adaptive enhancement of students' feedback literacy in the learning process empowered by artificial intelligence.

[0005] To achieve the above objectives, one aspect of this application proposes a multimodal large model feedback sentiment enhancement method, the method comprising the following steps:

[0006] Obtain feedback information from the current student, which includes multiple feedback literacy elements;

[0007] The feedback information is analyzed and quantified to obtain the feedback literacy score and feedback literacy level;

[0008] Obtain the first multimodal interaction data of the current student;

[0009] Emotion recognition is performed based on the first multimodal interaction data to obtain interaction emotion tags;

[0010] Obtain the current student's feedback history data;

[0011] The feedback history data is analyzed to obtain feedback history characteristics;

[0012] The feedback literacy score, the interaction sentiment label, and the feedback history features are input into the student feedback sentiment preference personalized model to obtain the current student's feedback sentiment preference label.

[0013] Obtain the topic of the feedback content to be generated for the current student;

[0014] The feedback sentiment preference tags, the topic of the feedback content to be generated, the feedback literacy level, and the interaction sentiment tags are input into a multimodal big data model for sentiment enhancement to obtain sentiment-enhanced feedback content.

[0015] The feedback content for enhanced emotion is optimized based on the human-machine symbiosis paradigm.

[0016] In some embodiments, the step of analyzing and quantifying the feedback information to obtain a feedback literacy score and a feedback literacy level includes:

[0017] The multiple feedback literacy elements are dimensionally divided;

[0018] The multiple feedback literacy elements are quantified to obtain the quantitative index scores of the feedback literacy elements.

[0019] The feedback literacy dimension score is calculated for each dimension based on the quantitative index scores of the feedback literacy elements.

[0020] The feedback literacy score is calculated based on the feedback literacy dimension score;

[0021] The current student's feedback literacy level is determined based on the feedback literacy score and the preset level range.

[0022] In some embodiments, the step of performing emotion recognition based on the first multimodal interaction data to obtain interaction emotion tags includes:

[0023] The first multimodal interaction data is preprocessed to obtain the second multimodal interaction data;

[0024] Feature extraction is performed on the second multimodal interaction data to obtain multimodal data features;

[0025] Based on the multimodal data features, emotion recognition is performed on the second multimodal interaction data to obtain the interaction emotion tag.

[0026] In some embodiments, the step of extracting features from the second multimodal interaction data to obtain multimodal data features includes:

[0027] The feature extractor extracts features of each modality in the second multimodal interaction data as unimodal features; wherein, the feature extractor performs optimization processing through modality similarity and contrast loss during feature extraction;

[0028] The fusion features of each modality are calculated based on the single-modal features and used as the multimodal data features.

[0029] In some embodiments, the step of performing emotion recognition on the second multimodal interaction data based on the multimodal data features to obtain the interaction emotion tag includes:

[0030] Calculate the similarity between each of the multimodal data features and the remaining multimodal data features;

[0031] Calculate the contrast loss based on the similarity;

[0032] Construct an objective function based on the contrastive loss;

[0033] The second multimodal interaction data is subjected to emotion recognition based on the objective function to obtain the interaction emotion tag.

[0034] In some embodiments, the step of inputting the feedback sentiment preference tag, the topic of the feedback content to be generated, the feedback literacy level, and the interaction sentiment tag into a multimodal large model for sentiment enhancement to obtain sentiment-enhanced feedback content includes:

[0035] A multimodal prompt template is constructed based on the topic of the feedback content to be generated, the feedback sentiment preference tag, the feedback literacy level, and the interaction sentiment tag;

[0036] Emotion control parameters are generated based on the feedback emotion preference tags, the feedback literacy level, and the interaction emotion tags;

[0037] A preset function is constructed based on the multimodal prompt template, the emotion control parameters, and the multimodal large model;

[0038] The emotion is enhanced according to the preset function, and the feedback content of the emotion enhancement is obtained.

[0039] In some embodiments, when inputting the feedback sentiment preference label, the topic of the feedback content to be generated, the feedback literacy level, and the interaction sentiment label into a multimodal large model for sentiment enhancement, the method further includes the following steps:

[0040] The feedback content for the emotion enhancement across different modalities is adjusted to ensure consistency.

[0041] To achieve the above objectives, another aspect of this application proposes a multimodal large model feedback emotion enhancement device, the device comprising:

[0042] The first module is used to obtain feedback information from the current student, and the feedback information includes multiple feedback literacy elements.

[0043] The second module is used to analyze and quantify the feedback information to obtain feedback literacy scores and feedback literacy levels.

[0044] The third module is used to acquire the first multimodal interaction data of the current student;

[0045] The fourth module is used to perform emotion recognition based on the first multimodal interaction data to obtain interaction emotion tags;

[0046] The fifth module is used to obtain the current student's feedback history data;

[0047] The sixth module is used to analyze the feedback history data to obtain feedback history characteristics;

[0048] The seventh module is used to input the feedback literacy score, the interaction sentiment label and the feedback history features into the student feedback sentiment preference personalized model to obtain the current student's feedback sentiment preference label;

[0049] The eighth module is used to obtain the topic of the feedback content to be generated for the current student;

[0050] The ninth module is used to input the feedback sentiment preference tags, the topic of the feedback content to be generated, the feedback literacy level, and the interaction sentiment tags into the multimodal big model for sentiment enhancement, so as to obtain sentiment-enhanced feedback content;

[0051] The tenth module is used to optimize the feedback content for enhanced emotion based on the human-machine symbiosis paradigm.

[0052] To achieve the above objectives, another aspect of this application provides an electronic device, comprising:

[0053] At least one processor;

[0054] At least one memory for storing at least one program;

[0055] When the at least one program is executed by the at least one processor, the at least one processor performs the method described above.

[0056] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0057] The embodiments of this application include at least the following beneficial effects: This application provides a multimodal large-scale model feedback emotion enhancement method, apparatus, device, and storage medium. This scheme analyzes and quantifies feedback information including multiple feedback literacy elements to obtain feedback literacy scores and feedback literacy levels. Simultaneously, it performs emotion recognition on the current student's multimodal interaction data to obtain interaction emotion tags, and analyzes the current student's feedback history data to obtain feedback history features. Then, the feedback literacy score, interaction emotion tags, and feedback history features are input into a personalized student feedback emotion preference model to obtain the current student's feedback emotion preference tags. The feedback emotion preference tags, the topic of the feedback content to be generated, the feedback literacy level, and the interaction emotion tags are then input into a multimodal large-scale model for emotion enhancement to obtain emotion-enhanced feedback content. Based on the human-machine symbiotic paradigm, the emotion-enhanced feedback content is optimized, thereby effectively improving the efficiency of AI-enabled adaptive enhancement of student learning feedback literacy. Attached Figure Description

[0058] Figure 1 This is a flowchart of the multimodal large model feedback emotion enhancement method provided in the embodiments of this application;

[0059] Figure 2 This is a schematic diagram of the structure of the multimodal large model feedback emotion enhancement device provided in the embodiments of this application;

[0060] Figure 3 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application.

[0062] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”

[0063] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0065] Before providing a detailed description of the embodiments of this application, some of the nouns and terms used in the embodiments of this application will be explained first. The nouns and terms used in the embodiments of this application shall be interpreted as follows:

[0066] The symbiotic paradigm is a system design and operation model that emphasizes collaboration, mutual benefit, and co-evolution between humans and artificial intelligence systems. In this paradigm, the relationship between humans and AI is not simply one of user and being used, but rather a mutually beneficial and co-growing ecosystem.

[0067] This application provides a method, apparatus, device, and storage medium for enhancing feedback emotion in a multimodal large-scale model. This application analyzes and quantifies feedback information including multiple feedback literacy elements to obtain feedback literacy scores and levels. Simultaneously, it performs emotion recognition on the current student's multimodal interaction data to obtain interaction emotion tags, and analyzes the current student's feedback history data to obtain feedback history features. Then, the feedback literacy score, interaction emotion tags, and feedback history features are input into a personalized student feedback emotion preference model to obtain the current student's feedback emotion preference tags. The feedback emotion preference tags, the topic of the feedback content to be generated, the feedback literacy level, and the interaction emotion tags are then input into a multimodal large-scale model for emotion enhancement to obtain emotion-enhanced feedback content. Based on a human-machine symbiotic paradigm, the emotion-enhanced feedback content is optimized, thereby effectively improving the efficiency of adaptive enhancement of student feedback literacy in the AI-enabled learning process.

[0068] The multimodal large model feedback sentiment enhancement method provided in this application relates to the field of sentiment data processing technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the multimodal large model feedback sentiment enhancement method, but is not limited to the above forms.

[0069] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0070] Figure 1 This is an optional flowchart of the multimodal large model feedback sentiment enhancement method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S00 to S90.

[0071] Step S100: Obtain feedback information from the current student. The feedback information includes multiple feedback literacy elements.

[0072] Step S110: Analyze and quantify the feedback information to obtain the feedback literacy score and feedback literacy level;

[0073] Step S120: Obtain the first multimodal interaction data of the current student;

[0074] Step S130: Perform emotion recognition based on the first multimodal interaction data to obtain interaction emotion tags;

[0075] Step S140: Obtain the current student's feedback history data;

[0076] Step S150: Analyze the historical feedback data to obtain historical feedback characteristics;

[0077] Step S160: Input the feedback literacy score, interaction sentiment label and feedback history features into the personalized student feedback sentiment preference model to obtain the current student's feedback sentiment preference label;

[0078] Step S170: Obtain the topic of the feedback content to be generated for the current student;

[0079] Step S180: Input the feedback sentiment preference tags, the topic of the feedback content to be generated, the feedback literacy level, and the interaction sentiment tags into the multimodal big data model for sentiment enhancement to obtain sentiment-enhanced feedback content;

[0080] Step S190: Optimize the feedback content for emotion enhancement based on the human-machine symbiosis paradigm.

[0081] In this embodiment, to monitor students' feedback literacy levels in real time and achieve adaptive enhancement of subsequent emotional feedback, this embodiment automatically analyzes and quantifies the feedback information after obtaining it. Here, feedback information refers to feedback data on student activity behavior. It is understood that the process of automatically analyzing and quantifying the feedback information in this embodiment includes, but is not limited to, the following steps:

[0082] Step S210: Divide multiple feedback literacy elements into dimensions;

[0083] Step S220: Quantify multiple feedback literacy elements to obtain quantitative index scores for the feedback literacy elements;

[0084] Step S230: Calculate the feedback literacy dimension score in each dimension based on the quantitative index scores of the feedback literacy elements;

[0085] Step S240: Calculate the feedback literacy score based on the feedback literacy dimension scores;

[0086] Step S250: Determine the current student's feedback literacy level based on the feedback literacy score and the preset level range.

[0087] In this embodiment, feedback information can be divided into four dimensions: feedback awareness, feedback strategy, feedback effect, and feedback emotion. Feedback awareness refers to students' understanding of the importance and value of feedback; feedback strategy refers to the methods and techniques students use in the feedback process; feedback effect refers to the impact of feedback on learning goals and grades; and feedback emotion refers to the positive or negative emotional reactions students experience during the feedback process. Each dimension contains several specific elements, as shown in Table 1.

[0088] Table 1

[0089]

[0090]

[0091] Furthermore, to transform feedback literacy elements into calculable data, this embodiment employs multiple data sources and data types to quantify each feedback literacy element. Specifically, these multiple data sources include, but are not limited to, student learning behavior data, academic performance data, learning emotion data, and learning feedback data. Learning behavior data refers to various operations and activities of students on the learning platform, such as login frequency, learning duration, learning progress, learning frequency, and learning interaction; academic performance data refers to the results of various assessments and tests conducted by students on the learning platform, such as homework scores, quiz scores, and exam scores; learning emotion data refers to various emotional expressions and feedback from students on the learning platform, such as emoticons, emotional words, and emotional evaluations; and learning feedback data refers to various feedback information received and given by students on the learning platform, such as feedback content, feedback format, feedback frequency, and feedback quality. Based on the above data, this embodiment combines a series of quantitative indicators to measure the quantitative indicator score of each feedback literacy element, as shown in Table 2:

[0092] Table 2

[0093]

[0094] It is understood that, in this embodiment, a weighted average method is used to calculate the score y of each feedback literacy dimension based on the quantitative indicators of each feedback literacy element. i Specifically, as shown in Formula 1:

[0095]

[0096] In the formula, w i It is the weight of the i-th feedback literacy element, x i is the quantitative score of the i-th feedback literacy element, and n is the number of elements under the feedback literacy dimension. Here, the weight w i The determination can be made based on expert scoring or data analysis methods.

[0097] After obtaining the feedback literacy dimension scores for each dimension, the feedback literacy score z is calculated using the analytic hierarchy process (AHP) according to Formula 2:

[0098]

[0099] In the formula, w j It is the weight of the j-th feedback literacy dimension, y j is the score of the j-th feedback literacy dimension, and m is the number of feedback literacy dimensions. Weight v j The determination can also be based on expert scoring or data analysis methods.

[0100] Then, based on the feedback literacy score and the preset level range, the current student's feedback literacy level L is determined. Specifically, this embodiment uses a five-level evaluation method to divide the feedback literacy level into excellent, good, average, poor, and insufficient, as shown in Table 3:

[0101] Table 3

[0102]

[0103] As shown in Table 3, feedback literacy reflects students' abilities and attitudes toward feedback in terms of need, trust, acceptance, seeking, utilization, reflection, satisfaction, improvement, achievement, and emotion, and has an important impact on the personalized generation of students' emotional feedback content.

[0104] In this embodiment of the application, the process of obtaining interaction emotion tags by performing emotion recognition based on the first multimodal interaction data includes, but is not limited to, the following steps:

[0105] Step S310: Preprocess the first multimodal interaction data to obtain the second multimodal interaction data;

[0106] Step S320: Extract features from the second multimodal interaction data to obtain multimodal data features;

[0107] Step S330: Perform emotion recognition on the second multimodal interaction data based on the multimodal data features to obtain the interaction emotion tag.

[0108] Understandably, the first type of multimodal interaction data can be collected in real time using various sensors and devices, such as cameras, microphones, keyboards, mice, and touchscreens, to capture multimodal data from students during the interactive feedback process. This includes, but is not limited to, visual data, speech data, text data, and behavioral data. Visual data refers to students' facial expressions, eye movements, and head posture; speech data refers to students' speech signals, intonation, and speech rate; text data refers to students' input content, language style, and emotional vocabulary; and behavioral data refers to students' operation frequency, operation duration, and operation type.

[0109] After obtaining the first multimodal interaction data, data cleaning, data alignment, and data standardization are performed on it. Data cleaning refers to removing invalid, noisy, and abnormal data; data alignment refers to synchronizing data from different modalities according to timestamps; and data standardization refers to converting data from different modalities into a unified format and scale.

[0110] After data preprocessing, feature extraction is performed on the second multimodal interaction data to obtain multimodal data features. Specifically, a feature extractor extracts features of each modality from the second multimodal interaction data as unimodal features, and then calculates the fused features of each modality based on the unimodal features as the multimodal data features. The feature extractor optimizes the feature extraction process using modality similarity and contrastive loss. It is understood that this embodiment can extract features helpful for emotion recognition from multimodal interaction data by utilizing a highly semantically centered cross-sample fusion network. Specifically, firstly, a unimodal feature extractor extracts features from each modality as unimodal features; then, highly semantically centered unimodal contrastive learning is used to preserve modality-specific information; finally, highly semantically centered cross-sample fusion is used to fuse the features of different modalities.

[0111] For example, the processing procedure in this embodiment can be adopted in the following manner:

[0112] First, extract single-modal features: For each modality m, use the feature extractor f. m From the second multimodal interaction data x m Extracting features z m , i.e. z m =f m (x m ).

[0113] Second, highly semantically centered single-modal contrastive learning: For each modality m, calculate its similarity s with highly semantic (e.g., text, meaningful numerical values) modalities. mt And use contrast loss L m To optimize the feature extractor f m ,Right now and Where d is the feature dimension and τ is the temperature parameter. This embodiment focuses on how to use textual information to guide feature extraction and alignment for other modalities, in order to preserve modality-specific information.

[0114] Third, highly semantically centered cross-sample fusion: calculate the fusion feature h for each modality m. m ,Right now in is the attention weight, and N is the number of samples.

[0115] This embodiment improves the accuracy and robustness of emotion feature representation learning by extracting features that are helpful for emotion recognition from multimodal interaction data while preserving modality-specific information.

[0116] After extracting multimodal data features, sentiment recognition is performed on the second multimodal interaction data based on these features. It is understood that this embodiment can utilize a highly semantically modal-centered cross-sample fusion network for sentiment recognition of multimodal interaction data, effectively integrating features from different modalities to construct a multimodal sentiment recognition model, thereby enabling real-time recognition of student interaction emotions. Specifically, bidirectional contrastive learning of the fusion modality is first used to optimize the fusion features, and then an objective function and prediction are used to achieve sentiment recognition. For example, the processing procedure in this embodiment can be as follows:

[0117] First, bidirectional contrastive learning of fused modalities: calculating the fused features h of each modality m. m Fusion features h with residual modes m′ similarity s mm′ And use contrast loss L mm′ To optimize fusion features, namely and Where d is the feature dimension and τ is the temperature parameter. This embodiment improves the accuracy and robustness of emotion recognition by effectively integrating features from different modalities.

[0118] Step 2, Objective Function and Prediction: In this embodiment, an objective function J is constructed based on the contrastive loss. The objective function J includes the sentiment classification loss L. c , contrast loss L mm′ And the regularization term R, i.e., J = L c +λL mm′ +γR, where λ and γ are weight parameters. In this embodiment, gradient descent is used to optimize the objective function to obtain the optimal model parameters. Then, based on this objective function, sentiment recognition is performed on multimodal data to obtain interactive sentiment labels E.

[0119] Specifically, the interactive emotion label E is a multivariate vector representing the student's emotional state during the interactive feedback process, such as happiness, sadness, anger, fear, surprise, and disgust. The interactive emotion label E reflects the student's emotional response during the interactive feedback process and has a significant impact on the personalized generation of the student's emotional feedback content. This embodiment further enhances the complementarity and synergy between modalities through a bidirectional comparative learning method that integrates modalities, enabling real-time identification of student interactive emotions.

[0120] Upon obtaining the interactive sentiment label E, this embodiment also analyzes the student's feedback history data to obtain each student's feedback history feature H, which serves as one of the inputs to the personalized model of student feedback sentiment preferences. Feedback history data refers to various feedback information received and given by students on the learning platform, such as feedback content, feedback format, feedback frequency, and feedback quality. The feedback history feature H is a multivariate vector representing the student's preferences, habits, and style regarding feedback content, such as the topic, length, tone, language, and emotion of the feedback content. The feedback history feature H reflects the student's personalized needs for feedback content and has a significant impact on the personalized generation of students' emotional feedback content.

[0121] Understandably, a personalized model of student feedback sentiment preferences can be constructed using a multilayer perceptron (MLP). This model P takes feedback literacy score z, interaction sentiment label E, and feedback history features H as input, and outputs the student's feedback sentiment preference label S = {s1, s2, ..., s...}. k Let k be the number of sentiment dimensions. The feedback sentiment preference label S serves as the input to multimodal AIGC, guiding the generation of multimodal content for enhanced sentiment feedback. Specifically, the feedback sentiment preference label S is a multivariate vector representing the student's sentiment preference for feedback content, such as positive, negative, neutral, motivating, encouraging, comforting, praising, or criticizing. The feedback sentiment preference label S is the core output of the personalized model of student feedback sentiment preferences and is key to achieving personalized generation of student sentiment feedback content.

[0122] In this embodiment, after obtaining the feedback sentiment preference label S, this embodiment can utilize the generation capabilities of a multimodal large model, combined with the student's feedback sentiment preferences, to generate multimodal feedback content with enhanced sentiment effects. It is understood that this embodiment includes, but is not limited to, the following steps:

[0123] Step S410: Construct a multimodal prompt template based on the topic T of the feedback content to be generated, the feedback sentiment preference label S, the feedback literacy level L, and the interaction sentiment label E; L∈{Excellent, Good, Average, Poor, Insufficient};

[0124] Step S420: Generate emotion control parameters based on feedback emotion preference tags, feedback literacy level, and interaction emotion tags;

[0125] Step S430: Construct a preset function based on the multimodal cue template, emotion control parameters, and multimodal large model;

[0126] Step S440: Perform emotion enhancement according to the preset function to obtain emotion-enhanced feedback content.

[0127] Specifically, the construction process of a multimodal prompt template can be achieved by constructing a function `ConstructEmotionalPrompt(T,S,L,E)`, which integrates the input information into a multimodal prompt template `Ψ`. The specific implementation is as follows: `Ψ = Concat(W` t ·T,W s ·S,W l ·L,W e ·E,Ψ base ), where W t W s W l W e These are the embedding matrices of the corresponding inputs, Ψ base It is a basic multimodal cue template, which ensures that emotional factors are effectively incorporated into the cue template.

[0128] The process of constructing emotion control parameters can be achieved by designing a function `GenerateEmotionControlParams(S,L,E)` to generate parameters Λ that control the intensity of emotions during the generation process. This function maps emotion preferences, feedback literacy levels, and interactive emotions to control parameters for each modality: Λ={λ text ,λ image ,λ audio} = f(S,L,E), where f is a mapping function that generates appropriate emotion control parameters based on the input.

[0129] Then, a predefined function `EmotionalMultiModalGenerate(M,Ψ,Λ)` is constructed based on the multimodal cue template Ψ, the emotion control parameter Λ, and the multimodal large model M. This function uses the multimodal large model M to generate emotion-enhanced feedback content C based on the cue template Ψ and the emotion control parameter Λ. This function generates text, image, and audio content respectively, and applies the corresponding emotion control parameters during the generation process to ensure that the generated content conforms to the required emotional characteristics.

[0130] In this embodiment, when performing sentiment enhancement using a multimodal large model, the consistency of the sentiment enhancement feedback content across different modalities is also adjusted. Specifically, this embodiment designs a function, AdjustModalConsistency(C), to ensure content consistency across different modalities. This function primarily performs the following operations: a) ensuring that the generated image content is semantically consistent with the text description; b) ensuring that the generated audio content is semantically consistent with the text content. This function ensures the semantic and emotional uniformity of the multimodal feedback content.

[0131] This embodiment fully utilizes the capabilities of a multimodal large model and designs prompts and control parameters to incorporate the required emotional features during the generation process. This achieves more accurate and efficient emotional feedback enhancement, improves the quality and relevance of feedback content, and provides new technical support for personalized learning and emotional intelligence education.

[0132] In this embodiment, after receiving emotionally enhanced feedback, the feedback content can be continuously optimized based on a human-machine symbiotic paradigm. Specifically, this embodiment can achieve continuous evolution through external knowledge base updates, prompt engineering optimization, and in-context learning, thereby continuously improving students' feedback literacy. It is understood that the continuous optimization process includes, but is not limited to, the following steps:

[0133] Step S510: Collect feedback literacy data: Collect key data on students' use of affective feedback.

[0134] Student feedback literacy score R = {r1, r2, ..., r n This corresponds to the feedback literacy elements defined in section 4.1;

[0135] Students' emotional response to generated content E={e1,e2,...,e m};

[0136] Student feedback behavior data B = {b1, b2, ..., b} k}, such as feedback frequency, feedback quality, etc.;

[0137] Step S520: Update the external knowledge base: Based on the collected data, update the external knowledge base K: K t+1 =f K (K t ,R,E,B), where f K This is the knowledge base update function, where t represents the current iteration number;

[0138] Step S530, Prompt Engineering Optimization: Based on the updated knowledge base, optimize the above multimodal prompt template Ψ: Ψ t+1 =f Ψ (Ψ f ,K t+1 ), where f Ψ It is a prompt template optimization function;

[0139] Step S540: Update In-context learning: Maintain a personalized context memory M for each student i. i And update: Where f M It is a context memory update function;

[0140] Step S550, Adaptive Feedback Generation: When generating new emotionally enhanced feedback, the optimized prompt template, updated knowledge base, and personalized contextual memory are combined. Where f F It is an adaptive feedback generation function, where T, S, L, and E represent the feedback topic, emotional preference, literacy level, and interactive emotion, respectively.

[0141] Step S560, Evaluate the effect of feedback literacy improvement: Evaluate the effect of affective feedback on improving students' feedback literacy: ΔR = f Δ (R before ,R after ), where f Δ It is a feedback-based competency improvement assessment function;

[0142] Step S570, Loop Optimization: Continuously execute steps 510 to S560 to form a closed-loop optimization process. Specifically, the iterative formula can be expressed as: (K,Ψ,M) i ,F) t+1 =f opt ((K,Ψ,M i ,F) t ,R,E,B), where f opt Represents the entire optimization loop;

[0143] Step S580, Tracking Long-Term Effects: Track and record the long-term trends in students' feedback literacy: T R =f T ({R t}), where f T It is a long-term progress tracking function, R t This represents the feedback literacy score at time point t.

[0144] This embodiment establishes a dynamically evolving external system (K,Ψ,M) i This approach enhances the emotional feedback capabilities of multimodal large models, enabling continuous optimization of the generated emotionally enhanced feedback content and continuous improvement of students' feedback literacy. This improves the adaptability and personalization of the system applied by the method in this application, and provides a new paradigm for sustainable development of AI-assisted education. Through human-machine symbiosis, it can learn and evolve together with students, continuously providing more accurate and effective emotionally enhanced feedback, thereby continuously improving students' feedback literacy.

[0145] In this embodiment of the application, the application of the human-machine symbiosis paradigm is mainly reflected in the following aspects:

[0146] Two-way interaction: The system used in this application method can not only provide students with emotionally enhanced feedback, but also learn and improve from students' responses.

[0147] Co-evolution: As students' feedback literacy improves, the system used in this application method is also constantly optimizing its feedback generation capabilities.

[0148] Personalized adaptation: The system used in this application method adapts to each student's unique needs and preferences through continuous learning.

[0149] Knowledge sharing: The external knowledge base of the system used in this application continuously accumulates new knowledge from student interactions, forming a dynamically updated knowledge ecosystem.

[0150] Complementary advantages: The system applied by the method in this application utilizes the computing power of AI and the creative thinking of humans, combining their strengths to achieve a 1+1>2 effect.

[0151] Continuous optimization: The method applied in this application uses a cyclical feedback mechanism, and both the system and the students involved in this application method are constantly improving and growing in this process.

[0152] Shared Goal: The system and students using this application method work together to improve learning outcomes and feedback skills, forming a symbiotic relationship with aligned goals.

[0153] This symbiotic paradigm transforms AI-assisted education systems from static tools into intelligent partners that grow and evolve alongside learners. This provides a new approach to the application of artificial intelligence in education and promises to bring about more personalized, intelligent, and sustainable learning experiences.

[0154] Reference Figure 2 This application provides a multimodal large model feedback emotion enhancement device, the device comprising:

[0155] The first module 600 is used to obtain feedback information from the current student, which includes multiple feedback literacy elements.

[0156] The second module 610 is used to analyze and quantify the feedback information to obtain the feedback literacy score and the feedback literacy level.

[0157] The third module, 620, is used to obtain the first multimodal interaction data of the current student;

[0158] The fourth module 630 is used to perform emotion recognition based on the first multimodal interaction data to obtain interaction emotion tags;

[0159] Module 5, number 640, is used to obtain the current student's feedback history data.

[0160] Module 650 is used to analyze historical feedback data to obtain historical feedback characteristics;

[0161] Module 7, 660, is used to input feedback literacy scores, interaction sentiment labels, and feedback history characteristics into the personalized student feedback sentiment preference model to obtain the current student's feedback sentiment preference label.

[0162] Module 8, number 670, is used to obtain the topic of the feedback content to be generated for the current student.

[0163] Module 9, 680, is used to input feedback sentiment preference tags, the topic of the feedback content to be generated, the feedback literacy level, and the interaction sentiment tags into the multimodal big model for sentiment enhancement, and obtain sentiment-enhanced feedback content.

[0164] Module 10, 690, is used to optimize the feedback content for emotion enhancement based on the human-machine symbiosis paradigm.

[0165] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0166] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described multimodal large-model feedback emotion enhancement method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0167] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0168] Please see Figure 3 , Figure 3 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0169] The processor 710 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0170] The memory 720 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 720 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 720 and called and executed by the processor 710 using the multimodal large-model feedback emotion enhancement method of the embodiments of this application.

[0171] The input / output interface 730 is used to implement information input and output;

[0172] The communication interface 740 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0173] Bus 750 transmits information between various components of the device (e.g., processor 710, memory 720, input / output interface 730, and communication interface 740);

[0174] The processor 710, memory 720, input / output interface 730 and communication interface 740 are connected to each other within the device via bus 750.

[0175] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described multimodal large model feedback emotion enhancement method.

[0176] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0177] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0178] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0179] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0180] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0181] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0182] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0183] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0184] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0185] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0186] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0187] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0188] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A multi-modal large model feedback emotion enhancement method, characterized in that, The method includes the following steps: Obtain feedback information from the current student, which includes multiple feedback literacy elements; The feedback information is analyzed and quantified to obtain the feedback literacy score and feedback literacy level; Obtain the first multimodal interaction data of the current student; Emotion recognition is performed based on the first multimodal interaction data to obtain interaction emotion tags; Obtain the current student's feedback history data; The feedback history data is analyzed to obtain feedback history characteristics; The feedback literacy score, the interaction sentiment label, and the feedback history features are input into the student feedback sentiment preference personalized model to obtain the current student's feedback sentiment preference label. Obtain the topic of the feedback content to be generated for the current student; The feedback sentiment preference tags, the topic of the feedback content to be generated, the feedback literacy level, and the interaction sentiment tags are input into a multimodal big data model for sentiment enhancement to obtain sentiment-enhanced feedback content. The feedback content for enhanced emotion is optimized based on the human-machine symbiosis paradigm; The step of analyzing and quantifying the feedback information to obtain feedback literacy scores and feedback literacy levels includes: The multiple feedback literacy elements are dimensionally divided; The multiple feedback literacy elements are quantified to obtain the quantitative index scores of the feedback literacy elements. The feedback literacy dimension score is calculated for each dimension based on the quantitative index scores of the feedback literacy elements. The feedback literacy score is calculated based on the feedback literacy dimension score; The current student's feedback literacy level is determined based on the feedback literacy score and the preset level range; The step of performing emotion recognition based on the first multimodal interaction data to obtain interaction emotion tags includes: The first multimodal interaction data is preprocessed to obtain the second multimodal interaction data; Feature extraction is performed on the second multimodal interaction data to obtain multimodal data features; Based on the multimodal data features, emotion recognition is performed on the second multimodal interaction data to obtain the interaction emotion tag; The step of extracting features from the second multimodal interaction data to obtain multimodal data features includes: The feature extractor extracts features of each modality in the second multimodal interaction data as unimodal features; wherein, the feature extractor performs optimization processing through modality similarity and contrast loss during feature extraction; The fusion features of each modality are calculated based on the single-modal features and used as the multimodal data features; The step of performing emotion recognition on the second multimodal interaction data based on the multimodal data features to obtain the interaction emotion tag includes: Calculate the similarity between each of the multimodal data features and the remaining multimodal data features; Calculate the contrast loss based on the similarity; Construct an objective function based on the contrastive loss; Emotion recognition is performed on the second multimodal interaction data according to the objective function to obtain the interaction emotion tag; The step of inputting the feedback sentiment preference tag, the topic of the feedback content to be generated, the feedback literacy level, and the interaction sentiment tag into a multimodal large model for sentiment enhancement to obtain sentiment-enhanced feedback content includes: A multimodal prompt template is constructed based on the topic of the feedback content to be generated, the feedback sentiment preference tag, the feedback literacy level, and the interaction sentiment tag; Emotion control parameters are generated based on the feedback emotion preference tags, the feedback literacy level, and the interaction emotion tags; A preset function is constructed based on the multimodal prompt template, the emotion control parameters, and the multimodal large model; The emotion is enhanced according to the preset function, and the feedback content of the emotion enhancement is obtained.

2. The method of claim 1, wherein, When inputting the feedback sentiment preference label, the topic of the feedback content to be generated, the feedback literacy level, and the interaction sentiment label into a multimodal large model for sentiment enhancement, the method further includes the following steps: The feedback content for the emotion enhancement across different modalities is adjusted to ensure consistency.

3. A multi-modal large model feedback sentiment enhancement device, characterized in that, The device includes: The first module is used to obtain feedback information from the current student, and the feedback information includes multiple feedback literacy elements. The second module is used to analyze and quantify the feedback information to obtain feedback literacy scores and feedback literacy levels. The third module is used to acquire the first multimodal interaction data of the current student; The fourth module is used to perform emotion recognition based on the first multimodal interaction data to obtain interaction emotion tags; The fifth module is used to obtain the current student's feedback history data; The sixth module is used to analyze the feedback history data to obtain feedback history characteristics; The seventh module is used to input the feedback literacy score, the interaction sentiment label and the feedback history features into the student feedback sentiment preference personalized model to obtain the current student's feedback sentiment preference label; The eighth module is used to obtain the topic of the feedback content to be generated for the current student; The ninth module is used to input the feedback sentiment preference tags, the topic of the feedback content to be generated, the feedback literacy level, and the interaction sentiment tags into the multimodal big model for sentiment enhancement, so as to obtain sentiment-enhanced feedback content; The tenth module is used to optimize the feedback content for enhanced emotion based on the human-machine symbiosis paradigm; The step of analyzing and quantifying the feedback information to obtain feedback literacy scores and feedback literacy levels includes: The multiple feedback literacy elements are dimensionally divided; The multiple feedback literacy elements are quantified to obtain the quantitative index scores of the feedback literacy elements. The feedback literacy dimension score is calculated for each dimension based on the quantitative index scores of the feedback literacy elements. The feedback literacy score is calculated based on the feedback literacy dimension score; The current student's feedback literacy level is determined based on the feedback literacy score and the preset level range; The step of performing emotion recognition based on the first multimodal interaction data to obtain interaction emotion tags includes: The first multimodal interaction data is preprocessed to obtain the second multimodal interaction data; Feature extraction is performed on the second multimodal interaction data to obtain multimodal data features; Based on the multimodal data features, emotion recognition is performed on the second multimodal interaction data to obtain the interaction emotion tag; The step of extracting features from the second multimodal interaction data to obtain multimodal data features includes: The feature extractor extracts features of each modality in the second multimodal interaction data as unimodal features; wherein, the feature extractor performs optimization processing through modality similarity and contrast loss during feature extraction; The fusion features of each modality are calculated based on the single-modal features and used as the multimodal data features; The step of performing emotion recognition on the second multimodal interaction data based on the multimodal data features to obtain the interaction emotion tag includes: Calculate the similarity between each of the multimodal data features and the remaining multimodal data features; Calculate the contrast loss based on the similarity; Construct an objective function based on the contrastive loss; Emotion recognition is performed on the second multimodal interaction data according to the objective function to obtain the interaction emotion tag; The step of inputting the feedback sentiment preference tag, the topic of the feedback content to be generated, the feedback literacy level, and the interaction sentiment tag into a multimodal large model for sentiment enhancement to obtain sentiment-enhanced feedback content includes: A multimodal prompt template is constructed based on the topic of the feedback content to be generated, the feedback sentiment preference tag, the feedback literacy level, and the interaction sentiment tag; Emotion control parameters are generated based on the feedback emotion preference tags, the feedback literacy level, and the interaction emotion tags; A preset function is constructed based on the multimodal prompt template, the emotion control parameters, and the multimodal large model; The emotion is enhanced according to the preset function, and the feedback content of the emotion enhancement is obtained.

4. An electronic device, comprising: include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method as described in any one of claims 1 to 2.

5. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 4. When the computer program is executed by a processor, it implements the method of any one of claims 1 to 2.

Citation Information

Patent Citations

  • Affective interaction systems, devices, and methods based on affective computing user interface

    US11226673B2

  • KR20230024095A