A multi-dimensional emotion perception and intelligent eloquence teaching system

Through a multi-dimensional emotional perception and intelligent eloquence teaching system, combining visual and auditory emotional information, a user knowledge graph is built, the relationship between emotions and eloquence skills is analyzed, and teaching strategies is dynamically adjusted, which solves the problem of insufficient emotion recognition and teaching strategies in the existing technology, and achieves high-precision and personalized teaching support.

CN118313680BActive Publication Date: 2025-08-08CHINA NEW LINE EDUCATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410413140.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-08
Publication Date
2025-08-08
Estimated Expiration
2044-04-08

AI Technical Summary

Technical Problem

The existing technology has shortcomings in the depth of emotion perception, the breadth of multimodal data fusion, the pertinence of personalized teaching strategies, the dynamic update and adaptability of the system, and contextualized teaching support, resulting in limited emotional recognition accuracy and teaching strategy accuracy, and it is impossible to provide personalized and real-time teaching support.

Method used

A multi-dimensional emotional perception and intelligent eloquence teaching system is adopted to obtain visual and auditory emotional information through the emotion recognition and perception module, and a user knowledge graph is built in combination with the user portrait construction module. The emotional computing and response module are used to analyze the impact relationship between emotional state and eloquence skills. The personalized strategy generation module dynamically adjusts the teaching strategy, and real-time optimization is carried out through the intelligent feedback and adjustment module.

Benefits of technology

It improves the accuracy of emotion recognition and the pertinence of teaching strategies, realizes dynamic updates and adaptability of personalized teaching, ensures that the teaching content is synchronized with the student status, and provides instant response and personalized teaching experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118313680B_ABST
    Figure CN118313680B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-dimensional emotion perception and intelligent eloquence teaching system, including: an emotion recognition and perception module, a user portrait construction module, an emotion calculation and response module, a personalized strategy generation module and an intelligent feedback and adjustment module; the emotion recognition and perception module is used to obtain the user's physical information and voice information to obtain a multimodal feature vector; the user portrait construction module is used to obtain user feature data, construct the user's knowledge graph, and thus generate a user portrait; the emotion calculation and response module is used to adjust teaching suggestions; the personalized strategy generation module is used to adjust and optimize the initial teaching strategy. The intelligent feedback and adjustment module is used to further adjust and optimize the initial teaching strategy. The present invention solves the technical problems in the prior art in terms of the depth of emotion perception, the breadth of multimodal data fusion, the pertinence of personalized teaching strategies, the dynamic update and adaptability of the system, and the support for contextualized teaching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic information technology, and in particular to a multi-dimensional emotion perception and intelligent eloquence teaching system. Background Art

[0002] While progress has been made in areas such as sentiment analysis, personalized learning, intelligent teaching assistance, and eloquence training, some limitations remain. Sentiment analysis systems primarily focus on identifying emotions through multimodal data, but often fail to directly link sentiment analysis to specific skills, such as eloquence. Personalized learning systems, on the other hand, focus on academic performance and learning efficiency, rarely considering the impact of emotional states on learning. While intelligent teaching assistants excel in teaching assistance, they often lack emotional perception capabilities and specialized training functions for eloquence skills. While eloquence training software provides simulated speech environments and assessment functions, it often overlooks the role of emotional factors and lacks personalized teaching strategies supported by deep learning and multimodal data analysis.

[0003] Existing sentiment analysis systems primarily identify basic emotional states but fail to delve into how emotions impact learning or skill performance, limiting the effectiveness of personalized instruction. Furthermore, many systems rely on a single data source and fail to integrate multiple data types, including text, speech, and visual data, limiting the accuracy and depth of emotion recognition. While personalized learning systems offer customized content, they often overlook students' emotional states and eloquence skills, resulting in teaching strategies that fail to effectively address specific challenges and insufficient personalized instruction. Furthermore, existing educational technology systems lack dynamic update mechanisms and are unable to reflect users' latest progress in real time, limiting their ability to provide effective long-term support. They also fail to fully utilize complex mathematics and advanced machine learning algorithms to extract deep insights from multimodal data, impacting the accuracy of instructional content and strategies. Eloquence training software and intelligent teaching assistants lack a deep understanding of teaching scenarios and are unable to provide contextualized learning experiences, impacting learning outcomes. Summary of the Invention

[0004] The present invention provides a multi-dimensional emotion perception and intelligent eloquence teaching system to solve the technical problems in the existing technology that are obviously insufficient in the depth of emotion perception, the breadth of multimodal data fusion, the pertinence of personalized teaching strategies, the dynamic update and adaptability of the system, and contextualized teaching support.

[0005] In order to solve the above technical problems, the embodiment of the present invention provides a multi-dimensional emotion perception and intelligent eloquence teaching system, which includes: an emotion recognition and perception module, a user portrait construction module, an emotion calculation and response module, a personalized strategy generation module and an intelligent feedback and adjustment module;

[0006] The emotion recognition and perception module is used to obtain the user's body information and voice information, and obtain visual emotion information and auditory emotion information based on the body information and the voice information, respectively, and then fuse the visual emotion information and the auditory emotion information to obtain a multimodal feature vector;

[0007] The user portrait construction module is used to obtain user feature data and perform language, audio and visual analysis on the user feature data, facial information and voice information respectively to obtain user language features, user audio features and user visual features, and then perform user clustering analysis based on the user language features, user audio features and user visual features to construct the user's knowledge graph, thereby generating a user portrait;

[0008] The emotion calculation and response module is used to calculate the user's current emotional state based on the multimodal feature vector, and analyze the influence relationship between the user's current emotional state and eloquence skill performance based on preset scenarios, and then adjust the teaching suggestions based on the influence relationship;

[0009] The personalized strategy generation module is used to obtain the user's performance information and match the initial teaching strategy from a preset resource library based on the performance information and the user portrait, so as to continuously monitor the user's performance information when the initial teaching strategy is executed, and then adjust and optimize the initial teaching strategy.

[0010] The intelligent feedback and adjustment module is used to perform big data analysis on the user's multimodal feature vector, user portrait and performance information, generate learning feedback, and adjust and optimize the initial teaching strategy based on the learning feedback.

[0011] As a preferred solution, the emotion recognition and perception module includes: a visual acquisition submodule, an audio acquisition submodule and a multimodal fusion submodule;

[0012] The visual acquisition submodule is used to obtain the user's facial expressions and body language, and extract facial muscle movement feature information of the facial expressions based on the face detection module, so as to identify the user's visual emotional information based on the facial muscle movement feature information; wherein the body information includes facial expressions and body language;

[0013] The audio acquisition submodule is used to obtain the user's voice information, and perform voice feature extraction and natural language processing on the voice information to obtain language information and audio information; and identify the user's auditory emotion information based on the audio information and the language information;

[0014] The multimodal fusion submodule is used to align and match the visual emotion information and the auditory emotion information, thereby fusing the aligned and matched visual emotion information and auditory emotion information to obtain a multimodal feature vector.

[0015] As a preferred solution, the multimodal fusion submodule includes: a data integration layer, an emotion fusion layer and a decision output interface;

[0016] The data integration layer is used to receive the visual emotion information and the auditory emotion information, and align and match the visual emotion information and the auditory emotion information on a preset time axis;

[0017] The emotion fusion layer is used to fuse the aligned and matched visual emotion information and auditory emotion information to obtain a multimodal feature vector;

[0018] The decision output interface is used to send the multimodal feature vector to the emotion calculation and response module, the personalized strategy generation module or the intelligent feedback and adjustment module for teaching generation.

[0019] As a preferred solution, the user portrait construction module includes: a data collection submodule, an analysis and modeling submodule, and a portrait generation and update submodule;

[0020] The data collection submodule is used to obtain user characteristic data; wherein the user characteristic data includes learning record behavior, test scores and user feedback information;

[0021] The analysis and modeling submodule is used to perform in-depth mining on the user feature data, facial information, and voice information to obtain user language features, user audio features, and user visual features, and to construct corresponding feature models based on the user language features, user audio features, and user visual features, and then to construct the user's knowledge graph through the feature models;

[0022] The portrait generation and update submodule is used to generate or update the user portrait based on the user's knowledge graph.

[0023] As a preferred solution, the analysis and modeling submodule includes: a feature extraction engine, an associated model layer and a model library;

[0024] The feature extraction engine is used to perform in-depth mining on the user feature data, facial information, voice information, and acquired user eloquence data, thereby extracting user language features, user audio features, user visual features, and user eloquence features;

[0025] The association model layer is used to obtain the corresponding emotional state based on the user's language features, user audio features, and user visual features, and then construct an association feature model corresponding to the emotional state and the user's eloquence features, and then construct the user's knowledge graph through the association feature model;

[0026] The model library is used to store the associated feature models generated by users and their corresponding knowledge graphs.

[0027] As a preferred solution, the portrait generation and updating submodule includes: a user attribute layer, an eloquence dimension evaluation matrix, an emotional response curve, and a portrait dynamic update layer;

[0028] The user attribute layer is used to obtain the user's basic information and the user's eloquence characteristics;

[0029] The eloquence dimension evaluation matrix is used to form a visual evaluation matrix based on the user's eloquence characteristics and display it;

[0030] The emotional response curve is used to obtain the user's emotions in different situations, generate an emotional response curve, and combine the user's eloquence characteristics to obtain the influence of emotions on the user's eloquence;

[0031] The dynamic portrait update layer is used to generate or update the user portrait based on the user's basic information, evaluation matrix, the influence of emotions on the user's eloquence and the user's knowledge graph.

[0032] As a preferred solution, the emotion calculation and response module includes: an emotion data acquisition submodule, an emotion analysis and modeling submodule, a situational awareness and understanding submodule, a feedback generation and optimization submodule, and a real-time interactive teaching decision submodule;

[0033] The emotion data collection submodule is used to obtain the user's historical emotion information;

[0034] The sentiment analysis modeling submodule is used to form a sentiment model based on the user's historical sentiment information, and to calculate the user's current quantified emotional state by combining the multimodal feature vector and the user portrait;

[0035] The contextual awareness and understanding submodule is used to analyze the influence relationship between the user's current emotional state and eloquence skill performance based on a preset scenario and the user's current quantified emotional state;

[0036] The feedback generation optimization submodule is used to generate appropriate motivational feedback or teaching suggestions based on the influence relationship;

[0037] The real-time interactive teaching decision-making submodule is used to adjust the course content, rhythm and presentation method according to the motivational feedback or teaching suggestions.

[0038] As a preferred solution, the context awareness and understanding submodule is further configured to:

[0039] Conduct in-depth analysis of the current teaching scene and extract the characteristics of the teaching scene;

[0040] According to the teaching scene characteristics, the user's emotional response based on the teaching scene characteristics is associated with his or her specific eloquence performance;

[0041] Form a contextual cognitive map based on the user's emotional response and specific eloquence performance based on the characteristics of the teaching scenario;

[0042] Based on the contextualized cognitive map, deep learning optimization of the association model is performed, and the cognitive map is continuously adjusted and updated according to the latest real-time data and deep learning feedback, thereby showing the dynamic interaction of user emotions and oral ability in different teaching scenarios.

[0043] As a preferred solution, the feedback generation optimization submodule is further used to:

[0044] Generate personalized feedback using a preset rule library and natural language processing technology based on the emotional state and eloquence performance obtained by the sentiment analysis modeling submodule;

[0045] At the same time, based on the user's current emotional state and eloquence characteristics, corresponding teaching activities and training plans are designed and updated, so as to automatically generate optimization suggestions for the user's emotional state and eloquence shortcomings based on the results output by the emotional data collection submodule, the emotional analysis modeling submodule, and the situational perception and understanding submodule.

[0046] As a preferred solution, the real-time interactive teaching decision submodule further includes:

[0047] Continuously receive and analyze user emotional state data, and dynamically adjust teaching content and methods based on the user's emotional state to match the user's learning needs and psychological state;

[0048] Dispatching corresponding teaching resources based on real-time analysis of users' learning progress and emotional responses;

[0049] Adjust the difficulty level of course tasks according to users' emotional reactions and ability assessments;

[0050] Generate adaptive teaching decisions based on real-time analytics, emotional responses, and ability assessments.

[0051] As a preferred solution, the personalized strategy generation module includes: a user evaluation submodule, a recommendation customization submodule and a feedback optimization submodule;

[0052] The user evaluation submodule is used to obtain the user's performance information through test and homework completion;

[0053] The recommendation customization submodule is used to match an initial teaching strategy from a preset resource library based on the performance information and the user profile; wherein the initial teaching strategy includes teaching and learning materials and homework tests;

[0054] The feedback optimization submodule is used to continuously monitor the user's performance information during the learning process of teaching learning materials and homework tests, and feed it back to the recommendation customization submodule so that the recommendation customization submodule can adjust and optimize the initial teaching strategy based on the feedback information.

[0055] As a preferred solution, the user assessment submodule includes: a language skills assessment layer, a logical thinking and content organization layer, and a non-verbal communication ability layer;

[0056] The language skills assessment layer is used to evaluate the user's pronunciation accuracy, pitch variation, speech speed control and vocabulary richness based on the user's voice information collected during tests and assignments;

[0057] The logical thinking and content organization layer is used to evaluate the user's logical thinking ability and content organization ability based on the voice information collected by the user during tests and assignments;

[0058] The non-verbal communication capability layer is used to evaluate the quality of the user's non-verbal signals based on the user's physical information collected during tests and assignments.

[0059] As a preferred solution, the recommendation customization submodule includes: a problem diagnosis layer and a strategy generation layer;

[0060] The problem diagnosis layer is used to conduct in-depth analysis based on the results output by the user evaluation submodule to obtain the shortcomings and deficiencies shown by the user in eloquence training;

[0061] The strategy generation layer is used to dynamically build a personalized learning path model based on the shortcomings and deficiencies shown by the user in eloquence training, and the recommendation and customization engine uses reinforcement learning and collaborative filtering learning technologies.

[0062] As a preferred solution, the intelligent feedback and adjustment module includes: a data monitoring submodule, a data processing and analysis submodule, a feedback generation submodule and an adjustment execution submodule;

[0063] The data monitoring submodule is used to monitor and obtain the user's multimodal feature vector, user portrait and performance information;

[0064] The data processing and analysis submodule is used to perform big data mining and analysis on the user's multimodal feature vector, user portrait and performance information;

[0065] The feedback generation submodule is used to generate learning feedback based on the results of big data mining and analysis;

[0066] The adjustment execution submodule is used to adjust and optimize the initial teaching strategy and further adjust and optimize the initial teaching strategy according to the learning feedback.

[0067] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0068] The technical solution of the present invention combines visual, voice, and text data to comprehensively identify and analyze emotional states to improve the accuracy of emotion recognition in eloquence teaching scenarios. At the same time, it uses advanced algorithms to process multimodal data, including data fusion, correlation analysis between emotions and eloquence performance, and the construction and updating of user profiles, thereby deeply exploring the insights behind the data. It also dynamically generates and adjusts teaching strategies based on students' emotions, eloquence abilities, and needs, including an intelligent recommendation system and a dynamic teaching resource library to optimize the learning experience. It then constructs a cognitive map that reflects the relationship between students' emotions, eloquence performance, and teaching scenarios to generate more targeted teaching strategies and feedback. As user training and feedback accumulate, user profiles are updated in real time to ensure that the teaching content stays synchronized with the students' latest status. The accuracy of user profiles is maintained through efficient data processing and machine learning models. The teaching content and rhythm are dynamically adjusted based on real-time emotion analysis and changes in user profiles, emphasizing the system's immediate responsiveness and adaptability to provide a more personalized teaching experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 : A structural diagram of a multi-dimensional emotion perception and intelligent eloquence teaching system provided by an embodiment of the present invention;

[0070] Figure 2 : A specific structural diagram of the multi-dimensional emotion perception and intelligent eloquence teaching system provided by an embodiment of the present invention;

[0071] Figure 3 : This is the system operation process provided by the embodiment of the present invention. DETAILED DESCRIPTION

[0072] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0073] Example 1

[0074] Please refer to Figure 1 , a multi-dimensional emotion perception and intelligent eloquence teaching system provided by an embodiment of the present invention, includes: an emotion recognition and perception module 101, a user portrait construction module 102, an emotion calculation and response module 103, a personalized strategy generation module 104 and an intelligent feedback and adjustment module 105.

[0075] The emotion recognition and perception module is used to obtain the user's body information and voice information, and obtain visual emotion information and auditory emotion information based on the body information and the voice information respectively, and then fuse the visual emotion information and the auditory emotion information to obtain a multimodal feature vector.

[0076] As a preferred solution of this embodiment, the emotion recognition and perception module includes: a visual acquisition submodule, an audio acquisition submodule and a multimodal fusion submodule.

[0077] The visual acquisition submodule is used to obtain the user's facial expressions and body language, and extract facial muscle movement feature information of the facial expressions based on the face detection module, so as to identify the user's visual emotional information based on the facial muscle movement feature information; wherein the body information includes facial expressions and body language.

[0078] In this embodiment, the visual acquisition submodule includes a visual sensor unit, an expression recognition engine, and a body language analysis module. The visual sensor unit utilizes a high-precision camera to capture subtle changes in students' facial expressions and body language, including but not limited to key features reflecting emotion, such as eyebrow movement, eye posture, and mouth curves. The facial expression recognition engine, based on deep learning-based facial key point detection technology, extracts and encodes facial muscle movement information to identify basic emotions such as joy, anger, sadness, and happiness, as well as more subtle emotional expressions. The body language analysis module uses a posture recognition algorithm to analyze non-verbal signals such as students' head posture, hand gestures, and body inclinations. These signals help determine the student's emotional engagement and confidence level during the presentation.

[0079] In this embodiment, a visual sensor unit, preferably a high-precision camera, is used as the front-end device to capture real-time student facial expressions and body language changes. The camera can record the video stream at a specific frame rate and continuously track the student in the frame. These cameras are typically equipped with high-definition resolution, wide dynamic range, and good low-light performance, ensuring that rich details can be captured in various conditions.

[0080] In this embodiment, for facial expression recognition, the facial expression recognition engine within the system utilizes deep learning techniques such as convolutional neural networks (CNNs) or facial key point detection algorithms to process captured video images. First, facial detection technology is used to locate the student's facial region. Next, a pre-trained model is applied within this region to extract dozens or even hundreds of key points, including the specific coordinate information for areas such as eyebrows, eyes, nose, and mouth corners. After obtaining these key points, the system further calculates and encodes the changing patterns of facial muscles, such as frowning, blinking frequency, and the curved shape of the mouth corners when smiling or feeling sad. By quantifying and pattern-analyzing these features, the emotion recognition engine can identify and classify basic emotions, such as joy, anger, sadness, and happiness, as well as more complex and subtle emotional states, such as surprise, disgust, fear, or neutrality.

[0081] In this embodiment, the body language analysis module uses a gesture recognition algorithm to analyze non-verbal signals such as head tilt, hand gestures (waving, nodding, shaking the head), body posture, and body movements to infer a student's attention, emotional engagement, and confidence during a speech or interactive situation. This component may employ 3D gesture estimation or 2D joint recognition to map the positional relationships of various human joints into quantifiable data, thereby interpreting the underlying emotional and attitude information.

[0082] The audio acquisition submodule is used to obtain the user's voice information, and perform voice feature extraction and natural language processing on the voice information to obtain language information and audio information; and identify the user's auditory emotion information based on the audio information and the language information.

[0083] In this embodiment, the audio acquisition submodule includes but is not limited to a high-definition microphone array, a speech signal preprocessing module, an emotional speech feature extractor, and a natural language processing unit. The high-definition microphone array is used to capture and enhance the student's voice signal in 360 degrees to ensure clear and distortion-free speech quality; the speech signal preprocessing module performs a series of operations on the audio, such as noise reduction, framing, and endpoint detection, to provide high-quality speech signal input for subsequent analysis; the emotional speech feature extractor uses a variety of methods such as spectrum analysis, MFCC (Mel-Frequency Cepstral Coefficients), and LPC (Linear Predictive Coding) to extract emotional features in the speech, such as intonation, volume, speaking speed, and rhythm; the natural language processing unit, in conjunction with speech-to-text technology, converts speech content into analyzable text data, providing further insight into the emotional color contained in vocabulary selection and sentence structure.

[0084] In this embodiment, a high-definition microphone array is the first step in the speech signal processing and understanding subsystem. It is designed to capture 360-degree student voices in a classroom or seminar room. Through multi-microphone collaboration and sound source localization technology, it effectively reduces ambient noise interference, enhances the target speech signal, and ensures clear, distortion-free audio information collected from all directions, providing high-quality raw speech input for subsequent emotion recognition. The speech signal preprocessing module is responsible for preliminary processing of the collected sound signal. First, it performs noise reduction to remove background noise and highlight the speech content. Second, it performs framing, dividing the continuous speech stream into a series of short speech frames to facilitate speech feature analysis within a small time window. Endpoint detection is used to determine the start and end points of speech, thereby eliminating invalid silence segments and retaining only the sections containing valid speech information. After preprocessing, the emotional speech feature extractor utilizes a variety of advanced signal processing methods and techniques to mine emotional cues, including spectral analysis, MFCC (Mel-Frequency Cepstral Coefficient) analysis, and LPC (Linear Predictive Coding) analysis. Spectral analysis reveals the frequency distribution of speech signals, reflecting variations in sound quality. MFCC (Mel-Frequency Cepstral Coefficient) analysis captures spectral characteristics important to human speech perception. LPC (Linear Predictive Coding) analysis extracts speech rhythm and formant characteristics by analyzing the correlation between the current signal sample and past samples. Feature extraction techniques aim to quantify factors such as intonation (pitch), volume (loudness), speech rate (rhythm), and more complex prosodic patterns, all of which are important dimensions of emotional expression in speech.

[0085] Furthermore, the natural language processing unit combines this with speech-to-text technology to convert the preprocessed speech signal into text. Through text analysis, the system can further understand language-level information such as vocabulary choice, syntactic structure, and rhetorical techniques, all of which may carry significant emotional overtones. For example, word choice can reflect the speaker's positive or negative attitude, while sentence length and complexity can reflect their fluency and psychological state.

[0086] The multimodal fusion submodule is used to align and match the visual emotion information and the auditory emotion information, thereby fusing the aligned and matched visual emotion information and auditory emotion information to obtain a multimodal feature vector.

[0087] As a preferred solution of this embodiment, the multimodal fusion submodule includes: a data integration layer, an emotion fusion layer and a decision output interface;

[0088] The data integration layer is used to receive the visual emotion information and the auditory emotion information, and align and match the visual emotion information and the auditory emotion information on a preset time axis;

[0089] The emotion fusion layer is used to fuse the aligned and matched visual emotion information and auditory emotion information to obtain a multimodal feature vector;

[0090] The decision output interface is used to send the multimodal feature vector to the emotion calculation and response module, the personalized strategy generation module or the intelligent feedback and adjustment module for teaching generation.

[0091] In this embodiment, the data integration layer simultaneously receives emotional cues from both visual and auditory channels and accurately aligns and matches them on a timeline. The emotion fusion layer is also used to build a deep neural network model. By combining visual and speech features and training a large number of labeled data samples, it can achieve accurate real-time recognition and comprehensive assessment of students' emotions. It then fuses and processes the aligned and matched visual and auditory emotional information to obtain a multimodal feature vector. The decision output interface is also used to generate corresponding teaching strategy instructions based on the identified emotional state, which are transmitted to the virtual tutor's teaching behavior control module for timely teaching intervention.

[0092] In this embodiment, the time series data alignment formula is:

[0093] D ts (i,j)=min{D ts (i-1,j-1)+δ(t i ,t j ),D ts (i-1,j),Dts (i,j-1)}

[0094] Among them, D ts (i, j) is the cumulative distance matrix in the dynamic time warping algorithm, δ(t i ,t j ) is the time point t i and t j The distance between them.

[0095] The formula for extracting the eloquence dimension feature is:

[0096]

[0097] Among them, F rhetoric is the comprehensive eloquence characteristic, ρ n is the weight of the nth eloquence feature, F n is the nth eloquence feature (such as speaking speed, intonation change, language richness, etc.), and N is the total number of eloquence features.

[0098] The emotion-eloquence fusion model is:

[0099] R fusion =σ(κ e ·E integrated +κ r ·F rhetoric )

[0100] Among them, R fusion is the fused emotion-eloquence response, σ is the activation function (such as sigmoid or ReLU), κe is the emotion data weight, and κr is the eloquence data weight.

[0101] The decision-making formula for multi-dimensional teaching strategies is:

[0102] T strategy =NN deep (R fusion ,C context )

[0103] Among them, T strategy For teaching strategies based on deep neural networks, NN deep is a deep neural network model, C context Provides contextual information for teaching.

[0104] The parameter optimization and regularization formula is

[0105]

[0106] Among them, Ω opt is the optimized parameter set, L is the loss function, λ is the regularization parameter, ‖θ‖ 2is the L2 norm of the model parameters.

[0107] In this embodiment, time series data alignment requires synchronous processing of data from the student's visual and auditory channels. These data may contain different timestamps, so they need to be processed through a dynamic time warping algorithm D ts (i, j) to align these time series data, so as to ensure that the visual and auditory data correspond in time, so as to perform accurate sentiment analysis. In the eloquence dimension feature extraction, natural language processing technology and speech analysis technology are used to extract multiple features of students' oral expression, such as speaking speed, intonation changes, language richness, etc. These features are weighted summed To form a comprehensive eloquence feature score, it is possible to quantify the student's oral expression ability as an important factor in subsequent decision-making. The emotion-eloquence fusion model combines the integrated emotion data E integrated and eloquence trait score F rhetoric By integrating the formula R fusion =σ(κ e ·E integrated +κ r ·F rhetoric ) are combined to form a comprehensive emotion-eloquence response score, and then a comprehensive score that includes both emotion and eloquence dimensions is created to more comprehensively understand the student's status. Regarding multi-dimensional teaching strategy decision-making, based on the deep neural network model NN deep , the system uses the comprehensive score R obtained in the previous step fusion and teaching context information C context To generate teaching strategies T strategy , so that appropriate teaching strategies can be dynamically generated according to students’ emotional state and eloquence ability. Finally, through parameter optimization and regularization, during the model training process, the system will minimize the loss function L and the regularization term λ·‖θ‖ 2 To adjust and optimize the model parameters Ω opt , thereby optimizing the performance of the model and ensuring that teaching strategies can be generated accurately and effectively in practical applications.

[0108] It is understandable that by continuously recording and analyzing students' audio-visual signals during eloquence training, their instantaneous emotional changes and fluctuations in eloquence skills can be quickly captured. At the same time, through advanced visual and voice processing technologies, complex emotional information is quantified into computable data, accurately identifying students' emotional categories and intensities, while evaluating their specific performance in eloquence dimensions, such as fluency and logical coherence. In addition, the results of emotion recognition are combined with eloquence dimension indicators to form in-depth feedback information, helping virtual tutors understand students' psychological needs and ability shortcomings in specific situations, thereby formulating targeted teaching strategies. Based on the above functions, the module can adjust the virtual tutor's teaching methods and content in real time, enabling it to respond to students' emotions and eloquence performance immediately and appropriately, improving the effectiveness and efficiency of personalized teaching.

[0109] The user portrait construction module is used to obtain user feature data, and perform language, audio and visual analysis on the user feature data, facial information and voice information respectively, to obtain user language features, user audio features and user visual features, and then perform user clustering analysis based on the user language features, user audio features and user visual features to construct the user's knowledge graph, thereby generating a user portrait.

[0110] As a preferred solution of this embodiment, the user portrait construction module includes: a data collection submodule, an analysis and modeling submodule, and a portrait generation and update submodule;

[0111] The data collection submodule is used to obtain user characteristic data; wherein the user characteristic data includes learning record behavior, test scores and user feedback information;

[0112] In this embodiment, the data collection submodule includes an eloquence data collector and an emotion recognition sensor. The eloquence data collector is responsible for collecting users' actual performance in various eloquence training scenarios, including but not limited to multimedia information such as voice recordings and video capture, for use in evaluating various eloquence indicators such as vocabulary usage, speech rate control, pitch variation, language logic, and expression fluency. The emotion recognition sensor uses built-in AI facial expression recognition and voice emotion analysis technology to capture users' emotional reactions in real time, such as joy, tension, anxiety, and confidence, and quantify them into analyzable data.

[0113] In this embodiment, the eloquence ability data collector monitors eloquence training scenarios. Specifically, the system is deployed in various eloquence training application scenarios, such as online courses, simulated speech practice, and real-time interactive communication platforms. Through its integrated multimedia data capture function, it automatically records the user's voice and video information. Voice recording analysis is then performed, recording the user's speech audio and converting it into text using speech recognition technology. Metrics such as vocabulary diversity, speaking rate, pause duration, and coherence are then extracted. Furthermore, tonal variations, such as stress distribution and intonation, are analyzed to assess the artistry and logic of the language expression. Finally, video capture and motion analysis are performed. Video capture technology is used to observe and record non-verbal expressions such as the user's body language, facial expressions, and eye contact. These factors also reflect the expressiveness and persuasiveness of eloquence.

[0114] In this embodiment, the emotion recognition sensor uses facial expression recognition, that is, an AI facial expression recognition algorithm is used to capture a sequence of facial images of the user through a camera. These images are then analyzed in real time using a deep learning model (such as a convolutional neural network) to decode six basic emotions (joy, sadness, surprise, anger, fear, and disgust) or other more complex emotional states, and quantify them into numerical indicators. Voice emotion analysis is then performed by combining speech signal processing technology with machine learning models (such as long short-term memory networks (LSTMs) or deep learning-based emotion recognition models) to extract features such as prosody, intensity, intonation, and rhythm from the user's voice to judge and quantify the user's emotional response, such as tension and confidence level. It is understandable that the data acquisition submodule plays a key role in the entire user profile construction process. It continuously extracts rich data on the user's verbal ability and emotional response from a variety of interactive scenarios, providing basic input for subsequent steps such as data cleaning, feature engineering, and model training, thereby helping to build a comprehensive and accurate user profile to facilitate the formulation and implementation of personalized teaching strategies.

[0115] The analysis and modeling submodule is used to perform in-depth mining on the user feature data, facial information, and voice information to obtain user language features, user audio features, and user visual features, and to construct corresponding feature models based on the user language features, user audio features, and user visual features, and then to construct the user's knowledge graph through the feature models;

[0116] As a preferred solution of this embodiment, the analysis and modeling submodule includes: a feature extraction engine, an associated model layer and a model library;

[0117] The feature extraction engine is used to perform in-depth mining on the user feature data, facial information, voice information, and acquired user eloquence data, thereby extracting user language features, user audio features, user visual features, and user eloquence features;

[0118] The association model layer is used to obtain the corresponding emotional state based on the user's language features, user audio features, and user visual features, and then construct an association feature model corresponding to the emotional state and the user's eloquence features, and then construct the user's knowledge graph through the association feature model;

[0119] The model library is used to store the associated feature models generated by users and their corresponding knowledge graphs.

[0120] In this embodiment, the analysis and modeling submodule includes: a feature extraction engine, an association model layer, and a model library. Among them, the feature extraction engine is an eloquence feature extraction engine, which is used to perform in-depth processing and feature extraction on the collected basic eloquence ability data to form representative eloquence dimension data points. The association model layer is an emotion-eloquence association model, which is used to establish an association model between emotional state and eloquence performance based on statistics and machine learning methods, and analyze how emotions affect the user's eloquence in different situations. The model library is a personalized growth model library, which is used to dynamically generate a personalized eloquence growth model based on the user's historical behavior data, eloquence ability, and emotional response.

[0121] In this embodiment, the multi-dimensional characterization formula is:

[0122]

[0123] in, The extended eloquence feature vector. i : The weight of the i-th extended eloquence dimension. φ: Nonlinear mapping function used to extract complex features. T i (x): The i-th eloquence dimension feature extracted from the original data x. N: The total number of extended eloquence dimensions.

[0124] The advanced feature optimization technology is kernel principal component analysis (Kernel PCA), which enhances feature expression capabilities by extracting key information in nonlinear feature space.

[0125] The emotion-eloquence association model is:

[0126]

[0127] Among them, Y complex : Evaluation of eloquence performance in complex situations. p , Extended emotional and eloquence characteristics. ·Ψ(E,Fext ) r : The interaction term between emotion and eloquence characteristics. β p ,γ q ,δ r : Correlation coefficient. P, Q, R: Number of emotion features, eloquence features, and interaction terms. ∈: Error term.

[0128] The advanced model optimization strategy is Elastic Net Regularization, which can balance model complexity and prediction accuracy.

[0129] The dynamic growth prediction model is:

[0130]

[0131] in, Dynamic personalized growth prediction. history : User historical behavior data. Current extended eloquence features. P profile : User profile. RNN: Recurrent neural network model for processing time series data.

[0132] Model adaptive adjustment technology combines transfer learning and online learning to ensure that the model can self-adjust and optimize according to new data, thereby improving the accuracy and personalization of predictions.

[0133] In this embodiment, the multi-dimensional characterization process is to use the eloquence feature extraction engine to collect the user's basic eloquence ability data. Then, through the formula Deeply process and extract features from these data to form representative eloquence dimension data points, providing accurate feature representation for subsequent analysis. Feature extraction optimization is then performed, and advanced techniques such as kernel principal component analysis (Kernel PCA) are applied to further refine and optimize eloquence features, thereby increasing the nonlinearity and complexity of feature expression and ensuring that features are more recognizable and relevant. Emotion-eloquence correlation analysis is also performed, using complex correlation analysis formulas. To analyze the relationship between emotional state and eloquence, we can understand how emotions affect users' eloquence in different situations and provide a basis for personalized recommendations. The optimization of the association model uses methods such as Elastic Net Regularization to optimize the association model, which can balance the accuracy and generalization ability of the model and reduce the risk of overfitting. We also build a personalized growth model and use a dynamic growth prediction model. Based on a user's historical behavioral data, current speaking ability, and profile, a personalized speaking growth model is generated. This dynamically reflects the user's growth trajectory and potential, providing customized development recommendations. Finally, the model is adaptively adjusted using transfer learning and online learning techniques to continuously adjust and optimize the growth model based on new user data, ensuring that the growth model can adapt to changes in user behavior and abilities, providing long-term, effective, and personalized support.

[0134] The portrait generation and update submodule is used to generate or update the user portrait based on the user's knowledge graph.

[0135] As a preferred solution of this embodiment, the portrait generation and updating submodule includes: a user attribute layer, an eloquence dimension evaluation matrix, an emotional response curve, and a portrait dynamic update layer;

[0136] The user attribute layer is used to obtain the user's basic information and the user's eloquence characteristics;

[0137] The eloquence dimension evaluation matrix is used to form a visual evaluation matrix based on the user's eloquence characteristics and display it;

[0138] The emotional response curve is used to obtain the user's emotions in different situations, generate an emotional response curve, and combine the user's eloquence characteristics to obtain the influence of emotions on the user's eloquence;

[0139] The dynamic portrait update layer is used to generate or update the user portrait based on the user's basic information, evaluation matrix, the influence of emotions on the user's eloquence and the user's knowledge graph.

[0140] In this embodiment, the portrait generation and update submodule includes a user attribute layer, an eloquence dimension evaluation matrix, an emotional response curve, and a portrait dynamic update layer. Among them, the user attribute layer is the user's core attribute area, which contains basic information (such as age, gender, educational background, etc.) and characteristic labels extracted from eloquence training (such as strong persuasiveness, quick thinking, etc.). The eloquence dimension evaluation matrix is used to display various eloquence ability indicators in a multi-dimensional manner to form a visual evaluation matrix, which intuitively displays the user's strengths and room for improvement. The emotional response curve is used to depict the characteristics of the user's corresponding emotional fluctuations in different eloquence activities, and to assist in understanding the impact of emotions on eloquence skills. The portrait dynamic update layer adopts a portrait dynamic update mechanism. As the user participates in more training and feedback, the system continuously iterates and optimizes the user portrait to ensure that it always reflects the latest development of eloquence ability and progress in emotional management ability.

[0141] In this embodiment, the user attribute feature vectorization formula is:

[0142] V core=Vec(a,g,e,…,T L )

[0143] Among them, V core : Feature vector of user attributes. Vec: Feature vectorization function. a, g, e: Represent the user's age, gender, educational background and other numerical features respectively. T L : Feature label vector obtained based on text analysis.

[0144] The eloquence dimension evaluation matrix - the advanced evaluation formula for eloquence ability is:

[0145]

[0146] in, The user's advanced rating on the i-th eloquence dimension and the j-th ability indicator. ik : The weight coefficient of the i-th dimension. S jk (D): Score function of data D based on the j-th capability indicator. K: Number of evaluation indicators.

[0147] The emotional response curve - emotional fluctuation model formula is:

[0148]

[0149] Among them, E m (t): composite score of emotion at time t. n : The weight of the nth sentiment indicator. φ n : Conversion function of sentiment indicator. Emo n (t): The nth sentiment indicator data at time t. N: The total number of sentiment indicators.

[0150] The dynamic update mechanism of the portrait - the composite function of the portrait update is:

[0151] P update =Dynamic Update (P old ,D new ,Θ)

[0152] Among them, P update :Updated user portrait. old :Original user portrait. D new : Newly collected user data. Θ: Update the parameter set of the model. Dynamic Update : Dynamically update the function and adjust the portrait features based on machine learning and user feedback.

[0153] The portrait optimization and iteration formula is:

[0154] P opt =Optimize(Pupdate ,Feedback,α)

[0155] Among them, P opt : The optimized user profile. Feedback: Feedback data from users. α: The parameter to be optimized. Optimize: The optimization function responsible for adjusting the profile based on feedback and historical data.

[0156] In this embodiment, the processing of the user core attribute area is performed by using the feature vectorization formula V core =Vec(a,g,e,…,T L ) converts the user's basic information (such as age, gender, educational background, etc.) and the characteristic labels extracted from eloquence training into numerical feature vectors, thereby creating a comprehensive user core attribute feature set to better understand the user's basic situation and personal characteristics. Then, the eloquence dimension evaluation matrix is generated, and the advanced evaluation formula is used. To evaluate the user's performance in various oral ability indicators, and to present these scores in a multi-dimensional way, forming a visual evaluation matrix, so that the user's oral ability can be intuitively displayed, including strengths and areas for improvement. And to draw an emotional response curve, using the emotional fluctuation model formula To analyze the user's emotional fluctuations in different eloquence activities, and draw an emotional response curve, and then depict the user's emotional change trend, to help understand how emotions affect the performance of eloquence skills. At the same time, the dynamic update mechanism of the portrait is implemented, through the portrait update composite function P update =Dynamic Update (P old ,D new ,Θ) to integrate the newly collected user data and the old user portraits, so as to realize the dynamic update of the user portrait and ensure that the user portrait can always reflect the user's latest oral ability development and emotional management ability. Finally, the portrait is optimized and iterated, using the portrait optimization and iteration formula P opt =Optimize(P update ,Feedback,α) optimizes and adjusts the updated portrait, considering the user’s feedback and historical data, which can further improve the accuracy and personalization of the user portrait and better serve the user’s training and development needs.

[0157] It is understandable that the system can comprehensively evaluate the user's oral ability, from basic language skills to complex situational coping strategies, provide a detailed ability distribution report, and monitor and accurately identify the user's emotional changes during the oral expression process in real time, revealing the interactive relationship between emotions and oral output, so as to better understand and guide the user's emotional regulation. At the same time, it can automatically generate targeted teaching plans and training tasks based on the oral dimension scores and emotional response characteristics in the user portrait, including special improvement exercises, practical simulation challenges and psychological counseling strategies. It can also flexibly adjust the teaching content and difficulty based on the user's progress and immediate emotional feedback during the training process, ensuring that each stage of training is close to the user's personalized needs. Through in-depth interpretation of user portraits, it provides the R&D team with valuable insights into user needs, course design effectiveness and system improvement suggestions.

[0158] The emotion calculation and response module is used to calculate the user's current emotional state based on the multimodal feature vector, and analyze the influence relationship between the user's current emotional state and eloquence skill performance based on the preset scenario, and then adjust the teaching suggestions based on the influence relationship.

[0159] As a preferred solution of this embodiment, the emotion calculation and response module includes: an emotion data acquisition submodule, an emotion analysis and modeling submodule, a situational awareness and understanding submodule, a feedback generation and optimization submodule, and a real-time interactive teaching decision submodule;

[0160] The emotion data collection submodule is used to obtain the user's historical emotion information;

[0161] In this embodiment, the emotion calculation and response module includes a multi-faceted data input interface, from which the emotion data collection submodule collects multimodal information about the user's eloquence by integrating multiple sensors and algorithms. This includes, but is not limited to, facial expression recognition technology (such as deep learning-based facial landmark detection), speech emotion analysis tools (for extracting features such as pitch, speaking speed, and volume), and natural language processing components (for capturing the emotional tone of the speech content).

[0162] In this embodiment, the multi-integrated data input interface can simultaneously process multimodal user emotional information from different sensing devices and software algorithms. The core function of the data input interface is to synchronously integrate data streams from multiple sources to ensure that the system can capture and understand the user's full range of emotional reactions when expressing eloquence in real time. Among them, the facial expression recognition subsystem is based on deep learning technology. It captures the user's facial image or video stream through a camera or other visual sensor, detects facial key points, and uses a pre-trained convolutional neural network (such as Dlib, Face++ or OpenFace framework) to analyze the image in real time, locate the key feature points of the face, thereby determining the position changes of parts such as eyes, eyebrows, and mouth, and performing expression classification and quantification, so as to further analyze the change patterns of these key points, and combine with pre-defined expression models to convert the user's facial expressions into a variety of basic or compound emotion categories such as anger, joy, sadness, surprise, fear, etc., and quantify them into numerical representations. Finally, voice emotion analysis tools utilize advanced audio signal processing technology and machine learning algorithms to conduct in-depth analysis of the user's voice signal. This includes: Feature extraction: Extracting characteristic parameters such as pitch, speaking rate, volume, prosody, and non-verbal sound characteristics (such as laughter and sighs). Emotion recognition: Based on these features, an emotion model is constructed. This may use support vector machines (SVMs), recurrent neural networks (RNNs), convolutional neural networks (CNNs), or other deep neural network models to determine and quantify the user's voice emotional state, such as nervousness, excitement, confidence, or indecision.

[0163] In this embodiment, the natural language processing component mainly analyzes the user's speech content to obtain the emotional color at the text level, including: speech content analysis, using NLP technology to perform word segmentation, part-of-speech tagging, syntactic analysis and semantic understanding on the text spoken by the user, and identify the emotional vocabulary and emotional intensity expressed; emotional tendency analysis: applying emotional analysis algorithms (such as rule-based methods, dictionary methods or deep learning models) to analyze the emotional polarity of the text (positive, negative, neutral) and more subtle emotional dimensions, such as approval, dissatisfaction, surprise or anxiety.

[0164] The sentiment analysis modeling submodule is used to form a sentiment model based on the user's historical sentiment information, and to calculate the user's current quantified emotional state by combining the multimodal feature vector and the user portrait;

[0165] In this embodiment, the core of the sentiment analysis modeling submodule is to conduct in-depth mining and comprehensive analysis of collected sentiment data. Specifically, customized sentiment feature models are constructed for eloquence. These models can quantify users' eloquence-related emotional states, such as confidence, persuasiveness, and likeability, from different perspectives. These states are then tracked and predicted in real time using machine learning or deep learning methods.

[0166] In this embodiment, the multimodal information obtained from the emotion data acquisition submodule is preprocessed through data preprocessing and feature extraction, such as audio signal noise reduction, speech segmentation into frames, video image standardization, etc. For facial expression recognition data, the coordinates of key points are further refined and geometric features are calculated, such as the degree of eye opening and closing, the amplitude of the corners of the mouth, etc. For speech emotion data, acoustic features such as pitch, speaking speed, volume, and emotional vocabulary and syntactic structure features based on NLP are extracted. For text content, the text is converted into a vectorized representation using a bag-of-words model, TF-IDF or deep learning embedding (such as Word2Vec, BERT). Then, emotional features customized for the eloquence dimension are constructed, and emotional feature models for specific dimensions such as self-confidence, persuasiveness, and affinity are constructed in combination with the characteristics of eloquence expression. For example: Confidence: This can be measured through indicators such as volume, speaking speed stability, eye contact, and the frequency of using affirmative words; Persuasiveness: This can be measured by analyzing factors such as language logic, sufficiency of arguments, appropriateness of pauses and emphasis, and audience feedback; Approachability: This can be measured by considering the frequency of facial smiles, the openness of body movements, the use of polite language, and the degree of empathy. Machine learning and deep learning model training utilizes labeled datasets and employs machine learning algorithms (such as SVM, decision trees, random forests, etc.) or deep learning models (such as convolutional neural networks, recurrent neural networks, and long short-term memory networks, etc.) to train emotion recognition models for different eloquence dimensions. These models aim to output a user's emotional state score for each eloquence dimension based on the input emotional feature data, and can track and predict the user's emotional changes in real time. Finally, emotional fusion and comprehensive evaluation are carried out to integrate the emotional feature results from different modalities (visual, auditory, and language) to form a multi-dimensional emotional vector to comprehensively reflect the user's current eloquence-related emotional state. Through weight distribution or other fusion strategies, the emotional scores of each dimension are comprehensively evaluated to generate a comprehensive eloquence performance evaluation index.

[0167] The contextual awareness and understanding submodule is used to analyze the influence relationship between the user's current emotional state and eloquence skill performance based on a preset scenario and the user's current quantified emotional state;

[0168] As a preferred solution of this embodiment, the context awareness and understanding submodule is further configured to:

[0169] Conduct in-depth analysis of the current teaching scene and extract the characteristics of the teaching scene;

[0170] According to the teaching scene characteristics, the user's emotional response based on the teaching scene characteristics is associated with his or her specific eloquence performance;

[0171] Form a contextual cognitive map based on the user's emotional response and specific eloquence performance based on the characteristics of the teaching scenario;

[0172] Based on the contextualized cognitive map, deep learning optimization of the association model is performed, and the cognitive map is continuously adjusted and updated according to the latest real-time data and deep learning feedback, thereby showing the dynamic interaction of user emotions and oral ability in different teaching scenarios.

[0173] In this embodiment, the contextual awareness and understanding submodule analyzes the current teaching scenario and the user's expression. Combining this with a pre-defined eloquence dimension indicator system, it links the user's emotional response with their specific eloquence performance, forming a contextualized cognitive map. This helps the system accurately determine how the user's emotions affect their performance on specific eloquence tasks.

[0174] In this embodiment, the extended context awareness unit-teaching scene depth analysis formula is:

[0175]

[0176] in, The feature vector of the teaching scene after in-depth analysis. p : The weight of the P-th scene feature. The pth raw data component in the teaching scenario. Θ p : The parameter set of the p-th feature extraction function. P: The total number of scene components.

[0177] The formula for high-level semantic analysis of user expression content is:

[0178]

[0179] in, High-level semantic analysis results of user expressions. Deep learning models for semantic analysis.

[0180] The high-level association model of emotion and eloquence is:

[0181]

[0182] in, Results of advanced correlation analysis between emotions and eloquence performance. The user's extended sentiment feature vector. Extended eloquence dimension data. h: Advanced emotion and eloquence correlation analysis function.

[0183] In this embodiment, a complex contextual cognitive map is constructed using advanced graph theory algorithms, including multiple levels of emotions, eloquence, and scene elements. Deep learning optimization of the association model uses advanced deep learning techniques (such as variational autoencoders or generative adversarial networks) to optimize the high-level association analysis function h, thereby further improving the accuracy and complexity of the analysis of the association between emotions and eloquence. The dynamic cognitive map is updated in real time, continuously adjusting and updating the cognitive map based on real-time data and deep learning feedback, ensuring that the cognitive map accurately reflects the user's emotions and eloquence in different situations at any time.

[0184] In this embodiment, the depth analysis of the teaching scene uses the formula In-depth analysis of the current teaching scene, extraction of key features, and accurate analysis of each component of the teaching scene provide a basis for subsequent analysis of emotions and eloquence. At the same time, the advanced semantic analysis of user expression content, using formula Perform advanced semantic analysis on the user's expression content, understand and analyze the user's eloquence expression, and extract the key indicators and characteristics of the user's eloquence ability. At the same time, perform emotion-eloquence advanced correlation analysis and apply formula Analyze the complex relationship between the user's emotional state and eloquence performance, so as to gain a deep understanding of how emotions affect the user's performance on specific eloquence tasks and identify the key connections between emotions and eloquence skills. It also constructs a complex contextual cognitive map, using advanced graph theory algorithms to build a complex contextual cognitive map covering emotions, eloquence performance, and scene elements, thus forming a comprehensive view that shows the dynamic interaction between user emotions and eloquence ability in different teaching scenarios. Deep learning optimization of the association model, using advanced deep learning technology to optimize the advanced association analysis function h to improve the accuracy and depth of the analysis, ensuring high precision and reliability of the analysis of the association between emotions and eloquence performance. Real-time update of the dynamic cognitive map, constantly adjusting and updating the cognitive map based on the latest real-time data and deep learning feedback, can maintain the real-time and accuracy of the cognitive map, ensuring that it always reflects the user's current emotional and eloquence performance status.

[0185] The feedback generation optimization submodule is used to generate appropriate motivational feedback or teaching suggestions based on the influence relationship;

[0186] As a preferred solution of this embodiment, the feedback generation optimization submodule is further used to:

[0187] Generate personalized feedback using a preset rule library and natural language processing technology based on the emotional state and eloquence performance obtained by the sentiment analysis modeling submodule;

[0188] At the same time, based on the user's current emotional state and eloquence characteristics, corresponding teaching activities and training plans are designed and updated, so as to automatically generate optimization suggestions for the user's emotional state and eloquence shortcomings based on the results output by the emotional data collection submodule, the emotional analysis modeling submodule, and the situational perception and understanding submodule.

[0189] In this embodiment, the feedback generation and optimization submodule, based on the analysis results from the previous layers, automatically generates optimization suggestions tailored to the user's emotional state and eloquence shortcomings, such as adjusting tone, controlling speech rate, and improving nonverbal behavior. Furthermore, this layer also includes a dynamic strategy generation module, which can design targeted teaching activities and training plans.

[0190] In this embodiment, the results are integrated and analyzed based on the user's emotional state data and eloquence dimension scores obtained from the "Sentiment Analysis and Modeling Layer." This layer analyzes the user's performance in different eloquence dimensions, such as confidence, persuasiveness, and likeability, and identifies potential shortcomings or areas for improvement.

[0191] In this embodiment, targeted feedback is generated. Based on the above analysis results, the feedback generator uses a preset rule library and natural language processing technology to generate personalized feedback. For example, in the case of low self-confidence, the system may suggest that the user increase the volume, use more firm language, or provide psychological construction techniques to enhance self-confidence. For the problem of speaking too fast, the system will remind the user to slow down appropriately to ensure clarity and audience understanding. For non-verbal behaviors, such as unnatural facial expressions or insufficient openness of body language, the system will give adjustment suggestions, such as increasing eye contact, improving the amplitude of smiles, etc.

[0192] In this embodiment, dynamic strategy generation designs and updates corresponding teaching activities and training plans based on the user's current emotional state and eloquence. For example, if the system detects that a user's emotional management skills are weak in a specific situation, it will develop corresponding situational simulation training tasks to help the user learn how to control and regulate their emotional expression in practice. If the user demonstrates low persuasiveness on a certain topic, the system may recommend participating in debate or speech practice and provide relevant materials and preparation strategies.

[0193] In this embodiment, continuous evaluation and optimization not only generates immediate feedback but also tracks and evaluates users' actual performance after adopting the suggestions, continuously optimizing subsequent recommendations. When users follow the system's guidance and make improvements, the system collects sentiment data again, compares the changes before and after, and adjusts the optimization algorithm to make the recommendations even more accurate and effective.

[0194] The real-time interactive teaching decision-making submodule is used to adjust the course content, rhythm and presentation method according to the motivational feedback or teaching suggestions.

[0195] As a preferred solution of this embodiment, the real-time interactive teaching decision submodule further includes:

[0196] Continuously receive and analyze user emotional state data, and dynamically adjust teaching content and methods based on the user's emotional state to match the user's learning needs and psychological state;

[0197] Dispatching corresponding teaching resources based on real-time analysis of users' learning progress and emotional responses;

[0198] Adjust the difficulty level of course tasks according to users' emotional reactions and ability assessments;

[0199] Generate adaptive teaching decisions based on real-time analytics, emotional responses, and ability assessments.

[0200] In this embodiment, the real-time interactive teaching decision-making submodule presents all of the above calculation results to the user in a user-friendly and interactive manner, while also updating the teaching path in real time to ensure the implementation of the personalized teaching plan. This submodule involves functions such as teaching resource scheduling and task difficulty adjustment, allowing the virtual tutor to make relevant teaching decisions in real time based on the student's emotional state.

[0201] In this embodiment, real-time emotional state monitoring and feedback is presented by continuously receiving and analyzing the user's (student's) emotional data, including but not limited to facial expressions, voice intonation, and body posture. This data is input into an emotional computing model for real-time processing and interpretation, deriving the user's current emotional state and presenting it to the user in a user-friendly and interactive manner, such as through charts, animations, or text prompts.

[0202] In this embodiment, personalized teaching content is adjusted according to the user's emotional state. This layer will dynamically adjust the teaching content and methods to match the user's learning needs and psychological state. For example, when a user feels frustrated or confused, the system may provide more detailed explanations, examples, or guided learning materials to enhance understanding and confidence. At the same time, teaching resource scheduling can be based on real-time analysis of the user's learning progress and emotional response. This layer is responsible for intelligent scheduling of various teaching resources, such as recommending teaching videos, case studies, or exercises that are suitable for the user's current emotional characteristics. If it is detected that the user has a good grasp of a certain knowledge point, the module will quickly advance to the next stage; otherwise, it may repeat the explanation or switch to a different teaching method to help the user overcome difficulties.

[0203] In this embodiment, the system manages task difficulty and pacing. Based on the user's emotional response and ability assessment, it adjusts the difficulty level of course tasks appropriately, ensuring a balance between challenge and a sense of accomplishment, thereby stimulating positive emotions and learning motivation. If the user is stressed or fatigued, the module automatically slows down the pace of instruction and schedules relaxing review activities or breaks, thereby maintaining a positive learning experience and stamina.

[0204] In this embodiment, decision-making and execution are instantaneous. The virtual tutor makes adaptive teaching decisions based on these real-time calculations. For example, if the user is highly focused and in a positive mood, a challenging activity may be offered; if the user is experiencing negative emotions, incentives may be provided or teaching strategies may be adjusted to alleviate tension. These decisions are also reflected in the updated teaching path, allowing the entire teaching process to flexibly adapt to individual differences and achieve truly personalized teaching plans.

[0205] As you can understand, the real-time emotion recognition and tracking module can monitor and continuously track a user's emotional changes during a speech or communication. This isn't limited to basic emotion categories (such as happiness, sadness, and nervousness), but also includes subtle emotional fluctuations closely related to eloquence. Simultaneously, contextualized assessment of eloquence dimensions evaluates a user's performance across various eloquence dimensions by analyzing their emotions in specific eloquence training situations, revealing the inherent connection between emotion and eloquence skills. This module then provides personalized feedback and guidance, providing precise and constructive feedback based on the user's emotional state, helping them understand the impact of their emotions on their eloquence and providing specific suggestions for improvement. This dynamically generates teaching strategies and automatically adjusts teaching content and methods to suit the user's current emotional state. For example, it can push stress-relieving exercises when a user is feeling nervous, or assign more challenging speaking tasks when confidence is high. This leads to a closed-loop optimization process for emotional regulation and skill improvement. The entire module implements a closed-loop optimization process from emotional data acquisition and analysis, feedback and guidance, to dynamic teaching strategy adjustment, helping users better regulate their emotions and achieve efficient and targeted progress in eloquence skills.

[0206] The personalized strategy generation module is used to obtain the user's performance information and match the initial teaching strategy from a preset resource library based on the performance information and the user portrait, so as to continuously monitor the user's performance information when the initial teaching strategy is executed, and then adjust and optimize the initial teaching strategy.

[0207] As a preferred solution of this embodiment, the personalized strategy generation module includes: a user evaluation submodule, a recommendation customization submodule and a feedback optimization submodule;

[0208] The user evaluation submodule is used to obtain the user's performance information through test and homework completion;

[0209] As a preferred solution of this embodiment, the user assessment submodule includes: a language skills assessment layer, a logical thinking and content organization layer, and a non-verbal communication ability layer;

[0210] The language skills assessment layer is used to evaluate the user's pronunciation accuracy, pitch variation, speech speed control and vocabulary richness based on the user's voice information collected during tests and assignments;

[0211] The logical thinking and content organization layer is used to evaluate the user's logical thinking ability and content organization ability based on the voice information collected by the user during tests and assignments;

[0212] The non-verbal communication capability layer is used to evaluate the quality of the user's non-verbal signals based on the user's physical information collected during tests and assignments.

[0213] In this embodiment, the user evaluation submodule integrates multi-dimensional eloquence index analysis tools, such as speech recognition technology and video analysis algorithms, to comprehensively collect and quantify the learner's speech or expression. This includes but is not limited to the following aspects:

[0214] Evaluate pronunciation accuracy, intonation, speed control, vocabulary richness, etc.

[0215] Analyze whether the user's arguments are clear, the logic is tight, and the level of information organization;

[0216] Capture and evaluate the quality and effectiveness of non-verbal cues such as body language, facial expressions, and eye contact.

[0217] In this embodiment, the user evaluation submodule integrates multi-dimensional eloquence index analysis tools, such as voice recognition technology and video analysis algorithms installed on smart devices, to collect audio and video data generated by learners during speeches or expressions in real time or non-real time. The data collected includes but is not limited to learners' voice recordings, video recordings, and possible text records. The language skills assessment layer systematically quantifies the accuracy of learners' pronunciation by analyzing voice signals, such as detecting whether there are mispronunciations or inaccurate syllables. Using voice feature analysis technology, it assesses whether the learner's tone changes are natural and appealing, and whether the speed of speech is properly controlled to avoid being too fast or too slow to affect comprehension. At the same time, vocabulary and word richness are also part of the evaluation content. The system will count and analyze the diversity of vocabulary used by learners and its appropriateness. The logical thinking and content organization layer automatically translates the views and content expressed by learners into structured information. The system uses natural language processing (NLP) technology to analyze the clarity of the argument, that is, whether the viewpoint is clear and easy to understand. The system compares and relates the learners' various views, evaluates the tightness of their logical chain, and judges whether the argument is organized and coherent. The assessment of content organization involves analyzing the cohesion and transitions between paragraphs, the development of themes, and the strength of the conclusion. The nonverbal communication layer utilizes computer vision technology, allowing the system to capture nonverbal information such as the learner's body language, facial expressions, and eye movements. The quality of these nonverbal signals, such as whether body language is coordinated with verbal expression, whether facial expressions are consistent with the contextual emotional expression, and whether eye contact is effective, are analyzed to reflect the learner's confidence, sincerity, and ability to interact with the audience.

[0218] The recommendation customization submodule is used to match an initial teaching strategy from a preset resource library based on the performance information and the user profile; wherein the initial teaching strategy includes teaching and learning materials and homework tests;

[0219] As a preferred solution of this embodiment, the recommendation customization submodule includes: a problem diagnosis layer and a strategy generation layer;

[0220] The problem diagnosis layer is used to conduct in-depth analysis based on the results output by the user evaluation submodule to obtain the shortcomings and deficiencies shown by the user in eloquence training;

[0221] The strategy generation layer is used to dynamically build a personalized learning path model based on the shortcomings and deficiencies shown by the user in eloquence training, and the recommendation and customization engine uses reinforcement learning and collaborative filtering learning technologies.

[0222] In this embodiment, the recommendation customization sub-module uses advanced machine learning and data mining techniques to build a personalized learning path model based on the ability assessment results in each of the above dimensions. This subsystem comprises a problem diagnosis layer and a strategy generation layer. The problem diagnosis layer quickly identifies students' weaknesses in eloquence training and provides a detailed report on their shortcomings. The strategy generation layer automatically designs a set of targeted teaching strategies based on each student's specific situation, including but not limited to specific tasks such as theoretical knowledge learning, practical skills training, and simulated scenario application.

[0223] In this embodiment, the problem diagnosis layer conducts an in-depth analysis of the multi-dimensional evaluation results received from the user ability assessment subsystem. Through machine learning algorithms, such as cluster analysis, decision trees, or neural networks, pattern recognition and correlation research are performed on the student's various eloquence indicator data. Quickly locate the shortcomings and deficiencies shown by each student in eloquence training, such as low pronunciation accuracy, unclear logical thinking, poor non-verbal communication skills, and other specific problems. Based on this diagnostic information, the system generates a detailed ability shortcoming report, providing an accurate basis for the subsequent formulation of teaching strategies.

[0224] In this embodiment, the strategy generation layer is based on the results obtained by the problem diagnosis module, and the intelligent recommendation and customization engine uses reinforcement learning, collaborative filtering or other adaptive learning technologies to dynamically build a personalized learning path model. According to the specific needs and improvement directions of each student, the system will screen out the most suitable learning content and activities from the huge educational resource library. For example, for the problem of inaccurate pronunciation, corresponding pronunciation correction courses or phonetic symbol practice tasks are recommended; for weak logical thinking, theoretical knowledge learning materials and supporting exercises for logical structure analysis and argumentation may be recommended; for the improvement of non-verbal communication skills, practical role-playing and speech simulation scenario application training are designed. The system will also dynamically adjust the recommended content and training intensity according to the student's progress to ensure that the learning path always fits the student's actual progress curve and realizes precise and efficient personalized teaching.

[0225] In this embodiment, the personalized strategy generation module also includes a dynamic teaching resource library and a content management system; the system is equipped with a large amount of educational resources and has real-time updating and integration capabilities. It can screen and combine the most appropriate teaching materials from the resource library according to the needs of personalized teaching strategies to form a personalized course package.

[0226] In this embodiment, resource collection and integration first builds a dynamic teaching resource library containing diversified and multi-dimensional educational resources. This resource library continuously collects and updates various materials related to eloquence training, such as teaching materials, video tutorials, audio files, case studies, and practical simulations. Teaching resources are derived from authoritative publications, expert handouts, high-quality online resources, and original content created or reviewed by teachers, ensuring the richness and professionalism of the resources. Resource labeling and indexing: All stored teaching resources are standardized, including but not limited to classification, metadata information (such as difficulty level, knowledge point association, and student type suitability), and labeling for easy search and filtering. These labels help the system quickly identify and understand the content characteristics and applicable scenarios of each teaching resource. In real-time response to personalized needs, once the intelligent recommendation and customization engine develops a personalized teaching strategy based on the student's ability assessment results, the dynamic teaching resource library and content management system begin to operate. Based on the learning path and goals tailored for the student in the strategy, the system automatically selects the most matching teaching materials from the resource library through algorithms and organically combines them. Personalized course package generation: According to the needs of personalized teaching strategies, the system will organize the selected teaching resources into a structured form a series of highly targeted, coherent and orderly teaching units or complete course packages. At the same time, these course packages may include theoretical learning materials, practical exercise tasks, auxiliary tools and supporting assessments, etc., aiming to fully meet the students' specific learning needs.

[0227] The feedback optimization submodule is used to continuously monitor the user's performance information during the learning process of teaching learning materials and homework tests, and feed it back to the recommendation customization submodule so that the recommendation customization submodule can adjust and optimize the initial teaching strategy based on the feedback information.

[0228] As a preferred solution of this embodiment, the intelligent feedback and adjustment module includes: a data monitoring submodule, a data processing and analysis submodule, a feedback generation submodule and an adjustment execution submodule;

[0229] The data monitoring submodule is used to monitor and obtain the user's multimodal feature vector, user portrait and performance information;

[0230] The data processing and analysis submodule is used to perform big data mining and analysis on the user's multimodal feature vector, user portrait and performance information;

[0231] The feedback generation submodule is used to generate learning feedback based on the results of big data mining and analysis;

[0232] The adjustment execution submodule is used to adjust and optimize the initial teaching strategy and further adjust and optimize the initial teaching strategy according to the learning feedback.

[0233] In this embodiment, during the implementation of personalized teaching strategies by the feedback optimization submodule, the system continuously collects students' learning behavior data and performance feedback, and through real-time monitoring and intelligent analysis, continuously adjusts and improves subsequent teaching plans to achieve closed-loop learning optimization.

[0234] In this embodiment, through data collection and behavioral analysis, during the personalized teaching implementation phase, the system monitors and records students' learning behavior data in real time. This includes, but is not limited to, study time, speed and quality of task completion, frequency of interaction, knowledge mastery, test scores, and more. Through advanced data collection tools and technologies, the system is able to capture subtle changes in students' learning process and their responses to different subject content, learning resources, and teaching methods. In this embodiment, intelligent assessment and feedback generation, based on the large amount of collected behavioral data, utilizes big data analysis and artificial intelligence algorithms to quantitatively evaluate students' learning outcomes, identifying areas of strength, weaknesses, and potential learning barriers. Detailed personalized learning feedback reports are automatically generated and presented to teachers and students in an intuitive manner, helping them understand their learning progress, understand their own learning characteristics, and provide directions and suggestions for improvement. The teaching plan is then dynamically adjusted. Based on real-time feedback, the system can intelligently adjust subsequent teaching content, difficulty level, and teaching pace, such as adding teaching resources or exercises to target weak points or accelerating the teaching process of content that has already been mastered. Adjustments to the teaching plan are not limited to the course content itself, but also include optimization of teaching methods, learning paths, and recommended learning resources. Finally, closed-loop optimization and continuous iteration are carried out. Among them, this feedback and optimization cycle is an ongoing process. After each learning cycle, the system will re-enter the cycle based on the new feedback data, continuously updating and improving personalized teaching strategies to ensure that the teaching plan always keeps pace with the individual development needs and progress trajectory of the students. At the same time, as the learning deepens and the accumulated data increases, the system can achieve more accurate predictions and guidance, thereby improving teaching efficiency and students' learning outcomes.

[0235] Understandably, the personalized strategy generation module focuses on targeted enhancement of each student's abilities across various dimensions of eloquence, ensuring that every user receives the most appropriate improvement plan, avoiding the drawbacks of a "one-size-fits-all" approach. Through contextualized interactive practice, utilizing virtual reality or simulated scenarios, students are provided with abundant hands-on opportunities to hone and improve their eloquence in various speaking situations. They receive immediate feedback and guidance, and adaptive progress management is provided. Based on the learner's progress and individual differences during actual training, the system flexibly adjusts the teaching pace and difficulty to maintain a challenging and motivating environment, effectively promoting the accumulation of long-term learning outcomes. This integration of emotional intelligence incorporates emotion recognition and regulation into the teaching process, helping students better master effective expression in different emotional states, thereby comprehensively improving their overall eloquence. This then allows for visual growth tracking, providing comprehensive ability development maps and data analysis reports, enabling students to clearly understand their growth trajectory across various dimensions of eloquence and facilitating more targeted teaching interventions and support from teachers or coaches.

[0236] In this embodiment, the adjustment execution submodule includes a user capability assessment subsystem, an intelligent recommendation and customization engine, a dynamic teaching resource library and content management subsystem, and a feedback and optimization cycle mechanism.

[0237] In this embodiment, the user ability assessment subsystem comprehensively collects and quantifies a learner's speech or expression by integrating multi-dimensional eloquence indicator analysis tools, such as speech recognition technology and video analysis algorithms. This includes, but is not limited to, the following aspects: A language skills assessment module evaluates pronunciation accuracy, pitch variation, speech rate control, and vocabulary richness. A logical thinking and content organization module analyzes the clarity of the user's arguments, the tightness of their logical chain, and the hierarchical organization of their information. A non-verbal communication skills module captures and evaluates the quality and effectiveness of non-verbal signals such as body language, facial expressions, and eye contact.

[0238] In this embodiment, data collection and preprocessing integrate multi-dimensional eloquence index analysis tools, such as voice recognition technology and video analysis algorithms installed on smart devices, to collect audio and video data generated by learners when they speak or express themselves in real time or non-real time. The collected data includes but is not limited to learners' voice recordings, video recordings, and possible text records. Language skills assessment, through analysis of voice signals, the system quantifies the accuracy of learners' pronunciation, such as detecting whether there are mispronunciations or inaccurate syllables. Using voice feature analysis technology, it is evaluated whether the learner's tone changes are natural and contagious, and whether the speaking speed is properly controlled to avoid speaking too fast or too slowly to affect understanding. At the same time, vocabulary and word richness are also part of the evaluation content. The system will count and analyze the diversity of vocabulary used by learners and its appropriateness.

[0239] In this embodiment, logical thinking and content organization utilize the learner's expressed viewpoints and content, which are automatically translated into structured information. The system then uses natural language processing (NLP) technology to analyze the clarity of the arguments, specifically whether the viewpoints are clear and easy to understand. The system compares and relates the learner's various viewpoints, assessing the tightness of their logical chain and determining whether the argument is organized and coherent. The evaluation of the content organization hierarchy involves analyzing the transitions between paragraphs, the development of the theme, and the strength of the conclusion.

[0240] In this embodiment, nonverbal communication utilizes computer vision technology, allowing the system to capture nonverbal information such as the learner's body movements, facial expressions, and eye movements. The system analyzes the quality of these nonverbal signals, such as whether body language is coordinated with verbal expression, whether facial expressions are consistent with the contextual emotional expression, and whether eye contact is effective, to reflect the learner's confidence, sincerity, and ability to interact with the audience.

[0241] In this embodiment, the intelligent recommendation and customization engine uses advanced machine learning and data mining technologies to build a personalized learning path model based on the ability assessment results of the above dimensions. This includes:

[0242] Problem diagnosis: quickly locate students' weak links in eloquence training and provide a detailed report on their shortcomings. It first conducts an in-depth analysis of the multi-dimensional evaluation results received from the user ability assessment subsystem. Through machine learning algorithms, such as cluster analysis, decision trees or neural networks, pattern recognition and correlation research are conducted on the students' various eloquence indicator data. Quickly locate the shortcomings and deficiencies shown by each student in eloquence training, such as low pronunciation accuracy, unclear logical thinking, poor non-verbal communication skills and other specific problems. The system then generates a detailed report on the shortcomings of the ability based on this diagnostic information, providing an accurate basis for the formulation of subsequent teaching strategies.

[0243] Strategy Generation: Based on the specific circumstances of each student, a targeted teaching strategy is automatically designed, including but not limited to specific tasks such as theoretical knowledge learning, practical skills training, and simulated scenario applications. Based on the results obtained by the problem diagnosis module, the intelligent recommendation and customization engine uses reinforcement learning, collaborative filtering, or other adaptive learning technologies to dynamically build a personalized learning path model. Based on the specific needs and improvement directions of each student, the system will select the most suitable learning content and activities from the vast educational resource library. For example, for problems with inaccurate pronunciation, corresponding pronunciation correction courses or phonetic symbol practice tasks are recommended; for weak logical thinking, theoretical knowledge learning materials and supporting exercises on logical structure analysis and argumentation may be recommended; for improving non-verbal communication skills, practical role-playing and speech simulation scenario application training are designed. The system will also dynamically adjust the recommended content and training intensity based on the student's progress to ensure that the learning path always fits the student's actual progress curve, achieving precise and efficient personalized teaching.

[0244] In this embodiment, the dynamic teaching resource library and content management subsystem are equipped with a large amount of educational resources and have real-time updating and integration capabilities. They can screen and combine the most suitable teaching materials from the resource library according to the needs of personalized teaching strategies to form personalized course packages. Among them, resource collection and integration, the dynamic teaching resource library and content management subsystem first build a dynamic teaching resource library containing diversified and multi-dimensional educational resources. The resource library continuously collects and updates various materials related to eloquence training, such as teaching materials, video tutorials, audio files, case studies, and actual combat simulations. Teaching resources come from authoritative publications, expert handouts, high-quality online resources, and original content created or reviewed by teachers to ensure the richness and professionalism of resources. Resource labeling and indexing, all teaching resources in the library will undergo standardized processing, including but not limited to classification, annotation of metadata information (such as difficulty level, knowledge point association, suitable student type, etc.), and labeling for easy search and screening. These labels help the system quickly identify and understand the content characteristics and applicable scenarios of each teaching resource. Responding to personalized needs in real time, once the intelligent recommendation and customization engine develops a personalized teaching strategy based on the student's ability assessment results, the dynamic teaching resource library and content management system begin to operate. Based on the learning path and goals tailored for the student in the strategy, the system automatically selects the most matching teaching materials from the resource library through algorithms and organically combines them. Finally, a personalized course package is generated. According to the needs of the personalized teaching strategy, the system will structure the selected teaching resources to form a series of highly targeted, coherent and orderly teaching units or complete course packages. These course packages may include theoretical learning materials, practical exercises, auxiliary tools and supporting assessments, etc., aiming to fully meet the specific learning needs of students.

[0245] In this embodiment, the feedback and optimization cycle mechanism continuously adjusts and improves subsequent teaching plans through real-time monitoring and intelligent analysis to achieve closed-loop learning optimization. During the personalized teaching implementation stage, the system will monitor and record the students' learning behavior data in real time, including but not limited to learning time, speed and quality of completing tasks, interaction frequency, knowledge point mastery, test scores, etc. Through advanced data collection tools and technologies, the system can capture subtle changes in students during the learning process, as well as their responses to different subject content, learning resources and teaching methods. Based on the large amount of behavioral data collected, big data analysis and artificial intelligence algorithms are used to quantitatively evaluate the students' learning effects, identify areas of strength, weaknesses and possible learning obstacles, and automatically generate detailed personalized learning feedback reports, which are presented to teachers and students in an intuitive way to help them understand their learning progress, understand their own learning characteristics, and provide directions and suggestions for improvement.

[0246] In this embodiment, based on real-time feedback, the system can intelligently adjust subsequent teaching content, difficulty level, and teaching pace. For example, it can add teaching resources or exercises to target weak points or accelerate the teaching process for content that has already been mastered. Adjustments to the teaching plan are not limited to the course content itself, but also include optimization of teaching methods, learning paths, and recommended learning resources.

[0247] The intelligent feedback and adjustment module is used to perform big data analysis on the user's multimodal feature vector, user portrait and performance information, generate learning feedback, and adjust and optimize the initial teaching strategy based on the learning feedback.

[0248] In this embodiment, the intelligent feedback and adjustment module includes a data acquisition layer, a data processing and analysis layer, an intelligent feedback generation layer, and a dynamic adjustment execution layer. The data acquisition layer uses a high-precision microphone array through an audio capture unit to collect voice signals when students are speaking or expressing themselves. Through the visual perception unit, facial expression recognition technology and a full-body motion capture system are used to record students' non-verbal communication performance, such as facial expressions, body movements, and postures. Through the semantic analysis unit, it connects to the natural language processing (NLP) engine to analyze the logic, information content, and rhetoric of the students' speech content.

[0249] In this embodiment, a high-precision microphone array is positioned appropriately to capture and record the omnidirectional voice signals of students during speeches or oral presentations. Microphone array technology not only improves the quality of the sound signal but also uses algorithms such as beamforming to locate the sound source, effectively suppressing ambient noise and reverberation, ensuring the acquisition of clear and spatially directed speech information. The visual data collection component utilizes facial expression recognition technology and a full-body motion capture system. The facial expression recognition system monitors and records students' facial muscle movements in real time, analyzing their emotional state (such as joy, confusion, tension, etc.) and their level of focus. The full-body motion capture system, on the other hand, uses cameras and other sensors to track students' body movements and posture changes, revealing subtle nuances in their nonverbal expressions, such as gesture emphasis and body orientation, which indirectly reflect their comprehension and communication effectiveness. After receiving the student's speech, the system connects to a natural language processing (NLP) engine for in-depth analysis. This unit performs word segmentation, syntactic analysis, semantic understanding, and emotion recognition on the student's oral presentation, assessing its logical rigor, the adequacy of the information provided, and the appropriateness of the rhetorical techniques used. This series of analyses helps us understand students’ actual grasp of the course content, as well as the level of thinking and expression ability they demonstrate in communication.

[0250] In this embodiment, the data processing and analysis layer converts the collected audio and video data into quantifiable eloquence skill indicators, such as clarity of expression, agility of thinking, emotional appeal, etc., based on a series of preset eloquence evaluation models. It uses emotion recognition technology to analyze the students' emotional state and combines it with the eloquence expression effect to provide deeper support for teaching feedback.

[0251] In this embodiment, the received audio and video data are first processed preliminarily, such as noise reduction and framing of the voice signal, and extraction of key features (such as pitch, speaking speed, pause time, etc.), while facial expressions and body movement information are refined. Based on a series of preset eloquence evaluation models (such as logic evaluation model, clarity scoring model, emotional expression intensity model, etc.), these multimodal features are fused and converted into quantifiable eloquence skill indicators. For example, in terms of expression clarity, pronunciation accuracy, language fluency, etc. are evaluated through voice recognition technology; in terms of mental agility, it is measured based on the student's reaction speed, argument organization ability, and the richness of vocabulary; in terms of emotional appeal, a comprehensive evaluation is made based on factors such as facial expressions, body language, and voice emotion changes.

[0252] In this embodiment, the subsystem uses deep learning or machine learning algorithms to conduct in-depth analysis of the collected emotion-related data. Facial expression recognition technology is used to analyze students' basic emotional states such as joy, anger, sorrow, and happiness during the speech, as well as more subtle and complex emotional tendencies. At the same time, combined with voice emotion recognition technology, acoustic features such as tone, speaking speed, and volume are analyzed to judge students' emotional fluctuations when speaking, so as to fully understand their emotional state and how it affects their eloquence. Combining the above-mentioned eloquence dimension indicators and emotion recognition results, more accurate and in-depth support is provided for subsequent teaching feedback, allowing teachers or virtual tutors to give targeted advice and guidance on specific issues.

[0253] In this embodiment, the intelligent feedback generation layer generates objective and specific feedback based on the analysis results of the above dimensions. This includes text-based summary reports, visual charts, and voice suggestions using voiceprint simulation. Based on the student's progress and current weaknesses, it automatically generates personalized training plans and practice tasks tailored to their needs.

[0254] In this embodiment, these abstract data are integrated and analyzed through the quantitative results of the eloquence dimension indicators and the emotion recognition diagnostic information obtained. For each student's performance, the system will automatically generate an objective and accurate text summary report, for example: "Your expression clarity score is high, but there are slight emotional fluctuations during the speech, which may affect the audience's understanding of the content; in terms of visual chart display, the system can draw trend charts and comparative analysis charts of various skill indicators over time, so as to intuitively see the student's progress or areas for improvement in different dimensions; the voice suggestion function of voiceprint simulation uses advanced speech synthesis technology to imitate the voice of teachers or other authorities to provide students with more friendly and personalized audio feedback, such as simulating the voice of a real teacher to give "Try to maintain a stable emotion in your next speech, and increase pauses appropriately to enhance the expression effect.

[0255] In this embodiment, based on the analysis results provided by the real-time feedback generator and the students' long-term performance records, the dynamic training plan maker can intelligently analyze each student's strengths and weaknesses, taking into account their learning goals and personal characteristics. It then automatically designs and generates personalized training plans tailored to the student's current needs. For example, it provides pronunciation practice tasks for students with unclear language expression, and arranges emotion management courses or situational simulation training for students with poor emotional control. The training plan not only includes specific learning content, but also covers multiple aspects such as difficulty gradients, frequency settings, and practice opportunities, ensuring the scientific and effective nature of the training and helping students improve their oral skills in a targeted manner.

[0256] In this embodiment, the dynamic adjustment execution layer can adjust the virtual learning environment according to the students' emotional state, such as switching to a learning scene that is more suitable for relaxation or concentration, and continuously optimize and iterate the teaching strategy based on the students' feedback response speed and effect to ensure that the interaction process is highly interactive and effective.

[0257] In this embodiment, biometric technologies (such as facial expression recognition and heart rate monitoring) are integrated to monitor students' emotional states in real time, such as anxiety, focus, or relaxation. When a student's emotional changes are detected, the system automatically adjusts the virtual learning environment based on a pre-set psychological model and context-matching algorithm. For example, if a student exhibits high stress levels, the system may switch to a virtual space with soft colors, relaxing background music, and natural scenery to help them relieve their emotions and improve their learning efficiency. Conversely, if a student needs to improve their focus, the system may switch to a bright, concise, and distraction-free learning environment. Based on in-depth analysis of student behavior data, learning outcomes, and feedback response time generated by the intelligent feedback generation layer, the system dynamically evaluates the impact of current teaching methods and content on students. Based on these real-time data analysis results, the teaching strategy optimizer can quickly iterate and adjust teaching plans, such as changing the explanation rhythm, adding interactive elements, and adjusting the difficulty of knowledge points, to ensure that the teaching strategy matches the student's needs and abilities. This highly personalized teaching strategy update is designed to enhance the interactivity between students and the learning system, improve the quality of the learning experience, and ultimately promote students' knowledge acquisition and skill improvement.

[0258] The implementation of the above embodiment has the following effects:

[0259] The technical solution of the present invention combines visual, voice, and text data to comprehensively identify and analyze emotional states to improve the accuracy of emotion recognition in eloquence teaching scenarios. At the same time, it uses advanced algorithms to process multimodal data, including data fusion, correlation analysis between emotions and eloquence performance, and the construction and updating of user profiles, thereby deeply exploring the insights behind the data. It also dynamically generates and adjusts teaching strategies based on students' emotions, eloquence abilities, and needs, including an intelligent recommendation system and a dynamic teaching resource library to optimize the learning experience. It then constructs a cognitive map that reflects the relationship between students' emotions, eloquence performance, and teaching scenarios to generate more targeted teaching strategies and feedback. As user training and feedback accumulate, user profiles are updated in real time to ensure that the teaching content stays synchronized with the students' latest status. The accuracy of user profiles is maintained through efficient data processing and machine learning models. The teaching content and rhythm are dynamically adjusted based on real-time emotion analysis and changes in user profiles, emphasizing the system's immediate responsiveness and adaptability to provide a more personalized teaching experience.

[0260] Example 2

[0261] See also Figure 3 , which is the system operation process provided by the embodiment of the present invention, including: the startup stage, the operation of the emotion recognition and perception module, the operation of the user portrait construction module, the work of the emotion calculation and response module, the execution of the personalized teaching strategy generation module, the operation of the intelligent feedback and adjustment module, and the end stage.

[0262] During the startup phase, the system is initialized, the interfaces of each module are connected, and basic data and model parameters are loaded.

[0263] The emotion recognition and perception module operates through visual data acquisition, using cameras to capture non-verbal information such as students' facial expressions and body language, and performs real-time processing and emotion recognition. The speech signal processing and understanding module collects students' voice information through a microphone, analyzing the intonation, tone, and content to determine their emotional state and learning intention. This comprehensive analysis of visual and speech information forms preliminary emotional recognition results, triggering further teaching activities or strategy adjustments as needed.

[0264] The user portrait construction module operates by obtaining user characteristic data (such as learning behavior records, test scores, and user feedback), conducting in-depth mining of the collected data, and establishing multi-dimensional characteristic models such as user behavior patterns, interest preferences, and ability levels. Based on the above analysis results, the user portrait is constructed and dynamically updated to ensure the accuracy of personalized teaching strategies.

[0265] The emotional computing and response module works by continuously receiving emotion recognition data from a module and analyzing it in combination with historical emotional information in the user portrait, further refining the classification and quantification of emotional states, forming an emotional model, and interpreting students' emotional reactions in combination with current learning tasks and situations, and inferring their internal needs and motivations. Based on the students' emotional states, it generates appropriate motivational feedback or teaching suggestions to guide the generation of adaptive teaching strategies. Finally, it adjusts the course content, rhythm and presentation methods according to students' emotional feedback and learning status to achieve real-time interaction and personalized teaching.

[0266] The personalized teaching strategy generation module is executed to objectively evaluate students' abilities through tests and homework completion, and uses user portraits and ability assessment results to match the most appropriate learning materials and exercises from the dynamic teaching resource library and content management system. It also continuously collects user performance data during the learning process and feeds it back to each module for improving subsequent teaching strategies.

[0267] The intelligent feedback and adjustment module monitors all relevant data generated during the entire learning process, including learning outcomes, interaction frequency, satisfaction, etc. It conducts in-depth analysis of these data to identify possible problems and areas for improvement, and generates specific, clear, and personalized learning feedback based on the analysis results. This feedback can be in the form of text, images, or videos, allowing for immediate adjustments to the teaching plan based on intelligent feedback, such as increasing explanations of difficult points, providing additional tutoring, or adjusting the learning progress.

[0268] At the end of the learning process, the system summarizes the data from the learning session, updates the user profile and teaching profile, and develops a more precise and personalized plan for the next learning session. Simultaneously, the system self-learns and optimizes its algorithms to provide more efficient and user-friendly teaching services the next time the user is used.

[0269] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A multi-dimensional emotion perception and intelligent eloquence teaching system, characterized by: include: Emotion recognition and perception module, user portrait construction module, emotion calculation and response module, personalized strategy generation module and intelligent feedback and adjustment module; The emotion recognition and perception module is used to obtain the user's body information and voice information, and obtain visual emotion information and auditory emotion information based on the body information and the voice information, respectively, and then fuse the visual emotion information and the auditory emotion information to obtain a multimodal feature vector; The user portrait construction module is used to obtain user feature data and perform language, audio and visual analysis on the user feature data, facial information and voice information respectively to obtain user language features, user audio features and user visual features, and then perform user clustering analysis based on the user language features, user audio features and user visual features to construct the user's knowledge graph, thereby generating a user portrait; The emotion calculation and response module is used to calculate the user's current emotional state based on the multimodal feature vector, and analyze the influence relationship between the user's current emotional state and eloquence skill performance based on preset scenarios, and then adjust the teaching suggestions based on the influence relationship; wherein, the current teaching scene is deeply analyzed to extract teaching scene features; Based on the characteristics of the teaching scenario, the user's emotional response based on the characteristics of the teaching scenario is associated with their specific eloquence performance; a contextualized cognitive map is formed based on the user's emotional response based on the characteristics of the teaching scenario and their specific eloquence performance; based on the contextualized cognitive map, a deep learning optimization of the association model is performed, and the cognitive map is continuously adjusted and updated based on the latest real-time data and deep learning feedback, thereby showing the dynamic interaction between the user's emotions and eloquence ability in different teaching scenarios; The personalized strategy generation module is used to obtain the user's performance information and match the initial teaching strategy from a preset resource library based on the performance information and the user profile, so as to continuously monitor the user's performance information during the execution of the initial teaching strategy, and then adjust and optimize the initial teaching strategy; The intelligent feedback and adjustment module is used to perform big data analysis on the user's multimodal feature vector, user portrait and performance information, generate learning feedback, and adjust and optimize the initial teaching strategy based on the learning feedback.

2. A multi-dimensional emotion perception and intelligent eloquence teaching system according to claim 1, characterized in that: The emotion recognition and perception module includes: a visual acquisition submodule, an audio acquisition submodule and a multimodal fusion submodule; The visual acquisition submodule is used to obtain the user's facial expressions and body language, and extract facial muscle movement feature information of the facial expressions based on the face detection module, so as to identify the user's visual emotional information based on the facial muscle movement feature information; wherein the body information includes facial expressions and body language; The audio acquisition submodule is used to obtain the user's voice information, and perform voice feature extraction and natural language processing on the voice information to obtain language information and audio information; and identify the user's auditory emotion information based on the audio information and the language information; The multimodal fusion submodule is used to align and match the visual emotion information and the auditory emotion information, thereby fusing the aligned and matched visual emotion information and auditory emotion information to obtain a multimodal feature vector.

3. A multi-dimensional emotion perception and intelligent eloquence teaching system as claimed in claim 2, characterized in that: The multimodal fusion submodule includes: a data integration layer, an emotion fusion layer and a decision output interface; The data integration layer is used to receive the visual emotion information and the auditory emotion information, and align and match the visual emotion information and the auditory emotion information on a preset time axis; The emotion fusion layer is used to fuse the aligned and matched visual emotion information and auditory emotion information to obtain a multimodal feature vector; The decision output interface is used to send the multimodal feature vector to the emotion calculation and response module, the personalized strategy generation module or the intelligent feedback and adjustment module for teaching generation.

4. The multi-dimensional emotion perception and intelligent eloquence teaching system according to claim 1, characterized in that: The user portrait construction module includes: a data collection submodule, an analysis and modeling submodule, and a portrait generation and update submodule; The data collection submodule is used to obtain user characteristic data; wherein the user characteristic data includes learning record behavior, test scores and user feedback information; The analysis and modeling submodule is used to perform in-depth mining on the user feature data, facial information, and voice information to obtain user language features, user audio features, and user visual features, and to construct corresponding feature models based on the user language features, user audio features, and user visual features, and then to construct the user's knowledge graph through the feature models; The portrait generation and update submodule is used to generate or update the user portrait based on the user's knowledge graph.

5. A multi-dimensional emotion perception and intelligent eloquence teaching system as claimed in claim 4, characterized in that: The analysis and modeling submodule includes: a feature extraction engine, an associated model layer and a model library; The feature extraction engine is used to perform in-depth mining on the user feature data, facial information, voice information, and acquired user eloquence data, thereby extracting user language features, user audio features, user visual features, and user eloquence features; The association model layer is used to obtain the corresponding emotional state based on the user's language features, user audio features, and user visual features, and then construct an association feature model corresponding to the emotional state and the user's eloquence features, and then construct the user's knowledge graph through the association feature model; The model library is used to store the associated feature models generated by users and their corresponding knowledge graphs.

6. A multi-dimensional emotion perception and intelligent eloquence teaching system according to claim 5, characterized in that: The portrait generation and updating submodule includes: a user attribute layer, an eloquence dimension evaluation matrix, an emotional response curve, and a portrait dynamic update layer; The user attribute layer is used to obtain the user's basic information and the user's eloquence characteristics; The eloquence dimension evaluation matrix is used to form a visual evaluation matrix based on the user's eloquence characteristics and display it; The emotional response curve is used to obtain the user's emotions in different situations, generate an emotional response curve, and combine the user's eloquence characteristics to obtain the influence of emotions on the user's eloquence; The dynamic portrait update layer is used to generate or update the user portrait based on the user's basic information, evaluation matrix, the influence of emotions on the user's eloquence and the user's knowledge graph.

7. The multi-dimensional emotion perception and intelligent eloquence teaching system according to claim 1, characterized in that: The emotional calculation and response module includes: an emotional data acquisition submodule, an emotional analysis and modeling submodule, a situational perception and understanding submodule, a feedback generation and optimization submodule, and a real-time interactive teaching decision submodule; The emotion data collection submodule is used to obtain the user's historical emotion information; The sentiment analysis modeling submodule is used to form a sentiment model based on the user's historical sentiment information, and to calculate the user's current quantified emotional state by combining the multimodal feature vector and the user portrait; The contextual awareness and understanding submodule is used to analyze the influence relationship between the user's current emotional state and eloquence skill performance based on a preset scenario and the user's current quantified emotional state; The feedback generation optimization submodule is used to generate appropriate motivational feedback or teaching suggestions based on the influence relationship; The real-time interactive teaching decision-making submodule is used to adjust the course content, rhythm and presentation method according to the motivational feedback or teaching suggestions.

8. The multi-dimensional emotion perception and intelligent eloquence teaching system according to claim 7, characterized in that: The feedback generation optimization submodule is further used to: Generate personalized feedback using a preset rule library and natural language processing technology based on the emotional state and eloquence performance obtained by the sentiment analysis modeling submodule; At the same time, based on the user's current emotional state and eloquence characteristics, corresponding teaching activities and training plans are designed and updated, so as to automatically generate optimization suggestions for the user's emotional state and eloquence shortcomings based on the results output by the emotional data collection submodule, the emotional analysis modeling submodule, and the situational perception and understanding submodule.

9. The multi-dimensional emotion perception and intelligent eloquence teaching system according to claim 8, characterized in that: The real-time interactive teaching decision submodule further includes: Continuously receive and analyze user emotional state data, and dynamically adjust teaching content and methods based on the user's emotional state to match the user's learning needs and psychological state; Dispatching corresponding teaching resources based on real-time analysis of users' learning progress and emotional responses; Adjust the difficulty level of course tasks according to users' emotional reactions and ability assessments; Generate adaptive teaching decisions based on real-time analytics, emotional responses, and ability assessments.

10. The multi-dimensional emotion perception and intelligent eloquence teaching system according to claim 1, characterized in that: The personalized strategy generation module includes: a user evaluation submodule, a recommendation customization submodule, and a feedback optimization submodule; The user evaluation submodule is used to obtain the user's performance information through test and homework completion; The recommendation customization submodule is used to match an initial teaching strategy from a preset resource library based on the performance information and the user profile; wherein the initial teaching strategy includes teaching and learning materials and homework tests; The feedback optimization submodule is used to continuously monitor the user's performance information during the learning process of teaching learning materials and homework tests, and feed it back to the recommendation customization submodule so that the recommendation customization submodule can adjust and optimize the initial teaching strategy based on the feedback information.

11. The multi-dimensional emotion perception and intelligent eloquence teaching system according to claim 10, characterized in that: The user assessment submodule includes: a language skills assessment layer, a logical thinking and content organization layer, and a non-verbal communication ability layer; The language skills assessment layer is used to evaluate the user's pronunciation accuracy, pitch variation, speech speed control and vocabulary richness based on the user's voice information collected during tests and assignments; The logical thinking and content organization layer is used to evaluate the user's logical thinking ability and content organization ability based on the voice information collected by the user during tests and assignments; The non-verbal communication capability layer is used to evaluate the quality of the user's non-verbal signals based on the user's physical information collected during tests and assignments.

12. The multi-dimensional emotion perception and intelligent eloquence teaching system according to claim 10, characterized in that: The recommendation customization submodule includes: a problem diagnosis layer and a strategy generation layer; The problem diagnosis layer is used to conduct in-depth analysis based on the results output by the user evaluation submodule to obtain the shortcomings and deficiencies shown by the user in eloquence training; The strategy generation layer is used to dynamically build a personalized learning path model based on the shortcomings and deficiencies shown by the user in eloquence training, and the recommendation and customization engine uses reinforcement learning and collaborative filtering learning technologies.

13. The multi-dimensional emotion perception and intelligent eloquence teaching system according to claim 1, characterized in that: The intelligent feedback and adjustment module includes: a data monitoring submodule, a data processing and analysis submodule, a feedback generation submodule and an adjustment execution submodule; The data monitoring submodule is used to monitor and obtain the user's multimodal feature vector, user portrait and performance information; The data processing and analysis submodule is used to perform big data mining and analysis on the user's multimodal feature vector, user portrait and performance information; The feedback generation submodule is used to generate learning feedback based on the results of big data mining and analysis; The adjustment execution submodule is used to adjust and optimize the initial teaching strategy and further adjust and optimize the initial teaching strategy according to the learning feedback.

Citation Information

Patent Citations

  • English teaching man-machine interaction system and method thereof

    CN113837907A

  • Personalized learning platform based on self-expansion knowledge base and multi-modal portrait

    CN114372155A