A digital human-computer interaction method and system for an art education course
By performing multimodal analysis on the speech and facial expressions of the learners, digital humans can more accurately identify the learners' intentions and emotional states, adjust teaching strategies, solve the problem of rigid interactive behavior in existing systems, and achieve more natural emotional resonance and personalized art education experiences.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FOSHAN POLYTECHNIC
- Filing Date
- 2026-03-09
- Publication Date
- 2026-07-10
AI Technical Summary
Existing digital human teaching systems cannot accurately identify learners' real-time state and inner emotions in art education scenarios, resulting in mechanical and rigid interactive behaviors. They cannot dynamically adjust the teaching pace and expression methods according to students' emotional state and intentions, which affects the personalization and interactive experience of art education.
By acquiring the speech and facial expression information of the learners, we can conduct micro-feature analysis of emotional expression, spontaneous deviation analysis, cross-modal performance consistency analysis, and emotional lag analysis to identify learners' intentions and emotional tendencies. Based on this information, we can adjust the digital human's expression and teaching strategies to achieve more natural and emotionally resonant interactions.
It improves the naturalness of interaction and emotional resonance of digital humans in art teaching, provides a personalized and immersive art education experience, and significantly enhances teaching quality.
Smart Images

Figure CN122363496A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a digital human-computer interaction method and system for art education courses. Background Technology
[0002] With the rapid development of artificial intelligence and digital human technology, digital humans are increasingly being applied in the field of arts education, becoming an important tool to support teaching. Traditional arts education courses typically rely on face-to-face demonstrations and guidance from teachers, but limitations in teacher availability, time and space constraints, and insufficient personalized teaching resources make it difficult to meet the diverse learning needs of students. Digital humans can simulate the teaching behavior of real teachers, providing immersive and repeatable interactive teaching experiences. They have shown particular potential in arts disciplines such as vocal music, performance, and dance, which require real-time feedback and emotional expression, offering a new technological path for the popularization and personalization of arts education.
[0003] However, existing digital human teaching systems have significant limitations in art education scenarios. Most systems can only output information one-way based on preset fixed processes, lacking effective perception and response to learners' real-time status and inner emotions. These systems often fail to simultaneously capture and integrate learners' voice content and facial expressions, making it difficult to accurately identify their learning intentions and emotional tendencies. This results in mechanical and rigid interactive behavior of digital humans, unable to dynamically adjust teaching pace and expression according to students' emotional states and intentions like a real teacher. Consequently, this affects the effectiveness of emotional resonance and personalized guidance in art education, limiting the practical role of digital humans in improving the quality of art teaching and interactive experiences. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides a digital human-computer interaction method and system for art education courses, enabling digital humans to achieve more natural and emotionally resonant interactions, providing a personalized and immersive art education experience, and significantly improving the practical role and teaching quality of digital humans in art teaching.
[0005] To address the aforementioned technical problems, this invention provides a digital human-computer interaction method for art education courses, the method comprising: The system acquires the vocal and facial expression information of the teaching subjects in art education courses, and performs micro-feature analysis of emotional expression based on the vocal and facial expression information to obtain micro-feature information of emotional expression. Based on the aforementioned voice and facial expression information, spontaneous deviation analysis of the learning subjects during the learning process is performed to obtain spontaneous deviation information; Cross-modal performance consistency analysis is performed based on the aforementioned voice performance information and facial performance information to obtain cross-modal performance consistency information; Based on the aforementioned voice and facial expression information, analyze the emotional lag of the teaching subjects in response to the preset stimuli; Based on the aforementioned micro-feature information of emotional expression, spontaneous deviation information, cross-modal performance consistency information, and emotional lag information, the intention of the teaching object is analyzed to obtain the target intention information. Furthermore, based on the aforementioned micro-feature information of emotional expression, spontaneous deviation information, cross-modal performance consistency information, and emotional lag information, the degree of deviation from the performance mode is used to analyze the emotional tendency of the teaching object to obtain the target emotional tendency. Based on the target intent information and target emotional tendency, determine the digital human's expressive information, determine the digital human's teaching adjustment strategy based on the target intent information and target emotional tendency, and drive the digital human to interact with the teaching object based on the expressive information and teaching adjustment strategy.
[0006] Optionally, the step of performing micro-feature analysis of emotional expression based on the voice performance information and facial performance information to obtain micro-feature information of emotional expression includes: The speech performance information is analyzed for minute speech change features to obtain minute speech change feature information; The facial expression information is analyzed for movement changes to obtain movement change features. Based on the subtle changes in speech and movement changes, the micro-features of emotional expression are analyzed to obtain micro-features of emotional expression.
[0007] Optionally, the step of analyzing the spontaneous deviation of the learning object during the learning process based on the voice performance information and facial performance information to obtain spontaneous deviation information includes: Detail feature extraction is performed on the speech performance information to obtain speech detail feature information; Detailed feature extraction is performed on the facial expression information to obtain facial detail feature information; Analyze the first deviation of speech detail feature information from standard speech detail, and analyze the second deviation of facial detail feature information from standard facial detail; Based on the first deviation and the second deviation, a spontaneous deviation analysis of the learning subjects is performed to obtain spontaneous deviation information.
[0008] Optionally, the step of performing cross-modal performance consistency analysis based on the voice performance information and facial performance information to obtain cross-modal performance consistency information includes: Feature vectors are extracted from the speech performance information to obtain speech feature vectors, and feature vectors are extracted from the facial performance information to obtain facial feature vectors. Based on the speech feature vector and facial feature vector, a performance consistency index is calculated, and cross-modal performance consistency analysis is performed based on the performance consistency index to obtain cross-modal performance consistency information.
[0009] Optionally, the step of extracting feature vectors from the speech performance information to obtain speech feature vectors, and extracting feature vectors from the facial performance information to obtain facial feature vectors, includes: Based on the fundamental frequency extraction algorithm and the time-domain sliding window, the vocal cord tremor frequency and amplitude are extracted from the speech performance information, and a speech feature vector is generated based on the vocal cord tremor frequency and amplitude. Facial muscle stiffness analysis was performed on the facial expression information to obtain facial muscle stiffness information. Based on the Hough circle transform, blink frequency and pupil diameter changes are extracted from the facial expression information, and facial feature vectors are generated based on the facial muscle stiffness information, blink frequency, and pupil diameter changes.
[0010] Optionally, the step of analyzing the teaching object's intent based on the micro-features of emotional expression, spontaneous deviation information, cross-modal consistency information, and emotional lag information to obtain target intent information includes: Based on the cross-modal performance consistency information and emotional lag information, pattern logic judgment is performed to obtain the pattern logic judgment result; Based on the micro-feature information of the emotional expression, an emotional resonance depth analysis is performed to obtain the target emotional resonance depth; Based on the spontaneous deviation information and the depth of target emotional resonance, an internal state assessment is performed to obtain the internal state assessment results. Based on the pattern logic judgment results and the internal state assessment results, the intention analysis of the teaching object is performed to obtain the target intention information.
[0011] Optionally, the step of analyzing the emotional tendency of teaching subjects based on the degree of deviation from the performance mode, using the micro-feature information of emotional expression, spontaneous deviation information, cross-modal performance consistency information, and emotional lag information to obtain the target emotional tendency, includes: Obtain the historical performance pattern sequence of teaching subjects in art education courses, and analyze the evolution trend of performance patterns based on the historical performance pattern sequence to obtain information on the evolution trend of historical performance patterns. Based on the voice performance information and facial performance information, the current performance pattern information of the teaching object is generated, and the degree of deviation between the current performance pattern information and the historical performance pattern evolution trend information is analyzed. Based on the aforementioned micro-features of emotional expression, spontaneous deviation information, cross-modal consistency information, and emotional lag information, the degree of deviation is used to analyze the emotional tendencies of the teaching subjects and obtain the target emotional tendencies.
[0012] Optionally, the analysis of the deviation between the current performance pattern information and the historical performance pattern evolution trend information includes: The historical performance pattern sequence is decomposed into a time series to obtain the decomposed historical performance pattern sequence. A historical benchmark model is constructed based on the historical performance pattern sequence after time series decomposition. The deviation between the current performance pattern information and the historical performance pattern evolution trend information is analyzed based on the historical benchmark model.
[0013] Optionally, determining the digital human's expressive information based on the target intent information and target emotional tendency includes: Based on the target intention information and target emotional tendency, several emotional levels of the teaching object are determined; Determine the intensity weight of each emotional level, and select several target non-verbal expression elements based on the intensity weight using a non-verbal expression element library; Determine the animated action expression information of the digital human based on several target non-verbal expression elements; Based on the target intent information and target emotional tendency, language expression elements are determined, and based on the animation action expression information and language expression elements, the digital human's expression information is determined.
[0014] In addition, the present invention also provides a digital human-computer interaction system for art education courses, the system comprising: Feature analysis module: used to acquire the speech and facial expression information of the teaching subjects in the art education course, and to perform micro-feature analysis of emotional expression based on the speech and facial expression information to obtain micro-feature information of emotional expression; Learning Deviation Analysis Module: Used to analyze the spontaneous deviation of the learning object during the learning process based on the voice performance information and facial performance information, and obtain spontaneous deviation information; Performance consistency analysis module: used to perform cross-modal performance consistency analysis based on the voice performance information and facial performance information to obtain cross-modal performance consistency information; Emotional lag analysis module: used to analyze the emotional lag information of the teaching object to the preset stimuli based on the voice performance information and facial performance information; The intention and sentiment analysis module is used to analyze the intention of the teaching object based on the micro-feature information of the emotional expression, spontaneous deviation information, cross-modal performance consistency information and emotional lag information to obtain the target intention information. Based on the micro-feature information of the emotional expression, spontaneous deviation information, cross-modal performance consistency information and emotional lag information, the module uses the degree of deviation of the performance mode to analyze the sentiment tendency of the teaching object to obtain the target sentiment tendency. Interaction module: used to determine the digital human's expressive information based on the target intention information and target emotional tendency, determine the digital human's teaching adjustment strategy based on the target intention information and target emotional tendency, and drive the digital human to interact with the teaching object based on the expressive information and teaching adjustment strategy.
[0015] In this embodiment of the invention, vocal and facial expression information of the learners in art education courses is acquired. Based on this information, micro-feature analysis of emotional expression is performed to more accurately capture subtle emotional changes in the learners. Spontaneous deviation analysis of the learners during the learning process, based on vocal and facial expression information, allows for a more objective assessment of their focus and engagement. Cross-modal consistency analysis, based on vocal and facial expression information, provides a more comprehensive understanding of their inner emotional state, avoiding the limitations of single-modal analysis. Analysis of the learners' emotional lag in response to pre-set stimuli is conducted based on vocal and facial expression information. Intention analysis of the learners is performed based on micro-feature information of emotional expression, spontaneous deviation information, cross-modal consistency information, and emotional lag information. Furthermore, analysis of the learners' emotional tendencies is conducted using the degree of deviation from performance patterns, based on these factors, enabling a comprehensive and in-depth perception of the learners' real-time state and inner emotions. This accurately identifies the learners' learning intentions and emotional states. Based on the target intention information and target emotional tendency, the expressive information of the digital human and the teaching adjustment strategy are determined. Based on the expressive information and teaching adjustment strategy, the digital human interacts with the teaching object, enabling the digital human to achieve more natural and emotionally resonant interaction, providing a personalized and immersive art education experience, and significantly improving the actual role and teaching quality of the digital human in art teaching. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating a digital human-computer interaction method for art education courses according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating a digital human-computer interaction method for art education courses according to another embodiment of the present invention. Figure 3 This is a schematic diagram of the structural composition of a digital human-computer interaction system for art education courses according to an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] Example 1 Please see Figure 1 , Figure 1 This is a flowchart illustrating a digital human-computer interaction method for art education courses according to an embodiment of the present invention. The method includes: S11: Obtain the speech and facial expression information of the teaching subjects in the art education course, and perform micro-feature analysis of emotional expression based on the speech and facial expression information to obtain micro-feature information of emotional expression. In the specific implementation of this invention, the speech and facial expression information of the teaching subjects in art education courses are acquired. The speech expression information is analyzed for subtle changes in speech characteristics to obtain subtle change feature information. The facial expression information is analyzed for movement changes to obtain movement change features. Based on the subtle change feature information of speech and movement changes, the micro-feature analysis of emotional expression is performed to obtain micro-feature information of emotional expression. This more accurately captures the subtle emotional changes of the teaching subjects, providing more precise input for subsequent emotional tendency and intention analysis, thereby improving the accuracy of the digital human's perception of the emotional state of the teaching subjects.
[0020] S12: Based on the spoken and facial expression information, perform spontaneous deviation analysis on the learning process of the teaching object to obtain spontaneous deviation information; In the specific implementation of this invention, detailed features are extracted from the speech performance information to obtain speech detail feature information; detailed features are extracted from the facial performance information to obtain facial detail feature information; the first deviation between the speech detail feature information and standard speech details is analyzed, and the second deviation between the facial detail feature information and standard facial details is analyzed; based on the first and second deviations, spontaneous deviation analysis of the learning object is performed to obtain spontaneous deviation information; by comparing the deviation of the learning object's speech and facial details from standard details, its spontaneous deviation in the learning process is quantified, thereby more objectively assessing the learning object's focus and engagement, and providing a basis for adjusting teaching strategies for digital humans.
[0021] S13: Perform cross-modal performance consistency analysis based on the aforementioned voice performance information and facial performance information to obtain cross-modal performance consistency information; In the specific implementation of this invention, feature vectors are extracted from the speech performance information to obtain speech feature vectors, and feature vectors are extracted from the facial performance information to obtain facial feature vectors. Based on the speech feature vectors and facial feature vectors, a performance consistency index is calculated, and cross-modal performance consistency analysis is performed based on the performance consistency index to obtain cross-modal performance consistency information. By calculating the performance consistency index of speech and facial feature vectors, the coordination of the teaching object's different modal expressions can be effectively evaluated, thereby gaining a more comprehensive understanding of its inner emotional state and avoiding the limitations of single-modal analysis.
[0022] S14: Analyze the emotional lag information of the teaching object to the preset stimuli based on the voice performance information and facial performance information; In the specific implementation of this invention, the emotional lag information of the teaching object to the preset stimuli is analyzed based on the voice performance information and facial performance information, so as to provide more comprehensive data support for subsequent intention analysis and emotional tendency analysis.
[0023] S15: Based on the micro-feature information of emotional expression, spontaneous deviation information, cross-modal performance consistency information and emotional lag information, perform intention analysis on the teaching object to obtain target intention information, and based on the micro-feature information of emotional expression, spontaneous deviation information, cross-modal performance consistency information and emotional lag information, use the degree of deviation of performance mode to perform emotional tendency analysis on the teaching object to obtain target emotional tendency. In the specific implementation of this invention, pattern logic judgment is performed based on the cross-modal performance consistency information and emotional lag information to obtain pattern logic judgment results; emotional resonance depth analysis is performed based on the emotional expression micro-feature information to obtain the target emotional resonance depth; internal state assessment is performed based on the spontaneous deviation information and the target emotional resonance depth to obtain internal state assessment results; and the intention analysis of the teaching object is performed based on the pattern logic judgment results and the internal state assessment results to obtain target intention information, thereby more accurately identifying the deep learning intention of the teaching object and providing more precise teaching guidance for digital humans. The system acquires a sequence of historical performance patterns of students in art education courses, and analyzes the evolution trend of these patterns to obtain historical performance pattern evolution trend information. Based on the spoken and facial performance information, it generates current performance pattern information for the students and analyzes the degree of deviation between the current performance pattern information and the historical performance pattern evolution trend information. Using the degree of deviation, it analyzes the students' emotional tendencies based on the micro-features of emotional expression, spontaneous deviation information, cross-modal performance consistency information, and emotional lag information to obtain target emotional tendencies. This allows for a more dynamic and personalized assessment of the students' emotional tendencies, enabling the digital human to better understand the students' long-term emotional changes and provide more targeted teaching feedback.
[0024] S16: Determine the digital human's expressive information based on the target intention information and target emotional tendency, determine the digital human's teaching adjustment strategy based on the target intention information and target emotional tendency, and drive the digital human to interact with the teaching object based on the expressive information and teaching adjustment strategy.
[0025] In the specific implementation of this invention, several emotional levels of the teaching object are determined based on the target intention information and target emotional tendency; the intensity weight of each emotional level is determined, and several corresponding target non-verbal expression elements are selected based on the intensity weight using a non-verbal expression element library; the animated action expression information of the digital human is determined based on the several target non-verbal expression elements; language expression elements are determined based on the target intention information and target emotional tendency, and the expression information of the digital human is determined based on the animated action expression information and language expression elements; the teaching adjustment strategy of the digital human is determined based on the target intention information and target emotional tendency, and the digital human interacts with the teaching object based on the expression information and teaching adjustment strategy. By refining the emotional levels, determining the intensity weights, and combining non-verbal and verbal expression elements, the expression information of the digital human becomes more layered and personalized, thereby improving the naturalness of the digital human interaction and the emotional resonance effect.
[0026] In this embodiment of the invention, vocal and facial expression information of the learners in art education courses is acquired. Based on this information, micro-feature analysis of emotional expression is performed to more accurately capture subtle emotional changes in the learners. Spontaneous deviation analysis of the learners during the learning process, based on vocal and facial expression information, allows for a more objective assessment of their focus and engagement. Cross-modal consistency analysis, based on vocal and facial expression information, provides a more comprehensive understanding of their inner emotional state, avoiding the limitations of single-modal analysis. Analysis of the learners' emotional lag in response to pre-set stimuli is conducted based on vocal and facial expression information. Intention analysis of the learners is performed based on micro-feature information of emotional expression, spontaneous deviation information, cross-modal consistency information, and emotional lag information. Furthermore, analysis of the learners' emotional tendencies is conducted using the degree of deviation from performance patterns, based on these factors, enabling a comprehensive and in-depth perception of the learners' real-time state and inner emotions. This accurately identifies the learners' learning intentions and emotional states. Based on the target intention information and target emotional tendency, the expressive information of the digital human and the teaching adjustment strategy are determined. Based on the expressive information and teaching adjustment strategy, the digital human interacts with the teaching object, enabling the digital human to achieve more natural and emotionally resonant interaction, providing a personalized and immersive art education experience, and significantly improving the actual role and teaching quality of the digital human in art teaching.
[0027] Example 2 Please see Figure 2 , Figure 2 This is a flowchart illustrating a digital human-computer interaction method for art education courses according to another embodiment of the present invention, the method comprising: S201: Obtain the speech and facial expression information of the teaching object in the art education course, and perform micro-feature analysis of emotional expression based on the speech and facial expression information to obtain micro-feature information of emotional expression; In the specific implementation of this invention, the step of performing micro-feature analysis of emotional expression based on the voice performance information and facial performance information to obtain micro-feature information of emotional expression includes: performing micro-voice change feature analysis on the voice performance information to obtain micro-voice change feature information; performing motion change analysis on the facial performance information to obtain motion change features; and performing micro-feature analysis of emotional expression based on the micro-voice change feature information and motion change features to obtain micro-feature information of emotional expression.
[0028] Specifically, this involves acquiring vocal and facial expression information of students in arts education courses. Vocal expression information can include acoustic characteristics such as pitch, volume, speech rate, and timbre when students are speaking, singing, or reciting. Facial expression information can include visual characteristics such as facial expressions, eye contact, and head posture when students are performing or expressing emotions. This information can be collected using various sensor devices, such as microphones and high-definition cameras. For example, in vocal lessons, microphones can capture students' singing voices in real time, while cameras can record changes in students' facial expressions while singing.
[0029] The speech performance information is analyzed for subtle changes to obtain characteristic information. This analysis involves a deep dive into the subtle, often imperceptible, changes in the speech performance of the learners, such as minute fluctuations in parameters like pitch, volume, speech rate, and timbre, as well as subtle adjustments to linguistic features like pauses and stress. These subtle changes often reflect deeper emotional states and psychological activities of the learners in arts education courses. Specifically, these subtle speech changes can be quantified using acoustic feature extraction algorithms, such as Mel-frequency cepstral coefficients, fundamental frequency change rate, and energy change rate, thereby obtaining the characteristic information. The aim is to capture subtle clues in the learners' emotional expression, providing more refined data support for subsequent emotion analysis.
[0030] The facial expression information is analyzed for movement changes to obtain movement change features. Movement change analysis refers to capturing and quantifying subtle facial muscle movements, instantaneous changes in expression, and dynamic eye movements. Examples include a slight raising of the eyebrows, a subtle twitch of the corner of the mouth, rapid eye movement, or a fixed gaze. These subtle changes in facial movements are also important components of emotional expression. Specifically, computer vision technologies, such as facial landmark detection and expression unit recognition, can be used to analyze the movement patterns and intensity of facial muscles to obtain movement change features. The aim is to capture the subtle nuances of nonverbal emotional expression from the visual modality. Based on the subtle changes in speech and movement features, micro-feature analysis of emotional expression is performed to obtain micro-feature information. A fusion analysis based on these two modalities' micro-features is then conducted to more comprehensively and accurately understand the micro-features of the student's emotional expression. This fusion analysis can employ multimodal data fusion techniques, such as feature-level fusion or decision-level fusion, to integrate micro-features from different modalities, thereby obtaining richer and more accurate micro-feature information of emotional expression. As a result, the obtained micro-feature information of emotional expression is richer and more authentic, providing a more solid data foundation for subsequent analysis of the teaching subjects' intentions and emotional tendencies. This helps the digital human to more accurately understand the internal state of the teaching subjects and adjust interaction strategies accordingly, thereby enhancing the personalization and effectiveness of art education courses.
[0031] S202: Based on the aforementioned voice performance information and facial performance information, perform spontaneous deviation analysis on the learning process of the teaching object to obtain spontaneous deviation information; In a specific implementation of this invention, the step of performing spontaneous deviation analysis on the learning process of the teaching object based on the speech performance information and facial performance information to obtain spontaneous deviation information includes: extracting detailed features from the speech performance information to obtain speech detail feature information; extracting detailed features from the facial performance information to obtain facial detail feature information; analyzing a first deviation degree between the speech detail feature information and standard speech details, and analyzing a second deviation degree between the facial detail feature information and standard facial details; and performing spontaneous deviation analysis on the learning process of the teaching object based on the first deviation degree and the second deviation degree to obtain spontaneous deviation information.
[0032] Specifically, extracting detailed features from the aforementioned speech performance information to obtain speech detail feature information refers to extracting subtle features related to learning status, emotion, or cognitive load from the speech data of students in art education courses using acoustic analysis techniques. These features include speech rate, tone, volume, changes in intonation, pause duration, stress distribution, and pronunciation clarity. These speech detail features can serve as indicators reflecting the students' internal learning status and emotional fluctuations.
[0033] Extracting detailed facial features from the aforementioned facial expression information involves using image processing and computer vision technologies to identify and quantify the intensity and duration of facial movement units related to emotion, attention, and cognitive state, as well as subtle changes in facial muscles, such as eyebrow raising or lowering, slight eye movements, mouth corner curvature, facial muscle tension, and blinking frequency, from facial video or image data of the learner. This detailed facial feature information can reveal the learner's emotional responses, level of focus, and comprehension of the learning content during the learning process.
[0034] The first deviation of analyzing speech detail features from standard speech details refers to comparing the speech detail features currently extracted by the learner with a pre-established standard speech detail model representing a normal or ideal learning state, and quantifying the degree of difference. For example, standard speech details might be characteristics such as a steady speaking speed, moderate intonation, and no obvious pauses or stuttering in a focused and active learning state. The higher the first deviation, the greater the deviation of the learner's speech performance from a normal learning state. The second deviation of analyzing facial detail features from standard facial details refers to comparing the facial detail features currently extracted by the learner with a pre-established standard facial detail model representing a normal or ideal learning state, and quantifying the degree of difference. For example, standard facial details might be characteristics such as focused eye contact, relaxed facial expression, and no frequent blinking or frowning in a focused and active learning state. The higher the second deviation, the greater the deviation of the learner's facial expression from a normal learning state.
[0035] Analyzing spontaneous deviations in the learning process based on the first and second deviations to obtain spontaneous deviation information involves comprehensively considering the degree of deviation from both vocal and facial modalities. This is achieved through a fusion analysis model (such as a multimodal neural network or decision tree) to determine whether the learner exhibits spontaneous deviation. Spontaneous deviation information can include the type of deviation (e.g., inattention, confusion, boredom, anxiety), the degree of deviation (e.g., slight, moderate, severe deviation), and the timing and duration of the deviation. This allows for earlier and more accurate identification of potential spontaneous deviation behaviors such as inattention, emotional fluctuations, or comprehension difficulties. This detailed analysis helps digital humans adjust teaching strategies more promptly and provide personalized interventions, such as adjusting the pace of instruction, providing additional explanations, or offering emotional support, thereby effectively improving the teaching outcomes and learning experience of arts education courses.
[0036] S203: Perform cross-modal performance consistency analysis based on the aforementioned voice performance information and facial performance information to obtain cross-modal performance consistency information; In a specific implementation of this invention, the step of performing cross-modal performance consistency analysis based on the speech performance information and facial performance information to obtain cross-modal performance consistency information includes: extracting feature vectors from the speech performance information to obtain speech feature vectors; extracting feature vectors from the facial performance information to obtain facial feature vectors; calculating a performance consistency index based on the speech feature vectors and facial feature vectors; and performing cross-modal performance consistency analysis based on the performance consistency index to obtain cross-modal performance consistency information.
[0037] Specifically, feature vectors are extracted from the speech performance information to obtain speech feature vectors, aiming to quantify the acoustic features of the speech of the teaching subject. Speech feature vectors may include, but are not limited to, parameters such as fundamental frequency, formant frequency, energy, speech rate, timbre, and pitch variation rate. These features can be obtained using digital signal processing techniques, such as Fourier transform and Mel-frequency cepstral coefficient extraction. Feature vectors are also extracted from the facial performance information to obtain facial feature vectors, aiming to capture the dynamic and static features of the teaching subject's face. Facial feature vectors may include the intensity and duration of facial action units, the movement trajectory of facial key points, eye movement patterns (such as blink frequency and pupil size changes), and head posture changes. These features can be extracted using computer vision techniques, such as facial key point detection, expression recognition algorithms, and eye-tracking technology.
[0038] The performance consistency index is calculated based on the aforementioned speech and facial feature vectors. This index is a quantitative measure of the synchronicity, coordination, and consistency of a learner's expression across both speech and facial modalities. For example, it can be obtained by calculating the cross-correlation coefficient between speech and facial feature vector sequences, dynamic time-warped distance, or by using a consistency score output by a multimodal fusion model (such as a deep learning model). This index reflects whether a learner's speech and facial expressions mutually support and corroborate each other when expressing a particular emotion or intention. Cross-modal performance consistency analysis is then performed based on this index to obtain cross-modal performance consistency information. This analysis aims to assess whether inconsistencies exist between a learner's speech and facial expressions in an arts education course. For example, when a learner verbally expresses a positive emotion but their facial expression appears stiff or negative, the performance consistency index will be low, indicating cross-modal performance inconsistency. This inconsistency information is crucial for understanding a learner's true emotional state, cognitive load, or potential psychological activities. The explicit feature vector extraction and performance consistency index calculation make the analysis process more transparent and quantifiable. This not only improves the accuracy and reliability of cross-modal performance consistency information, but also provides richer and more precise input for subsequent intent analysis and sentiment analysis, thereby helping digital humans to more accurately understand the real state of the learners and develop more targeted interaction strategies and teaching adjustment strategies.
[0039] Furthermore, the step of extracting feature vectors from the speech performance information to obtain speech feature vectors and extracting feature vectors from the facial performance information to obtain facial feature vectors includes: extracting vocal cord tremor frequency and amplitude from the speech performance information based on a fundamental frequency extraction algorithm and a time-domain sliding window, and generating speech feature vectors based on the vocal cord tremor frequency and amplitude; performing facial muscle stiffness analysis on the facial performance information to obtain facial muscle stiffness information; and extracting blink frequency and pupil diameter changes from the facial performance information based on Hough circle transform, and generating facial feature vectors based on the facial muscle stiffness information, blink frequency, and pupil diameter changes.
[0040] Specifically, based on a fundamental frequency extraction algorithm and a time-domain sliding window, the vocal cord micro-vibration frequency and amplitude are extracted from the speech performance information. The fundamental frequency extraction algorithm is used to identify the fundamental frequency in the speech signal, that is, the basic frequency of vocal cord vibration, which is closely related to the speaker's emotional state and physiological tension. The time-domain sliding window allows for local analysis of the speech signal, thereby capturing subtle changes such as vocal cord micro-vibration frequency and amplitude. Vocal cord micro-vibration frequency refers to the small, rapid frequency fluctuations that occur in the vocal cords during phonation, while vocal cord micro-vibration amplitude refers to the intensity of these fluctuations. These micro-vibration features are considered physiological indicators reflecting an individual's emotional state, such as tension, excitement, or fatigue. A speech feature vector is generated based on the vocal cord micro-vibration frequency and amplitude. By combining these extracted vocal cord micro-vibration frequencies and amplitudes, a speech feature vector that can represent the micro-features of speech emotion can be generated.
[0041] Facial muscle stiffness analysis was performed on the facial expression information to obtain facial muscle stiffness information, which can be understood as the degree of tension or relaxation of facial muscles. This information is closely related to an individual's emotional state, particularly stress, anxiety, or relaxation. For example, when an individual feels tense, their facial muscles may exhibit higher stiffness.
[0042] Based on the Hough circle transform, blink frequency and pupil diameter changes were extracted from the facial expression information. The Hough circle transform is an image processing technique commonly used to identify circular structures in images; here, it was used to accurately detect the eye region and then extract blink frequency and pupil diameter changes. Blink frequency refers to the number of blinks per unit time, and its changes may reflect an individual's fatigue level, concentration, or emotional fluctuations. Pupil diameter changes are related to the activity of the autonomic nervous system and can reflect an individual's emotional arousal level and cognitive load. A facial feature vector was generated based on the facial muscle stiffness information, blink frequency, and pupil diameter changes. By combining these micro-facial features, a facial feature vector representing the micro-features of facial emotions can be generated. These specific and detailed feature extraction methods provide high-quality, high-dimensional input data for subsequent cross-modal performance consistency analysis, enabling a deeper and more comprehensive understanding of the emotional expression of the learners.
[0043] S204: Analyze the emotional lag information of the teaching object to the preset stimuli based on the voice performance information and facial performance information; In the specific implementation of this invention, the emotional lag information of the learning object to the preset stimuli is analyzed based on the aforementioned voice performance information and facial performance information. Emotional lag analysis aims to assess the time delay in the appearance of the learning object's emotional response after receiving external stimuli (such as questions, demonstrations, or feedback from a digital human). For example, when a digital human poses a challenging question, the learning object may need some time to express feelings of thinking, confusion, or excitement. Emotional lag can reflect the learning object's reaction speed, comprehension ability, or emotional regulation ability.
[0044] S205: Based on the aforementioned micro-feature information of emotional expression, spontaneous deviation information, cross-modal performance consistency information, and emotional lag information, perform intention analysis on the teaching subjects to obtain target intention information; In the specific implementation of this invention, the step of analyzing the teaching object's intent based on the micro-feature information of emotional expression, spontaneous deviation information, cross-modal performance consistency information, and emotional lag information to obtain target intent information includes: performing pattern logic judgment based on the cross-modal performance consistency information and emotional lag information to obtain pattern logic judgment results; performing emotional resonance depth analysis based on the micro-feature information of emotional expression to obtain target emotional resonance depth; performing internal state assessment based on the spontaneous deviation information and target emotional resonance depth to obtain internal state assessment results; and performing teaching object intent analysis based on the pattern logic judgment results and internal state assessment results to obtain target intent information.
[0045] Specifically, pattern logic judgment is performed based on the cross-modal performance consistency information and emotional lag information to obtain pattern logic judgment results. Pattern logic judgment refers to inferring the internal thinking patterns and behavioral logic of the learner by analyzing the consistency of the learner's performance across different modalities (e.g., voice and facial expressions) and the lag in emotional responses to external stimuli. For example, if a learner displays positive emotions in their voice but appears stiff or sluggish in their facial expressions, this may indicate inconsistency in their emotional expression, requiring further investigation into their true intentions. Emotional lag information can reflect the learner's reaction speed and depth to stimuli; for example, a slow response to a work of art may indicate insufficient understanding or resonance.
[0046] Based on the aforementioned micro-features of emotional expression, an emotional resonance depth analysis is conducted to obtain the target emotional resonance depth. This analysis can be understood as assessing the degree of resonance the learner experiences with art education course content or digital human interaction by deeply analyzing the micro-features of their emotional expression. These micro-features can include subtle changes in speech (e.g., speech rate, pitch fluctuations) and subtle changes in facial movements (e.g., slight raising of eyebrows, subtle twitching of the corners of the mouth). These subtle features often more realistically reflect the learner's emotional state and resonance depth. For example, when a learner is appreciating a painting, slight sighs or relaxed facial muscles in their speech may indicate a deeper emotional resonance.
[0047] An internal state assessment is conducted based on the spontaneous deviation information and the depth of target emotional resonance. Specifically, this assessment comprehensively judges the learner's internal learning state and psychological activities by integrating spontaneous deviation information and the depth of target emotional resonance. Spontaneous deviation information reflects whether the learner exhibits unexpected behaviors such as inattention or shifting interest during the learning process, while the depth of emotional resonance reveals their level of engagement with the learning content. By combining these two aspects, a more comprehensive assessment can be made of whether the learner is actively engaged, confused, or inattentive. For example, if a learner exhibits high spontaneous deviation (e.g., frequently switching topics or wandering eyes) but simultaneously has a low depth of emotional resonance, it may indicate a lack of interest in or difficulty understanding the current learning content. Based on the pattern logic judgment results and the internal state assessment results, the learner's intention is analyzed to obtain target intention information. Combining this with the spontaneous deviation information and the depth of emotional resonance in the internal state assessment allows for a comprehensive judgment of the learner's learning state and psychological activities. It is precisely because of this hierarchical and multi-dimensional integrated analysis method that the final intent analysis results are more comprehensive, in-depth and accurate, avoiding the one-sidedness that may be brought about by single-dimensional analysis.
[0048] S206: Obtain the historical performance pattern sequence of the teaching objects in the art education course, and analyze the evolution trend of the performance pattern based on the historical performance pattern sequence to obtain information on the evolution trend of the historical performance pattern. In the specific implementation of this invention, a historical performance pattern sequence of the teaching subjects in art education courses is obtained. This historical performance pattern sequence refers to the system continuously recording the vocal and facial expressions of the teaching subjects at different times and in different course segments, integrating and abstracting them into a series of performance patterns with a temporal order. These performance patterns can be a quantitative description of the teaching subjects' comprehensive state, including emotional expression, learning engagement, and concentration, within a specific time period. The purpose is to provide a long-term, dynamic reference benchmark for subsequent emotional tendency analysis. Based on the historical performance pattern sequence, an evolution trend analysis of the performance patterns is performed to obtain historical performance pattern evolution trend information. This can be understood as processing the acquired historical performance pattern sequence using time series analysis, pattern recognition, or machine learning to reveal the patterns and trends of the teaching subjects' performance patterns changing over time in art education courses. For example, the frequency of emotional fluctuations, the cycle of concentration changes, and the evolution path of learning enthusiasm can be analyzed. The purpose is to establish a dynamic model for predicting or evaluating the future performance of teaching subjects, thereby better understanding the deeper meaning of their current performance.
[0049] S207: Generate the current performance pattern information of the teaching object based on the voice performance information and facial performance information, and analyze the degree of deviation between the current performance pattern information and the historical performance pattern evolution trend information; In the specific implementation of this invention, the analysis of the deviation between the current performance pattern information and the historical performance pattern evolution trend information includes: performing time series decomposition on the historical performance pattern sequence to obtain the time series decomposed historical performance pattern sequence; constructing a historical benchmark model based on the time series decomposed historical performance pattern sequence; and analyzing the deviation between the current performance pattern information and the historical performance pattern evolution trend information based on the historical benchmark model.
[0050] Specifically, generating the current performance pattern information of the learning subject based on the aforementioned voice and facial expression information refers to constructing the learning subject's current comprehensive performance state during real-time interaction by utilizing voice and facial expression information collected at the current moment or recently, through a processing method similar to that used to generate historical performance pattern sequences. For example, this may include current emotional intensity, focus level, and engagement. Its purpose is to provide an immediate data point for comparison with historical trends.
[0051] Time-series decomposition of the historical performance pattern sequence yields a time-series decomposed historical performance pattern sequence. This involves breaking down the data on the historical performance patterns of students in art education courses—such as changes in tone of voice, facial expressions, and body movements—over time into different components. These components typically include trend components, seasonal components, and residual components. The trend component reflects the long-term direction of change in the performance pattern, the seasonal component reveals periodic fluctuation patterns, and the residual component represents random, unexplained fluctuations. The aim is to gain a deeper understanding of the internal structure and evolutionary laws of historical performance patterns, remove noise and periodic interference, and thus obtain purer and more representative evolutionary trend information.
[0052] Constructing a historical benchmark model based on historical performance pattern sequences after time series decomposition refers to using decomposed historical performance pattern data to establish a model that can describe or predict the typical performance patterns of learners through statistical methods or machine learning algorithms (such as autoregressive integral moving average models, exponential smoothing models, recurrent neural networks, etc.). This historical benchmark model aims to provide a stable, reliable, and personalized reference system for measuring the degree of deviation of learners' current performance patterns.
[0053] Analyzing the deviation between current performance patterns and historical performance pattern evolution trends based on the historical benchmark model involves comparing the performance patterns of students in the current art education curriculum with the typical performance patterns predicted or described by the aforementioned historical benchmark model, and quantifying the differences between the two. For example, Euclidean distance, Mahalanobis distance, and other indicators can be calculated to represent the degree of deviation. The aim is to accurately identify the differences between students' current performance and their historical norms, providing accurate and detailed input for subsequent sentiment analysis. By comparing current performance pattern information with this refined historical benchmark model, subtle deviations in students' performance patterns can be detected more accurately and sensitively, thus providing more reliable data support for sentiment analysis.
[0054] For example, suppose a student's historical performance pattern sequence in an art education course includes data on their tone of voice, facial expression activity, and body movement amplitude at different time periods. First, this historical performance pattern sequence is decomposed into a time series, for example, using the STL (Seasonal-Trend decomposition using Loess) method, into trend components, seasonal components, and residual components. The trend component can reveal the student's long-term progress or regression in artistic performance; the seasonal component may reflect periodic performance characteristics at different course stages or under different emotional states; and the residual component represents random, unpredictable fluctuations.
[0055] Next, based on the decomposed trend and seasonal components, a historical baseline model can be constructed, such as using an ARIMA model or a long short-term memory network, to predict the normal performance pattern of the learner in a given context. When the learner performs the current artistic expression, the system generates information about their current performance pattern. Subsequently, this current performance pattern information is compared with the normal performance pattern predicted by the historical baseline model, and the degree of deviation between the two is calculated. For example, if the current performance pattern is significantly lower than the predicted value in terms of vocal intonation and facial expression activity, a large negative deviation can be identified, which may indicate that the learner is currently in a low mood or lacks interest.
[0056] S208: Based on the aforementioned micro-feature information of emotional expression, spontaneous deviation information, cross-modal performance consistency information, and emotional lag information, the degree of deviation is used to analyze the emotional tendency of the teaching object to obtain the target emotional tendency; In the specific implementation of this invention, the emotional tendency analysis of the teaching object is conducted based on the aforementioned micro-feature information of emotional expression, spontaneous deviation information, cross-modal performance consistency information, and emotional lag information, using the degree of deviation to obtain the target emotional tendency. This means that the deviation degree obtained from the above analysis is used as one of the important input features, combined with the emotional expression micro-feature information, spontaneous deviation information, cross-modal performance consistency information, and emotional lag information obtained from the basic scheme. Through multimodal fusion, machine learning models, and other methods, the current target emotional tendency of the teaching object is comprehensively judged. The purpose is to provide a comprehensive, accurate, and historically contextualized emotional assessment result. The introduction of this deviation degree makes the judgment of emotional tendency more personalized and dynamic, effectively distinguishing between occasional emotional fluctuations and continuous changes in emotional tendency, thereby providing digital humans with more refined and targeted interaction strategies and teaching adjustment suggestions, ultimately improving the teaching effectiveness of art education courses and the learning experience of the teaching object.
[0057] For example, suppose a student is learning to play the piano. The system continuously records the student's vocal (e.g., humming, sighing, talking to themselves) and facial (e.g., expressions, eye contact, head posture) performance over the past few weeks, generating a historical performance pattern sequence. For instance, this sequence might show that the student typically exhibits high focus and enthusiasm in the early stages of mastering a new piece, but when encountering technical difficulties, their facial expressions become slightly tense, and their voice occasionally reveals frustration. By analyzing the evolution trend of this historical sequence, the system can identify the student's typical emotional patterns along the learning curve. When the student's current vocal and facial performance information is collected during a particular practice session, the system generates their current performance pattern information. For example, if the current performance pattern shows that the student's focus is significantly lower than the historical average when facing a piece they have already mastered, and their facial expressions show signs of impatience, the system will analyze that there is a significant negative deviation between the current performance pattern information and the historical performance pattern evolution trend information. Based on this degree of deviation, and combined with data obtained from micro-features of emotional expression, spontaneous deviation information, cross-modal performance consistency information, and emotional lag information, the system will comprehensively judge whether the student currently has a target emotional tendency of learning burnout or loss of interest in the current practice content.
[0058] S209: Determine the digital human's expressive information based on the target intention information and target emotional tendency, determine the digital human's teaching adjustment strategy based on the target intention information and target emotional tendency, and drive the digital human to interact with the teaching object based on the expressive information and teaching adjustment strategy.
[0059] In the specific implementation of this invention, determining the digital human's expressive information based on the target intention information and target emotional tendency includes: determining several emotional levels of the teaching object based on the target intention information and target emotional tendency; determining the intensity weight of each emotional level, and selecting several corresponding target non-verbal expression elements based on the intensity weight using a non-verbal expression element library; determining the digital human's animated action expressive information based on the several target non-verbal expression elements; determining language expression elements based on the target intention information and target emotional tendency, and determining the digital human's expressive information based on the animated action expressive information and language expression elements.
[0060] Specifically, determining several emotional levels of the learner based on the target intention information and target emotional tendency refers to mapping the learner's internal state to a series of predefined emotional categories or dimensions based on the learner's target intention information and target emotional tendency. These emotional levels can include, but are not limited to, broad categories such as positive, negative, and neutral, and can also be subdivided into more specific emotional states such as joy, sadness, anger, curiosity, confusion, and focus. In this way, the learner's emotions can be identified and classified in a multi-dimensional and refined manner, providing a foundation for the subsequent generation of digital human expression.
[0061] Determining the intensity weight of each emotional level and selecting corresponding target nonverbal expression elements from a nonverbal expression element library based on these intensity weights involves assessing the salience or influence of each emotional level after identifying the learner's emotional level and assigning it a corresponding intensity weight. For example, if the learner exhibits high curiosity and slight confusion, the intensity weight of curiosity will be higher than that of confusion. Subsequently, based on these emotional levels and their intensity weights, the nonverbal elements that best match the current emotional state are selected from a pre-constructed nonverbal expression element library. This nonverbal expression element library can include various facial expressions (such as smiling, frowning, and surprise), body movements (such as nodding, shaking the head, and shrugging), and eye contact (such as direct eye contact and avoidance), to ensure that the digital human can express itself through rich nonverbal forms.
[0062] Determining the animated action expression information of a digital human based on several target nonverbal expression elements refers to converting the selected target nonverbal expression elements into an executable animation sequence for the digital human. This includes adjusting the digital human model's posture, facial expressions, gestures, etc., so that it can vividly and naturally display the selected nonverbal expressions. For example, if the selected nonverbal elements are "smiling" and "nodding," an animation sequence of the digital human smiling accompanied by a nodding action will be generated.
[0063] Determining language expression elements based on the target intent information and target emotional tendency refers to selecting or generating corresponding language content from a language expression database based on the target intent information and target emotional tendency of the learner. These language expression elements can be preset phrases and sentences, or text generated in real time through natural language generation technology. The aim is to ensure that the digital human's language feedback accurately responds to the learner's intent and matches their emotional tendency. For example, when a learner shows confusion, the digital human might generate a question like, "Do you have any questions about this concept?" Determining the digital human's expressive information based on the animated action expression information and language expression elements refers to integrating the generated animated action expression information with the language expression elements to form the digital human's complete, multimodal expressive information. This integration ensures that the digital human's verbal and nonverbal expressions are synchronized in time and consistent in content, thus presenting a natural and coherent interactive experience. By integrating animated action expression information with language expression elements, cross-modal consistency and coordination of the digital human's expression are achieved, making the digital human's interactive behavior more natural, fluent, and expressive.
[0064] Based on the target intent information and target emotional tendency, the teaching adjustment strategy of the digital human is determined, and the digital human interacts with the learner based on the expressed information and the teaching adjustment strategy. For example, when the learner shows confusion, the digital human can use a gentler tone and more patient expressions to explain. The teaching adjustment strategy can include adjusting the difficulty, pace, and interaction method of the teaching content. For example, when the learner shows high engagement, the digital human can speed up the teaching pace or provide more challenging tasks; when the learner shows frustration, the digital human can provide encouraging feedback or switch to easier exercises. Finally, based on these expressed information and teaching adjustment strategies, the digital human interacts with the learner, thereby achieving a personalized and intelligent art education experience. This achieves accurate capture of the learner's deep emotions and learning state. This deep perception and intelligent response mechanism greatly enhances the effectiveness of the digital human in emotional resonance and personalized guidance in art education, providing a more advanced technological path for the popularization and personalization of art education.
[0065] In this embodiment of the invention, vocal and facial expression information of the learners in art education courses is acquired. Based on this information, micro-feature analysis of emotional expression is performed to more accurately capture subtle emotional changes in the learners. Spontaneous deviation analysis of the learners during the learning process, based on vocal and facial expression information, allows for a more objective assessment of their focus and engagement. Cross-modal consistency analysis, based on vocal and facial expression information, provides a more comprehensive understanding of their inner emotional state, avoiding the limitations of single-modal analysis. Analysis of the learners' emotional lag in response to pre-set stimuli is conducted based on vocal and facial expression information. Intention analysis of the learners is performed based on micro-feature information of emotional expression, spontaneous deviation information, cross-modal consistency information, and emotional lag information. Furthermore, analysis of the learners' emotional tendencies is conducted using the degree of deviation from performance patterns, based on these factors, enabling a comprehensive and in-depth perception of the learners' real-time state and inner emotions. This accurately identifies the learners' learning intentions and emotional states. Based on the target intention information and target emotional tendency, the expressive information of the digital human and the teaching adjustment strategy are determined. Based on the expressive information and teaching adjustment strategy, the digital human interacts with the teaching object, enabling the digital human to achieve more natural and emotionally resonant interaction, providing a personalized and immersive art education experience, and significantly improving the actual role and teaching quality of the digital human in art teaching.
[0066] Example 3 Please see Figure 3 , Figure 3 This is a schematic diagram illustrating the structural composition of a digital human-computer interaction system for art education courses according to an embodiment of the present invention. The system includes: Feature analysis module 31: used to acquire the speech and facial expression information of the teaching object in the art education course, and to perform micro-feature analysis of emotional expression based on the speech and facial expression information to obtain micro-feature information of emotional expression; Learning deviation analysis module 32: used to perform spontaneous deviation analysis of the learning object in the learning process based on the voice performance information and facial performance information, and obtain spontaneous deviation information; Performance consistency analysis module 33: used to perform cross-modal performance consistency analysis based on the voice performance information and facial performance information to obtain cross-modal performance consistency information; Emotional lag analysis module 34: used to analyze the emotional lag information of the teaching object to the preset stimulus based on the voice performance information and facial performance information; Intent-sentiment analysis module 35: is used to analyze the intent of the teaching object based on the micro-feature information of emotional expression, spontaneous deviation information, cross-modal performance consistency information and emotional lag information to obtain target intent information, and to analyze the emotional tendency of the teaching object based on the degree of deviation of performance mode based on the micro-feature information of emotional expression, spontaneous deviation information, cross-modal performance consistency information and emotional lag information to obtain target emotional tendency. Interaction module 36: used to determine the digital human's expressive information based on the target intention information and target emotional tendency, determine the digital human's teaching adjustment strategy based on the target intention information and target emotional tendency, and drive the digital human to interact with the teaching object based on the expressive information and teaching adjustment strategy.
[0067] In the specific implementation of this invention, the specific implementation methods of the system items can be referred to the implementation methods of the above-mentioned method items, and will not be repeated here.
[0068] In this embodiment of the invention, vocal and facial expression information of the learners in art education courses is acquired. Based on this information, micro-feature analysis of emotional expression is performed to more accurately capture subtle emotional changes in the learners. Spontaneous deviation analysis of the learners during the learning process, based on vocal and facial expression information, allows for a more objective assessment of their focus and engagement. Cross-modal consistency analysis, based on vocal and facial expression information, provides a more comprehensive understanding of their inner emotional state, avoiding the limitations of single-modal analysis. Analysis of the learners' emotional lag in response to pre-set stimuli is conducted based on vocal and facial expression information. Intention analysis of the learners is performed based on micro-feature information of emotional expression, spontaneous deviation information, cross-modal consistency information, and emotional lag information. Furthermore, analysis of the learners' emotional tendencies is conducted using the degree of deviation from performance patterns, based on these factors, enabling a comprehensive and in-depth perception of the learners' real-time state and inner emotions. This accurately identifies the learners' learning intentions and emotional states. Based on the target intention information and target emotional tendency, the expressive information of the digital human and the teaching adjustment strategy are determined. Based on the expressive information and teaching adjustment strategy, the digital human interacts with the teaching object, enabling the digital human to achieve more natural and emotionally resonant interaction, providing a personalized and immersive art education experience, and significantly improving the actual role and teaching quality of the digital human in art teaching.
[0069] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.
[0070] Furthermore, the above provides a detailed description of a digital human-computer interaction method and system for art education courses provided by the embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A digital human-computer interaction method for art education courses, characterized in that, The method includes: The system acquires the vocal and facial expression information of the teaching subjects in art education courses, and performs micro-feature analysis of emotional expression based on the vocal and facial expression information to obtain micro-feature information of emotional expression. Based on the aforementioned voice and facial expression information, spontaneous deviation analysis of the learning subjects during the learning process is performed to obtain spontaneous deviation information; Cross-modal performance consistency analysis is performed based on the aforementioned voice performance information and facial performance information to obtain cross-modal performance consistency information; Based on the aforementioned voice and facial expression information, analyze the emotional lag of the teaching subjects in response to the preset stimuli; Based on the aforementioned micro-feature information of emotional expression, spontaneous deviation information, cross-modal performance consistency information, and emotional lag information, the intention of the teaching object is analyzed to obtain the target intention information. Furthermore, based on the aforementioned micro-feature information of emotional expression, spontaneous deviation information, cross-modal performance consistency information, and emotional lag information, the degree of deviation from the performance mode is used to analyze the emotional tendency of the teaching object to obtain the target emotional tendency. Based on the target intent information and target emotional tendency, determine the digital human's expressive information, determine the digital human's teaching adjustment strategy based on the target intent information and target emotional tendency, and drive the digital human to interact with the teaching object based on the expressive information and teaching adjustment strategy.
2. The digital human-computer interaction method for art education courses according to claim 1, characterized in that, The step of performing micro-feature analysis of emotional expression based on the voice performance information and facial performance information to obtain micro-feature information of emotional expression includes: The speech performance information is analyzed for minute speech change features to obtain minute speech change feature information; The facial expression information is analyzed for movement changes to obtain movement change features. Based on the subtle changes in speech and movement changes, the micro-features of emotional expression are analyzed to obtain micro-features of emotional expression.
3. The digital human-computer interaction method for art education courses according to claim 1, characterized in that, The analysis of spontaneous deviations of the learning subjects during the learning process based on the voice and facial expression information, to obtain spontaneous deviation information, includes: Detail feature extraction is performed on the speech performance information to obtain speech detail feature information; Detailed feature extraction is performed on the facial expression information to obtain facial detail feature information; Analyze the first deviation of speech detail feature information from standard speech detail, and analyze the second deviation of facial detail feature information from standard facial detail; Based on the first and second deviations, a spontaneous deviation analysis of the learning subjects is conducted to obtain spontaneous deviation information.
4. The digital human-computer interaction method for art education courses according to claim 1, characterized in that, The cross-modal performance consistency analysis based on the voice performance information and facial performance information to obtain cross-modal performance consistency information includes: Feature vectors are extracted from the speech performance information to obtain speech feature vectors, and feature vectors are extracted from the facial performance information to obtain facial feature vectors. Based on the speech feature vector and facial feature vector, a performance consistency index is calculated, and cross-modal performance consistency analysis is performed based on the performance consistency index to obtain cross-modal performance consistency information.
5. The digital human-computer interaction method for art education courses according to claim 4, characterized in that, The step of extracting feature vectors from the speech performance information to obtain speech feature vectors, and extracting feature vectors from the facial performance information to obtain facial feature vectors, includes: Based on the fundamental frequency extraction algorithm and the time-domain sliding window, the vocal cord tremor frequency and amplitude are extracted from the speech performance information, and a speech feature vector is generated based on the vocal cord tremor frequency and amplitude. Facial muscle stiffness analysis was performed on the facial expression information to obtain facial muscle stiffness information. Based on the Hough circle transform, blink frequency and pupil diameter changes are extracted from the facial expression information, and facial feature vectors are generated based on the facial muscle stiffness information, blink frequency, and pupil diameter changes.
6. The digital human-computer interaction method for art education courses according to claim 1, characterized in that, The intention analysis of the learning subjects based on the micro-features of emotional expression, spontaneous deviation information, cross-modal consistency information, and emotional lag information, to obtain target intention information, includes: Based on the cross-modal performance consistency information and emotional lag information, pattern logic judgment is performed to obtain the pattern logic judgment result; Based on the micro-feature information of the emotional expression, an emotional resonance depth analysis is performed to obtain the target emotional resonance depth; Based on the spontaneous deviation information and the depth of target emotional resonance, an internal state assessment is performed to obtain the internal state assessment results. Based on the pattern logic judgment results and the internal state assessment results, the intention analysis of the teaching object is performed to obtain the target intention information.
7. The digital human-computer interaction method for art education courses according to claim 1, characterized in that, The method of analyzing the emotional tendencies of teaching subjects based on the micro-features of emotional expression, spontaneous deviation information, cross-modal performance consistency information, and emotional lag information, and obtaining the target emotional tendency, includes: Obtain the historical performance pattern sequence of teaching subjects in art education courses, and analyze the evolution trend of performance patterns based on the historical performance pattern sequence to obtain information on the evolution trend of historical performance patterns. Based on the voice performance information and facial performance information, the current performance pattern information of the teaching object is generated, and the degree of deviation between the current performance pattern information and the historical performance pattern evolution trend information is analyzed. Based on the aforementioned micro-features of emotional expression, spontaneous deviation information, cross-modal consistency information, and emotional lag information, the degree of deviation is used to analyze the emotional tendencies of the teaching subjects and obtain the target emotional tendencies.
8. The digital human-computer interaction method for art education courses according to claim 7, characterized in that, The analysis of the deviation between current performance pattern information and historical performance pattern evolution trend information includes: The historical performance pattern sequence is decomposed into a time series to obtain the decomposed historical performance pattern sequence. A historical benchmark model is constructed based on the historical performance pattern sequence after time series decomposition. The deviation between the current performance pattern information and the historical performance pattern evolution trend information is analyzed based on the historical benchmark model.
9. The digital human-computer interaction method for art education courses according to claim 1, characterized in that, The process of determining the digital human's expressive information based on the target intent information and target emotional tendency includes: Based on the target intention information and target emotional tendency, several emotional levels of the teaching object are determined; Determine the intensity weight of each emotional level, and select several target non-verbal expression elements based on the intensity weight using a non-verbal expression element library; Determine the animated action expression information of the digital human based on several target non-verbal expression elements; Based on the target intent information and target emotional tendency, language expression elements are determined, and based on the animation action expression information and language expression elements, the digital human's expression information is determined.
10. A digital human-computer interaction system for art education courses, characterized in that, The system includes: Feature analysis module: used to acquire the speech and facial expression information of the teaching subjects in the art education course, and to perform micro-feature analysis of emotional expression based on the speech and facial expression information to obtain micro-feature information of emotional expression; Learning Deviation Analysis Module: Used to analyze the spontaneous deviation of the learning object during the learning process based on the voice performance information and facial performance information, and obtain spontaneous deviation information; Performance consistency analysis module: used to perform cross-modal performance consistency analysis based on the voice performance information and facial performance information to obtain cross-modal performance consistency information; Emotional lag analysis module: used to analyze the emotional lag information of the teaching object to the preset stimuli based on the voice performance information and facial performance information; The intention and sentiment analysis module is used to analyze the intention of the teaching object based on the micro-feature information of the emotional expression, spontaneous deviation information, cross-modal performance consistency information and emotional lag information to obtain the target intention information. Based on the micro-feature information of the emotional expression, spontaneous deviation information, cross-modal performance consistency information and emotional lag information, the module uses the degree of deviation of the performance mode to analyze the sentiment tendency of the teaching object to obtain the target sentiment tendency. Interaction module: used to determine the digital human's expressive information based on the target intention information and target emotional tendency, determine the digital human's teaching adjustment strategy based on the target intention information and target emotional tendency, and drive the digital human to interact with the teaching object based on the expressive information and teaching adjustment strategy.