Immersive language learning and evaluation system fusing multi-modal interaction

Through a multimodal interactive immersive language learning system, the core elements of cognition and aesthetic education in classical poetry are deconstructed, lightweight immersive scenarios are constructed, and multimodal data is collected and integrated for evaluation. This solves the problems of single evaluation dimensions and insufficient feedback in existing technologies, and realizes personalized learning path optimization and ability enhancement.

CN121935577AInactive Publication Date: 2026-04-28YONGZHOU VOCATIONAL & TECH COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YONGZHOU VOCATIONAL & TECH COLLEGE
Filing Date
2025-12-11
Publication Date
2026-04-28
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing language learning systems cannot fully cover the core aesthetic education needs of language content such as classical poetry. Their assessment dimensions are singular and lack multimodal data fusion, resulting in one-sided assessment results and a lack of precise support for personalized feedback and learning path optimization.

Method used

Design an immersive language learning and assessment system that integrates multimodal interaction. By defining modules to decompose the core elements of language content in terms of cognition and aesthetic education, construct a suitable lightweight immersive scenario, collect multimodal data and perform weighted fusion and dimensional cross-validation, and generate targeted learning feedback and path optimization.

Benefits of technology

It achieves precise capture of learners' rhythmic expression, emotional resonance, and ability to interpret artistic conception, enhances learning immersion and device compatibility, provides personalized and accurate assessment and dynamic learning path adjustment, and solves the problems of one-sided assessment dimensions and weak feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935577A_ABST
    Figure CN121935577A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of education, in particular to an immersive language learning and evaluation system fusing multi-modal interaction, which comprises a definition module, a scene construction module, a preprocessing module and a feedback module. The system collects multi-modal data of characters, voices, facial expressions and limb postures, and after validity verification and feature extraction, comprehensive evaluation is achieved through a weighted fusion and dimension cross validation model. According to the method, evaluation standards are set according to differences of different literary and sports types, an immersive scene adaptive to a mobile terminal and a lightweight VR device is constructed, and scene modes can be dynamically switched. The system solves the problems of single evaluation dimension, insufficient immersion and weak feedback pertinence in the prior art, not only can cover double requirements of cognition and beauty, but also can provide personalized feedback by dynamically adjusting a learning path, improve the evaluation accuracy and learning effect, and assist the improvement of the comprehensive ability of a learner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of educational technology, and more specifically, to an immersive language learning and assessment system that integrates multimodal interaction. Background Technology

[0002] Current language learning systems primarily focus on assessing memorization and cognitive levels, using single formats such as dictation and multiple-choice questions. This fails to address the core aesthetic needs of language content, particularly classical poetry, and cannot effectively capture learners' deeper understanding and expression of rhythm, emotional resonance, and evocative imagery. Existing systems suffer from limited interaction, lacking lightweight immersive scenarios adapted for mobile devices and lightweight VR, and failing to achieve comprehensive assessment through multimodal data fusion. Their assessment models are often based on single dimensions, neglecting to consider heterogeneous data such as text, speech, facial expressions, and body posture, resulting in biased assessments and a lack of precise support for personalized feedback and learning path optimization. Furthermore, current technologies do not differentiate assessment standards for different literary genres (such as bold and unrestrained versus delicate and refined poetry), failing to meet the personalized learning needs at the aesthetic level. Therefore, there is an urgent need for an immersive language learning and assessment system that integrates multimodal interaction to address the problems of limited assessment dimensions, insufficient immersion, and weak targeted feedback in existing technologies, thereby achieving a comprehensive improvement in cognitive and aesthetic abilities. Summary of the Invention

[0003] In view of this, the present invention addresses the shortcomings of the prior art by proposing an immersive language learning and assessment system that integrates multimodal interaction, aiming to solve at least one of the problems mentioned in the background art.

[0004] This invention provides an immersive language learning and assessment system that integrates multimodal interaction, including: a definition module, configured to decompose the cognitive, aesthetic core elements and immersion trigger points of the language learning content based on the genre and aesthetic expression requirements of the target language learning content, and determine the appropriate multimodal interaction dimensions; The scene building module is configured to build lightweight immersive scenes containing core imagery and interactive points of language learning content based on device type and hardware performance. The preprocessing module is configured to collect learners’ multimodal data through the adapter terminal and perform validity verification and feature extraction based on preset rules. The multimodal data includes text, speech, facial expressions, and body posture. The feedback module is configured to use a weighted fusion and dimensional cross-validation model to comprehensively evaluate the preprocessed data and generate targeted learning feedback and path optimization schemes based on the evaluation results.

[0005] In some embodiments, the target language learning content is classical poetry, the core cognitive elements are specifically imagery, allusions, and tonal rules, the core aesthetic education elements are specifically recitation rhythm, emotional tone, and body-fitting movements, the immersion trigger point is a scene interaction trigger point strongly related to the core imagery of the poem, and the multimodal interaction dimensions are text modality, voice modality, facial expression modality, and body posture modality. If the emotional tone of the poem is determined to be bold and unrestrained, then the speech modality speed value is set to be greater than or equal to the preset first speech speed threshold, the fundamental frequency fluctuation amplitude is set to be greater than or equal to the preset first fluctuation amplitude threshold, and the body modality is bound to large and sweeping movements. If the emotional tone of the poem is determined to be gentle and subtle, the speech modality speed value is set to be less than or equal to the preset second speech speed threshold, the number of pauses is greater than or equal to the preset first threshold times / sentence, and the body modality is bound to introverted actions.

[0006] In some embodiments, the core imagery interaction points are constructed based on the cognitive core elements of language learning content, and the scene modes include AR mode, fully immersive mode, and minimalist scene mode; If the device type is detected as a mobile device, AR mode is automatically enabled, which uses the camera to identify planes in the real scene and overlay virtual images; If the device type is detected as VR glasses, then the full immersion mode is enabled, and 360° scene rendering is performed; If a learner fails to trigger any core imagery interaction points for 30 consecutive seconds, it is determined that there is insufficient immersion, and a guidance prompt will pop up. If the scene loading time is detected to be longer than the preset time, it is determined that the hardware performance is insufficient, and the scene will be automatically switched to a minimalist scene mode to retain the core imagery and reduce the rendering precision.

[0007] In some embodiments, the multimodal data acquisition is achieved through a mobile camera, a built-in microphone, and a touchscreen; The validity verification and feature extraction based on preset rules include: The text modality uses OCR to recognize the input content and compares it with the standard answer corresponding to the core cognitive elements to determine the text matching degree. Speech modality extraction features include speech rate, fundamental frequency, and pause location. Feature points are extracted from facial expression modalities and mapped to preset emotion types; Extracting skeletal key points from limb posture modalities; If the text matching degree is less than the preset first matching degree threshold, it is determined that the understanding of the core cognitive elements is insufficient, and the text modality supplementary learning resources are pushed. If the deviation between the speech modal features and the standard template is greater than the preset deviation threshold, it is further determined that the rhythm control or emotional rhythm in the core elements of aesthetic education is insufficient. If the match between the emotional tone of facial expressions and the emotional tone of language learning content is less than a preset second matching threshold, the association is determined to be insufficient emotional resonance in the core elements of aesthetic education.

[0008] In some embodiments, during speech modality preprocessing, the extracted speech features are compared with a standard prosodic template of the target language learning content, wherein the standard prosodic template is set according to the emotional tone of the language learning content. If the speech rate deviates from the standard value by more than the first deviation threshold, and the fundamental frequency fluctuation amplitude is less than the preset first fluctuation amplitude threshold, it is comprehensively judged as a double deficiency in rhythm and emotional expression in the core elements of aesthetic education. If the pause position differs from the standard template by a number of preset second thresholds and is concentrated in the core emotional sentences of the poem, it is determined that the emotional rhythm is severely lacking, and rhythm training resources are pushed.

[0009] In some embodiments, a preset number of feature points are extracted in the facial expression modality preprocessing, mapped to four types of emotions: joy, sadness, excitement, and resentment, and the matching degree is calculated with the emotional tone of the language learning content. If the emotional matching degree is less than the preset second matching degree threshold and the emotional intensity fluctuation is less than the preset first emotional intensity fluctuation threshold, it is determined that there is neither emotional resonance nor deliberate expression, triggering the combination of emotional guidance and authentic expression training resources. If the emotional matching degree is greater than the second matching degree threshold, but the emotional intensity fluctuation is less than the first emotional intensity fluctuation threshold, it is determined that the emotional cognition meets the standard but the expression is unnatural, and training resources for natural emotional expression are pushed.

[0010] In some embodiments, in the preprocessing of body posture modalities, a preset number of skeletal key points are extracted and compared with the motion trajectories in a predefined contextual motion library, which is constructed based on the cognitive core elements and aesthetic education core elements of the target language learning content. If the motion matching degree is less than the preset third matching degree threshold, and the fluctuation amplitude of the skeletal key points is less than the preset first key point fluctuation amplitude threshold, it is judged that the artistic expression is insufficient and the limbs are stiff, triggering motion decomposition teaching and flexibility training guidance. If the motion matching degree is greater than or equal to the third matching degree threshold, but the fluctuation amplitude of the skeletal key points is less than the first key point fluctuation amplitude threshold, then it is determined that the motion adaptation meets the standard but the expression is stiff, and a natural performance demonstration video is pushed.

[0011] In some embodiments, the comprehensive evaluation of the preprocessed data using a weighted fusion and dimensional cross-validation model includes: If the speech prosody score is less than the preset first score threshold and the facial emotion score is less than the preset second score threshold, then it is cross-judged as an overall lack of emotional expression in the core elements of aesthetic education, and the proportion of emotional training in the learning path is increased. If the text recognition score is greater than or equal to the preset third score threshold, but the physical performance score is less than the preset fourth score threshold, it is determined that the mastery of the core cognitive elements is disconnected from the expression of the core aesthetic elements, and training resources related to imagery and physical expression are pushed.

[0012] In some embodiments, the step of using a weighted fusion and dimensional cross-validation model to comprehensively evaluate the preprocessed data and generating targeted learning feedback and path optimization schemes based on the evaluation results includes: If a modality corresponding to a core element meets the standard in two consecutive assessments, its training frequency should be reduced. If the target is not met twice in a row, its training percentage will be increased to the total training volume, and the learning path will be dynamically adjusted based on the results of multiple assessments.

[0013] In some embodiments, if the consistency between the system score and the expert score is less than 80% in the algorithm test, the modal weights are adjusted. If the data acquisition accuracy of a certain device is less than 90% during hardware adaptation testing, then the feature extraction algorithm should be optimized.

[0014] Compared with existing technologies, the beneficial effects of this invention are as follows: It innovatively integrates multimodal data including text, speech, facial expressions, and body postures, breaking through the limitations of traditional single-cognitive assessment systems. It comprehensively covers the dual cognitive and aesthetic needs of language learning, accurately capturing learners' rhythmic expression, emotional resonance, and ability to interpret imagery, thus solving the problem of one-sided assessment dimensions. It constructs a lightweight immersive scene adapted to mobile devices and lightweight VR devices, combining core imagery interactive design and dynamic mode switching to significantly improve learning immersion and device compatibility, lowering the barrier to entry. It sets differentiated assessment standards and interaction rules for different literary genres (such as bold and unrestrained poetry and delicate and graceful poetry), and uses weighted fusion and dimensional cross-validation models to achieve personalized and accurate assessment. Based on the assessment results, it dynamically adjusts the learning path, strengthening training in weak dimensions and optimizing the frequency of achieving target dimensions, providing targeted feedback and resource recommendations. This effectively solves problems such as the disconnect between cognition and expression, and insufficient emotional expression, helping learners achieve a comprehensive improvement in cognitive and aesthetic abilities.

[0015] The above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure.

[0016] Other features and aspects of this disclosure will become clearer from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is a functional block diagram of an immersive language learning and assessment system that integrates multimodal interaction, provided in an embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0020] In the description of this application, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0021] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0022] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0023] See Figure 1As shown, an immersive language learning and assessment system integrating multimodal interaction according to an embodiment of this application includes: The definition module is configured to break down the cognitive aspects of language learning content, the core elements of aesthetic education, and the immersion triggers based on the genre and aesthetic expression requirements of the target language learning content, and to determine the appropriate multimodal interaction dimensions. The scene building module is configured to build lightweight immersive scenes containing core imagery and interactive points of language learning content based on device type and hardware performance. The preprocessing module is configured to collect learners’ multimodal data through the adapter terminal and perform validity verification and feature extraction based on preset rules. The multimodal data includes text, speech, facial expressions, and body posture. The feedback module is configured to use a weighted fusion and dimensional cross-validation model to comprehensively evaluate the preprocessed data and generate targeted learning feedback and path optimization schemes based on the evaluation results.

[0024] It should be understood that the definition module, as the "demand center" of the system, has the core task of completing "element decomposition" and "dimensional matching" based on the essential attributes (literary genre) of the target language learning content and the aesthetic education goals. Among these, "cognitive core elements" refer to the basic comprehension level of language learning (such as the imagery and allusions of classical poetry), "aesthetic education core elements" focus on the level of emotional expression and artistic conception (such as the rhythm of recitation and body-fitting movements), and "immersion trigger points" are the key interactive carriers connecting cognition and aesthetic education (such as scene elements strongly related to the imagery of poetry). This module analyzes the characteristics of the language content and selects four multimodal interaction dimensions: text, speech, facial expressions, and body posture. These four dimensions respectively cover text comprehension, audio rhythm, visual emotion, and body performance, forming a full-dimensional coverage of "cognition + aesthetic education," overcoming the limitations of traditional systems that only focus on textual cognition.

[0025] Scene Construction Module: This module acts as an "immersive carrier," with its core logic being "adaptive construction." This means it doesn't rely on expensive professional equipment, but rather dynamically generates lightweight immersive scenes based on the user's terminal type (mobile / lightweight VR glasses) and hardware performance (such as processor computing power and network speed). The core of the scene is the "core imagery interaction point," retaining only interactive elements strongly related to core cognitive elements (such as "river water" and "moonlight" from "Spring River Flower Moon Night"), avoiding ineffective interactions that interfere with learning objectives.

[0026] The preprocessing module, acting as the "data hub," is responsible for the "collection-verification-extraction" of multimodal raw data. The collection phase utilizes common mobile hardware (camera, microphone, touchscreen) to lower the barrier to entry. The core of the preprocessing phase is "standardization transformation," converting unstructured raw data (text input, audio, facial images, body videos) into computable structured data (matching degree, feature parameters, status labels), and verifying data validity through preset rules (such as removing blurry facial images and noisy audio), providing reliable input for subsequent evaluation.

[0027] Feedback Module: Serving as the "evaluation and optimization hub," it employs a dual-core model of "weighted fusion + dimensional cross-validation." "Weighted fusion" ensures that the evaluation weights of each modality match the learning objectives (with higher weights for aesthetic education-related modalities), while "dimensional cross-validation" avoids misjudgments from a single modality (such as combining speech and facial expressions to judge emotional expression). Ultimately, it generates "targeted feedback + dynamic path optimization," achieving a closed loop of "evaluation-feedback-improvement."

[0028] After the system starts, the definition module first breaks down the core elements and interaction dimensions of the target language learning content; then the scene construction module generates an adapted immersive scene based on the user's device; learners engage in multimodal interaction in the scene (such as memorizing text, reciting poems, and performing physical actions), and the preprocessing module collects and processes the data in real time; finally, the feedback module completes a comprehensive evaluation based on the processed data, pushes personalized resources and adjusts the learning path, forming a complete learning and evaluation closed loop.

[0029] In some specific embodiments, the target language learning content is classical poetry, the core cognitive elements are imagery, allusions, and tonal rules, the core aesthetic education elements are recitation rhythm, emotional tone, and body-fitting movements, the immersion trigger point is a scene interaction trigger point strongly related to the core imagery of the poem, and the multimodal interaction dimensions are text modality, voice modality, facial expression modality, and body posture modality. If the emotional tone of the poem is determined to be bold and unrestrained, then the speech modality speed value is set to be greater than or equal to the preset first speech speed threshold, the fundamental frequency fluctuation amplitude is set to be greater than or equal to the preset first fluctuation amplitude threshold, and the body modality is bound to large and sweeping movements. If the emotional tone of the poem is determined to be gentle and subtle, the speech modality speed value is set to be less than or equal to the preset second speech speed threshold, the number of pauses is greater than or equal to the preset first threshold times / sentence, and the body modality is bound to introverted actions.

[0030] It should be understood that "heroic" poetry is explicitly defined as classical poetry whose themes include frontier warfare, historical reflection, and the expression of lofty aspirations, with emotions characterized by grandeur, passion, and magnificence. Examples include "Bring in the Wine," "To the Tune of 'Breaking the Enemy's Formation' - A Heroic Poem for Chen Tongfu," and "To the Tune of 'Nian Nu Jiao' - Reminiscences of Chibi." Its core characteristic is "expressive emotion and magnificent momentum," which dictates that the corresponding multimodal interaction rules must match this feature.

[0031] The "graceful and restrained" style of poetry is clearly defined as classical poetry whose themes include "love and longing in one's boudoir, sorrow of parting, and description of scenery and expression of feelings, with emotions that are delicate, subtle, melancholic, gentle, and beautiful in their imagery." Examples include "Sheng Sheng Man - Searching and Searching," "Rainy Alley," and "Magpie Bridge Fairy - Delicate Clouds Weaving Patterns." Its core characteristics are "restrained emotions and beautiful imagery," and the corresponding multimodal interaction rules must conform to these traits.

[0032] Large, sweeping movements: These are defined as "movements with a large range of motion, a brisk rhythm, and a graceful posture," specifically including raising a glass and waving an arm, walking with head held high, waving a hand to create momentum, turning around and extending an arm, clenching a fist and shaking an arm, etc. The core characteristics of these movements are "large spatial span, obvious force exertion, and externalization of emotion," which are suitable for the emotional expression needs of bold and unrestrained poetry.

[0033] Subtle and gentle movements: Replacing the original "reserved movements" to clarify the scope of protection, these movements are limited to "movements with small range of motion, slow rhythm, and reserved posture." Specifically, they include slow walking, looking down in thought, gently gathering sleeves, bowing the head, and lightly brushing the fingertips. The core characteristics of these movements are "small spatial span, gentle force, and restrained emotion," which are suitable for the artistic expression of graceful poetry.

[0034] Preset first speaking speed threshold: 120 words / minute (referencing the recitation speed of professional reciters for bold and unrestrained poems to ensure an emotionally stirring rhythm). Preset first fluctuation amplitude threshold: 30Hz (fundamental frequency fluctuation reflects emotional fluctuations, and this threshold ensures the emotional tension when reciting bold and unrestrained poems). Preset second speech rate threshold: 80 words / minute (to suit the rhythmic needs of delicate expressions in graceful poetry); Preset first threshold number of times: 5 times / sentence (to enhance the melancholy or tender emotions of graceful poetry through reasonable pauses); Entry-level weighting thresholds: Text modality weighting ≥ 40%, Body / facial modality weighting ≤ 15% (Focusing on cognitive foundations and lowering the threshold for aesthetic expression education); Advanced level weight thresholds: text modality weight ≤ 20%, body / facial expression modality weight ≥ 25% (strengthening aesthetic expression and improving comprehensive ability requirements).

[0035] The definition module first identifies the type of the target classical poetry: by analyzing the subject matter (e.g., "frontier warfare" is classified as heroic) and emotional expression (e.g., "delicate and subtle" is classified as graceful), the type is categorized. Then, multimodal interaction rules are set according to the type differences; heroic poems need to highlight "exhilaration", so the speech rate and fundamental frequency fluctuation threshold are increased, and large-scale actions are bound; graceful poems need to highlight "delicacy", so the speech rate is reduced, pause requirements are increased, and subtle and gentle actions are bound. At the same time, the weight allocation of each modality is dynamically adjusted according to the difficulty level selected by the learner to ensure that the learning objectives match the difficulty and achieve "personalized teaching".

[0036] In some specific embodiments, the core imagery interaction points are constructed based on the cognitive core elements of language learning content, and the scene modes include AR mode, fully immersive mode, and minimalist scene mode; If the device type is detected as a mobile device, AR mode is automatically enabled, which uses the camera to identify planes in the real scene and overlay virtual images; If the device type is detected as VR glasses, then the full immersion mode is enabled, and 360° scene rendering is performed; If a learner fails to trigger any core imagery interaction points for 30 consecutive seconds, it is determined that there is insufficient immersion, and a guidance prompt will pop up. If the scene loading time is detected to be longer than the preset time, it is determined that the hardware performance is insufficient, and the scene will be automatically switched to a minimalist scene mode to retain the core imagery and reduce the rendering precision.

[0037] It should be understood that the preset time is 5 seconds (a reasonable smoothness threshold is set based on the scene loading performance of mainstream mobile devices and lightweight VR devices). Core imagery interaction points: Based on the core cognitive elements, for example, the core images of "Ascending the Heights" are "fallen leaves", "Yangtze River" and "giant cries". Then only these three types of interaction points are retained in the scene. Each interaction point triggers the corresponding cognitive content (such as touching "fallen leaves" to pop up the imagery analysis of "endless falling leaves rustling down"), avoiding ineffective interactions that distract attention.

[0038] The core logic of the scene construction module is "balancing adaptability and immersion". First, it detects the user's device type: if it is a mobile device (phone / tablet), it enables AR mode; it uses the mobile device's camera to identify flat surfaces in the real scene (such as desktops, walls), and overlays virtual elements related to the core imagery of the poem (such as displaying a dynamic effect of "river water" on the desktop) to achieve a lightweight immersive experience of "reality + virtuality"; if it is a lightweight VR headset, it enables full immersion mode, and restores the artistic conception of the poem through 360° scene rendering (such as the moonlit scene by the river in "Spring River Flower Moon Night"), allowing learners to be completely immersed in it.

[0039] Meanwhile, the module monitors user interaction status and device performance in real time: if a learner does not trigger any core imagery interaction points for 30 consecutive seconds, it indicates insufficient immersion, and the system will pop up a guidance prompt (such as "Try touching the Yangtze River in the scene to unlock the story behind the poem") to guide the user to participate more deeply; if the scene loading time exceeds 5 seconds, it is determined that the user's device performance is insufficient (such as an old mobile phone or a low-configuration VR device), and it will automatically switch to a minimalist scene mode—retaining core imagery interaction points, but reducing the scene rendering precision (such as reducing particle effects and simplifying model textures) to ensure smooth system operation and avoid affecting the learning experience due to performance issues.

[0040] In some specific embodiments, the multimodal data acquisition is achieved through a mobile terminal camera, a built-in microphone, and a touch screen; The validity verification and feature extraction based on preset rules include: The text modality uses OCR to recognize the input content and compares it with the standard answer corresponding to the core cognitive elements to determine the text matching degree. Speech modality extraction features include speech rate, fundamental frequency, and pause location. Feature points are extracted from facial expression modalities and mapped to preset emotion types; Extracting skeletal key points from limb posture modalities; If the text matching degree is less than the preset first matching degree threshold, it is determined that the understanding of the core cognitive elements is insufficient, and the text modality supplementary learning resources are pushed. If the deviation between the speech modal features and the standard template is greater than the preset deviation threshold, it is further determined that the rhythm control or emotional rhythm in the core elements of aesthetic education is insufficient. If the match between the emotional tone of facial expressions and the emotional tone of language learning content is less than a preset second matching threshold, the association is determined to be insufficient emotional resonance in the core elements of aesthetic education.

[0041] It should be understood that the preset first matching threshold is 60% (referring to the cognitive achievement standards for language learning, setting a basic threshold for text comprehension). Preset deviation threshold: 30% (based on the feature range of the standard prosodic template, setting a reasonable deviation threshold for speech expression); Preset second matching threshold: 50% (refer to the algorithm accuracy of emotion recognition to set the basic threshold for emotional resonance); Data acquisition equipment: Mobile device with built-in camera (supports 1080P resolution to ensure clear facial and body images), built-in microphone (supports noise reduction function to extract pure voice), and touch screen (supports handwriting and keyboard input to adapt to different input habits).

[0042] The core task of the preprocessing module is to "collect valid data and extract key features". The specific process is as follows: Data collection: The system collects learners' text input (handwriting or keyboard input) through a touch screen, their recitation voice through a microphone, and their facial expressions and body movements video through a camera. Feature extraction and validity verification: Text Modality: Utilizing OCR technology to recognize handwritten input (supporting common fonts such as regular script and running script), or directly reading keyboard input, the system compares the input text with the standard answers corresponding to the core cognitive elements (such as the original text of dictation questions and keywords for imagery analysis questions) to calculate the text matching degree. If the matching degree is less than 60%, it indicates that the learner's understanding of the core cognitive elements (such as imagery and allusions) is insufficient, and the system automatically triggers the push of supplementary learning resources in the text modality (such as short videos on imagery analysis and popular science articles on allusions). Speech modality: The core features of the speech (speech rate, fundamental frequency, and pause positions) are extracted through audio processing algorithms and compared with the standard rhyme template of the poem. If the deviation exceeds 30%, it indicates that the learner has not mastered the rhythm or emotional expression of the poem, and is further judged as "insufficient rhythm control" (such as speech rate too fast / too slow) or "lack of emotional rhythm" (such as incorrect pause positions). Facial expression modality: Facial feature points are extracted using image recognition algorithms and mapped to preset emotion types (joy / sadness / excitement / grief), which are then matched with the emotional tone of the poem. If the matching degree is less than 50%, it indicates that the learner has not generated emotional resonance and is judged as "insufficient emotional resonance".

[0043] In some specific embodiments, in speech modality preprocessing, the extracted speech features are compared with the standard prosodic template of the target language learning content, and the standard prosodic template is set according to the emotional tone difference of the language learning content. If the speech rate deviates from the standard value by more than the first deviation threshold, and the fundamental frequency fluctuation amplitude is less than the preset first fluctuation amplitude threshold, it is comprehensively judged as a double deficiency in rhythm and emotional expression in the core elements of aesthetic education. If the pause position differs from the standard template by a number of preset second thresholds and is concentrated in the core emotional sentences of the poem, it is determined that the emotional rhythm is severely lacking, and rhythm training resources are pushed.

[0044] It should be understood that the first deviation threshold is 30%. The preset first fluctuation amplitude threshold is 20Hz (lower than the fundamental frequency fluctuation threshold of bold and unrestrained poetry, used to determine insufficient emotional expression). The preset number of second thresholds is 2 (refer to the sentence and paragraph structure of poetry and set a reasonable threshold for pause differences). Standard rhythmic templates: recorded by more than 3 professional reciters, with different settings based on the emotional tone of the poems; the bold and unrestrained templates emphasize "excitement" (fast speech speed, large fundamental frequency fluctuation), while the graceful and restrained templates emphasize "soothing" (slow speech speed, small fundamental frequency fluctuation), ensuring the authority and suitability of the templates.

[0045] The preprocessing flow for speech modalities is "feature comparison - comprehensive judgment". First, the extracted speech features (speech rate, fundamental frequency, pause positions) are compared dimension-by-dimensionally with the corresponding standard prosodic template: If the speaking speed deviates from the standard value by more than 30% and the fundamental frequency fluctuation is less than 20Hz, it means that the learner has neither reached the standard speaking speed nor has a lack of emotional fluctuation. The overall judgment is "dual deficiency in rhythm and emotional expression". The system will push a combination of rhythm training and emotional guidance resources (such as speaking speed adjustment exercises and emotional expression demonstration audio). If the pauses differ from the standard template in two or more places, and these differences are concentrated in the core emotional segments of the poem (such as "searching and searching, cold and desolate" in "Sheng Sheng Man"), it indicates that the learner has not grasped the emotional rhythm of the poem and is judged as having "severe lack of emotional rhythm". In this case, rhythm training resources (such as sentence and paragraph pause markings and follow-up reading exercises) will be given priority.

[0046] This working principle avoids misjudgment based on a single feature by cross-comparing multiple features, ensuring the accuracy of speech modality assessment and aligning with the cultivation needs of "recitation rhythm" in the core elements of aesthetic education.

[0047] In some specific embodiments, a preset number of feature points are extracted in the facial expression modality preprocessing, mapped to four types of emotions: joy, sadness, excitement, and resentment, and the matching degree is calculated with the emotional tone of the language learning content. If the emotional matching degree is less than the preset second matching degree threshold and the emotional intensity fluctuation is less than the preset first emotional intensity fluctuation threshold, it is determined that there is neither emotional resonance nor deliberate expression, triggering the combination of emotional guidance and authentic expression training resources. If the emotional matching degree is greater than the second matching degree threshold, but the emotional intensity fluctuation is less than the first emotional intensity fluctuation threshold, it is determined that the emotional cognition meets the standard but the expression is unnatural, and training resources for natural emotional expression are pushed.

[0048] It should be understood that the preset number of feature points is 72 (using the standard feature point extraction scheme of Face++ API, covering key facial areas such as the corners of the eyes, corners of the mouth, and brow bones to ensure the accuracy of emotion recognition). Preset second matching threshold: 50% (to maintain consistent judgment criteria); The preset first emotional intensity fluctuation threshold is 0.2 (out of 1; the minimum fluctuation threshold for "natural expression" is set with reference to the intensity quantification standard of emotion recognition algorithms). Emotion mapping logic: By using the positional changes of 72 facial feature points (such as an upturned mouth corresponding to "joy" and a drooping eye corresponding to "sadness"), combined with a machine learning model, facial expressions are mapped to 4 preset emotions with a mapping accuracy of ≥90% (based on the technical parameters of Face++ API).

[0049] The core of facial expression modality preprocessing is "emotion matching and authenticity determination." First, 72 feature points are extracted from the learner's face, and the corresponding emotion type and emotion intensity fluctuation value are obtained through an emotion mapping algorithm. Then, the matching degree between the emotion type and the emotional tone of the poem is calculated. If the emotional matching degree is less than 50% and the emotional intensity fluctuation is less than 0.2, it means that the learner not only did not generate emotional resonance, but also had a stiff and deliberate expression (such as forcibly frowning to express "sadness"). It is judged as "neither emotional resonance nor deliberate expression". The system will push emotional guidance (such as explaining the background of the poet's creation to help understanding).

[0050] In some specific embodiments, in the preprocessing of body posture modality, a preset number of skeletal key points are extracted and compared with the motion trajectory in a predefined mood motion library, which is constructed based on the cognitive core elements and aesthetic education core elements of the target language learning content. If the motion matching degree is less than the preset third matching degree threshold, and the fluctuation amplitude of the skeletal key points is less than the preset first key point fluctuation amplitude threshold, it is judged that the artistic expression is insufficient and the limbs are stiff, triggering motion decomposition teaching and flexibility training guidance. If the motion matching degree is greater than or equal to the third matching degree threshold, but the fluctuation amplitude of the skeletal key points is less than the first key point fluctuation amplitude threshold, then it is determined that the motion adaptation meets the standard but the expression is stiff, and a natural performance demonstration video is pushed.

[0051] It should be understood that the "artistic imagery movement library" is a standardized set of movements constructed based on the core cognitive elements (imagery, allusions) and aesthetic education elements (emotional tone, physical adaptation movements) of the target classical poetry. For example, the movement library for the bold and unrestrained poem "Bring in the Wine" includes expansive movements such as "raising the cup and waving the arm," "raising the head and shaking the arm," and "turning around and stretching the arm"; while the movement library for the graceful and restrained poem "Slowly, Slowly" includes subtle and gentle movements such as "lightly gathering the sleeves," "lowering the head and nodding," and "lightly brushing the fingertips." Each movement is marked with the standard trajectory coordinates of 18 skeletal key points to ensure accurate comparison.

[0052] Preset number of skeletal key points: 18 (using the standard skeletal extraction scheme of the OpenPose open-source library, covering key limb parts such as head, neck, shoulder, elbow, wrist, hip, knee, and ankle, which can completely capture limb movement trajectory).

[0053] Preset third matching threshold: 40% (referring to the requirement of conformity between body movements and standard trajectories, setting the basic standard for artistic expression). The preset threshold for the first key point fluctuation amplitude is 0.1 (out of 1). Based on the fluctuation range of the human body's natural limb movement, a threshold is set to distinguish between "natural expression" and "stiff expression". A fluctuation amplitude lower than this value indicates that the limb movement lacks flexibility.

[0054] The core of preprocessing body posture modalities is "movement trajectory comparison and interpretation state determination," which fully serves the assessment needs of "body-adaptive movements" in the core elements of aesthetic education. The specific process is as follows: Skeletal key point extraction: The learner's limb movement video is collected by the mobile camera, and the coordinate data of 18 skeletal key points are extracted in real time using the OpenPose algorithm to generate continuous key point trajectories (such as the coordinate change sequence of the shoulder, elbow and wrist when the arm is waving). Action comparison and status determination: The extracted skeletal key point trajectories are compared frame by frame with the standard motion trajectories of the corresponding poems in the "Artistic Motion Library" to calculate the motion matching degree. If the movement matching degree is less than 40% and the fluctuation amplitude of the skeletal key points is less than 0.1, it means that the learner has neither mastered the appropriate limb movement nor has limb stiffness (such as mechanically imitating the "raising a cup" movement without natural fluctuation). It is judged as "insufficient artistic expression and limb stiffness". The system triggers the combination of resources of "movement decomposition teaching + flexibility training guidance" (such as step-by-step demonstration of the skeletal movement logic of the "wielding a brush" movement and push limb stretching exercise videos). If the movement matching degree is ≥40% but the fluctuation range of the skeletal key points is less than 0.1, it means that the learner has mastered the basic form of the movement, but the expression is stiff and unnatural (such as the limbs not swinging naturally when "walking slowly"). It is judged as "movement matching meets the standard but expression is stiff", and a natural performance demonstration video is pushed (such as the details of the body movements of professional actors performing the poem, with the fluctuation range of key points marked).

[0055] In some specific embodiments, the comprehensive evaluation of the preprocessed data using a weighted fusion and dimensional cross-validation model includes: If the speech prosody score is less than the preset first score threshold and the facial emotion score is less than the preset second score threshold, then it is cross-judged as an overall lack of emotional expression in the core elements of aesthetic education, and the proportion of emotional training in the learning path is increased. If the text recognition score is greater than or equal to the preset third score threshold, but the physical performance score is less than the preset fourth score threshold, it is determined that the mastery of the core cognitive elements is disconnected from the expression of the core aesthetic elements, and training resources related to imagery and physical expression are pushed.

[0056] It should be understood that the preset first score threshold is 12.5 points (the full score for the speech prosody module is 25 points, and this threshold is the passing score, corresponding to a 50% score rate). The preset second score threshold is 12.5 points (the facial emotion module has a maximum score of 25 points, consistent with the passing score for voice prosody, to ensure a unified emotional assessment standard). The preset third score threshold is 16 points (the full score for the text cognition module is 20 points, and this threshold corresponds to an 80% score rate, which is set as the standard for "mastery" of the core cognitive element). The preset fourth score threshold is 15 points (the full score for the physical performance module is 30 points, and this threshold is the passing score, corresponding to a 50% score rate). Emotional training percentage: Increased to 40% of total training (referring to the overall resource allocation of multimodal training to ensure that emotional weaknesses are addressed).

[0057] The single-dimensional score is calculated as "feature parameter score × module weight". For example, the text recognition module has a weight of 20%. If the learner's text matching degree is 80%, then the text recognition score is 80% × 20 = 16 points. The phonological prosody module has a weight of 25%. If the phonological features fit the standard template by 70%, then the phonological prosody score is 70% × 25 = 17.5 points. And so on. The total score is obtained by adding the four single-dimensional scores together.

[0058] The core of the feedback module's "multimodal fusion evaluation" is "single-dimensional weighted scoring + cross-dimensional cross-validation," which avoids the one-sidedness of single-module evaluation and ensures the accuracy of comprehensive evaluation. The specific logic is as follows: Cross-dimensional cross-validation judgment: If the speech prosody score is less than 12.5 (failing) and the facial emotion score is less than 12.5 (failing), it indicates that the learner has serious deficiencies in both the audio (prosody) and visual (facial expression) core emotional expression modalities. This is cross-judged as "overall lack of emotional expression in the core elements of aesthetic education". The system will automatically increase the proportion of emotional training in the learning path to 40% (such as increasing the frequency of push of resources such as emotional simulation training and emotional guidance for recitation) to specifically strengthen emotional expression ability. If the textual cognition score is ≥16 (proficient) but the physical performance score is less than 15 (failing), it means that the learner has fully understood the core cognitive elements of poetry (imagery, allusions, etc.), but cannot express the artistic conception through physical movements. This is judged as "a disconnect between the mastery of the core cognitive elements and the expression of the core aesthetic elements". The system will push "imagery-physical expression related training resources" (such as analyzing the logic of the physical stretching movements corresponding to the image of "Yangtze River" and providing "imagery-movement" comparison exercises) to break down the barriers between cognition and expression.

[0059] In some specific embodiments, the step of using a weighted fusion and dimensional cross-validation model to comprehensively evaluate the preprocessed data and generating targeted learning feedback and path optimization schemes based on the evaluation results includes: If a modality corresponding to a core element meets the standard in two consecutive assessments, its training frequency should be reduced. If the target is not met twice in a row, its training percentage will be increased to the total training volume, and the learning path will be dynamically adjusted based on the results of multiple assessments.

[0060] It should be understood that reducing the training frequency—adjusting it to 50% of the original frequency (for example, reducing the original 3 times per week for text recognition training to 1-2 times per week after reaching the target)—is necessary to avoid overtraining and wasting resources. Increase the training ratio: It is clearly defined as "increasing to 35% of the total training volume". This ratio ensures that weak dimensions are strengthened in a focused manner, while reserving enough training space for other modules to avoid excessive resource consumption by a single dimension.

[0061] This refers to three or more consecutive assessments (each assessment spaced 1-7 days apart, adapted to the learner's learning cycle) to ensure that the learning path adjustment is based on stable learning status data and to avoid misadjustment due to a single accidental performance.

[0062] This embodiment demonstrates the specific implementation logic of the feedback module "personalized feedback and path optimization," with its core being "dynamic resource allocation based on evaluation results," thus achieving a closed-loop learning system of "teaching according to aptitude." Targeted Modal Processing: If a modal corresponding to a core element (such as text recognition) achieves the target in two consecutive assessments (text recognition score ≥ 16 points), it indicates that the learner has mastered the ability in that dimension. The system will automatically reduce the training frequency to 50% of the original frequency and allocate the released training resources to other weaker dimensions. Weak modality handling: If a modality corresponding to a core element (such as physical performance) fails to meet the standard in two consecutive assessments (physical performance score less than 15 points), it indicates that this dimension is a core weak point. The system will increase its training proportion to 35% of the total training volume. For example, in 6 training sessions per week, 2 physical performance-specific training sessions will be arranged, and a weekly special improvement report will be generated (including training data, progress, and optimization suggestions). Dynamic adjustment of learning path: The system continuously tracks the assessment results of three or more consecutive assessments. If the weak modality reaches the standard after reinforcement training, the training ratio will be adjusted back. If other modalities show new signs of weakness, the resource allocation will be adjusted in a timely manner to ensure that the learning path always focuses on the learner's current ability shortcomings and maximizes learning efficiency.

[0063] In some specific embodiments, if the consistency between the system score and the expert score is less than 80% in the algorithm test, the modal weights are adjusted. If the data acquisition accuracy of a certain device is less than 90% during hardware adaptation testing, then the feature extraction algorithm should be optimized.

[0064] It should be understood that the algorithm testing consistency threshold is 80% (referring to the consistency standards between intelligent evaluation systems and expert evaluations in the industry, setting the minimum requirement for model reliability). Hardware compatibility accuracy threshold: 90% (Based on the general performance of mobile devices and lightweight VR devices, this sets an effective standard for data collection to ensure evaluation accuracy across different devices). User feedback percentage thresholds: "Insufficient immersion" feedback percentage > 30%, "Complex operation" feedback percentage > 20% (based on the statistical patterns of user survey sample size, setting optimization trigger conditions for scenario and interaction design).

[0065] Adjust modal weights: Based on the difference analysis between expert scores and system scores, for example, if experts give higher weight to emotional expression, while the system currently has a facial emotion weight of 25%, then appropriately increase it to 30% to ensure that the evaluation model matches professional judgment; Optimize feature extraction algorithms: For devices with a data acquisition accuracy of less than 90% (such as older mobile devices), simplify the feature extraction process (such as reducing the frame rate of skeletal key point extraction and optimizing the font adaptation logic of OCR recognition) to improve adaptability while ensuring accuracy. Optimize the design of scene interaction points: Add feedback sound effects (such as the wind sound effect triggered by touching "fallen leaves") and dynamic animations (such as "river water" ripples when touched) to enhance the sense of immersion; Simplify the interaction process: remove unnecessary confirmation steps (such as canceling the secondary confirmation of "whether to enter the scene") and optimize the operation entry point (such as setting the core interactive function to be triggered with one click).

[0066] This embodiment describes the system's "iterative optimization mechanism." Its core is to continuously improve the system's reliability, adaptability, and user experience through three types of testing, ensuring the system can adapt to different technical environments and user needs. The specific process is as follows: Algorithm testing iteration: The system's evaluation results for a batch of learners are compared with the human scores from more than 3 language education experts, and the consistency is calculated. If the consistency is less than 80%, it indicates that the weight allocation of the evaluation model is unreasonable (such as the weight of aesthetic education elements is too low). By analyzing the dimensional preferences of the expert scores, the scoring weights of each modality are adjusted until the consistency is ≥80%. Hardware compatibility testing iteration: Data collection tests are conducted on mainstream mobile devices (different brands, models, and system versions) and lightweight VR devices to statistically analyze the multimodal data collection accuracy of each device; if the accuracy of a certain device is less than 90%, the feature extraction algorithm is optimized for the hardware performance of that device (such as camera resolution and processor computing power) to ensure the effectiveness of data collection; User feedback testing iteration: Collect user feedback from a batch of users and count the percentage of issues such as "insufficient immersion" and "complex operation"; if the percentage of "insufficient immersion" is >30%, optimize the feedback design (sound effects, animations) of scene interaction points; if the percentage of "complex operation" is >20%, simplify the interaction process and remove redundant steps.

[0067] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. An immersive language learning and assessment system integrating multimodal interaction, characterized in that, include: The definition module is configured to break down the cognitive aspects of language learning content, the core elements of aesthetic education, and the immersion triggers based on the genre and aesthetic expression requirements of the target language learning content, and to determine the appropriate multimodal interaction dimensions. The scene building module is configured to build lightweight immersive scenes containing core imagery and interactive points of language learning content based on device type and hardware performance. The preprocessing module is configured to collect learners’ multimodal data through the adapter terminal and perform validity verification and feature extraction based on preset rules. The multimodal data includes text, speech, facial expressions, and body posture. The feedback module is configured to use a weighted fusion and dimensional cross-validation model to comprehensively evaluate the preprocessed data and generate targeted learning feedback and path optimization schemes based on the evaluation results.

2. The immersive language learning and assessment system integrating multimodal interaction according to claim 1, characterized in that, The target language learning content is classical poetry, the core cognitive elements are imagery, allusions, and tonal rules, the core aesthetic elements are recitation rhythm, emotional tone, and body-fitting movements, the immersion trigger points are scene interaction trigger points that are strongly related to the core imagery of the poetry, and the multimodal interaction dimensions are text modality, voice modality, facial expression modality, and body posture modality. If the emotional tone of the poem is determined to be bold and unrestrained, then the speech modality speech rate value is set to be greater than or equal to the preset first speech rate threshold, the fundamental frequency fluctuation amplitude is set to be greater than or equal to the preset first fluctuation amplitude threshold, and the body modality is bound to large and sweeping movements. If the emotional tone of the poem is determined to be gentle and subtle, the speech modality speed value is set to be less than or equal to the preset second speech speed threshold, the number of pauses is greater than or equal to the preset first threshold times / sentence, and the body modality is bound to introverted actions.

3. The immersive language learning and assessment system integrating multimodal interaction according to claim 2, characterized in that, The core imagery interaction points are constructed based on the cognitive core elements of language learning content, and the scene modes include AR mode, fully immersive mode, and minimalist scene mode; If the device type is detected as a mobile device, AR mode is automatically enabled, which uses the camera to identify planes in the real scene and overlay virtual images; If the device type is detected as VR glasses, then the full immersion mode is enabled, and 360° scene rendering is performed; If a learner fails to trigger any core imagery interaction points for 30 consecutive seconds, it is determined that there is insufficient immersion, and a guidance prompt will pop up. If the scene loading time is detected to be longer than the preset time, it is determined that the hardware performance is insufficient, and the scene will be automatically switched to a minimalist scene mode to retain the core imagery and reduce the rendering precision.

4. The immersive language learning and assessment system integrating multimodal interaction according to claim 3, characterized in that, The multimodal data acquisition is achieved through the mobile device's camera, built-in microphone, and touchscreen; The validity verification and feature extraction based on preset rules include: The text modality uses OCR to recognize the input content and compares it with the standard answer corresponding to the core cognitive elements to determine the text matching degree. Speech modality extraction features include speech rate, fundamental frequency, and pause location. Feature points are extracted from facial expression modalities and mapped to preset emotion types; Extracting skeletal key points from limb posture modalities; If the text matching degree is less than the preset first matching degree threshold, it is determined that the understanding of the core cognitive elements is insufficient, and the text modality supplementary learning resources are pushed. If the deviation between the speech modal features and the standard template is greater than the preset deviation threshold, it is further determined that the rhythm control or emotional rhythm in the core elements of aesthetic education is insufficient. If the match between the emotional tone of facial expressions and the emotional tone of language learning content is less than a preset second matching threshold, the association is determined to be insufficient emotional resonance in the core elements of aesthetic education.

5. The immersive language learning and assessment system integrating multimodal interaction according to claim 4, characterized in that, In speech modality preprocessing, the extracted speech features are compared with the standard prosodic template of the target language learning content, and the standard prosodic template is set according to the emotional tone difference of the language learning content. If the speech rate deviates from the standard value by more than the first deviation threshold, and the fundamental frequency fluctuation amplitude is less than the preset first fluctuation amplitude threshold, it is comprehensively judged as a double deficiency in rhythm and emotional expression in the core elements of aesthetic education. If the pause position differs from the standard template by a number of preset second thresholds and is concentrated in the core emotional sentences of the poem, it is determined that the emotional rhythm is severely lacking, and rhythm training resources are pushed.

6. The immersive language learning and assessment system integrating multimodal interaction according to claim 5, characterized in that, A predetermined number of feature points are extracted in the facial expression modality preprocessing, mapped to four types of emotions: joy, sadness, excitement, and resentment, and the matching degree is calculated with the emotional tone of the language learning content. If the emotional matching degree is less than the preset second matching degree threshold and the emotional intensity fluctuation is less than the preset first emotional intensity fluctuation threshold, it is determined that there is neither emotional resonance nor deliberate expression, triggering the combination of emotional guidance and authentic expression training resources. If the emotional matching degree is greater than the second matching degree threshold, but the emotional intensity fluctuation is less than the first emotional intensity fluctuation threshold, it is determined that the emotional cognition meets the standard but the expression is unnatural, and training resources for natural emotional expression are pushed.

7. The immersive language learning and assessment system integrating multimodal interaction according to claim 6, characterized in that, In the preprocessing of body posture modality, a preset number of skeletal key points are extracted and compared with the motion trajectory in a predefined contextual motion library. The contextual motion library is constructed based on the cognitive core elements and aesthetic education core elements of the target language learning content. If the motion matching degree is less than the preset third matching degree threshold, and the fluctuation amplitude of the skeletal key points is less than the preset first key point fluctuation amplitude threshold, it is judged that the artistic expression is insufficient and the limbs are stiff, triggering motion decomposition teaching and flexibility training guidance. If the motion matching degree is greater than or equal to the third matching degree threshold but the fluctuation amplitude of the skeletal key points is less than the first key point fluctuation amplitude threshold, it is determined that the motion adaptation meets the standard but the expression is stiff, and a natural performance demonstration video is pushed.

8. The immersive language learning and assessment system integrating multimodal interaction according to claim 7, characterized in that, The comprehensive evaluation of the preprocessed data using a weighted fusion and dimensional cross-validation model includes: If the speech prosody score is less than the preset first score threshold and the facial emotion score is less than the preset second score threshold, then it is cross-judged as an overall lack of emotional expression in the core elements of aesthetic education, and the proportion of emotional training in the learning path is increased. If the text recognition score is greater than or equal to the preset third score threshold, but the physical performance score is less than the preset fourth score threshold, it is determined that the mastery of the core cognitive elements is disconnected from the expression of the core aesthetic elements, and training resources related to imagery and physical expression are pushed.

9. The immersive language learning and assessment system integrating multimodal interaction according to claim 8, characterized in that, The method of using a weighted fusion and dimensional cross-validation model to comprehensively evaluate the preprocessed data, and generating targeted learning feedback and path optimization schemes based on the evaluation results, includes: If a modality corresponding to a core element meets the standard in two consecutive evaluations, its training frequency will be reduced. If the target is not met twice in a row, its training percentage will be increased to the total training volume, and the learning path will be dynamically adjusted based on the results of multiple assessments.

10. An immersive language learning and assessment system integrating multimodal interaction according to claim 9, characterized in that, If the consistency between the system score and the expert score is less than 80% in the algorithm test, the modal weights will be adjusted. If the data acquisition accuracy of a certain device is less than 90% during hardware adaptation testing, then the feature extraction algorithm should be optimized.