English immersion teaching system based on multi-modal interaction

By employing multimodal interaction and dual-path assessment, combined with learner state modeling, and dynamically adjusting feedback strategies, this approach addresses the issue of insufficient language creativity in existing English immersion teaching systems, thereby enhancing learners' language innovation capabilities and immersive learning experience.

CN121921147AInactive Publication Date: 2026-04-24ZHENGZHOU TECHN COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-13
Publication Date
2026-04-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing English immersion teaching systems focus too much on grammar and vocabulary standardization, which suppresses learners' language creativity and free and fluent communication skills, makes the interaction process rigid, and makes it difficult to achieve the leap from language knowledge to free communication ability.

Method used

Employing a multimodal interactive interface, an immersive scene engine, a learner state modeling module, and a dual-path assessment and guidance module, this system constructs a parallel mechanism for both standardized and creative assessments through voice, visual, and text input collection and feedback. It dynamically adjusts feedback strategies based on learners' emotional and ability states to encourage learners' creative language attempts.

Benefits of technology

It effectively stimulated learners' creative language attempts while ensuring the standardization of basic language, enhanced learners' language innovation ability and participation in immersive learning, and cultivated strategic thinking and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921147A_ABST
    Figure CN121921147A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer-aided teaching, and particularly discloses an English immersive teaching system based on multi-modal interaction, which comprises a multi-modal interaction interface, an immersive scene engine, a learner state modeling module, a dual-path evaluation and guidance module and a dynamic content adaptation module. Wherein the dual-path evaluation and guidance module executes normative evaluation and creativity evaluation in parallel, and dynamically fuses and generates composite feedback according to the real-time state of a learner. Through the above scheme, the method can effectively stimulate the creative expression of a learner while guaranteeing the language normalization, and provides self-adaptive personalized immersive learning experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer-aided instruction technology, specifically relating to an English immersion teaching system based on multimodal interaction. Background Technology

[0002] The integration of artificial intelligence and educational technology is profoundly transforming the field of language learning, particularly the teaching model of English as a Second Language. Immersion learning, as a core concept in this field, aims to promote learners' language acquisition and comprehensive application abilities by simulating authentic language contexts.

[0003] Multimodal interactive English immersion teaching systems integrate text, speech, visual, and tactile interaction modalities to create a highly realistic language environment for learners, aiming to enhance learning engagement and effectiveness. Existing technologies generally employ pre-set dialogue scripts and standard answer libraries to drive interaction and rely on rule-based or statistical model-based error correction mechanisms to evaluate learners' language output.

[0004] Existing technologies focus excessively on the grammatical accuracy and lexical standardization of language expression. Their error correction and feedback mechanisms often directly point out errors and provide standard answers. This single-dimensional evaluation standard severely inhibits learners' desire for language creativity and personalized expression in immersive environments, leading to a rigid interaction process. Learners tend to use safe but uncreative simple syntax, making it difficult to achieve the leap from language knowledge to free and fluent communication skills. Summary of the Invention

[0005] The purpose of this invention is to provide an English immersion teaching system based on multimodal interaction, so as to solve the technical contradiction in the prior art that the excessive focus on grammar and vocabulary norms inhibits learners' language creativity, resulting in a rigid immersion interaction process and difficulty in cultivating free and fluent communication skills.

[0006] To achieve the above objectives, this invention provides an English immersion teaching system based on multimodal interaction. The system includes a multimodal interaction interface, an immersive scene engine, a learner state modeling module, a dual-path assessment and guidance module, and a dynamic content adaptation module.

[0007] The multimodal interaction interface is used to collect learners' multimodal input data and present the system's multimodal feedback. This interface integrates a voice acquisition unit, a visual acquisition unit, a text input unit, and a multi-channel feedback unit. The voice acquisition unit captures learners' voice input through a microphone array and performs noise reduction and enhancement processing. The visual acquisition unit simultaneously captures learners' facial expressions, body movements, and gestures through a depth camera and an RGB camera. The text input unit receives text information entered by learners via a keyboard or touchscreen. The multi-channel feedback unit integrates a graphics rendering engine, a spatial audio engine, and a haptic feedback device to generate and present visual scenes, 3D spatial audio, and haptic feedback simulating physical interactions.

[0008] The immersive scene engine is used to build and drive a dynamic, interactive virtual language environment. This engine includes a scene database, a physics and behavior rule library, and a real-time rendering core. The scene database stores 3D models of virtual scenes, character models, object attributes, and environmental sound effects for multiple themes. The physics and behavior rule library defines the interaction logic between objects in the virtual environment, the character behavior tree, and event triggering conditions. Based on the learner's interactive input and system decisions, the real-time rendering core calls the scene database and rule library to drive the state update of the virtual scene and sends rendering commands to the multi-channel feedback unit of the multimodal interaction interface.

[0009] The learner state modeling module is used to construct and update a multi-dimensional learner cognitive and emotional state model in real time. This module includes a language ability analysis submodule, an interaction behavior analysis submodule, and an emotional state inference submodule. The language ability analysis submodule parses the speech and text input from the multimodal interaction interface, extracting quantitative indicators such as lexical complexity, syntactic diversity, discourse coherence, and pronunciation accuracy. The interaction behavior analysis submodule analyzes body language data acquired by the visual acquisition unit to identify the learner's attention direction, engagement level, and nonverbal communicative intentions. The emotional state inference submodule integrates the prosodic features of speech, micro-expression recognition results of facial expressions, and interaction behavior patterns, employing a multimodal emotional fusion algorithm based on deep neural networks to infer the learner's current emotional state, including dimensions such as confidence, confusion, engagement, or anxiety. The learner state modeling module integrates the above analysis results into a structured state vector and updates it over time.

[0010] The dual-path assessment and guidance module is the core of this invention. It is used to assess learners' language output in a parallel and collaborative manner and generate composite feedback that integrates normative guidance and creative encouragement. This module consists of a normative assessment path, a creative assessment path, and a feedback synthesizer.

[0011] The normative evaluation path is used to assess the accuracy and normativity of language output. This path includes a grammatical error detection unit, a lexical suitability assessment unit, and a pragmatic appropriateness analysis unit. The grammatical error detection unit uses a hybrid algorithm based on dependency parsing and statistical language models to identify grammatical errors in sentences. The lexical suitability assessment unit, based on a context vector space model constructed from a large-scale corpus, evaluates the appropriateness of the vocabulary used in the current context. The pragmatic appropriateness analysis unit, based on a pre-defined pragmatic rule base, judges the politeness and functionality of language expressions in specific social contexts. The normative evaluation path outputs a normative evaluation report containing specific error types, locations, and suggested corrections.

[0012] The creativity assessment path evaluates the novelty, richness, and communicative effectiveness of language output. This path includes a unit for measuring expressive novelty, a unit for analyzing semantic richness, and a unit for assessing communicative intent achievement. The unit for measuring expressive novelty quantifies the unconventionality of the learner's output by calculating the difference between the probability distribution of the learner's output sentence's n-gram model and the probability distribution of the corresponding pattern in a large-scale standard corpus. The unit for analyzing semantic richness assesses the semantic load and information content of the expression by analyzing the number, diversity, and conceptual network density of content words in the sentence. The unit for assessing communicative intent achievement, combined with the current interactive context provided by the immersive scene engine, analyzes whether the learner's language output effectively advances the virtual dialogue or completes the preset communicative task, even if the expression is not grammatically perfect. The creativity assessment path outputs a creativity assessment report that includes a creativity score, a description of specific highlights, and potential areas for improvement.

[0013] The feedback synthesizer receives reports from both the prescriptive and creative assessment paths and dynamically assigns weights and fuses feedback information based on real-time state vectors from the learner state modeling module. The synthesizer has a pre-defined feedback strategy matrix that maps the learner's language proficiency, emotional state, and current interaction stage, outputting weight configuration strategies for prescriptive and creative feedback. For example, for learners in a high-anxiety state or at the language foundation stage, the synthesizer increases the weight of prescriptive feedback to ensure a solid foundation; for learners demonstrating high confidence and strong language skills, it significantly increases the weight of creative feedback to encourage more complex language attempts. The synthesizer ultimately generates a structured composite feedback instruction, which includes explicit correction prompts for key errors, positive reinforcement descriptions for creative expressions, and optional advanced expression suggestions based on the current context. This composite feedback instruction is sent to the immersive scene engine and multimodal interaction interface, transforming into concrete forms such as virtual character dialogue responses, visual highlighting prompts, and encouraging voice comments.

[0014] The dynamic content adaptation module dynamically adjusts the difficulty and content direction of immersive scenarios based on the output of the dual-path assessment and guidance module and the long-term evolution of the learner's state model. This module includes a difficulty adjuster and a narrative branch controller. The difficulty adjuster analyzes trends in normative errors and creative scores in recent interactions and, based on a pre-set adaptive algorithm, adjusts the complexity of target syntax, the professionalism of required vocabulary, and the challenge of communicative scenarios in subsequent interactive tasks. The narrative branch controller, based on the communicative strategies and value orientations reflected in the learner's language choices at key dialogue nodes, selects a matching subsequent plot development path from multiple pre-set narrative branches in the scenario database, thereby achieving a more personalized and engaging immersive experience.

[0015] Furthermore, the dynamic weight allocation process in the feedback synthesizer is as follows: The feedback synthesizer receives a state vector output by the learner state modeling module in real time. This vector contains a comprehensive language ability index and an emotional state code. A nonlinear mapping function is pre-set within the feedback synthesizer. This function takes the comprehensive language ability index and the emotional state code as input and outputs a normative feedback weight coefficient between 0 and 1. The creative feedback weight coefficient is 1 minus the normative feedback weight coefficient. When the emotional state code is identified as high anxiety or the comprehensive language ability index is lower than the first threshold, the nonlinear mapping function outputs a normative feedback weight coefficient higher than 0.7. When the emotional state code is identified as high confidence and the comprehensive language ability index is higher than the second threshold, the nonlinear mapping function outputs a normative feedback weight coefficient lower than 0.3.

[0016] Furthermore, the specific measurement process of the expression novelty measurement unit is as follows: First, the learner's output statement is processed by word segmentation and part-of-speech tagging; then, all consecutive 2-gram and 3-gram combinations in the statement are extracted; next, the frequency of each n-gram combination obtained from the pre-loaded large-scale standard English corpus is queried; the logarithmic mean of the frequencies of all n-gram combinations in the learner's statement is calculated and used as the regularity benchmark score of the statement; finally, through a preset transformation function, the regularity benchmark score is mapped to a novelty score between 0 and 100, where a higher score indicates that the expression deviates more from the common pattern and the stronger the novelty.

[0017] Furthermore, the semantic richness analysis unit performs the following analysis process: Semantic role labeling is performed on the input statement to identify the core predicate and its related semantic roles such as agent, patient, time, and place; the number of different semantic roles in the statement is counted, and the ratio of these roles to the total number of words in the statement is calculated as a semantic role density index; simultaneously, using a pre-trained word vector model, the average cosine similarity of all real word vectors in the statement is calculated as a lexical semantic concentration index; the semantic richness analysis unit performs a weighted sum of the semantic role density index and the lexical semantic concentration index, and after normalization, finally outputs a semantic richness index.

[0018] Furthermore, the difficulty adjuster of the dynamic content adaptation module adopts an adjustment mechanism based on reinforcement learning strategies. The difficulty adjuster treats each teaching interaction as a state, outputs the action of adjusting the difficulty level as a strategy, and uses the combination of a reduction in the learner's normative error rate and an increase in creative scores in subsequent interactions as a reward signal. Through continuous interaction, the strategy network in the difficulty adjuster is constantly updated to learn the optimal difficulty adjustment strategy for a specific learner state, thereby maximizing long-term teaching effectiveness.

[0019] Furthermore, the system operates within a hierarchical decision-making loop framework. This framework comprises a perception layer, a cognitive decision-making layer, and an execution layer. The perception layer corresponds to the multimodal interaction interface and learner state modeling module, responsible for raw data acquisition and initial state perception. The cognitive decision-making layer corresponds to the core logic of the dual-path evaluation and guidance module, responsible for parallel evaluation, strategy trade-offs, and composite decision generation. The execution layer corresponds to the immersive scene engine and dynamic content adaptation module, responsible for translating decisions into specific virtual environment changes and content presentation. This loop framework runs synchronously at a frequency of 30 times per second, ensuring the system's real-time response to learner interactions and the coherence of the environment.

[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention, by introducing a dual-path assessment and guidance module, constructs a collaborative mechanism that combines normative and creative assessments, fundamentally changing the single-dimensional evaluation criteria of immersive teaching systems. This system not only focuses on the accuracy of language form but also quantitatively assesses the novelty, richness, and communicative effectiveness of language expression. A feedback synthesizer dynamically integrates and weights the two assessment results. This design enables the system to effectively identify and motivate learners' creative language attempts while ensuring basic language standardization, thereby significantly stimulating learners' willingness to express themselves and their language innovation abilities.

[0021] 2. This invention achieves highly personalized adaptive feedback through deep coupling of the learner state modeling module and the feedback synthesizer. The system can perceive the learner's emotional state and ability level in real time and dynamically adjust the proportion of normative guidance and creative incentives in the feedback strategy accordingly. For anxious or weak learners, the system provides clearer and more supportive normative guidance to build confidence; for confident and capable learners, the system focuses on encouraging them to break through conventions and engage in more complex language construction. This real-time adjustment based on psychological and cognitive states makes teaching intervention more precise and humane, effectively reducing learners' emotional filtering and enhancing their psychological safety and long-term engagement in immersive learning.

[0022] 3. This invention, through a dynamic content adaptation module, transforms assessment results and long-term learning status into a dynamic basis for adjusting the immersive environment and content. The system not only provides feedback on individual interactions but also adaptively adjusts the difficulty and narrative direction of subsequent tasks based on learners' performance trends. This ability to combine micro-level interactive assessment with macro-level teaching path planning enables the system to provide each learner with a continuously challenging and highly relevant personalized learning journey. Control over narrative branches further enhances learners' sense of immersion and autonomy, deeply integrating language learning with meaningful virtual experiences. This, in turn, improves comprehensive language application skills while cultivating learners' strategic thinking and adaptability in real communicative situations. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the overall technical solution architecture of the present invention; Figure 2 This is a schematic diagram of the core principle framework of the dual-path evaluation and guidance module in this invention; Figure 3 This is a schematic diagram of the multi-level interaction relationship and data flow of learner state modeling and feedback synthesis in this invention; Figure 4 This is a logical flow diagram of the immersive scene engine and dynamic content adaptation module in this invention. Detailed Implementation

[0024] Example 1: This invention provides an English immersion teaching system based on multimodal interaction, the overall technical architecture of which is shown in the attached figure. Figures 1 to 4 As shown in the attached diagram. This system utilizes five core components working collaboratively: a multimodal interaction interface, an immersive scene engine, a learner state modeling module, a dual-path assessment and guidance module, and a dynamic content adaptation module. These components work together to construct an immersive English learning environment capable of real-time perception, intelligent assessment, dynamic feedback, and continuous evolution. The following will combine the attached diagram... Figure 1 To be continued Figure 4 The specific implementation methods of each module are described in detail.

[0025] The multimodal interaction interface, serving as the sole physical and informational communication channel between the system and the learner, bears the dual responsibility of input acquisition and output presentation. This interface comprises a voice acquisition unit, a visual acquisition unit, a text input unit, and a multi-channel feedback unit. The voice acquisition unit employs a circular array of eight high-sensitivity omnidirectional microphones, deployed on a horizontal plane 1.2 meters in front of the learner, with a sampling frequency of 48 kHz and a bit depth of 24 bits. The acquired raw audio signal first undergoes noise reduction processing based on adaptive spectral subtraction, followed by beamforming to enhance the voice signal from the learner's direction and suppress environmental reverberation and background noise. The processed voice stream is then sent to the subsequent language analysis module in 16 kHz mono format. The visual acquisition unit consists of a depth camera and a high-resolution RGB camera synchronized, with hardware triggering achieving frame-level synchronization at a frame rate of 30 frames per second. The depth camera acquires 3D point cloud data of the learner's upper body with millimeter-level accuracy for reconstructing limb posture; the RGB camera, with a resolution of 1920×1080, captures facial expression details. After the two video streams are timestamped, they are fed into the facial micro-expression recognition model and the full-body pose estimation algorithm, respectively.

[0026] The text input unit supports both standard QWERTY keyboard and capacitive touchscreen input. The received text character stream is encoded in UTF-8 and then appended with timestamps and input context identifiers to form structured text event packets. The multi-channel feedback unit integrates a graphics rendering engine, a spatial audio engine, and a haptic feedback device. The graphics rendering engine is built on the Vulkan API and supports real-time ray tracing and global illumination to generate high-fidelity virtual scenes. The spatial audio engine uses the HRTF (Head-Related Transfer Function) model to synthesize binaural audio signals with a sense of direction and distance in real time based on the coordinates of the virtual sound source in three-dimensional space. The haptic feedback device is a linear resonant actuator array installed on the armrests of the learner's chair, which can simulate physical interactive tactile sensations such as tapping, vibration, and sliding. Its driving signal is generated by the system based on the force, material, and duration parameters of the virtual interactive event.

[0027] The immersive scene engine is responsible for building and driving the operational logic and state evolution of the entire virtual language environment. Please refer to the appendix. Figure 4The engine comprises three subsystems: a scene database, a physics and behavior rule library, and a real-time rendering core. The scene database stores multiple themed scenes in asset packages, including typical social scenarios such as airport check-in counters, coffee shop ordering areas, and business meeting tables. Each scene asset package contains a complete 3D mesh model, PBR material textures, static lighting probes, environmental sound effect samples, and preset character animation sequences. All models have undergone LOD (Level of Detail) optimization to ensure a frame rate of over 30 frames per second at different rendering distances. The physics and behavior rule library adopts an ECS (Entity-Component-System) architecture, defining collision response rules between objects, state machines for interactive objects (such as the empty, half-full, and full states of a coffee cup), behavior trees for virtual characters (including nodes for greetings, questions, responses, and farewells), and event triggering conditions (such as triggering a coffee-making animation when a learner says the keyword "coffee" and approaches the coffee machine). The real-time rendering core operates in a 33-millisecond cycle. Within each cycle, it first retrieves the latest user input events from the multimodal interaction interface, then queries the latest state vector from the learner state modeling module, and finally combines this with the composite feedback instructions generated by the dual-path evaluation and guidance module. It then loads the corresponding resources from the scene database and updates the state of all virtual entities based on the physics and behavior rule base. Ultimately, the rendering core generates a complete scene frame and its corresponding audio mix, which is output through the multi-channel feedback unit.

[0028] The learner state modeling module is the core sensory hub for the system to achieve personalized instruction. Please refer to the attached document. Figure 3 This module constructs a multi-dimensional state vector that evolves over time by fusing heterogeneous data from multiple sources. It comprises a language ability analysis submodule, an interaction behavior analysis submodule, and an emotion state inference submodule. The language ability analysis submodule receives speech streams from the speech acquisition unit and text streams from the text input unit.

[0029] For speech streams, endpoint detection is first performed to segment effective speech segments. Then, a pre-trained end-to-end speech recognition model (based on the Transformer architecture) converts the speech into text, extracting prosodic features such as fundamental frequency, energy, and speech rate. The recognized text and the original text input are fed into a language feature extraction pipeline: lexical complexity is calculated by the TTR (type / form ratio) and average word length; syntactic diversity is measured by the number of different dependency relation types and their distribution entropy; discourse coherence is evaluated by calculating the variance of the semantic similarity between adjacent sentences (based on cosine similarity from Sentence-BERT embeddings); and pronunciation accuracy is calculated by forcibly aligning the speech with the reference text and calculating the phoneme-level error rate. All indicators are standardized and weighted to form a comprehensive language proficiency index.

[0030] The interaction behavior analysis submodule processes the posture and facial expression data output by the visual acquisition unit. It extracts 68 facial keypoints and 33 body keypoints using the MediaPipe framework, calculating features such as head tilt angle, gaze direction vector, and gesture openness. A pre-trained LSTM network is used to identify whether the learner's attention is focused on the virtual character (gaze angle less than 30 degrees and lasting more than 500 milliseconds). Engagement is quantified by the number of effective interactions per unit time (e.g., voice speech, gesture pointing), while nonverbal communicative intent is jointly determined by gesture type (e.g., pointing, shrugging, nodding) and body lean angle. The emotion state inference submodule employs a multimodal fusion strategy.

[0031] The input includes prosodic features of speech (fundamental frequency standard deviation, speech rate variation rate), facial expression AU (action unit) intensity vectors (obtained from face images via convolutional neural network regression), and interaction behavior patterns (such as prolonged silence, frequent corrections). These features are concatenated into a 128-dimensional fused feature vector, which is input into a three-layer fully connected neural network, outputting a four-dimensional emotional state code, corresponding to the probability distributions of four dimensions: confidence, confusion, engagement, and anxiety. The learner state modeling module updates the state vector every 100 milliseconds. This vector contains the outputs of all the above sub-modules and caches the data of the most recent 10 seconds in time series format for subsequent modules to perform trend analysis.

[0032] The dual-path evaluation and guidance module is the core innovation that distinguishes this invention from existing technologies. Please refer to the appendix. Figure 2 This module consists of three parts: a prescriptive evaluation path, a creative evaluation path, and a feedback synthesizer, enabling parallel, collaborative, and dynamic evaluation of learners' language output. The prescriptive evaluation path first preprocesses the input sentences, including word segmentation, part-of-speech tagging, and dependency parsing. The grammatical error detection unit employs a hybrid algorithm: on the one hand, it detects structural errors such as subject-verb agreement and tense conflicts using a rule set built on the Stanford dependency parser; on the other hand, it calculates the perplexity of sentences using a BERT language model trained on a corpus of billions of words. If the perplexity exceeds a threshold, it is marked as a potential grammatical anomaly. The lexical applicability judgment unit constructs a context vector space based on the GloVe word vector model trained on the CommonCrawl corpus. For each target word in a sentence, it calculates the weighted average cosine similarity between the target word and other word vectors within the context window; if the similarity is below 0.4, it is judged as inappropriate word use.

[0033] The pragmatic appropriateness analysis unit incorporates a knowledge base of 200 social pragmatic rules, such as using the modal verbs "could" or "would" when asking for help and avoiding abbreviations when speaking to strangers. The system matches learners' statements with the social roles in the current scenario (e.g., customer-waiter) to check for violations of relevant rules. The prescriptive assessment path ultimately outputs a structured report, including error location (identified by character offset), error type (e.g., subject-verb disagreement), suggested corrections (e.g., changing "is" to "are"), and a confidence score.

[0034] The creativity assessment approach quantifies creativity across three dimensions: novelty, richness, and communicative effectiveness. The expression novelty measurement unit executes the following process: after word segmentation and part-of-speech tagging of the input sentence, all consecutive 2-gram and 3-gram combinations are extracted; the frequency of each n-gram is queried in the pre-loaded Google Books Ngram corpus (covering publications from 1900 to 2019); the logarithmic mean of all n-gram frequencies is calculated and recorded as the conventional benchmark score. Finally, the transformation function is used to convert... Mapped to a novelty score of 0 to 100 : Where a and b are preset parameters, calibrated experimentally to a = -2.5 and b = 12.0. This formula ensures that when... At lower levels (i.e., when n-grammars are rare), The value approaches 100, and vice versa. The semantic richness analysis unit first labels the statements with semantic roles, identifies the core predicates and their arguments (agent, patient, instrument, location, time, etc.); and counts the number of different semantic roles. And calculate its relationship with the total number of words. The ratio of [value] to [value] is used as a semantic role density index. Simultaneously, the Word2Vec model is used to calculate the average cosine similarity of all content word (noun, verb, adjective, adverb) vectors. As an indicator of semantic concentration, the final semantic richness index... We obtain the following through weighted summation and normalization: in and Let the weighting coefficient be denoted as . =0.6, =0.4, emphasizing the importance of semantic role diversity. The communicative intent achievement assessment unit obtains the current interaction context from the immersive scene engine, including the preset task goal (such as asking for flight delay information), the history of previous conversations, and the expected response of the virtual character. The system identifies the learner's intent in the statement using a classification model (such as requesting information or expressing dissatisfaction) and compares it with the task goal. If the intent matches and contains the necessary information elements (such as flight number and time), it is considered successfully achieved, even if there are minor grammatical errors. The creative assessment path output includes... , Description of the signs and specific highlights of the achievement of the intention.

[0035] The feedback synthesizer receives the two reports mentioned above and dynamically fuses them with the real-time state vector from the learner state modeling module. The feedback synthesizer internally has a pre-defined nonlinear mapping function. This function uses a comprehensive language proficiency index. With emotional state coding (Anxiety dimension probability) is the input, and the output is the prescriptive feedback weight coefficient. : in For the Sigmoid function, and To adjust the parameters, and The threshold value is used. <0.4 (first threshold) or When >0.7, >0.7; when >0.8 (second threshold) and When <0.3, <0.3. Creative Feedback Weight The feedback synthesizer is based on... and The recommendations from the two reports were weighted and merged: normative errors with a confidence level higher than 0.8 and If the value is >0.5, a clear correction suggestion will be generated; creative highlights if >70 and If the value is greater than 0.6, a positive reinforcement statement is generated (e.g., "This expression is very creative!"). Simultaneously, the system retrieves three more complex expressions from a pre-set advanced expression library that match the current context as optional suggestions. Finally, the composite feedback instruction is encapsulated in JSON format, containing text content, voice emotion tags (e.g., encouraging, neutral), visual highlight area coordinates, and haptic feedback intensity parameters, and sent to the immersive scene engine.

[0036] The dynamic content adaptation module is responsible for the system's long-term adaptive evolution. Please refer to the appendix. Figure 4 This module includes a difficulty adjuster and a narrative branch controller. The difficulty adjuster maintains a sliding window that records the canonical error rate of the last 20 interactions. Average Creativity Score Based on this, the regulator calculates the difficulty of adjusting the gradient: if Five consecutive declines and The difficulty level increases with increasing difficulty and decreases with decreasing difficulty. The difficulty level affects three parameters of subsequent tasks: target syntactic complexity (e.g., from simple sentences to complex sentences), vocabulary specialization (e.g., from general vocabulary to industry-specific terms), and scenario challenge (e.g., from single-turn dialogue to multi-turn negotiation). The narrative branch controller monitors key decision points. For example, in an airport scenario, when a learner is asked "Why are you yo-ulate?", a responsible answer (e.g., "My alarm didn't go off, it's my fault") triggers a responsible narrative branch, leading to greater trust from subsequent characters; a deflective answer (e.g., "The traffic was terrible, not my fault") triggers a shirking branch, causing suspicion from the characters. All branches are preset in the scenario database, and the controller selects the most suitable path through semantic matching.

[0037] The entire system operates within a hierarchical decision loop framework, as shown in the attached diagram. Figure 1 As shown, the perception layer (multimodal interaction interface and learner state modeling module) collects and processes data at a frequency of 30 Hz; the cognitive decision-making layer (dual-path evaluation and guidance module) performs evaluation and feedback synthesis at a frequency of 10 Hz; and the execution layer (immersive scene engine and dynamic content adaptation module) drives environment rendering and content updates at a frequency of 30 Hz. These three layers achieve low-latency communication through shared memory and message queues, ensuring end-to-end response latency is less than 100 milliseconds.

[0038] Example 2: Building upon Example 1, this example deeply optimizes the creative assessment path in the dual-path assessment and guidance module to address the complex language phenomena of advanced learners. Specifically, the expression novelty measurement unit introduces a cross-modal metaphor detection mechanism. When learners use idioms such as "It's raining cats and dogs" or similes such as "It's as quiet as a library," the system no longer relies solely on n-gram frequency but activates a dedicated metaphor understanding submodule. This submodule contains a pre-built knowledge graph of common English metaphors, covering 5000 source-target domain mappings (e.g., time-money, argument-war). The system first identifies metaphorical structures (e.g., as...as, like, isa, etc.) through dependency parsing, extracting concepts from the source and target domains; then, it queries the knowledge graph for valid mappings; if a valid mapping exists, it is considered a valid metaphor, and an additional 10 points are added to the novelty score; if no valid mapping exists but the distance between the source and target domains in the semantic vector space is less than 0.5, it is considered a creative metaphor, and an additional 20 points are added to the novelty score. This effectively distinguishes between clichés and truly creative language constructions.

[0039] Furthermore, the semantic richness analysis unit incorporates discourse marker analysis in this embodiment. The system identifies and statistically analyzes the frequency and appropriateness of discourse markers (such as however, therefore, incontrast). By analyzing the logical relationships (cause and effect, contrast, progression) between the preceding and following sentences, if the marker matches the logical relationship, it is considered to have improved discourse coherence and is given increased weight in the semantic richness index. Simultaneously, the system records whether learners actively use multiple discourse marker variations (such as instead of / rather than / as an alternative); higher diversity results in a higher score.

[0040] In this embodiment, the feedback synthesizer adds a cognitive load dimension. The learner state modeling module adds a cognitive load inference submodule, which estimates the current cognitive load level by analyzing the frequency of speech pauses, the number of self-corrections, and the proportion of eye-tracking regressions. When the cognitive load is too high (>0.8), even if the learner is in a high-confidence state, the feedback synthesizer will temporarily reduce the complexity of creative feedback, prioritizing concise affirmations to avoid information overload. This mechanism ensures that creative incentives do not fail due to cognitive overload.

[0041] In this embodiment, the narrative branch controller of the dynamic content adaptation module supports dynamic branch generation. When learners use expressions that are not preset by the system but are semantically reasonable and creative at key nodes, the narrative branch controller can invoke a small language model to generate a new dialogue branch in real time based on the current context and character settings, and cache it in the scene database for later use. This gives the system limited open-ended narrative capabilities, further enhancing immersion and personalization.

[0042] With the above enhancements, this embodiment is particularly suitable for senior undergraduate students majoring in English or learners preparing for the IELTS or TOEFL exams, and can effectively support their development of rhetorical skills and discourse organization abilities in academic writing and advanced speaking.

Claims

1. An English immersion teaching system based on multimodal interaction, characterized in that, include: A multimodal interaction interface is used to collect learners' multimodal input data and present the system's multimodal feedback; An immersive scene engine for building and driving a dynamic, interactive virtual language environment; The learner state modeling module is used to build and update a multi-dimensional learner cognitive and emotional state model in real time. The dual-path assessment and guidance module is used to assess learners’ language output in a parallel and collaborative manner and generate composite feedback that integrates normative guidance and creative incentives. The dynamic content adaptation module is used to dynamically adjust the difficulty and content direction of the immersive scene based on the output of the dual-path evaluation and guidance module and the long-term evolution of the learner state model. The multimodal interaction interface integrates a voice acquisition unit, a visual acquisition unit, a text input unit, and a multi-channel feedback unit. The immersive scene engine includes a scene database, a physical and behavioral rule library, and a real-time rendering core; The learner state modeling module includes a language ability analysis submodule, an interaction behavior analysis submodule, and an emotional state inference submodule. The dual-path evaluation and guidance module consists of a normative evaluation path, a creative evaluation path, and a feedback synthesizer. The normative evaluation path includes a grammatical error detection unit, a lexical applicability judgment unit, and a pragmatic appropriateness analysis unit; The creative assessment path includes a unit for measuring the novelty of expression, a unit for analyzing semantic richness, and a unit for assessing the degree of achievement of communicative intent; The dynamic content adaptation module includes a difficulty adjuster and a narrative branch controller.

2. The English immersion teaching system based on multimodal interaction according to claim 1, characterized in that, The feedback synthesizer receives reports from the normative assessment path and the creative assessment path, and performs dynamic weight allocation and feedback information fusion based on the real-time state vector from the learner state modeling module. The feedback synthesizer has a preset feedback strategy matrix, which maps the learner's language proficiency level, emotional state and current interaction stage, and outputs a weight configuration strategy for normative feedback and creative feedback. The feedback synthesizer ultimately generates a structured composite feedback instruction, which is sent to the immersive scene engine and multimodal interaction interface to be transformed into a virtual character's dialogue response, visual highlighting prompts, or encouraging voice comments.

3. The English immersion teaching system based on multimodal interaction according to claim 2, characterized in that, The dynamic weight allocation process in the feedback synthesizer is as follows: The feedback synthesizer receives the state vector output by the learner state modeling module in real time. This vector contains the language proficiency index and the emotional state encoding. The feedback synthesizer has a pre-set nonlinear mapping function that takes the language proficiency index and emotional state code as input and outputs a normative feedback weight coefficient between 0 and 1; the creative feedback weight coefficient is 1 minus the normative feedback weight coefficient. When the emotional state code identifies high anxiety or the comprehensive language ability index is below the first threshold, the nonlinear mapping function outputs a normative feedback weight coefficient higher than 0.

7. When the emotional state encoding is identified as highly confident and the comprehensive language ability index is higher than the second threshold, the nonlinear mapping function outputs a normative feedback weight coefficient of less than 0.

3.

4. The English immersion teaching system based on multimodal interaction according to claim 1, characterized in that, The specific measurement process of the expression novelty measurement unit is as follows: First, the sentence output by the learner is processed by word segmentation and part-of-speech tagging; then, all consecutive 2-gram and 3-gram combinations in the sentence are extracted; then, the frequency of occurrence of each n-gram combination obtained from the statistics of the preloaded large-scale standard English corpus is queried. The logarithmic mean of the frequencies of all n-gram combinations in the learner’s sentence is calculated and used as the regularity benchmark score for that sentence. Finally, the regularity benchmark score is mapped to a novelty score between 0 and 100 through a predefined transformation function.

5. The English immersion teaching system based on multimodal interaction according to claim 1, characterized in that, The analysis process of the semantic richness analysis unit is as follows: Semantic role labeling is performed on the input statement to identify the core predicate and its related agent, patient, time, and place semantic roles; The number of different semantic roles in a sentence is counted, and the ratio of these roles to the total number of words in the sentence is calculated as a semantic role density index. At the same time, the mean cosine similarity of all real word vectors in the sentence is calculated using a pre-trained word vector model as a lexical semantic concentration index. The semantic richness analysis unit performs a weighted summation of the semantic role density index and the lexical semantic concentration index, and then outputs a semantic richness index through normalization.

6. The English immersion teaching system based on multimodal interaction according to claim 1, characterized in that, The difficulty adjuster of the dynamic content adaptation module adopts an adjustment mechanism based on reinforcement learning strategy; The difficulty adjuster treats each teaching interaction as a state, takes the action of adjusting the difficulty level as the strategy output, and takes the combination of the reduction of normative error rate and the improvement of creative score in subsequent interactions as the reward signal. Through continuous interaction, the policy network in the difficulty adjuster is constantly updated to learn the optimal difficulty adjustment strategy for a particular learner's state.

7. The English immersion teaching system based on multimodal interaction according to claim 1, characterized in that, The system operates within a hierarchical decision-making loop framework, which includes a perception layer, a cognitive decision-making layer, and an execution layer. The perception layer corresponds to the multimodal interaction interface and the learner state modeling module, and is responsible for raw data acquisition and primary state perception. The cognitive decision-making layer corresponds to the core logic of the dual-path evaluation and guidance module, which is responsible for parallel evaluation, strategy trade-offs, and compound decision generation; the execution layer corresponds to the immersive scene engine and dynamic content adaptation module, which is responsible for transforming decisions into specific virtual environment changes and content presentation; this loop framework runs synchronously at a frequency of 30 times per second.

8. The English immersion teaching system based on multimodal interaction according to claim 1, characterized in that, The voice acquisition unit in the multimodal interaction interface captures the learner's voice input through a microphone array and performs noise reduction and enhancement processing. The visual acquisition unit simultaneously captures the learner's facial expressions, body movements, and gestures using a depth camera and an RGB camera; the text input unit receives text information input by the learner via a keyboard or touchscreen. The multi-channel feedback unit integrates a graphics rendering engine, a spatial audio engine, and a haptic feedback device to generate and present visual scenes, 3D spatial audio, and haptic feedback that simulates physical interaction.

9. The English immersion teaching system based on multimodal interaction according to claim 1, characterized in that, The scene database in the immersive scene engine stores 3D models of virtual scenes, character models, object attributes, and environmental sound effects for multiple themes. The physics and behavior rule base defines the interaction logic between objects in the virtual environment, the character behavior tree, and the event triggering conditions; The real-time rendering core calls the scene database and rule base based on the learner's interactive input and system decisions, drives the state update of the virtual scene, and sends rendering instructions to the multi-channel feedback unit of the multimodal interaction interface.

10. The English immersion teaching system based on multimodal interaction according to claim 1, characterized in that, The language ability analysis submodule in the learner state modeling module parses the speech and text input from the multimodal interaction interface and extracts quantifiable indicators such as lexical complexity, syntactic diversity, discourse coherence, and pronunciation accuracy. The interactive behavior analysis submodule analyzes the body language data acquired by the visual acquisition unit to identify learners' attention direction, engagement level, and non-verbal communication intentions; The emotion state inference submodule integrates the prosodic features of speech, the micro-expression recognition results of facial expressions, and the interaction behavior patterns. It adopts a multimodal emotion fusion algorithm based on deep neural networks to infer the learner's current emotion state.