Spoken English situational simulation teaching method and system based on somatosensory interaction
By constructing a joint interactive feature map of body-sense and speech, and dynamically adjusting the rhythm of dialogue and role feedback, the problems of insufficient immersive interaction and multimodal information processing in English oral teaching are solved. It realizes the synchronous collection and deep integration of learners' language output and non-verbal actions, and improves the intelligence of the teaching system and the personalization of learning paths.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SCHOOL OF SCI & LITERATURE JIANGSU NORMAL UNIV
- Filing Date
- 2026-02-02
- Publication Date
- 2026-05-12
AI Technical Summary
Existing English oral teaching systems lack immersive interaction, have low student participation, cannot simulate real conversational atmospheres, and lack multimodal information processing capabilities, making it difficult to achieve synchronous collection and deep integration of language output and non-verbal actions, resulting in discontinuous learning paths and fragmented training.
By acquiring learners' action behavior data and English oral output content, a joint interactive feature map of body-sense and speech is constructed. Combined with a speech semantic analysis model, the dialogue rhythm and role feedback strategies are dynamically adjusted to generate personalized training paths, thereby achieving synchronous acquisition and deep integration of language output and non-verbal actions.
It significantly improved the teaching system's ability to understand learners' true expressions, established a dynamic task recommendation mechanism and personalized cyclical training paths, enhanced the adaptability of teaching content and the intelligence of interactive feedback, and promoted the personalization and data-driven nature of learning paths.
Smart Images

Figure CN122018691A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of language scenario simulation teaching technology, specifically to an English oral scenario simulation teaching method and system based on motion-sensing interaction. Background Technology
[0002] Currently, English oral teaching widely employs methods such as video instruction, situational dialogues, and AI speech recognition scoring. While these methods have yielded some results in training standard pronunciation, they generally suffer from a lack of immersive interaction, low student participation, and a lack of authentic scenarios for oral output. This is especially true in primary and junior high school, where students often rely heavily on rote memorization for English expression, exhibiting weak contextual understanding and immediate responsiveness, making them ill-equipped for real-world communication scenarios.
[0003] Existing technologies attempt to incorporate virtual reality (VR), augmented reality (AR), or speech recognition to assist English teaching, but the following problems still exist:
[0004] (1) The equipment is expensive and the operation is complicated, making it unsuitable for widespread adoption;
[0005] (2) Speech recognition has low tolerance for dialect, speech rate and intonation changes, resulting in distorted evaluation;
[0006] (3) Most systems lack guidance on non-verbal factors (such as body language, eye contact, etc.) and cannot simulate a real dialogue atmosphere;
[0007] (4) Learners lack feedback mechanisms in simulated situations, resulting in discontinuous learning paths and fragmented training.
[0008] Especially in "task-driven" oral training scenarios, such as airport information, shopping mall, and hospital registration, traditional technologies struggle to establish dynamic feedback mechanisms based on learners' body language interactions, language output status, and the system's built-in multidimensional interaction model. Consequently, they cannot effectively achieve the closed-loop process of "dialogue rhythm guidance - language output training - interactive feedback evaluation," making it difficult to improve students' practical language application abilities. Summary of the Invention
[0009] The purpose of this invention is to provide a method and system for simulating English oral communication scenarios based on motion-sensing interaction, so as to address the shortcomings of the prior art.
[0010] To achieve the above objectives, the present invention provides the following technical solution: a method for simulating English oral communication scenarios based on motion-sensing interaction, comprising:
[0011] Acquire learners' action and behavior data in the teaching scenario, including head rotation, body posture, facial orientation, and hand movements, to form an initial interaction feature sequence;
[0012] The system acquires learners' spoken English output and performs semantic tag classification, fluency assessment, and intonation recognition on the speech content based on a speech semantic analysis model, thereby constructing a speech feature matrix.
[0013] The interaction feature sequence and the speech feature matrix are time-aligned and feature-fused to form a joint haptic-speech interaction feature map.
[0014] Based on the aforementioned somatosensory-voice joint interaction feature map, corresponding simulated task nodes and NPC character scripts are selected, and the dialogue rhythm, contextual difficulty, and character feedback strategies are dynamically adjusted.
[0015] Based on the learner’s current task completion status and the somatosensory-voice joint interaction feature map, the task completion degree, language interaction response time and interaction authenticity are calculated, and a capability assessment report is generated.
[0016] Based on the low response indicators marked in the competency assessment report, customized training tasks that include setting scenarios, vocabulary topics, and behavioral interaction requirements are automatically recommended to form a personalized cyclical training path for learners.
[0017] Preferably, the process of forming the haptic-voice joint interaction feature map includes the following steps: aligning the initial interaction feature sequence with the voice feature matrix based on a unified timestamp; calculating the optimal matching path between the aligned feature sequences using a dynamic time warping algorithm to establish a response mapping relationship between voice and action; weighting and fusing the aligned features through a multi-channel attention mechanism to extract joint features that play a key role in task feedback in the joint behavior expression; and constructing a graph structure with nodes as units to generate the haptic-voice joint interaction feature map.
[0018] Preferably, the process of selecting the corresponding simulated task node and NPC character script based on the somatosensory-voice joint interaction feature map includes the following steps:
[0019] Extract the set of key nodes with high semantic relevance and behavioral continuity from the joint interaction feature map of somatosensory and speech, and use it as a representation subgraph of the current learner's behavioral state;
[0020] The node feature vectors in the expression subgraph are matched with the task templates in the preset task knowledge graph to determine the best-fit simulated task node.
[0021] Based on the contextual semantic tags of simulated task nodes, retrieve the corresponding non-player character behavior script library and select character response scripts with context adaptability;
[0022] The language expression and interaction triggering conditions of the role response script are dynamically replaced to match the current learner's state, thereby realizing personalized task generation.
[0023] Preferably, the process of dynamically adjusting the dialogue rhythm, contextual difficulty, and role feedback strategy includes the following steps:
[0024] Based on the speech fluency index and limb response delay time of nodes in the somatosensory-speech joint interaction feature map, the learner's current cognitive load level is calculated.
[0025] Based on the cognitive load level, a matching context complexity level is selected from a preset context parameter library, and the length of the background task instructions, the number of keywords, and the grammatical structure level are adjusted accordingly.
[0026] By combining the response time and semantic deviation rate in the learner's historical interaction characteristics, the language output speed, pause interval and guidance prompts of the non-player character are dynamically adjusted.
[0027] Preferably, the process of calculating task completion rate, language interaction response time, and interaction realism includes the following steps:
[0028] Extract the key behavioral paths corresponding to the current simulated task node from the joint haptic-voice interaction feature map, and use them as the task behavior trajectory;
[0029] The task completion rate is calculated based on the semantic tag matching rate and action execution completeness in the task behavior trajectory.
[0030] The average time interval between learner's voice output and non-player character's voice input is analyzed to obtain the language interaction response time;
[0031] The interaction authenticity score is calculated by comprehensively utilizing the temporal coherence of action features and speech features, and the consistency index of tone and emotion.
[0032] Preferably, the process of acquiring learner action behavior data includes using a depth camera, an infrared recognition device, and a three-dimensional pose estimation algorithm to acquire the learner's head pitch angle, yaw angle, roll angle, skeletal joint position coordinates, gesture morphology parameters, and facial orientation vector, and organizing them into an initial interaction feature sequence according to timestamps.
[0033] Preferably, the speech semantic analysis model is built on a deep learning structure and uses a bidirectional long short-term memory-conditional random field model or a Transformer model to perform semantic intent recognition, grammatical structure parsing and keyword extraction on the speech transcription content, and generates a speech feature matrix including semantic tags, fluency scores and intonation features.
[0034] This invention also provides an English oral communication scenario simulation teaching system based on motion-sensing interaction, comprising:
[0035] The motion-sensing behavior acquisition module acquires learners' motion behavior data in the teaching scenario, including head rotation, body posture, facial orientation, and hand movements, forming an initial interaction feature sequence.
[0036] Speech Feature Analysis Module: Acquires learners' spoken English output and performs semantic label classification, fluency assessment, and intonation recognition on the speech content based on the speech semantic analysis model, and constructs a speech feature matrix;
[0037] Multimodal feature fusion module: performs temporal alignment and feature fusion of the interaction feature sequence and the speech feature matrix to form a joint haptic-speech interaction feature map;
[0038] Interaction generation module: Based on the aforementioned haptic-voice joint interaction feature map, select the corresponding simulated task nodes and NPC character scripts, and dynamically adjust the dialogue rhythm, contextual difficulty, and character feedback strategy;
[0039] Performance Analysis Module: Based on the learner's current task completion status and the somatosensory-voice joint interaction feature map, calculates task completion degree, language interaction response time and interaction authenticity, and generates a capability assessment report;
[0040] Training recommendation module: Based on the low response indicators marked in the ability assessment report, it automatically recommends customized training tasks that include setting scenarios, vocabulary topics, and behavioral interaction requirements, forming a personalized cyclical training path for learners.
[0041] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0042] 1. This invention constructs a multimodal behavior perception system that combines haptic interaction with speech semantic analysis. This system enables the simultaneous acquisition and deep fusion of learners' verbal output and non-verbal actions during English oral communication, significantly improving the teaching system's ability to understand learners' authentic expressive states. By employing a haptic-speech joint interaction feature map modeling method, not only is the temporal continuity and semantic relevance of behavioral expressions preserved, but it also provides precise data support for subsequent task matching and interactive feedback, effectively compensating for the shortcomings of existing English oral teaching methods in processing multimodal information.
[0043] 2. This invention establishes a dynamic task recommendation mechanism and a personalized cyclical training path based on multi-dimensional indicators such as learners' cognitive load, task completion, and interaction quality. It can accurately recommend training content based on learners' weaknesses and gradually adjust the task difficulty, achieving a closed-loop teaching logic of "assessment-intervention-training-reassessment." This method not only improves the adaptability of teaching content and the intelligence of interactive feedback but also promotes the personalization and data-driven nature of learning paths, possessing broad application and promotion value and significant innovative value in educational technology. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0045] Figure 1 This is a flowchart of the method of the present invention.
[0046] Figure 2 This is a flowchart of the system modules of the present invention. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] Example 1, please refer to Figure 1 As shown in this embodiment, an English oral communication scenario simulation teaching method based on motion-sensing interaction includes:
[0049] Acquire learners' action and behavior data in the teaching scenario, including head rotation, body posture, facial orientation, and hand movements, to form an initial interaction feature sequence.
[0050] In this invention, the teaching system first collects learners' motion behavior perception data non-contactly through a motion-sensing interaction module set up in the teaching scenario. The motion-sensing interaction module preferably employs multimodal sensing devices, including a depth camera, an infrared recognition module, and a 3D pose estimation algorithm, to capture and analyze learners' motion characteristics in real time during oral interaction. Specifically, the system acquires learners' motion behavior data in the following dimensions through the motion-sensing device:
[0051] Head rotation data: The pitch, yaw, and roll angles of the learner's head are collected during the dialogue to determine their focus, interaction intention, and level of participation; Body posture data: The positional changes of the learner's shoulder, elbow, wrist, and knee joints are detected through skeletal point tracking algorithms to identify their standing posture, hand gestures, and other behaviors; Facial orientation information: The facial orientation and gaze direction are analyzed based on facial landmark detection algorithms to determine whether the learner is looking at the simulated dialogue character or task target; Hand movement features: Dynamic hand gestures such as open hands, clenched fists, and pointing are extracted using gesture recognition models (such as convolutional neural networks (CNN) or MediaPipe hand models) to assist in instruction input and expression support in contextual tasks.
[0052] After collection, the aforementioned multi-dimensional action and behavior data undergoes denoising, time synchronization, and feature standardization operations in a preprocessing module, ultimately forming an initial interaction feature sequence. This sequence is a set of multi-dimensional vectors arranged by timestamps, used to describe the learner's dynamic interaction behavior trajectory in a specific context.
[0053] The initial interaction feature sequence, as one of the inputs for subsequent speech-action fusion analysis, not only reflects the learner's non-verbal expression, but also reveals their comprehension state and emotional tendency to a certain extent, providing key behavioral basis for the system to carry out subsequent situational simulation scheduling, task feedback generation and personalized teaching recommendations.
[0054] In this embodiment, the implementation of this step overcomes the technical shortcomings of existing English oral teaching systems that "rely solely on voice input while ignoring non-verbal interaction," effectively enhancing learners' participation and authentic interactive experience in immersive task scenarios, and laying the foundation for building a "language-behavior" dual-channel linkage teaching system.
[0055] The system acquires learners' spoken English output and performs semantic tag classification, fluency assessment, and intonation recognition on the speech content based on a speech semantic analysis model, thereby constructing a speech feature matrix.
[0056] In this invention, the system acquires learners' spoken English output in real time during interactive teaching scenarios through an embedded voice acquisition module. The voice acquisition module includes a directional noise-canceling microphone array, a voice signal preprocessing unit, and an edge speech recognition engine, enabling high-fidelity voice signal capture even under various environmental interference conditions.
[0057] The acquired speech data first undergoes noise suppression, echo cancellation, and speech rate normalization processing by a speech signal preprocessing module to ensure the accuracy and stability of the speech analysis process. Subsequently, the preprocessed speech signal is input into the speech semantic analysis model constructed in this invention. This model, based on a deep learning speech recognition architecture (such as BiLSTM-CRF or Transformer structure), achieves multi-dimensional speech feature extraction and understanding. This step includes the following sub-processes:
[0058] Semantic tag classification: The speech-to-text content is compared with the pre-set teaching task semantic database to identify the category of semantic units (such as request, response, question, evaluation, etc.), and key entities (such as location, object, person) and grammatical structure are extracted to generate semantic tag vectors;
[0059] Fluency assessment: Based on parameters such as speech rate, pause frequency, repetition rate and speech continuity, a learner's oral fluency assessment model is constructed, and corresponding scoring indicators are output to measure the naturalness of expression and language mastery level.
[0060] Intonation recognition: By extracting pitch parameters such as fundamental frequency (F0), energy envelope, and formant features from the speech signal, it identifies whether learners' sentences contain intonation features such as rising interrogative sentences, stressed expressions, and emotional coloring.
[0061] The above analysis results are ultimately summarized into a multi-dimensional speech feature matrix, where each row represents a time segment, and the corresponding columns include multiple dimensions such as semantic category labels, speech fluency scores, intonation recognition parameters, and keyword extraction results. This speech feature matrix serves as a structured representation of user language output behavior and is directly used for subsequent action-speech feature fusion processing and contextual feedback scheduling.
[0062] Through the above processing steps, the system can not only accurately capture the learner's language output, but also comprehensively evaluate their oral proficiency from two dimensions: semantic depth and expression quality, achieving fine-grained tracking and dynamic adaptation of the learning process.
[0063] Compared with existing teaching systems that rely solely on keywords or speech recognition and transcription, this invention significantly improves the accuracy of teaching feedback and the intelligence level of the feedback mechanism by integrating multiple dimensions such as semantic tag classification, fluency assessment, and intonation recognition, providing key speech data support for personalized English oral teaching.
[0064] The interaction feature sequence and the speech feature matrix are time-aligned and fused to form a haptic-speech joint interaction feature map.
[0065] The initial interaction feature sequence and the speech feature matrix are time-aligned. First, a unified reference timeline is set, and a global timeline is constructed with a sampling interval of 10 milliseconds (i.e., 100 frames per second). It is assumed that the interaction feature sequence is generated by a depth camera at a rate of 30 frames per second, that is, each frame interval is approximately 33.3 milliseconds, and the speech feature matrix is obtained by capturing speech streams from a microphone and parsing them at a rate of 100 frames per second.
[0066] For each sampling point in the interaction feature sequence, linear interpolation is used to interpolate it to the three most recent speech time frames. For example, if the interaction feature of the first frame is located at time point t=33ms, and the speech frames exist at t=30ms and t=40ms, then interpolation is performed as follows: Calculate the weights. Where t1 = 30ms and t2 = 40ms; interpolation results F1 and F2 are the feature vectors corresponding to adjacent speech frames. After interpolating and aligning each interactive feature vector to the 100Hz time axis, an aligned interactive feature sequence with the same time length and dimension as the speech feature matrix is obtained, laying the foundation for feature fusion.
[0067] Dynamic Time Warping is performed on the time-aligned interaction feature sequence and the speech feature matrix to establish their time response mapping relationship. The algorithm implementation steps are as follows:
[0068] Define the interaction feature sequence A = {a1, a2, ..., am} and the speech feature sequence B = {b1, b2, ..., bn}, where m ≈ n; construct the cost matrix D with dimension m × n, where each element D(i, j) represents the Euclidean distance between feature vectors ai and bj; define the recursive formula: Where d(ai,bj)= Backtracking from D(m,n) to D(1,1), the minimum cost path is extracted, and the corresponding (i,j) index pairs represent the optimal match of two features on the time axis. The output is a set of matching time point pairs P={(i1,j1),(i2,j2),...,(ik,jk)}, which is used for subsequent feature fusion.
[0069] For each pair (ai, bj) in the set of matching time points, the steps to calculate the fused feature vector F(i, j) are as follows: Define a three-channel attention function:
[0070] Channel 1 (Semantic Channel): Calculates the keyword weights in speech features. ;
[0071] Channel 2 (Action Channel): Calculates the salient magnitude of interactive actions. ;
[0072] Channel 3 (Coupled Channel): Fusing Speech-Action Joint Correlation: ;in This represents vector concatenation, where W1~W3 and b1~b3 are trainable parameters.
[0073] The three attention scores are used as weight coefficients to calculate the weighted fusion features: All F(i,j) constitute a time-aligned sequence of joint feature vectors. The attention mechanism determines the weight contribution of each channel in a specific task scenario through training, and can be optimized based on supervised learning datasets, using loss functions such as cross-entropy to minimize the system's error in action recognition tasks.
[0074] Using each time-aligned joint feature vector F(i,j) as a graph node Vk, construct a time-directed graph G=(V,E). The specific steps are as follows:
[0075] Each node Vk has the following attributes: feature vector F(i,j); timestamp Tk; and task label Lk (e.g., "asking for directions", "shopping", "making an appointment"). For any two adjacent time nodes Vk and V(k+1), a directed edge is established. The edge weight is defined as: This is used to measure the continuity of feature representation. Graph embedding models (such as GCN or GAT) are used to encode the graph structure into a set of feature vectors for subsequent behavior recognition models.
[0076] The resulting graph structure preserves the temporal sequence and behavioral expression state, and is called the somatosensory-voice joint interaction feature map. This map not only records the learner's language and body language features in oral interaction scenarios, but also maps the continuity and semantic relationships between behaviors through the connection structure in the graph.
[0077] Based on the aforementioned haptic-voice joint interaction feature map, corresponding simulated task nodes and NPC character scripts are selected, and the dialogue rhythm, contextual difficulty, and character feedback strategies are dynamically adjusted.
[0078] The haptic-voice joint interaction feature map is represented as a directed graph G(V,E), where node Vk contains feature vector Fk (256 dimensions), timestamp Tk, and semantic label Lk.
[0079] Iterate through all nodes Vk and filter the set of key nodes S according to the following rules:
[0080] If the semantic similarity score1 of Vk is greater than or equal to 0.85 (calculated by cosine similarity between the node's semantic label and the current scene's semantic word vector); and the edge weight (i.e., action continuity score) between Vk and the previous node V(k-1) is greater than or equal to 0.9; and the time interval ΔT is less than or equal to 1.5 seconds; then the node is included in the key set S.
[0081] An index is created for the subgraph G'(S,E') formed by the node set S, which serves as the subgraph representing the current behavior state. The temporal structure and feature sequence are preserved for subsequent matching.
[0082] The task knowledge graph T is a set of task templates {T1,T2,...,Tn}, where each task template contains a task semantic label Lt, a standard behavioral feature vector Ft (256 dimensions), and a keyword set Wt. The feature center vector Fc of the representation subgraph G'(S,E') is calculated as follows: , k∈S, meaning the average of the feature vectors of all key nodes is taken. For each task template Ti, calculate its cosine similarity with Fc: The maximum value Simmax is taken and the corresponding task node Tmax is recorded. If Simmax ≥ 0.92, then Tmax is selected as the current best-fit simulation task node. Each task template Ti is bound to multiple script entities {Si1, Si2, ...}, and each script includes: a character language template (including variable slots); behavior triggering conditions (such as user actions, tone feedback); and a set of contextual tags (such as "request type", "interrogative sentence", "airport scene").
[0083] The semantic tags of the selected task nodes are cross-referenced with the tag set in the role script: if the number of overlapping tags is ≥2, it is determined as a candidate script; if only one tag overlaps, but the behavioral structure is completely matched (e.g., containing "confirm-response" action pairs), it can also be included as a candidate. Finally, the script with the highest semantic matching score and action structure fit is selected for personalized processing.
[0084] The keyword set Wc extracted from the expression subgraph, along with the fluency score Scoref and action reaction time Rtime, are used as the current learner's parameter set. The character script contains replaceable slot markers (e.g., [Destination], [Suggested Phrases], [Confirmation Expression]), which are filled in as follows: [Destination] → Select the word with the highest part of speech match for "location" from Wc; [Suggested Phrases] → If Scoref < 0.6, use a prompting statement template, such as "Maybe you can try…"; [Confirmation Expression] → If Rtime > 2 seconds, add a supplementary confirmation step: "Did you mean…?" After replacement, the final non-player character response script is generated, achieving personalized teaching content output.
[0085] Construct the learner interaction load vector Lf=[Ff,Rf], where: Ff is the speech fluency score (range 0~1, the higher the score, the more fluent the speech), which comes from the speech feature extraction module; Rf is the body action response time (in seconds), which is obtained by subtracting the previous speech trigger time from the action timestamp.
[0086] Train a support vector machine model with input features [Ff, Rf] and labels representing cognitive load levels (low, medium, high). Supervised learning training is performed using manually labeled real-world interaction data. Input the current Lf vector into the model, and it outputs the cognitive load level Lc for use in the next step.
[0087] Define the context parameter template: Low complexity: number of instruction words ≤ 10, no subordinate clauses, keywords ≤ 3; Medium complexity: number of instruction words ≤ 15, includes simple subordinate clauses, keywords ≤ 5; High complexity: number of instruction words > 15, includes compound sentences, keywords ≥ 6.
[0088] The selection is based on the Lc value: Lc=High → Select a low-complexity task template; Lc=Medium → Maintain medium complexity; Lc=Low → Proceed to a high-complexity template. Update the current task instruction structure and refresh the corresponding script instruction content.
[0089] Adjusting the feedback rhythm of non-player characters based on historical behavior includes: collecting the following parameters from the last 5 rounds of interaction: average response time Tavg (in seconds); semantic deviation rate Perr=1−(system-recognized semantics / standard semantic matching score);
[0090] Construct feedback regulation rules:
[0091] If Tavg ≥ 3 seconds and Perr ≥ 0.3: Set the character's speech rate to 90 to 100 words per minute; insert a cue-like waiting statement; extend the pause interval for non-player characters to 2 seconds;
[0092] If Tagg ≤ 1.5 seconds and Perr ≤ 0.1: Set speech rate ≥ 130 words per minute; disable prompting language, and the character should reply using natural speech.
[0093] All adjustment results are written to the script parameter configuration file and dynamically applied to the speech synthesis and interaction logic.
[0094] Based on the learner's current task completion status and the somatosensory-voice joint interaction feature map, the task completion degree, language interaction response time, and interaction authenticity are calculated, and an ability assessment report is generated.
[0095] To compute the learner's behavioral performance in the current task, key behavioral paths corresponding to the current simulated task node are first extracted from the haptic-voice joint interaction feature map, thus constructing the task behavior trajectory. The specific implementation method is as follows:
[0096] Given that the task identifier label of the current simulated task node is Lt (such as "asking for directions" or "shopping"), match it based on the completed task nodes;
[0097] In the joint interactive feature map G(V,E), search for all node subsequences containing Lt;
[0098] Semantic clustering is performed on each subsequence, and the longest and most similar subsequence is selected as the optimal behavioral path P={V1,V2,...,Vn}.
[0099] Each node Vk contains a somatosensory feature vector, a speech feature vector, a timestamp Tk, and behavioral annotations. The path order represents the learner's actual execution trajectory in the task.
[0100] This path, as a structural representation of the task behavior trajectory, is the core input for subsequent calculations of task completion and interaction metrics.
[0101] This invention defines task completion as the degree of consistency between the learner's behavioral trajectory and the expected standard task process, and comprehensively evaluates two dimensions: semantic expression and action execution.
[0102] Extract the standard behavioral path Q={q1,q2,...,qm} of the target task Tt from the task knowledge graph. Each qi is an expected behavior node, containing standard semantic label Li and action label Ai.
[0103] Compare the semantic labels Lk of each node in the learner's behavioral trajectory P with the labels Li in the standard path Q, and count the number of semantic matches S.
[0104] Use a cosine similarity of tag vectors ≥ 0.85 as the matching criterion;
[0105] Compare the structural similarity between the action features in each node and the expected action (e.g., angle amplitude difference ≤ 15°, duration error ≤ 1 second), and count the number of complete actions executed, A.
[0106] The final completion score D is calculated using the following formula: D = 0.6 × (S / m) + 0.4 × (A / m); where m is the standard task path length, D ∈ [0,1], and the closer to 1, the higher the completion score.
[0107] Language interaction response time reflects the learner's speed of understanding and responding to non-player character's voice input. In the task behavior trajectory P, the set of all nodes containing voice features is extracted; interaction round pairs are labeled: (Ui, Ri), where Ui is the non-player character's voice trigger time, and Ri is the learner's response time; timestamps are obtained from the speech transcription record in seconds; response intervals are calculated for all interaction rounds. If ΔTi ≥ 0.5 seconds and ≤ 10 seconds, it is considered a valid response; average response time , where N is the number of effective rounds. This value is used to measure language processing efficiency and participates in subsequent feedback adjustment and evaluation report output.
[0108] Interactive realism is used to evaluate whether learners' language and motor expressions in simulated tasks are natural and coordinated. Temporal coherence is calculated by defining the start time of language output as... The peak time of the action characteristics is ; calculate for each node If ΔT ≤ 1.0 seconds, it is considered time synchronization; the proportion of synchronized nodes is Ps = Ns / N, and the total number N is the number of trajectory nodes; the time-series coordination score is Cs = Ps. Intonation and emotion consistency score: Extract the fundamental frequency variation range (F0) and the amplitude of speech rate and intensity variation from speech features; construct an intonation and emotion classification model (based on random forest or convolutional neural network); compare the classification results with the node semantic labels (such as interrogative sentences, request sentences) to see if they are consistent; calculate the proportion of consistent nodes Pe = Ne / N, and obtain the intonation consistency score Ce = Pe. Interaction authenticity score Cr: Cr = 0.5 × Cs + 0.5 × Ce, with a value range of [0,1], the higher the value, the higher the naturalness of the interaction.
[0109] Construct a learner task performance evaluation vector V=[D,Tavg,Cr], and compare this vector with past evaluation records to assess the trend of ability change.
[0110] The final output is a structured competency assessment report, which includes the current task performance level, room for improvement, and corresponding training suggestions. The report structure supports PDF export or API output to the teacher's terminal for personalized teaching scheduling.
[0111] Based on the low response indicators marked in the competency assessment report, customized training tasks that include setting scenarios, vocabulary topics, and behavioral interaction requirements are automatically recommended to form a personalized cyclical training path for learners.
[0112] To improve learners' efficiency in enhancing their target abilities during English speaking training, this invention, based on a generated ability assessment report and low-response indicators marked in the report, automatically triggers a customized training task recommendation process, constructing a personalized cyclical training path with dynamic adaptability and continuous feedback mechanisms. The recommendation and path construction process specifically includes the following steps:
[0113] The generated capability assessment report includes the following three key indicators: task completion (D), average response time (Tavg), and interaction authenticity (Cr). The system presets three capability thresholds: D < 0.75 indicates low semantic completion; Tavg > 3.0 seconds indicates lag in language organization and output; and Cr < 0.7 indicates insufficient multimodal interaction coordination. If any indicator fails to reach a threshold, it is mapped to the corresponding capability dimension Ci ∈ {semantic expression, speech organization, interaction naturalness}, and its weakness label is recorded.
[0114] Establish a training task template library. The template structure includes the following three parts: setting scenario labels (such as "airport security check", "hospital appointment", "restaurant ordering"); target vocabulary themes (such as "occupational titles", "emotional adjectives", "locative prepositions"); and behavioral interaction requirements (such as "actively asking questions", "receiving instructions and executing them", "providing negative feedback").
[0115] Weakness labels are used as indexing criteria to filter the task combinations with the highest relevance from the training task template library. For example, if Tavg > 3.0 seconds, the corresponding ability label is "speech organization," and tasks involving high-frequency sentence reconstruction or contextual response training are prioritized. If Cr < 0.7, training tasks containing intonation imitation and synchronized action cues are recommended. Each recommended task is constructed with at least one explicit target sentence pattern, two scene keywords, and one multimodal interaction trigger requirement.
[0116] The current recommended task is added to the learner's training path sequence R={T1,T2,...,Tn}, and the path is organized according to the structure of "capability dimension → scenario adaptation → feedback verification". After each round of training tasks is completed, the system re-triggers the capability assessment and updates the latest indicator set V'=[D',Tavg',Cr']. If the same capability dimension fails to meet the standard twice in a row, the system will automatically increase the training intensity of the recommended task under that dimension, including: increasing instruction complexity; inserting interference items (such as time limits, task switching); and strengthening the behavioral feedback mechanism.
[0117] The training path features a closed-loop update mechanism, which continues until all capability indicators reach the preset level or remain stable at a high level for three consecutive times, at which point it is marked as a stage achieved.
[0118] In this embodiment, this step ensures that learners form a structured, responsive, and traceable learning path during actual training, significantly improving their language proficiency and interactive response in English-speaking contexts, and effectively solving the technical problems of "insufficient feedback," "fragmented content," and "vague goals" in traditional teaching methods.
[0119] Example 2, please refer to Figure 2 As shown in this embodiment, an English oral communication scenario simulation teaching system based on motion-sensing interaction includes:
[0120] The motion-sensing behavior acquisition module acquires learners' motion behavior data in the teaching scenario, including head rotation, body posture, facial orientation, and hand movements, forming an initial interaction feature sequence.
[0121] Speech Feature Analysis Module: Acquires learners' spoken English output and performs semantic label classification, fluency assessment, and intonation recognition on the speech content based on the speech semantic analysis model, and constructs a speech feature matrix;
[0122] Multimodal feature fusion module: performs temporal alignment and feature fusion of the interaction feature sequence and the speech feature matrix to form a joint haptic-speech interaction feature map;
[0123] Interaction generation module: Based on the aforementioned haptic-voice joint interaction feature map, select the corresponding simulated task nodes and NPC character scripts, and dynamically adjust the dialogue rhythm, contextual difficulty, and character feedback strategy;
[0124] Performance Analysis Module: Based on the learner's current task completion status and the somatosensory-voice joint interaction feature map, calculates task completion degree, language interaction response time and interaction authenticity, and generates a capability assessment report;
[0125] Training recommendation module: Based on the low response indicators marked in the ability assessment report, it automatically recommends customized training tasks that include setting scenarios, vocabulary topics, and behavioral interaction requirements, forming a personalized cyclical training path for learners.
[0126] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for simulating English oral communication scenarios based on motion-sensing interaction, characterized in that: include: Acquire learners' action and behavior data in the teaching scenario, including head rotation, body posture, facial orientation, and hand movements, to form an initial interaction feature sequence; The system acquires learners' spoken English output and performs semantic tag classification, fluency assessment, and intonation recognition on the speech content based on a speech semantic analysis model, thereby constructing a speech feature matrix. The interaction feature sequence and the speech feature matrix are time-aligned and feature-fused to form a joint haptic-speech interaction feature map. Based on the aforementioned somatosensory-voice joint interaction feature map, corresponding simulated task nodes and NPC character scripts are selected, and the dialogue rhythm, contextual difficulty, and character feedback strategies are dynamically adjusted. Based on the learner’s current task completion status and the somatosensory-voice joint interaction feature map, the task completion degree, language interaction response time and interaction authenticity are calculated, and a capability assessment report is generated. Based on the low response indicators marked in the competency assessment report, customized training tasks that include setting scenarios, vocabulary topics, and behavioral interaction requirements are automatically recommended to form a personalized cyclical training path for learners.
2. The English oral communication scenario simulation teaching method based on motion-sensing interaction according to claim 1, characterized in that: The process of forming a joint haptic-voice interaction feature map includes the following steps: The initial interaction feature sequence and the speech feature matrix are time-aligned based on a unified timestamp; the optimal matching path between the aligned feature sequences is calculated using a dynamic time warping algorithm to establish a response mapping relationship between speech and action; the aligned features are weighted and fused using a multi-channel attention mechanism to extract joint features that play a key role in task feedback in the joint behavior expression; the fused joint features are used as nodes to construct a graph structure to generate a somatosensory-speech joint interaction feature map.
3. The English oral communication scenario simulation teaching method based on motion-sensing interaction according to claim 1, characterized in that: Based on the aforementioned haptic-voice joint interaction feature map, the process of selecting the corresponding simulated task node and NPC character script includes the following steps: Extract the set of key nodes with high semantic relevance and behavioral continuity from the joint interaction feature map of somatosensory and speech, and use it as a representation subgraph of the current learner's behavioral state; The node feature vectors in the expression subgraph are matched with the task templates in the preset task knowledge graph to determine the best-fit simulated task node. Based on the contextual semantic tags of simulated task nodes, retrieve the corresponding non-player character behavior script library and select character response scripts with context adaptability; The language expression and interaction triggering conditions of the role response script are dynamically replaced to match the current learner's state, thereby realizing personalized task generation.
4. The English oral communication scenario simulation teaching method based on motion-sensing interaction according to claim 3, characterized in that: The process of dynamically adjusting the dialogue pace, contextual difficulty, and role feedback strategies includes the following steps: Based on the speech fluency index and limb response delay time of nodes in the somatosensory-speech joint interaction feature map, the learner's current cognitive load level is calculated. Based on the cognitive load level, a matching context complexity level is selected from a preset context parameter library, and the length of the background task instructions, the number of keywords, and the grammatical structure level are adjusted accordingly. By combining the response time and semantic deviation rate in the learner's historical interaction characteristics, the language output speed, pause interval and guidance prompts of the non-player character are dynamically adjusted.
5. The English oral communication scenario simulation teaching method based on motion-sensing interaction according to claim 1, characterized in that: The process of calculating task completion rate, language interaction response time, and interaction realism includes the following steps: Extract the key behavioral paths corresponding to the current simulated task node from the joint haptic-voice interaction feature map, and use them as the task behavior trajectory; The task completion rate is calculated based on the semantic tag matching rate and action execution completeness in the task behavior trajectory. The average time interval between learner's voice output and non-player character's voice input is analyzed to obtain the language interaction response time; The interaction authenticity score is calculated by comprehensively utilizing the temporal coherence of action features and speech features, and the consistency index of tone and emotion.
6. The English oral communication scenario simulation teaching method based on motion-sensing interaction according to claim 1, characterized in that: The process of acquiring learner action behavior data includes using depth cameras, infrared recognition devices, and 3D pose estimation algorithms to obtain the learner's head pitch angle, yaw angle, roll angle, skeletal joint position coordinates, gesture morphology parameters, and facial orientation vector, and then organizing them into an initial interaction feature sequence according to timestamps.
7. The English oral communication scenario simulation teaching method based on motion-sensing interaction according to claim 1, characterized in that: The speech semantic analysis model is built on a deep learning structure. It uses a bidirectional long short-term memory-conditional random field model or a Transformer model to perform semantic intent recognition, grammatical structure parsing and keyword extraction on the speech transcription content, and generates a speech feature matrix including semantic tags, fluency scores and intonation features.
8. A motion-sensing interactive English oral scenario simulation teaching system, used to implement the motion-sensing interactive English oral scenario simulation teaching method according to any one of claims 1-7, characterized in that: include: The motion-sensing behavior acquisition module acquires learners' motion behavior data in the teaching scenario, including head rotation, body posture, facial orientation, and hand movements, forming an initial interaction feature sequence. Speech Feature Analysis Module: Acquires learners' spoken English output and performs semantic label classification, fluency assessment, and intonation recognition on the speech content based on the speech semantic analysis model, and constructs a speech feature matrix; Multimodal feature fusion module: performs temporal alignment and feature fusion of the interaction feature sequence and the speech feature matrix to form a joint haptic-speech interaction feature map; Interaction generation module: Based on the aforementioned haptic-voice joint interaction feature map, select the corresponding simulated task nodes and NPC character scripts, and dynamically adjust the dialogue rhythm, contextual difficulty, and character feedback strategy; Performance Analysis Module: Based on the learner's current task completion status and the somatosensory-voice joint interaction feature map, calculates task completion degree, language interaction response time and interaction authenticity, and generates a capability assessment report; Training recommendation module: Based on the low response indicators marked in the ability assessment report, it automatically recommends customized training tasks that include setting scenarios, vocabulary topics, and behavioral interaction requirements, forming a personalized cyclical training path for learners.