English teaching method based on somatosensory interaction
By collecting users' body movements and voice data in real time, and combining motion-sensing interaction with spatiotemporal graph convolutional networks, personalized teaching tasks are generated and errors are corrected in real time. This solves the problem of insufficient personalization and interactivity in traditional English teaching methods, and improves learning effectiveness and efficiency.
Patent Information
- Application Number
- CN202510992544.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional English teaching methods lack personalization and interactivity, making it difficult to continuously improve learning outcomes, and error correction is slow.
By collecting users' body movements and voice data in real time, and using motion-sensing interaction technology and spatiotemporal graph convolutional networks to analyze user behavior, personalized teaching tasks are generated, and errors are corrected in real time and the difficulty is dynamically adjusted.
It has enabled a personalized and highly interactive language learning experience, improved learning outcomes and user engagement, and optimized learning efficiency and language proficiency.
Smart Images

Figure CN120872151A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interactive teaching, and in particular to an English teaching method based on motion-sensing interaction. Background Technology
[0002] With the rapid development of information technology, traditional English teaching methods are facing increasing challenges, especially in improving students' language practice skills and interactive experiences. Motion-sensing interactive technology, with its ability to interact with virtual environments through body movements, has become an innovative application in education. Motion-sensing interactive English teaching methods can provide students with an immersive learning experience. Students can not only interact with English learning content in real time through physical movements, but also improve their practical language application skills in contextualized environments.
[0003] Compared to traditional English teaching methods, most current English teaching methods on the market have the following disadvantages: First, traditional methods typically rely on static textbooks and standardized courses, lacking personalization and unable to flexibly adjust teaching content according to each learner's progress, interests, and abilities. Second, traditional teaching focuses primarily on language output, lacking sufficient interactivity and immersion, making students easily bored and tired, hindering sustained learning progress. Furthermore, existing language learning platforms are relatively lagging in error correction and feedback mechanisms, making it difficult to promptly identify students' specific problems in grammar, pronunciation, etc., resulting in inaccurate correction of learning outcomes. Summary of the Invention
[0004] To improve existing methods, this paper proposes a motion-sensing interactive English teaching method. This method collects users' body movements and voice data in real time, combines dynamic tasks and an intelligent error correction system to achieve a personalized and highly interactive language learning experience, effectively improving learning outcomes and user engagement.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A motion-sensing interactive English teaching method includes: The system simultaneously collects user's physical and vocal data using a camera and microphone. Based on the teaching topic selected by the user, the system matches the corresponding grammatical structure and core vocabulary set from the database to generate dynamic teaching task instructions. Construct a body language-English intention mapping library to map real-time collected body movements into English communicative intention data; Users perform physical actions that match the target grammatical structure and core vocabulary according to the teaching task instructions, and simultaneously output the corresponding English speech. Based on spatiotemporal graph convolutional network, continuous frame action sequences of user body behavior data are analyzed, action spatiotemporal features are extracted and body behavior intention identifiers are output, and speech intention identifiers are generated through speech recognition and semantic parsing. Compare the consistency between the body language intent identifier and the voice intent identifier. If they match, the interaction is considered correct. If there is a conflict, the error correction process is triggered and a conflict label is generated. Based on data with conflict labels, the body behavior data and voice data are compared with the English communicative intent of the target grammatical structure and vocabulary set in the teaching task instructions, and correction instructions are generated based on the conflict type according to the inconsistency of intent. The teaching tasks are dynamically adjusted based on the accuracy of users' historical interactions. By adding visual interference variables and speech noise, single-action tasks are upgraded to complex communicative behavior tasks until all teaching task instructions are completed.
[0006] Preferably, the step of matching corresponding grammatical structures and core vocabulary sets from the database based on the user-selected teaching topic and generating dynamic teaching task instructions specifically includes: Based on the teaching topic selected by the user, three layers of related data are extracted from the database, including grammar layer, vocabulary layer and scenario layer; The grammar layer contains frequently used grammatical structures for the teaching topics selected by the user, the vocabulary layer contains a core vocabulary set related to the topic, and the scenario layer contains templates for typical communicative behaviors. Based on the acquired grammatical structure and core vocabulary set, teaching task instructions for various scenarios are generated, and these instructions are updated in real time.
[0007] Preferably, the construction of the body language-English intention mapping library, which maps real-time collected body movements to English communicative intention data, specifically includes: Based on historical data of interactive English teaching, a behavior-English intention mapping library was constructed. The collected body movements are used to construct a mapping model based on the relationship between the characteristics of the body movements and the English intentions, thus linking the body movements with the English intentions. The mapping model is trained based on historical data, and the trained model is used to analyze body movements in real time and provide English intentions.
[0008] Preferably, the user performs physical actions that match the target grammatical structure and core vocabulary according to the teaching task instructions, and simultaneously outputs the corresponding English speech, specifically including: Based on the generated dynamic teaching task instructions, the instructions are conveyed to the user through the display screen; Users perform corresponding physical actions based on the grammatical structures and core vocabulary required in the teaching task instructions, and verbally output English sounds that match the actions. The camera and microphone collect user's body movement and voice data in real time.
[0009] Preferably, the step of analyzing continuous frame action sequences of user body behavior data based on spatiotemporal graph convolutional networks, extracting spatiotemporal features of actions and outputting body behavior intention identifiers, and generating speech intention identifiers through speech recognition and semantic parsing specifically includes: Based on the acquired user body behavior data, a spatiotemporal graph convolutional network model is constructed and trained. Each frame's limb movement is represented as a graph structure, with the key points of the skeleton as nodes in the graph, and the connections between nodes representing the relationships between body parts. Spatiotemporal graph convolutional layers are used to capture spatial and temporal information. Spatial convolution learns the local spatial features of each node, while temporal convolution learns the dynamic changes between consecutive frames. Based on the acquired spatiotemporal features, the model is trained, and the trained spatiotemporal graph convolutional network model is used to identify the intention of body behavior and classify the intention of each continuous frame sequence. Based on the model's output, generate a label of body movement intention; Based on the acquired user voice data, the voice data is converted into text data through speech recognition, and the text data is parsed through natural language processing technology to extract the core semantics and obtain the voice intent identifier.
[0010] Preferably, the comparison of the consistency between the body language intent identifier and the voice intent identifier, if they match, is determined to be a correct interaction; if a conflict exists, the error correction process is triggered and a conflict label is generated, specifically including: Consistency is determined based on the acquired body language intent identifiers and voice intent identifiers. If the body language intent matches the voice intent... Figure 1 The interaction was deemed correct. If the intention of the body language is inconsistent with the intention of the voice, it is judged as a conflict, and an intention type label and a conflict label are added to the body language intention label and the voice intention label. The error correction process is triggered based on body language and voice data with added conflict tags.
[0011] Preferably, the step of comparing body language data and speech data with the English communicative intent of the target grammatical structure and vocabulary set in the teaching task instructions based on conflict-labeled data, and generating correction instructions based on conflict type according to inconsistencies in intent, specifically includes: Based on teaching task instructions with conflicting labels, the English communicative intent of the teaching task instructions is obtained through a body behavior-English intent mapping library. Acquire body action data and voice data with conflicting labels, including body action intent identifiers and voice intent identifiers; By comparing body language intent markers, voice intent markers, and English communicative intent, intent markers that do not match English communicative intent are obtained; Based on the mismatched intent identifiers, corrective instructions are generated to guide users to make corrections in their physical actions or speech.
[0012] Preferably, the step of dynamically adjusting the teaching task based on the user's historical interaction accuracy, by adding visual interference variables and speech noise, upgrades the single-action task into a complex communicative behavior task until all teaching task instructions are completed, specifically includes: Based on the user's past interactions, the accuracy, response time, and error types of task completion data are obtained. Based on the user's interaction accuracy, multi-level task difficulty is set by introducing different amounts of visual interference variables and voice noise; By setting complex communicative behavioral tasks, users can be taught through multi-channel interactive methods. During the teaching process of multi-level task difficulty, the accuracy and reaction time of users are monitored in real time, and the difficulty and complexity of the tasks are dynamically adjusted. Teaching is carried out through gradual difficulty adjustment. When the user successfully completes all task instructions, provide visual and verbal feedback to the user and collect the user's teaching feedback.
[0013] Compared with the prior art, the advantages of the present invention are: By collecting users' body language and speech data in real time, the system can tailor dynamic learning tasks for each user, making the learning process more aligned with their individual learning pace and needs. Combined with spatiotemporal graph convolutional network technology, this method can accurately analyze users' body language and speech output, identify and correct errors in learning, and ensure accuracy in grammar and vocabulary. Furthermore, by analyzing historical interaction data, the system dynamically adjusts the difficulty of tasks, gradually increasing the learning challenge for users and helping them master English at different levels. This multi-channel, interactive teaching method not only enhances the immersion in learning but also optimizes learning outcomes through continuous feedback and error correction mechanisms, improving learning efficiency and users' language proficiency. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the method proposed in this invention; Figure 2 This is a schematic diagram of the instructions for generating dynamic teaching tasks proposed in this invention; Figure 3 This is a schematic diagram illustrating the English communicative intent proposed in this invention; Figure 4This is a schematic diagram illustrating the user output body movements and English speech proposed in this invention; Figure 5 This is a schematic diagram illustrating the generation of limb behavior intent identifiers and voice intent identifiers proposed in this invention; Figure 6 This is a schematic diagram of the triggering error correction process proposed in this invention; Figure 7 This is a schematic diagram of the generation and correction instructions proposed in this invention; Figure 8 This is a schematic diagram of the instructions for completing teaching tasks proposed in this invention. Detailed Implementation
[0015] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.
[0016] See Figure 1 As shown, an English teaching method based on motion-sensing interaction includes: Step 1: Simultaneously collect the user's body movement data and voice data using a camera and microphone; Step 2: Based on the teaching topic selected by the user, match the corresponding grammatical structure and core vocabulary set from the database to generate dynamic teaching task instructions; Step 3: Construct a body language-English intention mapping library to map real-time collected body movements into English communicative intention data; Step 4: The user performs physical actions that match the target grammatical structure and core vocabulary according to the teaching task instructions, and simultaneously outputs the corresponding English speech; Step 5: Analyze the continuous frame action sequence of user body behavior data based on spatiotemporal graph convolutional network, extract spatiotemporal features of the action and output body behavior intention identifier, and generate speech intention identifier through speech recognition and semantic parsing; Step 6: Compare the consistency between the body language intent identifier and the voice intent identifier. If they match, the interaction is considered correct. If there is a conflict, the error correction process is triggered and a conflict label is generated. Step 7: Based on the data with conflict labels, compare the body behavior data and voice data with the English communicative intent of the target grammatical structure and vocabulary set in the teaching task instructions, and generate correction instructions based on the conflict type according to the inconsistency of intent. Step 8: Dynamically adjust the teaching tasks based on the accuracy of users' historical interactions. By adding visual interference variables and speech noise, upgrade the single-action task to a complex communicative behavior task until all teaching task instructions are completed.
[0017] See Figure 2As shown, based on the teaching topic selected by the user, the system matches the corresponding grammatical structures and core vocabulary sets from the database to generate dynamic teaching task instructions, specifically including: Based on the teaching topic selected by the user, three layers of related data are extracted from the database, including grammar layer, vocabulary layer and scenario layer; The grammar layer contains frequently used grammatical structures for the teaching topics selected by the user, the vocabulary layer contains a core vocabulary set related to the topic, and the scenario layer contains templates for typical communicative behaviors. Based on the acquired grammatical structure and core vocabulary set, teaching task instructions for various scenarios are generated, and these instructions are updated in real time.
[0018] Specifically, three layers of data related to the user-selected teaching topic are extracted from the database: grammar layer, vocabulary layer, and scenario layer. Based on this data, teaching task instructions for different scenarios are generated. The task instructions should be dynamically designed according to grammatical structure, vocabulary set, and communicative behavior. Based on student feedback, learning progress, and needs, teaching task instructions are updated in real time. If the user has mastered the basic quoting syntax, more complex tasks can be added. The difficulty of teaching tasks is automatically adjusted based on the user's errors in the tasks. If the user is more interested in a certain scenario, corresponding task instructions are recommended based on the student's interests.
[0019] See Figure 3 Constructing a body language-English intention mapping library, which maps real-time collected body movements to English communicative intention data, specifically includes: Based on historical data of interactive English teaching, a behavior-English intention mapping library was constructed. The collected body movements are used to construct a mapping model based on the relationship between the characteristics of the body movements and the English intentions, thus linking the body movements with the English intentions. The mapping model is trained based on historical data, and the trained model is used to analyze body movements in real time and provide English intentions.
[0020] Specifically, students' body movement data are collected through sensors, cameras or other motion-sensing devices. Common body movements include gestures, body postures and head movements. Each movement should have specific characteristics, such as the type, direction, speed and amplitude of the movement. Body language can represent specific intentions or emotional states, thereby influencing language output. Pointing to something may indicate "inquiry" or "request," while waving may indicate "greeting" or "saying goodbye." Based on historical data and linguistic research, we construct the relationship between body language and English intentions. A data-driven approach is used to build a mapping model between body language and English intentions. The mapping model is trained using historical data and then used to analyze students' body language in real time and predict their English intentions.
[0021] See Figure 4 As shown, users perform physical actions that match the target grammatical structure and core vocabulary according to the teaching task instructions, and simultaneously output the corresponding English speech. Specifically, this includes: Based on the generated dynamic teaching task instructions, the instructions are conveyed to the user through the display screen; Users perform corresponding physical actions based on the grammatical structures and core vocabulary required in the teaching task instructions, and verbally output English sounds that match the actions. The camera and microphone collect user's body movement and voice data in real time.
[0022] Specifically, based on the teaching topic selected by the user, the system extracts data from the grammar, vocabulary and scenario layers from the database, constructs appropriate task instructions, and presents the task instructions to the user through the display screen. The instructions should be clear and specify the task's objectives, grammar requirements and vocabulary requirements. Users provide feedback through body language according to the requirements of the task instructions. For example, based on the instructions in a shopping scenario, users may point to the product or make a questioning gesture. In addition to body language, users also need to verbally output English sentences that match the actions. The task requires users to use specified grammatical structures and core vocabulary. Sensors or cameras are used to capture users' body movement data. Cameras can track users' movements in real time and convert them into motion feature vectors. Microphones are used to collect users' spoken output in real time and convert the speech signal into text.
[0023] See Figure 5 As shown, based on spatiotemporal graph convolutional networks, continuous frame action sequences of user body behavior data are analyzed to extract spatiotemporal features of the actions and output body behavior intention identifiers. Specifically, speech intention identifiers are generated through speech recognition and semantic parsing, including: Based on the acquired user body behavior data, a spatiotemporal graph convolutional network model is constructed and trained. Each frame's limb movement is represented as a graph structure, with the key points of the skeleton as nodes in the graph, and the connections between nodes representing the relationships between body parts. Spatiotemporal graph convolutional layers are used to capture spatial and temporal information. Spatial convolution learns the local spatial features of each node, while temporal convolution learns the dynamic changes between consecutive frames. Based on the acquired spatiotemporal features, the model is trained, and the trained spatiotemporal graph convolutional network model is used to identify the intention of body behavior and classify the intention of each continuous frame sequence. Based on the model's output, generate a label of body movement intention; Based on the acquired user voice data, the voice data is converted into text data through speech recognition, and the text data is parsed through natural language processing technology to extract the core semantics and obtain the voice intent identifier.
[0024] Specifically, the limb movements in each frame are represented as a graph structure, where each key point of the skeleton is a node in the graph, and the connections between nodes represent the relationships between different body parts; the skeletal movements in each frame are mapped as a spatiotemporal graph, and the limb movements at each moment contain not only spatial information but also temporal information. The core idea of spatiotemporal graph convolutional networks is to capture both spatial and temporal information simultaneously. Spatial convolution learns the local spatial features of each node, while temporal convolution learns the dynamic changes between consecutive frames. For each time step, spatial convolution is used to aggregate information from neighboring nodes, learning the local spatial features of each skeletal keypoint. The formula is as follows: in, Let be the output feature of the k-th layer, representing the spatial features of node t. Let be the set of neighboring nodes of node t. Let be the weight matrix of the k-th layer. For bias terms, For activation functions; By combining spatial and temporal convolution, the spatiotemporal features of the nodes are obtained. The spatial and temporal features are then combined through simple concatenation or weighted fusion operations to serve as the final node feature representation. By using a pre-trained spatiotemporal graph convolutional network, each action sequence is classified into intentions. Based on the category output by the model, a corresponding intention label is generated for the user's physical behavior, with each category corresponding to a specific action intention. The microphone is used to collect the user's voice data, which is then converted into text data through speech recognition. Natural language processing technology is used to parse the voice text and extract the core semantics. Named entity recognition, part-of-speech tagging and other methods are used to extract key intent information. The results of body language intent recognition and voice intent recognition are combined to confirm the final intent.
[0025] See Figure 6 As shown, the consistency between body language intent identifiers and voice intent identifiers is compared. If they match, the interaction is considered correct. If a conflict exists, an error correction process is triggered, and a conflict label is generated. Specifically, this includes: Consistency is determined based on the acquired body language intent identifiers and voice intent identifiers. If the body language intent matches the voice intent... Figure 1 The interaction was deemed correct. If the intention of the body language is inconsistent with the intention of the voice, it is judged as a conflict, and an intention type label and a conflict label are added to the body language intention label and the voice intention label. The error correction process is triggered based on body language and voice data with added conflict tags.
[0026] Specifically, in cases of conflict, conflict labels are added to both the physical intention and the voice intention to indicate that the intentions are inconsistent. For conflicting intention pairs, an error correction process needs to be triggered to check whether the voice intention has sufficient contextual support, whether the conflict is caused by environmental noise or speech recognition errors, whether the same physical action may correspond to multiple intentions, resulting in inconsistency with the voice intention, and to check the errors of the speech recognition or physical action recognition models.
[0027] See Figure 7 As shown, based on data with conflict labels, body language data and speech data are compared with the English communicative intent of the target grammatical structures and vocabulary sets in the teaching task instructions. Corrective instructions are generated based on the conflict type when the intents are inconsistent. Specifically, these include: Based on teaching task instructions with conflicting labels, the English communicative intent of the teaching task instructions is obtained through a body behavior-English intent mapping library. Acquire body action data and voice data with conflicting labels, including body action intent identifiers and voice intent identifiers; By comparing body language intent markers, voice intent markers, and English communicative intent, intent markers that do not match English communicative intent are obtained; Based on the mismatched intent identifiers, corrective instructions are generated to guide users to make corrections in their physical actions or speech.
[0028] Specifically, for data with conflicting labels, the body language intent and voice intent identifiers for each data point are obtained. Using a mapping library, the communicative intent related to body language and voice signals is parsed out. By comparing the body language intent identifiers and voice intent identifiers with the predetermined English communicative intent, mismatched intent identifiers are identified. Based on the mismatched intent identifiers, correction instructions are generated to guide users to correct their physical behavior or voice. If a voice intent error is identified, it can be corrected through voice re-recognition, contextual understanding, or manual confirmation. If a physical intent error is identified, it can be corrected by increasing the sensitivity of the action or analyzing the details of the action. The error correction results are fed back to the system to optimize the model and recognition accuracy. The generated correction instructions are based on mismatched intent identifiers and help users make corrections through reasonable guidance, ensuring that the final body language and speech are consistent with the target English communicative intent.
[0029] See Figure 8 As shown, the teaching tasks are dynamically adjusted based on the user's historical interaction accuracy. By adding visual interference variables and speech noise, the single-action task is upgraded to a complex communicative behavior task until all teaching task instructions are completed. Specifically, this includes: Based on the user's past interactions, the accuracy, response time, and error types of task completion data are obtained. Based on the user's interaction accuracy, multi-level task difficulty is set by introducing different amounts of visual interference variables and voice noise; By setting complex communicative behavioral tasks, users can be taught through multi-channel interactive methods. During the teaching process of multi-level task difficulty, the accuracy and reaction time of users are monitored in real time, and the difficulty and complexity of the tasks are dynamically adjusted. Teaching is carried out through gradual difficulty adjustment. When the user successfully completes all task instructions, provide visual and verbal feedback to the user and collect the user's teaching feedback.
[0030] Specifically, the execution results of each task are verified, the correct completion rate is calculated, the time required for the user to complete the task from the start of execution is recorded, and the errors are classified and statistically analyzed according to different error types during task execution. Introduce varying amounts of visual interference into the task to increase its difficulty. This can be achieved by altering the complexity of images, animations, colors, or backgrounds in the task to impair the user's visual processing ability. Adding background noise to the task can simulate interfering sounds in the real environment. Speech noise can increase the difficulty of the task, especially for tasks that require processing voice commands. To make the tasks more relevant to real-world applications, we designed composite communicative behavior tasks. These tasks require interaction using multiple channels of information, such as voice, vision, and motion. For example, users need to complete instructions using both voice and motion in the presence of noise and visual interference. The difficulty of the task is adjusted in real time based on the user's accuracy and response time. For example, if the user has a high accuracy rate on a certain task, the intensity of visual interference or voice noise is increased to increase the difficulty of the task; conversely, the interference is reduced or the complexity of the task is decreased. As users improve their task completion rate, the difficulty of the tasks is gradually increased to ensure that users are constantly challenged during the learning process without it becoming too difficult. Whenever a user successfully completes a task, positive visual or auditory feedback is given to enhance the learning effect. After the user completes all task instructions, collect user feedback on the task to further optimize the design.
[0031] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0032] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0033] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An English teaching method based on motion-sensing interaction, characterized in that, include: The system simultaneously collects user's physical and vocal data using a camera and microphone. Based on the teaching topic selected by the user, the system matches the corresponding grammatical structure and core vocabulary set from the database to generate dynamic teaching task instructions. Construct a body language-English intention mapping library to map real-time collected body movements into English communicative intention data; Users perform physical actions that match the target grammatical structure and core vocabulary according to the teaching task instructions, and simultaneously output the corresponding English speech. Based on spatiotemporal graph convolutional network, continuous frame action sequences of user body behavior data are analyzed, action spatiotemporal features are extracted and body behavior intention identifiers are output, and speech intention identifiers are generated through speech recognition and semantic parsing. Compare the consistency between the body language intent identifier and the voice intent identifier. If they match, the interaction is considered correct. If there is a conflict, the error correction process is triggered and a conflict label is generated. Based on data with conflict labels, the body behavior data and voice data are compared with the English communicative intent of the target grammatical structure and vocabulary set in the teaching task instructions, and correction instructions are generated based on the conflict type according to the inconsistency of intent. The teaching tasks are dynamically adjusted based on the accuracy of users' historical interactions. By adding visual interference variables and speech noise, single-action tasks are upgraded to complex communicative behavior tasks until all teaching task instructions are completed.
2. The English teaching method based on motion-sensing interaction according to claim 1, characterized in that, The process of generating dynamic teaching task instructions based on user-selected teaching topics by matching corresponding grammatical structures and core vocabulary sets from the database specifically includes: Based on the teaching topic selected by the user, three layers of related data are extracted from the database, including grammar layer, vocabulary layer and scenario layer; The grammar layer contains frequently used grammatical structures for the teaching topics selected by the user, the vocabulary layer contains a core vocabulary set related to the topic, and the scenario layer contains templates for typical communicative behaviors. Based on the acquired grammatical structure and core vocabulary set, teaching task instructions for various scenarios are generated, and these instructions are updated in real time.
3. The English teaching method based on motion-sensing interaction according to claim 1, characterized in that, The construction of the body language-English intention mapping library, which maps real-time collected body movements to English communicative intention data, specifically includes: Based on historical data of interactive English teaching, a behavior-English intention mapping library was constructed. The collected body movements are used to construct a mapping model based on the relationship between the characteristics of the body movements and the English intentions, thus linking the body movements with the English intentions. The mapping model is trained based on historical data, and the trained model is used to analyze body movements in real time and provide English intentions.
4. The English teaching method based on motion-sensing interaction according to claim 1, characterized in that, The user, based on the teaching task instructions, performs physical actions that match the target grammatical structure and core vocabulary, and simultaneously outputs corresponding English speech. Specifically, this includes: Based on the generated dynamic teaching task instructions, the instructions are conveyed to the user through the display screen; Users perform corresponding physical actions based on the grammatical structures and core vocabulary required in the teaching task instructions, and verbally output English sounds that match the actions. The camera and microphone collect user's body movement and voice data in real time.
5. The English teaching method based on motion-sensing interaction according to claim 4, characterized in that, The method of analyzing continuous frame action sequences of user body behavior data based on spatiotemporal graph convolutional networks, extracting spatiotemporal features of actions and outputting body behavior intention identifiers, and generating speech intention identifiers through speech recognition and semantic parsing specifically includes: Based on the acquired user body behavior data, a spatiotemporal graph convolutional network model is constructed and trained. Each frame's limb movement is represented as a graph structure, with the key points of the skeleton as nodes in the graph, and the connections between nodes representing the relationships between body parts. Spatiotemporal graph convolutional layers are used to capture spatial and temporal information. Spatial convolution learns the local spatial features of each node, while temporal convolution learns the dynamic changes between consecutive frames. Based on the acquired spatiotemporal features, the model is trained, and the trained spatiotemporal graph convolutional network model is used to identify the intention of body behavior and classify the intention of each continuous frame sequence. Based on the model's output, generate a label of body movement intention; Based on the acquired user voice data, the voice data is converted into text data through speech recognition, and the text data is parsed through natural language processing technology to extract the core semantics and obtain the voice intent identifier.
6. The English teaching method based on motion-sensing interaction according to claim 1, characterized in that, The process of comparing the consistency between the body language intent identifier and the voice intent identifier, and determining that the interaction is correct if they match, and triggering an error correction process and generating a conflict label if there is a conflict, specifically includes: Consistency is determined based on the acquired body movement intent identifier and voice intent identifier. If the body movement and voice intent are consistent, the interaction is considered correct. If the intention of the body language is inconsistent with the intention of the voice, it is judged as a conflict, and an intention type label and a conflict label are added to the body language intention label and the voice intention label. The error correction process is triggered based on body language and voice data with added conflict tags.
7. The English teaching method based on motion-sensing interaction according to claim 1, characterized in that, The data based on conflict labels compares body language and speech data with the English communicative intent of the target grammatical structures and vocabulary sets in the teaching task instructions, and generates correction instructions based on the conflict type when the intents are inconsistent. Specifically, this includes: Based on teaching task instructions with conflicting labels, the English communicative intent of the teaching task instructions is obtained through a body behavior-English intent mapping library. Acquire body action data and voice data with conflicting labels, including body action intent identifiers and voice intent identifiers; By comparing body language intent markers, voice intent markers, and English communicative intent, intent markers that do not match English communicative intent are obtained; Based on the mismatched intent identifiers, corrective instructions are generated to guide users to make corrections in their physical actions or speech.
8. The English teaching method based on motion-sensing interaction according to claim 1, characterized in that, The method of dynamically adjusting teaching tasks based on the accuracy of users' historical interactions, and upgrading single-action tasks into complex communicative behavior tasks by adding visual interference variables and speech noise, until all teaching task instructions are completed, specifically includes: Based on the user's past interactions, the accuracy, response time, and error types of task completion data are obtained. Based on the user's interaction accuracy, multi-level task difficulty is set by introducing different amounts of visual interference variables and voice noise; By setting complex communicative behavioral tasks, users can be taught through multi-channel interactive methods. During the teaching process of multi-level task difficulty, the accuracy and reaction time of users are monitored in real time, and the difficulty and complexity of the tasks are dynamically adjusted. Teaching is carried out through gradual difficulty adjustment. When the user successfully completes all task instructions, provide visual and verbal feedback to the user and collect the user's teaching feedback.
Citation Information
Cited By
Digital human interaction method and system based on gesture and body movement recognition
CN122152124A