Multi-modal English learning interaction system and vocabulary memory training method

Through the multimodal English learning interactive system, collect and analyze learners' multimodal behavior data, group division and teaching strategy adjustments, and push customized resources, solving the problem of insufficient personalized teaching and learning interaction in traditional English learning technology, and improving learning effect and participation.

CN120580898AInactive Publication Date: 2025-09-02XINXIANG VOCATIONAL & TECHN COLLEGE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510756863.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-02
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional English learning technology has shortcomings in multimodal learning needs, personalized teaching and learning interactivity, and cannot effectively meet the personalized needs of learners. It lacks collection and analysis of multimodal learning behavior data, lacks targeted teaching resource push, and lacks learners' participation and enthusiasm.

Method used

A multimodal English learning interactive system is designed to collect learners' multimodal behavior data through the data acquisition module, use group intelligence algorithms to perform group division and behavioral pattern analysis, automatically adjust teaching strategies, push customized multimodal learning resources, and enhance learner participation through the interactive learning module.

Benefits of technology

The teaching strategy adjustments are achieved according to the learners' personalized needs, the learning effect and participation are improved, a closed-loop learning optimization process is formed, and the learners' vocabulary memory ability is continuously improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580898A_ABST
    Figure CN120580898A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal English learning interaction system and a vocabulary memory training method, and relates to the technical field of English learning, the system comprises the following components: a data acquisition module, a data analysis module, a strategy adjustment module, a resource push module and an interaction learning module; multi-modal learning behavior data, including text input, voice reading, handwritten notes, video learning behaviors, interactive operation and the like, of learners are collected through the data acquisition module, the learners are subjected to group division by applying a group intelligent algorithm, and the behavior pattern and performance of each group in vocabulary learning are analyzed for each group, so that the learning efficiency of the learners is improved. Based on the analysis, the system can automatically adjust teaching strategies and push customized multi-modal learning resources and training methods, so that personalized requirements of different learners are met, and the learning effect and experience are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of English learning, and in particular to a multimodal English learning interactive system and a vocabulary memory training method. Background Art

[0002] With the acceleration of globalization and the popularization of English as an international language, the demand for English learning is growing, especially in vocabulary memorization. Traditional English learning methods often focus on single-modal learning, such as memorizing vocabulary only through book reading or listening practice. This method ignores individual differences among learners and multi-sensory participation in the learning process. In recent years, with the rapid development of information technology, the concept of multimodal learning has gradually been introduced into the field of education, emphasizing the integration of multiple modalities such as text, voice, images, and video to provide a richer and more interactive learning experience to stimulate learners' interest and improve learning outcomes.

[0003] However, traditional English learning technologies seem to be unable to cope with the needs of multimodal learning. First, traditional technologies often lack the ability to comprehensively collect and analyze learners' multimodal learning behavior data, and cannot accurately grasp learners' learning styles, interest preferences and ability levels, resulting in a lack of targeted push of teaching resources. Secondly, traditional teaching strategy adjustments mostly rely on teachers' experience and judgment, lack scientific data support, and it is difficult to achieve personalized teaching. Thirdly, the traditional learning system lacks interactivity, and learners are often in a state of passively accepting knowledge and lack opportunities for active exploration and communication, which affects their learning enthusiasm and participation.

[0004] In summary, traditional English learning technology has obvious shortcomings in meeting multimodal learning needs, realizing personalized teaching, and enhancing learning interactivity. Therefore, it is particularly important to develop a multimodal English learning interactive system and vocabulary memory training method. Summary of the Invention

[0005] The purpose of the present invention is to make up for the shortcomings of the existing technology and provide a multimodal English learning interactive system and vocabulary memory training method. It can collect learners' multimodal learning behavior data through a data acquisition module, use a swarm intelligence algorithm to divide learners into groups, and analyze the behavior patterns and performance of each group in vocabulary learning. Based on these analyses, the system can automatically adjust teaching strategies and push customized multimodal learning resources and training methods to meet the personalized needs of different learners. At the same time, the system also enhances learners' learning enthusiasm and participation through interactive learning modules.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: On the one hand, a multimodal English learning interactive system includes the following components: a data acquisition module, a data analysis module, a strategy adjustment module, a resource push module and an interactive learning module;

[0007] The data collection module is used to collect multimodal learning behavior data of all learners, including but not limited to text input data, voice reading data, handwritten note data, video learning behavior data, and interactive operation data with learning content;

[0008] The data analysis module uses a swarm intelligence algorithm to analyze the collected multimodal learning behavior data. This module first divides learners into groups based on their age, learning level, and learning style. Then, for each group, it analyzes their behavior patterns and performance in learning various types of vocabulary, thereby identifying common problems that exist in certain vocabulary learning categories.

[0009] The strategy adjustment module automatically adjusts the teaching strategy when the data analysis module finds that a certain group has common problems in learning a certain type of vocabulary. The adjustment of the teaching strategy includes but is not limited to adjusting the presentation of vocabulary learning content, the difficulty and sequence of learning tasks, and the arrangement of learning time.

[0010] The resource push module pushes customized multimodal learning resources and training methods to groups with common problems based on the adjusted teaching strategy. The customized multimodal learning resources include English vocabulary learning materials in various forms such as text, pictures, audio, and video. The training methods include diversified exercise forms to meet the learning needs of different groups of learners and improve vocabulary memorization effects.

[0011] The interactive learning module supports interaction between learners and between learners and the system. Learners can conduct online discussions and share learning experiences and skills with each other in the system. The system can also provide real-time feedback and evaluation based on the learners' learning situation and interactive performance, thereby enhancing learners' learning enthusiasm and participation.

[0012] Furthermore, the multimodal learning behavior data collected by the data acquisition module include text input data, voice reading data, handwritten note data, video learning behavior data and interactive operation data with learning content, wherein the text input data is obtained by the learner entering English text in the system input box, and the handwritten text is converted into analyzable character data using optical character recognition technology. The voice reading data is collected through a microphone and analyzed after being converted into text by voice recognition technology. The handwritten note data records the coordinates, pressure, and writing speed information of the handwriting through an electronic handwriting device. The video learning behavior data includes the video viewing time, the number of pauses and time points, and fast forward and rewind operations, which are recorded through the operation log of the video playback module. The interactive operation data with the learning content, such as clicking on learning materials and answering questions, are recorded in real time through the system event listener. The system preprocesses the collected raw data, including noise removal and data standardization operations to ensure data quality.

[0013] Furthermore, the swarm intelligence algorithm used in the data analysis module is specifically an improved dynamic ant colony-particle swarm hybrid algorithm. This algorithm combines the positive feedback mechanism of the ant colony algorithm and the fast convergence characteristics of the particle swarm algorithm to divide learners into groups and analyze common problems in vocabulary learning. In the algorithm, let the number of ant colonies be m, the number of particle swarms be n, and the dimension of learning behavior data be d. In the group division stage, define the position vector X of each ant ij =(x ij1 ,x ij2 ,…,x ijd ) represents the value of the i-th ant on the j-th dimension data. Its initial position is randomly generated. During the search process, the ant moves according to the pheromone concentration τ ij and heuristic information η ij Select the next position, the transition probability formula is:

[0014]

[0015] Among them, α and β are parameters that control the relative importance of pheromone and heuristic information, which are determined by training and optimization of historical learning data. k represents the next position set that ant k can choose. At the same time, the speed update formula of the particle swarm algorithm is introduced:

[0016] v ij (t+1)=ωv ij (t)+c1r 1j (t)(p ij -x ij (t))+c2r 2j (t)(g j -x ij (t))

[0017] Where ω is the inertia weight, which is dynamically adjusted with the number of iterations. The initial value is set to 0.9 and is reduced by 0.01 every 10 iterations until it reaches 0.4. c1 and c2 are learning factors, which are determined by cross-validation to be c1 = 1.5, c2 = 1.5, r 1j (t) and r 2j (t) is a random number in the interval [0, 1], p ij is the historical optimal position of particle i, g i For the global optimal position of the entire particle swarm, in the phase of analyzing common problems in vocabulary learning, based on the group division results, the mean and variance of each group on various vocabulary learning indicators are calculated to determine the common problems.

[0018] Furthermore, when the data analysis module uses the improved dynamic ant colony-particle swarm hybrid algorithm, it also dynamically and adaptively adjusts the algorithm. During the execution of the algorithm, it monitors the accuracy indicators of group division and problem analysis in real time. When the accuracy indicator increases by less than 5% for five consecutive iterations, the number of ant colonies m and the number of particle swarms n are automatically increased by 10% of the current number. At the same time, the values ​​of α and β are adjusted, α is increased by 0.1 and β is decreased by 0.1 to enhance the algorithm's processing ability for complex learning behavior data. In addition, through cluster analysis of historical learning data, an association model between different learning behavior patterns and algorithm parameters is established. When a new learning behavior pattern is detected, the algorithm parameters are automatically adjusted according to the association model to improve the adaptability and analysis accuracy of the algorithm.

[0019] Furthermore, when adjusting the teaching strategy, the strategy adjustment module adopts a strategy optimization model based on a Bayesian network. The model uses the age, learning level, and learning style attributes of the learner group as nodes, and common vocabulary learning problems and teaching strategy adjustment plans as nodes to construct a Bayesian network structure. The Bayesian network is trained through historical learning data and teaching strategy adjustment effect data to determine the conditional probability distribution between each node. When it is found that a certain group has common problems in a certain type of vocabulary learning, the attribute information and problem information of the group are input into the Bayesian network, and the probabilities of different teaching strategy adjustment plans are calculated. The plan with the highest probability is selected as the adjusted teaching strategy. At the same time, the model also updates the parameters of the Bayesian network based on the feedback information of the learners in the subsequent learning process, and continuously optimizes the accuracy of the teaching strategy adjustment.

[0020] Furthermore, when pushing customized multimodal learning resources, the resource push module adopts a resource recommendation algorithm based on semantic association. The algorithm first semantically annotates the multimodal learning resources, uses natural language processing technology to extract keywords and themes from text resources, extracts visual features of picture resources through image recognition technology and converts them into semantic descriptions, and extracts semantic information after voice-to-text conversion of audio and video resources. Then, a semantic association network is constructed, with words as nodes and the semantic similarity between words and the association between resources and words as edge weights. When resources need to be pushed to a certain group, relevant resources are searched in the semantic association network based on the common vocabulary learning problems and learning goals of the group. Let the target word be v, and the semantic association calculation formula between resource r and v is:

[0021]

[0022] Where k is the dimension of semantic association, λ i is the weight of each dimension, which is determined by regression analysis of learners’ resource usage feedback data. i (v, r) is the semantic similarity between resource r and vocabulary v in the i-th dimension. Resources are pushed to the group in descending order of semantic relevance, and the order and content of resource push are dynamically adjusted according to the group's learning progress and feedback.

[0023] Furthermore, the interactive learning module supports multiple forms of interaction, including online discussion areas, learning group tasks, and real-time voice and video communication. In the online discussion area, the system uses sentiment analysis technology to judge the emotional tendency of learners' speeches. If negative emotional speeches are detected, encouraging information and relevant learning resources are automatically pushed. For learning group tasks, a collaborative learning algorithm based on task allocation optimization is adopted. This algorithm considers learners' learning ability and interests and hobbies, decomposes group tasks into subtasks, and allocates tasks through the Hungarian algorithm to minimize the total cost of task allocation. Suppose the learner set is L = {l1,l2,…,l n}, the subtask set is T = {t1, t2, ..., t m}, learner l i Complete subtask t j The cost is c ij , then the task allocation problem is transformed into solving The smallest allocation solution where x ij is a 0-1 variable. When the learner l i Assign to subtask t j When x ij =1, otherwise x ij =0. At the same time, the system evaluates and rewards based on the completion of group tasks and member contributions, encouraging learners to actively participate in dynamic learning.

[0024] Furthermore, the system also has a personalized learning path planning function. This function uses a reinforcement learning algorithm to plan a personalized learning path based on the learner's individual learning behavior data, learning goals, and group learning situation. The state space is defined as the learner's current learning state, and the action space is defined as the optional learning resources and training methods. The reward function is designed according to the learner's performance in the learning process. By continuously interacting with the learning environment, the optimal action is selected to maximize the cumulative reward, thereby planning a personalized learning path for the learner. During the learning process, the learning path is dynamically adjusted according to the learner's actual learning situation to ensure that the learning path always meets the learner's needs and learning progress.

[0025] On the other hand, a vocabulary memory training method for a multimodal English learning interactive system comprises the following specific steps:

[0026] S1. Data collection step: Continuously collect multimodal learning behavior data of all learners through the data collection module;

[0027] S2. Group division and data analysis: The data analysis module uses swarm intelligence algorithms to process the collected data. First, learners are divided into groups based on factors such as age and learning level. Then, the behavior patterns and performance of each group in learning various vocabulary items are analyzed to identify common problems.

[0028] S3, teaching strategy adjustment step: When a group of students is found to have common problems in learning a certain type of vocabulary, the strategy adjustment module automatically adjusts the teaching strategy to determine the presentation method, learning task difficulty and sequence that are suitable for the group;

[0029] S4. Resource and method push step: The resource push module pushes customized multimodal learning resources and training methods to the group based on the adjusted teaching strategy, guiding learners to conduct vocabulary memory training;

[0030] S5. Learning effect evaluation and feedback steps: The system collects learners' performance and feedback information during the training process through the interactive learning module, evaluates the learning effect, and further optimizes the teaching strategy and learning resource push based on the evaluation results, forming a closed-loop learning optimization process.

[0031] Compared with the existing technology, this multimodal English learning interactive system and vocabulary memory training method have the following beneficial effects:

[0032] 1. The system collects learners' multimodal learning behavior data through the data acquisition module, including text input, voice reading, handwritten notes, video learning behavior and interactive operations, and uses swarm intelligence algorithms to divide learners into groups. It then analyzes the behavior patterns and performance of each group in vocabulary learning. Based on these analyses, the system can automatically adjust teaching strategies and push customized multimodal learning resources and training methods to meet the personalized needs of different learners, significantly improving learning outcomes and experience.

[0033] Second, the system collects learners' performance and feedback information during training through interactive learning modules and evaluates learning outcomes. Based on the evaluation results, the system can further optimize teaching strategies and learning resource delivery, forming a closed-loop learning optimization process. This continuous optimization mechanism helps learners continuously identify their own shortcomings, adjust their learning strategies, and achieve continuous improvement in vocabulary memory ability.

[0034] Other advantages, objects and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art based on an examination of the following or may be learned from the practice of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0036] Figure 1 This is a process diagram of a multimodal English learning interactive system;

[0037] Figure 2 This is a flowchart of a vocabulary memory training method for a multimodal English learning interactive system. DETAILED DESCRIPTION

[0038] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.

[0039] Example 1

[0040] The data collection module continuously captures students' multi-dimensional behaviors in English learning:

[0041] In writing exercises, the text data entered by students through the keyboard were recorded, such as the frequency of incorrect use of “lookdeliciously” in the composition;

[0042] Using classroom microphones to collect speech reading data, we analyzed the pronunciation accuracy of the word "sound" in different contexts (e.g., intonation differences between interrogative and declarative sentences);

[0043] Using an electronic writing board to track handwritten note data, we found that students frequently altered the collocations of "taste" when copying verb phrases (e.g., repeatedly erasing "tastesweetly" and correcting it to "tastesweet");

[0044] The video playback module logs captured students' operational traces when watching verb explanation videos—62% of students paused more than three times in the "Grammar Rules of Sensory Verbs" segment, and 48% fast-forwarded to the "Example Sentence Demonstration" section. The raw data was pre-processed through system denoising and standardization to form a structured data set suitable for analysis.

[0045] The data analysis module starts the improved dynamic ant colony-particle swarm hybrid algorithm. The formula is: let the number of ant colonies be m, the number of particle swarms be n, the dimension of learning behavior data be d, and in the group division stage, define the position vector X of each ant ij =(x ij1 ,x ij2 ,…,x ijd ) represents the value of the i-th ant on the j-th dimension data. Its initial position is randomly generated. During the search process, the ant moves according to the pheromone concentration τ ij and heuristic information η ij Select the next position, the transition probability formula is:

[0046]

[0047] Where α and β are parameters that control the relative importance of pheromone and heuristic information, allowed k represents the set of next positions that ant k can choose. Students are divided into three subgroups based on age, learning level (based on the scores of the last three tests), and learning style (visual and auditory preferences identified through questionnaires):

[0048] Weak foundation group (about 25%): high spelling error rate and obvious speech delay;

[0049] Comprehension deviation group (about 30%): can recognize the meaning of words but cannot apply them correctly in context;

[0050] Memory confusion group (approximately 45%): This group showed confusion in the meaning of "sensory verbs," such as mixing up "smellgood" and "smellwell." Further analysis of the "memory confusion group" revealed an error rate of 68% in the "sensory verb + adjective" structure, and a lack of awareness of the implicit emotional color of verbs (e.g., failing to distinguish the semantic difference between "lookworried" and "lookworrying").

[0051] The strategy adjustment module uses the Bayesian network strategy optimization model to generate a three-tiered optimization strategy based on the group attributes (around 14 years old, medium learning level, and 58% visual learners) and the question type (word meaning confusion):

[0052] Innovation in presentation: Converting static grammar tables into dynamic animations—for example, using animated characters tasting fruit to demonstrate the correct use of "taste," with color changes highlighting the position of adjectives (e.g., the incorrect usage of "apple tastes red" versus the correct usage of "apple tastes sweet").

[0053] Restructuring the task sequence: Design a three-level task chain: "Word meaning matching → contextual filling-in-the-blank → scenario sentence creation." First, strengthen visual memory through picture matching, then deepen comprehension through short passage filling-in-the-blank, and finally create sentences based on a given scenario (e.g., "Describe a garden after rain").

[0054] Adjust the time rhythm: embed verb-specific training into the 10 minutes before morning self-study (taking advantage of the golden period of short-term memory), and push a 5-minute fun animation to review key knowledge points during the lunch break.

[0055] The resource push module relies on the semantic association resource recommendation algorithm. The formula is: Let the target word be v, and the semantic association calculation formula between resource r and v is:

[0056]

[0057] Where k is the dimension of semantic association, λ is the weight of each dimension, and s i (v, r) is the semantic similarity between resource r and vocabulary v in the i-th dimension. Resources are pushed to the group in descending order of semantic relevance. The order and content of resource push are dynamically adjusted based on the group's learning progress and feedback. A three-dimensional learning package is customized for the "memory confusion group":

[0058] Text resources: "A Handbook of Commonly Confused Sensory Verbs" includes analysis of frequently used errors (e.g., the difference between the physical and emotional perceptions of "feel" and "touch") and includes mnemonics ("sensory verb + adjective, modifying the subject, remember it");

[0059] Audio resources: A "Verb Pronunciation Analysis" micro-lecture recorded by a foreign teacher, comparing the stress differences between "sound" in "sound interesting" and "sound anoise";

[0060] Video resource: A clip from the campus sitcom "English Mini-Theatre," in which actors interpret scenes like "looktired" and "smellnice" through facial expressions and actions, while corresponding English sentences scroll in real time at the bottom of the screen;

[0061] Training method: Adopt the "flash card + context challenge" mode - the front of the flash card is a verb picture, and the back is a Chinese and English example sentence. The system dynamically adjusts the frequency of card appearance based on the student's correct answer rate. Cards with a high error rate will be repeated every 3 minutes.

[0062] The interactive learning module builds a "verb exploration community" to stimulate group collaboration and individual growth:

[0063] Group Task Design: Using a collaborative learning algorithm based on task allocation optimization, the task of "co-building a sensory verb knowledge base" was broken down into "drawing word meaning illustrations" (for visually oriented students), "recording example sentences" (for auditory oriented students), and "organizing error cases" (for analytical oriented students). The Hungarian algorithm was then used to achieve the optimal match between tasks and abilities.

[0064] Emotional intelligence intervention: The system uses sentiment analysis technology to scan discussion forum posts. When negative sentiment is detected, it automatically pushes an encouraging message, "You've mastered 80% of the verbs, just one step left!" and attaches a link to a simplified memory game.

[0065] Results tracking and iteration: Two weeks later, testing showed that the verb confusion error rate for this group had dropped to 22%. The system further optimized video resources based on the test-taking data, expanding the scene demonstrations of "abstract emotional verbs" from indoors to outdoors, thereby enriching the memory cues.

[0066] Example 2

[0067] The data collection module comprehensively records students’ actual combat performance:

[0068] In the simulated negotiation phase, the conversation content was transcribed using speech recognition technology to analyze the frequency and tone of use of key words such as "quotation" and "delivery date";

[0069] Tracking student behavior while watching the video "International Business Negotiation Case Library" revealed that 83% of students fast-forwarded through the segment "Customer Complaints about Quality Issues" and spent less than two minutes on the segment "Interpretation of Contract Terms."

[0070] When collecting electronic note data, we found that when students recorded terms such as "negotiation strategy," they only annotated the Chinese definitions and lacked scenario-based application examples.

[0071] Interactive operation data shows that the average number of times students click on the "Business Terminology Colloquial Conversion Tool" is only 1.2 times per day, and the usage rate is significantly low.

[0072] The data analysis module uses an improved dynamic ant colony-particle swarm hybrid algorithm. The formula is: let the number of ant colonies be m, the number of particle swarms be n, and the dimension of learning behavior data be d. In the group division stage, define the position vector X of each ant ij =(x ij1 ,x ij2 ,…,x ijd ) represents the value of the i-th ant on the j-th dimension data. Its initial position is randomly generated. During the search process, the ant moves according to the pheromone concentration τ ij and heuristic information η ij Select the next position, the transition probability formula is:

[0073]

[0074] Where α and β are parameters that control the relative importance of pheromone and heuristic information, allowed k The set of possible next positions for ant k is divided into groups based on professional experience (≥5 years), learning goals (business negotiation practice), and learning style (80% are practical learners):

[0075] The terminology group (55%) is used to translating written language directly, such as directly translating "underseparate cover" as "under a separate cover" instead of "sent in a separate letter";

[0076] Logic confusion group (45%): When responding to customer objections, they lack a structured expression framework of "acknowledge → explanation → solution".

[0077] Focusing on the "stiff terminology group", it was found that in the "price negotiation" scenario, the frequency of use of colloquial variants of terms such as "offer" and "counter-offer" (such as "Let's make an offer" and "Do you think this price is acceptable") was less than 20%, resulting in a stiff communication atmosphere.

[0078] The strategy adjustment module relies on the Bayesian network strategy optimization model, combines group characteristics and problem pain points, and formulates scenario-based strategies:

[0079] Transformation of content presentation: Introducing "Comparative Analysis of Negotiation Recordings"—playing both written and spoken communication from the same scenario ("We understand your concerns, but we really can't do this right now"), and highlighting the role of modal particles (such as "hmm" and "Is that so") in improving communication fluency.

[0080] Reconstructing the task scenario: Designing an "elevator negotiation" micro-task—requiring students to use colloquial language to introduce the advantages of a new product to a "customer" (a virtual character in the system) within 90 seconds, strengthening their ability to express themselves accurately in a short period of time;

[0081] Time management optimization: Take advantage of the gaps in fragmented sales work (such as the 15 minutes when customers are waiting for a reply) to push "business speech blind boxes" - each time randomly presenting three examples of spoken expressions for a scenario (such as "politely decline a price reduction request").

[0082] The resource push module uses the semantic association resource recommendation algorithm. The formula is: Let the target word be v, and the semantic association calculation formula between resource r and v is:

[0083]

[0084] Where k is the dimension of semantic association, λ i is the weight of each dimension, s i (v, r) is the semantic similarity between resource r and vocabulary v in the i-th dimension. Resources are pushed to the group in descending order of semantic relevance. The order and content of resource push are dynamically adjusted based on the group's learning progress and feedback. A scenario-based learning package is distributed to the "terminology-stupid group":

[0085] Audio resources: "100 Practical Examples of Cross-Border Sales Talk" includes high-frequency dialogues from real negotiations (e.g., "How to respond to a customer's 'price is too high'"), with emphasis and pause rhythms (e.g., "We really can't lower the price any further, but we can optimize the delivery time for you");

[0086] Video resources: A clip from "On-site Business Negotiation Recordings" shows how experienced salespeople use open-ended expressions like "Do you think this is feasible?" to guide clients toward consensus, with subtitles highlighting colloquial terminology.

[0087] Text resources: The "Business English Colloquial Conversion Guide" provides a "written language → spoken language" comparison table (e.g., "attached please find" corresponds to "there is in the attachment") and a "scenario quick reference table" (categorized by "quotation - objection handling - transaction");

[0088] Training method: Using an "instant feedback role-playing" model, trainees simulate negotiation responses through voice input. The system uses voice recognition technology to analyze the naturalness of terminology. If written expressions (such as "based on company regulations") are detected, a prompt will pop up immediately: "Try using 'Our current policy is', which is more colloquial."

[0089] The interactive learning module creates a "virtual negotiation hall" to simulate real business scenarios:

[0090] Intelligent collaboration mechanism: Using a collaborative learning algorithm based on task allocation optimization, we assigned roles for the group task "Planning a negotiation for a new product launch." Participants skilled in data analysis were assigned to organize product specifications (subtask 1), participants with fluent oral communication served as the lead negotiator (subtask 2), and participants with clear logic controlled the conversation logic (subtask 3). The Hungarian algorithm was used to ensure the shortest possible task duration.

[0091] Dual monitoring of emotions and efficiency: The system uses sentiment analysis technology to identify participants' anxiety signals during simulations and automatically delivers audio guidance on "deep breathing before negotiations." Points are awarded based on group task completion (e.g., achieving pre-set sales figures, terminology accuracy), which can be redeemed for opportunities to simulate offline business social scenarios.

[0092] Personalized path planning: Through reinforcement learning algorithms, a dynamic learning map is generated for each student. When the "stiff terminology" problem is improved to an error rate of less than 10%, the "Cross-Cultural Negotiation Strategy" module is automatically unlocked, and historical data is combined to recommend cases suitable for their negotiation style (for example, students who prefer a gentle communication style will prioritize studying "Southeast Asian Market Negotiation Records").

[0093] Example 3

[0094] Panoramic collection of multimodal data: The system records students' learning trajectory throughout the entire process: the text input module captures annotation data during literature reading, the voice report module uses voice recognition technology to analyze the pauses in the speech (such as the incoherent pronunciation of "convolutional neural network"), and the scanning of handwritten mind maps shows confusion in the conceptual hierarchy (such as the incorrect classification of "supervised learning" and "reinforcement learning" in the same category). The video learning module records the time nodes when students repeatedly watch the "Transformer architecture" demonstration video (an average of 4 minutes and 20 seconds each time), and the forum interaction data shows that "how to understand transfer learning" is a high-frequency question.

[0095] Accurate stratification of academic groups: The data analysis module uses an improved dynamic ant colony-particle swarm hybrid algorithm to divide students into four groups: "Computer major-graduate entrance examination-deep analysis type" and "Cross-disciplinary-employment-application-oriented type" according to their professional foundation (the top 40% are computer majors, and the bottom 30% are cross-disciplinary students), learning goals (academic English is required for postgraduate entrance examinations, and employment focuses on technical applications), and cognitive style (70% are logical analysis type). For the "cross-disciplinary-intermediate foundation-graduate entrance examination" group (28 people), analysis found that their professional vocabulary listening error rate reached 38%, the term collocation errors in the writing of literature abstracts accounted for 25%, and the compliance rate of English naming conventions in code comments was only 60%.

[0096] Intelligent adaptation of academic strategies: The strategy adjustment module is based on the Bayesian network model. Combined with the characteristics of this group (cross-disciplinary, needing to cope with postgraduate entrance examination professional English courses) and the type of questions (shallow understanding of terminology), it recommends the "academic scenario immersion + cross-modal linkage" strategy. The system adjusts the content presentation method to a "vocabulary-code-chart" trinity display (such as simultaneously presenting mathematical formula charts and Python code snippets when explaining "gradient descent"). The task difficulty is set as a progressive training of "paper abstract writing → academic report → code comment optimization". The learning time is scheduled for 3 hours of concentrated in-depth learning every Saturday morning.

[0097] Deep integration of academic resources: The resource push module builds an "artificial intelligence terminology knowledge network" through a semantic association recommendation algorithm to push:

[0098] Text resources: A parsed version of the abstract of the ACM conference paper "Neural Machine Translation" (highlighting high-frequency terms and including a "term-author perspective" correlation map);

[0099] Audio resources: Excerpts from the MIT Open Course "Frontiers of Deep Learning" lecture (with synchronized subtitles and terminology follow-up);

[0100] Video resources: A video of a "Vocabulary Speed ​​Challenge" simulating an international academic conference, focusing on the scenario-based applications of hot terms such as "Transformer" and "BERT";

[0101] Interactive tool: A code editor with integrated terminology recommendation. When students enter "model", it automatically prompts professional collocations such as "neuralnetworkmodel" and "pre-trainedmodel", and annotates the frequency of use and context examples.

[0102] Academic interaction and ability advancement: During real-time voice communication, the system provides real-time semantic correction for incorrect terminology used in student reports (such as misusing "reinforcement learning" in a supervised learning scenario), and pops up a prompt box saying "This term is more suitable for dynamic decision-making scenarios." Group collaboration tasks require the joint completion of the "AI terminology knowledge graph" construction. Using a task allocation model based on the Hungarian algorithm, "theoretical terminology organization" is assigned to computer science students, "interdisciplinary application case collection" is assigned to cross-disciplinary students, and "visualization design" is assigned to students who are good at chart making. The system generates an "academic ability radar chart" for each member based on the conceptual relevance of the graph (such as whether the subordinate relationship of "transfer learning-small sample learning" is correctly labeled) and the accuracy of the code examples.

[0103] Dynamic calibration of academic paths: Through analysis of student data using a reinforcement learning algorithm, it was found that after four weeks of training, the accuracy of terminology in literature abstracts for this group increased from 55% to 78%. However, the use of logical cohesive words in academic reports was insufficient (for example, the frequency of transition words such as "however" and "therefore" was 40% lower than the standard value). The system adjusted the learning path accordingly: a new special module "Logical Chain of Academic Writing" was added, and the structured expression task of "argument-evidence-conclusion" was incorporated into vocabulary training. Relevant chapters in "Academic Writing for Graduate Students" were recommended as expansion resources.

[0104] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as above in terms of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can, without departing from the scope of the technical solution of the present invention, make some changes or modifications to equivalent embodiments using the technical contents disclosed above. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A multimodal English learning interactive system, characterized by: The system includes the following components: data collection module, data analysis module, strategy adjustment module, resource push module and interactive learning module: The data collection module is used to collect multimodal learning behavior data of all learners, including but not limited to text input data, voice reading data, handwritten note data, video learning behavior data, and interactive operation data with learning content; The data analysis module uses a swarm intelligence algorithm to analyze the collected multimodal learning behavior data. This module first divides learners into groups based on their age, learning level, and learning style. Then, for each group, it analyzes their behavior patterns and performance in learning various types of vocabulary, thereby identifying common problems that exist in certain vocabulary learning categories. The strategy adjustment module automatically adjusts the teaching strategy when the data analysis module finds that a certain group has common problems in learning a certain type of vocabulary. The adjustment of the teaching strategy includes but is not limited to adjusting the presentation of vocabulary learning content, the difficulty and sequence of learning tasks, and the arrangement of learning time. The resource push module pushes customized multimodal learning resources and training methods to groups with common problems based on the adjusted teaching strategy. The customized multimodal learning resources include English vocabulary learning materials in various forms such as text, pictures, audio, and video, and the training methods include a variety of exercises. The interactive learning module supports interaction between learners and between learners and the system. Learners can conduct online discussions and share learning experiences and skills with each other in the system. The system can also provide real-time feedback and evaluation based on the learners' learning situation and interactive performance, thereby enhancing learners' learning enthusiasm and participation.

2. A multimodal English learning interactive system according to claim 1, characterized in that: The multimodal learning behavior data collected by the data acquisition module include text input data, voice reading data, handwritten note data, video learning behavior data and interactive operation data with learning content, among which, the text input data is obtained by the learner entering English text in the system input box, and the handwritten text is converted into analyzable character data using optical character recognition technology. The voice reading data is collected through a microphone and analyzed after being converted into text using voice recognition technology. The handwritten note data records the coordinates, pressure, and writing speed information of the handwriting through an electronic handwriting device. The video learning behavior data includes the video viewing time, the number of pauses and time points, and fast forward and rewind operations, which are recorded through the operation log of the video playback module. The interactive operation data with the learning content, such as clicking on learning materials and answering questions, are recorded in real time through the system event listener. The system preprocesses the collected raw data, including noise removal and data standardization operations.

3. A multimodal English learning interactive system according to claim 1, characterized in that: The swarm intelligence algorithm used in the data analysis module is specifically an improved dynamic ant colony-particle swarm hybrid algorithm. This algorithm combines the positive feedback mechanism of the ant colony algorithm and the fast convergence characteristics of the particle swarm algorithm to divide learners into groups and analyze common problems in vocabulary learning. In the algorithm, let the number of ant colonies be m, the number of particle swarms be n, and the dimension of learning behavior data be d. In the group division stage, the position vector X of each ant is defined as ij =(x ij1 ,x ij2 ,…,x ijd ) represents the value of the i-th ant on the j-th dimension data. Its initial position is randomly generated. During the search process, the ant moves according to the pheromone concentration τ ij and heuristic information η ij Select the next position, the transition probability formula is: Where α and β are parameters that control the relative importance of pheromone and heuristic information, allowerd k represents the next position set that ant k can choose. At the same time, the speed update formula of the particle swarm algorithm is introduced: v ij (t+1)=ωv ij (t)+c1r 1j (t)(p ij -x ij (t)+c2r 2j (t)(g j -x ij (t)) Where ω is the inertia weight, c1 and c2 are learning factors, r 1j (t) and r 2j (t) is a random number in the interval [0, 1], p ij is the historical optimal position of particle i, g i For the global optimal position of the entire particle swarm, in the phase of analyzing common problems in vocabulary learning, based on the group division results, the mean and variance of each group on various vocabulary learning indicators are calculated to determine the common problems.

4. A multimodal English learning interactive system according to claim 1, characterized in that: When using the improved dynamic ant colony-particle swarm hybrid algorithm, the data analysis module also dynamically and adaptively adjusts the algorithm. During the execution of the algorithm, the accuracy indicators of group division and problem analysis are monitored in real time. When the accuracy indicator increases by less than 5% for five consecutive iterations, the number of ant colonies m and the number of particle swarms n are automatically increased by 10% of the current number. At the same time, the values ​​of α and β are adjusted, α is increased by 0.1, and β is decreased by 0.

1. In addition, through cluster analysis of historical learning data, an association model between different learning behavior patterns and algorithm parameters is established. When a new learning behavior pattern is detected, the algorithm parameters are automatically adjusted according to the association model.

5. The multimodal English learning interactive system according to claim 1, characterized in that: When adjusting the teaching strategy, the strategy adjustment module adopts a strategy optimization model based on a Bayesian network. The model uses the age, learning level, and learning style attributes of the learner group as nodes, and common vocabulary learning problems and teaching strategy adjustment plans as nodes to construct a Bayesian network structure. When it is found that a certain group has common problems in a certain type of vocabulary learning, the group's attribute information and problem information are input into the Bayesian network, and the probabilities of different teaching strategy adjustment plans are calculated. The plan with the highest probability is selected as the adjusted teaching strategy. At the same time, the model also updates the parameters of the Bayesian network based on feedback information from learners in the subsequent learning process.

6. A multimodal English learning interactive system according to claim 1, characterized in that: When pushing customized multimodal learning resources, the resource push module adopts a resource recommendation algorithm based on semantic association. The algorithm first semantically annotates the multimodal learning resources, uses natural language processing technology to extract keywords and themes from text resources, extracts visual features of picture resources through image recognition technology and converts them into semantic descriptions, and extracts semantic information after performing speech-to-text conversion on audio and video resources. Then, a semantic association network is constructed, with words as nodes and the semantic similarity between words and the association between resources and words as edge weights. When resources need to be pushed to a certain group, relevant resources are searched in the semantic association network based on the common vocabulary learning problems and learning goals of the group. Let the target word be v, and the semantic association calculation formula between resource r and v is: Where k is the dimension of semantic association, λ i is the weight of each dimension, s i (v, r) is the semantic similarity between resource r and vocabulary v in the i-th dimension. Resources are pushed to the group in descending order of semantic relevance, and the order and content of resource push are dynamically adjusted according to the group's learning progress and feedback.

7. The multimodal English learning interactive system according to claim 1, characterized in that: The interactive learning module supports multiple forms of interaction, including online discussion areas, learning group tasks, and real-time voice and video communication. In the online discussion area, the system uses sentiment analysis technology to judge the emotional tendency of learners' speeches. If negative emotional speeches are detected, encouraging information and relevant learning resources are automatically pushed. For learning group tasks, a collaborative learning algorithm based on task allocation optimization is adopted. This algorithm considers learners' learning ability and interests and hobbies, decomposes group tasks into subtasks, and allocates tasks through the Hungarian algorithm to minimize the total cost of task allocation. Suppose the learner set is L = {l1,l2,…,l n }, the subtask set is T = {t1, t2, ..., t m }, learner l i Complete subtask t j The cost is c ij , then the task allocation problem is transformed into solving The smallest allocation solution where x ij is a 0-1 variable. When the learner l i Assign to subtask t j When x ij =1, otherwise x ij =0. At the same time, the system evaluates and rewards based on the completion of group tasks and member contributions, encouraging learners to actively participate in dynamic learning.

8. The multimodal English learning interactive system according to claim 1, characterized in that: The system also has a personalized learning path planning function. This function uses a reinforcement learning algorithm to plan a personalized learning path based on the learner's individual learning behavior data, learning goals, and group learning situation. The state space is defined as the learner's current learning state, and the action space is defined as the selectable learning resources and training methods. The reward function is designed based on the learner's performance in the learning process. By continuously interacting with the learning environment, the optimal action is selected to maximize the cumulative reward, thereby planning a personalized learning path for the learner. During the learning process, the learning path is dynamically adjusted according to the learner's actual learning situation.

9. A vocabulary memory training method for a multimodal English learning interactive system, applicable to a multimodal English learning interactive system according to any one of claims 1 to 8, characterized in that: The specific steps of this method are: S1. Data collection step: Continuously collect multimodal learning behavior data of all learners through the data collection module; S2. Group division and data analysis: The data analysis module uses swarm intelligence algorithms to process the collected data. First, learners are divided into groups based on factors such as age and learning level. Then, the behavior patterns and performance of each group in learning various vocabulary items are analyzed to identify common problems. S3, teaching strategy adjustment step: when a group of students is found to have common problems in learning a certain type of vocabulary, the strategy adjustment module automatically adjusts the teaching strategy; S4. Resource and method push step: The resource push module pushes customized multimodal learning resources and training methods to the group based on the adjusted teaching strategy, guiding learners to conduct vocabulary memory training; S5. Learning effect evaluation and feedback steps: The system collects learners' performance and feedback information during the training process through the interactive learning module, evaluates the learning effect, and further optimizes the teaching strategy and learning resource push based on the evaluation results, forming a closed-loop learning optimization process.

Citation Information

Cited By

  • Network course teaching effect analysis and optimization method and system based on big data

    CN121031992A

  • English auxiliary teaching software multi-modal optimization method based on embedded system

    CN121456799A