A word learning data generation method, device, equipment and storage medium

CN122551633APending Publication Date: 2026-08-11GUANGDONG XIAOTIANCAI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-08-11

Smart Images

  • Figure CN122551633A_ABST
    Figure CN122551633A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and storage medium for generating vocabulary learning data, including: acquiring user information; determining multiple semantically related words and practice modes corresponding to the user information; playing video animations corresponding to each word; retrieving target exercises from a preset exercise bank according to the playback progress and practice mode of each video animation; and collecting answer operation data for each target exercise; performing data analysis and processing on each answer operation data to generate learning result data. This improves the adaptability of word display methods and practice modes to users, enhances the diversity of word content presentation and practice modes, and makes learning content and learning modes more targeted and engaging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device, and storage medium for generating word learning data. Background Technology

[0002] English vocabulary learning is a core foundation of English education and a key starting point for students to build an English language system and improve their comprehensive listening, speaking, reading and writing skills. With the widespread adoption of educational informatization and intelligent educational terminals, lightweight and interactive vocabulary learning using digital tools such as mobile phones, tablets and learning machines has gradually replaced the traditional paper-based memorization model and has become the mainstream trend in primary school English teaching and home-based self-study.

[0003] In the process of teaching vocabulary through digital tools, related technologies typically mechanically list words in the order of textbook words and use pictures to show the meaning of each word. During vocabulary practice, a fixed basic vocabulary practice mode is used. The way words are displayed and the practice mode cannot adapt to the different learning abilities and progress of different users. The presentation of vocabulary content and the practice mode are monotonous, lacking interest and relevance. Summary of the Invention

[0004] This application provides a method, apparatus, device, and storage medium for generating vocabulary learning data. It addresses the problems of inconsistent word display and practice modes in digital vocabulary learning, which fail to adapt to different users' learning abilities and progress, and result in monotonous word content presentation and practice modes lacking interest and relevance. The method accurately matches and extracts multiple words from a word database based on user information and determines practice modes suitable for that user. By retrieving and playing video animations corresponding to each word, and selecting target exercises from a preset question bank based on the playback progress and practice mode of each video animation, it improves the adaptability of word display and practice modes to users, enhances the diversity of word content presentation and practice modes, and makes learning content and modes more targeted and engaging.

[0005] In a first aspect, embodiments of this application provide a method for generating word learning data, comprising: Obtain user information and identify multiple semantically related words and practice patterns corresponding to the user information; Play the video animation corresponding to each of the aforementioned words, retrieve the target exercises from the preset exercise bank according to the playback progress of each of the aforementioned video animations and the exercise mode, and collect the answer operation data of each of the aforementioned target exercises; The data from each of the question-answering operations is analyzed and processed to generate learning result data.

[0006] Optionally, determining the multiple semantically related words corresponding to the user information includes: Based on the user information, the number of target words is determined. According to the semantic tags of each word in the preset word database, words with the same part-of-speech tag, the same logical tag, and the same scene tag are semantically aggregated and associated to obtain multiple word combinations with different semantic types. The semantic tags include part-of-speech tag, logical tag, and scene tag. Based on the target number of words, extract multiple semantically related words from the same word combination.

[0007] Optionally, retrieving target exercises from a preset exercise bank based on the playback progress of each video animation and the exercise mode includes: The training level is determined based on the playback progress of each video animation. The exercises in the preset exercise library are matched with the words corresponding to the video animation and the training level. The target exercise is determined based on the matching result and extracted from the preset exercise library.

[0008] Optionally, determining the training level based on the playback progress of each of the video animations includes: If the playback progress of any of the video animations is at a preset progress value, the training level is determined to be the first training level; If the playback progress of each of the aforementioned video animations is a preset progress value, then the training level is determined to be the second training level.

[0009] Optionally, after collecting the answer operation data for each of the target exercises, the method further includes: If the answer operation data corresponding to each of the first training levels meets the preset challenge conditions, the training level is determined to be the third training level, and the answer operation data is real-time collected data. According to the practice mode, the challenge questions corresponding to the third training level are retrieved from the preset question bank, and the answer operation data of the challenge questions are collected in real time. Accordingly, the data analysis and processing of each of the answer operation data includes: Data analysis and processing are performed on the answer operation data of the target exercises and the answer operation data of the challenge exercises.

[0010] Optionally, after collecting the answer operation data of the challenge exercises, the method further includes: Based on the answer operation data of the target exercises and the answer operation data of the challenge exercises, identify the incorrect exercises and count the number of incorrect exercises; If there are multiple incorrect exercises, the multiple incorrect exercises are reset sequentially, and the incorrect exercises are retrieved in the reset order.

[0011] Optionally, before obtaining user information, the method further includes: The pre-stored words are processed to decompose their pronunciations, generating pronunciation teaching analysis data corresponding to each of the pre-stored words. The pronunciation teaching analysis data, spelling form, Chinese explanation and application scenario corresponding to each of the pre-stored words are input into the contextual animation resource model to obtain the video animation corresponding to each of the pre-stored words. The video animation includes multiple example sentences of different scenarios corresponding to the pre-stored words.

[0012] In a second aspect, embodiments of this application provide a word learning data generation apparatus, comprising: The data acquisition module is used to acquire user information; A word determination module is used to determine multiple semantically related words corresponding to the user information; The practice mode determination module is used to determine the practice mode corresponding to the user information; The learning data retrieval module is used to play the video animation corresponding to each of the aforementioned words, and retrieve the target exercises from the preset exercise bank according to the playback progress of each of the aforementioned video animations and the exercise mode. The answer operation data acquisition module is used to collect answer operation data for each of the target exercises. The learning result generation module is used to perform data analysis and processing on the answer operation data to generate learning result data.

[0013] In a third aspect, embodiments of this application provide an electronic device, the device comprising: one or more processors; and a storage device configured to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the word learning data generation method described in the first aspect.

[0014] In a fourth aspect, embodiments of this application provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the word learning data generation method as described in the first aspect.

[0015] This application embodiment obtains user information and a word database, extracts multiple semantically related words from the word database based on the user information, and determines the practice mode corresponding to the user information; it retrieves and plays the video animation corresponding to each word, and retrieves target exercises from a preset exercise bank according to the playback progress and practice mode of each video animation, and collects the answer operation data of each target exercise in real time; it performs data aggregation and analysis on each answer operation data to generate learning report data for multiple words. The above-mentioned method can accurately match and extract multiple words from the word database based on user information, determine the practice mode adapted to the user information, and improve the adaptability of word display and practice mode to users by retrieving and playing the video animation corresponding to each word and retrieving target exercises from a preset exercise bank according to the playback progress and practice mode of each video animation. This enhances the diversity of word content presentation and practice mode, making the learning content and learning mode more targeted and interesting. Attached Figure Description

[0016] Figure 1 This is a flowchart of a word learning data generation method provided in an embodiment of this application; Figure 2 This is a flowchart of another word learning data generation method provided in the embodiments of this application; Figure 3 This is a flowchart of a video animation generation method including pronunciation instruction provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a word learning data generation device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a word learning data generation device provided in an embodiment of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. It should also be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0018] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0019] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0020] The word learning data generation method, apparatus, device, and medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0021] The word learning data generation method provided in this application is used in scenarios of intelligent word learning on smart learning terminals. Based on the above application scenario, it can be understood that the executing entity of each step can be a computer device. This computer device refers to any electronic device with data computing, processing, and storage capabilities, such as mobile phones, PCs (Personal Computers), tablet computers, and other terminal devices. This application does not limit this.

[0022] Figure 1 This is a flowchart of a word learning data generation method provided in an embodiment of this application, such as... Figure 1 As shown, it includes: S101. Obtain user information and identify multiple semantically related words and practice modes corresponding to the user information.

[0023] User information can refer to basic data used for personalized learning, such as a student's grade level, learning ability, past mistakes, and learning progress. The vocabulary database can refer to a pre-stored collection of English vocabulary from different grade levels, including complete data such as spelling, phonetic symbols, definitions, semantic tags, difficulty, applicable grade level, and contextual attributes. Semantically related words can refer to words with the same part of speech, similar context, or similar logic. Practice modes can refer to learning and answering methods automatically matched according to the student's grade level, which can include animated guided modes for lower grades and lightweight self-practice modes for higher grades.

[0024] In one embodiment, user information can be obtained by the system calling a local storage interface or a network interface to read and load pre-entered and persistently stored user basic data, learning behavior data, etc., and can also read a pre-stored structured dataset of English vocabulary in the curriculum standards.

[0025] In one embodiment, the method for determining multiple semantically related words and practice modes corresponding to the user information can be as follows: parse feature data such as grade, learning ability, and current learning unit in the user information, determine the upper limit of the number of words and the difficulty range for this learning, then perform a full search of the word database to extract a set of candidate words that match the current learning unit, calculate the semantic similarity value between each word in the candidate word set using a semantic similarity calculation model, determine multiple words with similarity values ​​greater than a preset threshold as a group of semantically related words, and extract a corresponding number of words from this group of semantically related words as target words for this learning based on the determined upper limit of the number of words for this learning. Meanwhile, the acquired user information is parsed to extract key features such as grade, learning stage, historical answer accuracy, and interaction preferences. Users are divided into lower or higher grade categories according to preset hierarchical matching rules. Then, based on the grade tag, the corresponding interaction form, exercise display method, guidance logic, and difficulty parameters are retrieved from the mode configuration library. Lower grade automatically matches the practice mode with animated guidance, graphic assistance, and simplified interaction, while higher grade automatically matches the efficient practice mode with lightweight exercises, self-answering, and no unnecessary animations.

[0026] Optionally, determining multiple semantically related words corresponding to the user information includes: determining the target number of words based on the user information; performing semantic aggregation and association on words with the same part-of-speech tag, the same logical tag, and the same scene tag according to the semantic tags of each word in the preset word database to obtain multiple word combinations with different semantic types, wherein the semantic tags include part-of-speech tag, logical tag, and scene tag; and extracting multiple semantically related words from the same word combination according to the target number of words.

[0027] The target number of words can refer to the total number of words that should be mastered in a single learning session, determined based on the user's grade level and cognitive level. Semantic tags refer to classification labels used to describe the semantic attributes of words, which can include part-of-speech tags, logical tags, and context tags. Part-of-speech tags can be labels that mark the grammatical attributes of words, such as nouns, verbs, adjectives, and adverbs. Logical tags can be labels that mark the inherent relationships between words, such as synonyms, antonyms, subordination, causality, and parallelism. Context tags can be labels that mark the practical usage environments of words, such as daily life, learning, and social interactions. Semantic aggregation and association can refer to the clustering calculation process performed based on semantic tags. Word combinations can refer to a set of words formed after semantic aggregation, where the internal words have strong relationships; this is the basic unit for dynamic grouping.

[0028] In one embodiment, determining the target number of words based on user information can be achieved by first parsing the user information to extract key features such as grade level, learning stage, historical learning ability, and answer accuracy rate. Then, a pre-defined mapping rule library between grade and word count is invoked, mapping lower grades (grades 1-3) to 5-6 words and higher grades (grades 4-6) to 7-8 words. Alternatively, the number of words can be fine-tuned by combining the user's historical learning rate with the currently determined number of words to ultimately determine the total number of words to be learned. For example, if the default number of words is 5 for lower elementary grades and 7 for higher elementary grades, and if reading the user's historical answer data shows that the user has repeatedly achieved an accuracy rate above 90%, indicating a fast learning rate, then 1 word is added to the base mapping number. If the accuracy rate is below 60%, indicating a slow learning rate, then 1 word is removed from the base number, ultimately yielding the total number of words to be learned.

[0029] In one embodiment, the method of semantically aggregating and associating words with the same part-of-speech tag, logical tag, and scene tag based on the semantic tags of each word in the word database to obtain multiple word combinations with different semantic types can be as follows: Traverse the word database, read the three types of semantic tags (part-of-speech tag, logical tag, and scene tag) attached to each word, use the tags as clustering features, and group words with completely identical tags using a matching algorithm. Words with the same scene, part of speech, and logical association are automatically aggregated together to form multiple word sets with clear semantic themes. For example, pen, pencil, ruler, book, and bag with the tag "noun + school supplies + classroom scene" are aggregated into a school supplies word combination, and run, jump, swim, sing, and dance with the tag "verb + action + daily behavior" are aggregated into an action word combination, ultimately obtaining multiple groups of word combinations divided by semantic type.

[0030] In one embodiment, extracting multiple semantically related words from the same word combination based on the target word count can be achieved by: using the determined target word count as the extraction limit, traversing the currently semantically aggregated word combinations, and randomly extracting the target word count from each word combination to form the final semantically related word set for learning. For example, if a fourth-grade student has a target word count of 7, the system randomly extracts 7 words from the already aggregated "school supplies" word combination (containing 8 words: pen, pencil, ruler, book, bag, eraser, crayon, pencil-box) as the semantically related word group for this learning session. Alternatively, the words in the word combination can be sorted in ascending order of length to obtain a word sequence. Using the target word count as the upper limit, the first 7 words in the word sequence are extracted, and these 7 words are used as the semantically related word group for this learning session. Optionally, a difficulty gradient grouping mode can also be used, dividing the words into groups based on phonetic difficulty, spelling length, and semantic complexity, and determining the number of words in each group based on the target word count.

[0031] The above method determines the target number of words to be matched based on user information, and then performs precise semantic aggregation and association of words in the word database based on three types of semantic tags: part of speech, logic, and scenario. This can form highly related and logically unified word combinations. Then, words are extracted from the same word combination according to the target number. This can realize intelligent, dynamic, and semantic adaptation of word grouping, avoid the memory fragmentation problem caused by random or mechanical grouping, make word grouping more in line with the user's cognitive level, and significantly improve the coherence and memory efficiency of word learning.

[0032] S102. Play the video animation corresponding to each word, retrieve the target exercises from the preset exercise bank according to the playback progress and practice mode of each video animation, and collect the answer operation data of each target exercise in real time.

[0033] Among these, video animations can refer to pre-produced and stored vocabulary teaching animation resources, including visual teaching content such as word pronunciation, form, meaning, scenarios, pronunciation breakdown, and example sentences in multiple scenarios, used to dynamically display vocabulary learning information. Playback progress can refer to the real-time recording of the playback status of each word animation, including progress indicators such as not playing, playing, and playing completed. Pre-built exercise bank can refer to a pre-built and stored standardized exercise set, which can include various question types such as phonetic-Chinese translation, listening and word selection, word spelling, and comprehensive matching. Target exercises can refer to exclusive exercises that are precisely matched and selected from the exercise bank based on animation playback progress, practice mode, word content, and grade level, for the current learning stage. Answer operation data can refer to behavioral data collected in real time during the user's answering process, including raw interaction data such as answering time, option content, correct and incorrect results, click trajectory, submission status, and incorrect question markings.

[0034] In one embodiment, the method for retrieving and playing the video animation corresponding to each word can be as follows: based on the determined target word set, each word is traversed sequentially and its unique identifier is extracted. The corresponding pre-generated 3D context animation data is retrieved from the animation resource library through resource index matching rules, and the animation data is loaded into the rendering buffer. Then, according to the word learning order and preset playback duration control logic, the animation is decoded, rendered and played frame by frame on the display interface of the learning terminal, and audio data is output synchronously. At the same time, the playback status and progress of the current animation are recorded in real time until the animations corresponding to all words in the group have been played.

[0035] In one embodiment, the method of retrieving target exercises from a preset exercise bank based on the playback progress and practice mode of each video animation can be as follows: real-time reading and judgment of the playback completion progress of each word animation; after each video animation is completed, retrieving candidate exercises containing the corresponding word from the preset exercise bank; determining the question type, difficulty level, etc., based on the practice mode and performing multi-condition matching and filtering on the candidate exercises; filtering out exercises that do not conform to the grade, question type, or knowledge point range; and then extracting target exercises according to a preset number and question order.

[0036] Optionally, target exercises can be retrieved from a preset exercise bank based on the playback progress and practice mode of each video animation, including: determining the training level based on the playback progress of each video animation, matching the exercises in the preset exercise bank with the words and training levels corresponding to the video animations, determining the target exercises based on the matching results, and retrieving the target exercises from the preset exercise bank.

[0037] The training levels can refer to the practice stages divided according to the learning process and increasing difficulty, which may include a basic vocabulary semantics level, a comprehensive phrase application level, and a challenging and intensive level. The matching results refer to the selected practice questions that meet the criteria of word association, level requirements, and difficulty suitability after comparison.

[0038] In one embodiment, the training level can be determined based on the playback progress of each video animation as follows: real-time monitoring and statistics of the completion status of the video animations corresponding to all words in the current group; when only a single word animation is completed, the training level is set to the basic level for single-point consolidation of the sound, form, and meaning of that word; when two complete word animations are played, the training level is set to the sum application level of the two learned words; when all word animations in the current group are completed, the training level is set to the reinforcement level, which is to train on exercises that randomly combine all words in the group.

[0039] In one embodiment, AI can be used to dynamically generate question types, and the difficulty and style of the questions can be adjusted in real time based on the students' correct answer rate, so as to intelligently customize personalized questions to suit students with different learning abilities.

[0040] Optionally, the training level can be determined based on the playback progress of each video animation, including: determining the training level as the first training level when the playback progress of any video animation is at a preset progress value; and determining the training level as the second training level when the playback progress of each video animation is at a preset progress value.

[0041] The preset progress value refers to a pre-defined threshold for video playback nodes, representing the progress scale of a single animation reaching a predetermined learning completion standard, such as playing to the end or playing for a specified duration. The first training level refers to the practice level activated after a single word animation has reached its target, providing targeted practice for basic memorization, pronunciation, and definition of that word. The second training level refers to the practice level activated after all word animations in the entire group have reached their target, providing integrated practice for the association and comprehensive application of multiple words within the group.

[0042] In one embodiment, determining the training level as the first training level when the playback progress of any video animation reaches a preset progress value can be achieved by: real-time polling and monitoring the playback parameters of the video animation corresponding to each word in the group, continuously comparing the real-time playback progress with the system's pre-configured progress threshold, and when the playback progress of any video animation reaches the preset progress value, determining the current practice mode as the first training level according to the preset level judgment rules, and simultaneously determining the corresponding basic exercises for the word in that video animation. For example, if the preset progress value is set to 100% video playback, during the playback of animal-related word groups, when the animation corresponding to "bird" reaches 100% playback progress, regardless of the playback status of other animations in the group, the system will directly switch to the first training level to conduct listening and spelling exercises for that word.

[0043] In one embodiment, determining the training level as the second training level when the playback progress of each video animation is at a preset progress value can be achieved by: continuously collecting the real-time playback progress of all word video animations within the group, comparing and verifying the progress data one by one with a preset progress threshold, and once it is detected that the playback progress of all video animations has reached the preset progress value, setting the training level to the second training level, and loading the practice rules and question types for the second training level. For example, if the preset progress value is set to the last 98% of the video playback, when the progress of all animations corresponding to cat, dog, and rabbit within the animal word group reaches the target, switching to the second training level to conduct comprehensive exercises such as word classification, synonym differentiation, and short sentence usage within the group.

[0044] As described above, by activating the first training level when a single animation reaches the preset progress and switching to the second training level after all animations have completed the preset progress, the training can be automatically graded according to the learning progress, achieving dynamic adaptation of practice difficulty and learning pace, conducting training in layers, ensuring a gradual learning process, and effectively improving the rationality of training and learning effectiveness.

[0045] In one embodiment, matching exercises in a preset exercise bank with the words and training levels corresponding to video animations can be achieved by: first, extracting the target word identifier associated with the current video animation and the identified training level identifier; using the target word identifier and training level identifier as matching conditions; traversing the preset exercise bank; and performing conditional retrieval and bidirectional verification on the associated word attributes and training level attributes bound to each exercise. For example, if the word corresponding to the current animation is "apple" and the training level is the first training level (word semantic consolidation), then exercises bound to the word "apple" and labeled as basic consolidation type, such as listening comprehension and picture recognition exercises, are selected from the exercise bank.

[0046] In one embodiment, the method for determining the target exercise based on the matching results and extracting the target exercise from the preset exercise library can be as follows: Iterate through the matching results of the target exercise and each preset exercise, determine the exercise with the highest matching degree as the target exercise, and if there are multiple target exercises, select multiple target exercises that match the learning stage and user ability according to the priority order of matching degree from high to low, and accurately extract such exercises from the preset exercise library and push them to the learning terminal. For example, the current word is "cat", the current training level is level 1, and the matching results are: ① Listen and select "cat" (word "cat" + level 1) → matching value 1.0 (target exercise); ② Make a sentence using "cat" (word "cat" + level 2) → matching value 0.8; ③ Fill in the blank with the word "car" (word "car" + level 1) → matching value 0.5.

[0047] In one embodiment, the method for collecting the answer operation data of each target exercise in real time can be as follows: continuously capture each step of the user's operation behavior through the event listening interface of the terminal interface, record the user's original interaction information such as option selection, text input, click position, answer duration, submission action, and interruption operation for each exercise in real time, and bind the above data with dimension information such as current exercise ID, word identifier, practice stage, and user ID, and update it in real time in the local cache.

[0048] As described above, by automatically dividing training levels according to the playback progress of video animations, accurately matching exercises with corresponding words and training levels, and extracting target exercises, dynamic adaptation of learning progress and practice content can be achieved. This makes exercise delivery more targeted and coherent, avoids disconnect between practice and learned words and training stages, and ensures that the difficulty of exercises increases gradually, effectively improving the accuracy and learning efficiency of word practice.

[0049] S103. Perform data analysis and processing on the data from each answer operation to generate learning result data.

[0050] The analysis and processing can refer to the process of cleaning, classifying, statistically analyzing, and calculating answer data from multiple questions and stages. This can include data operations such as accuracy statistics, incorrect question classification, weak point identification, learning progress integration, animation viewing time, and pronunciation follow-up data. The learning outcome data can refer to the standardized learning outcome data output after summary and analysis, including vocabulary mastery, exercise accuracy, weak knowledge points, incorrect question list, learning progress, and overall rating.

[0051] In one embodiment, the method for generating learning result data by analyzing and processing the data of each question-answering operation can be as follows: First, all collected question-answering operation data is uniformly collected and cleaned to remove invalid, duplicate, and abnormal operation records. Then, the data is classified, statistically analyzed, and quantitatively calculated according to dimensions such as word ID, question type, correctness of answer, answering time, and frequency of wrong answers to identify the user's mastery of each word, weak knowledge points, and question-answering habits. The above statistical results are then structured, integrated, and encapsulated with information such as learning progress, accuracy rate, and weaknesses to finally generate standardized learning report data containing information such as word mastery, quality of exercise completion, distribution of wrong answers, and learning suggestions.

[0052] In one embodiment, after generating learning report data for multiple words, the learning report data is pushed to the learning terminal and parent monitoring terminals, teacher management terminals, etc., which are bound to the user's login account in the learning terminal, via wireless network or local area network.

[0053] In one possible embodiment, pre-learning exercises can also be added to conduct a simple assessment before vocabulary learning, determine the student's basic ability, and determine the practice mode or training level based on the student's basic ability.

[0054] This application embodiment obtains user information and a word database, extracts multiple semantically related words from the word database based on the user information, and determines the practice mode corresponding to the user information; it retrieves and plays the video animation corresponding to each word, and retrieves target exercises from a preset exercise bank according to the playback progress and practice mode of each video animation, and collects the answer operation data of each target exercise in real time; it performs data aggregation and analysis on each answer operation data to generate learning report data for multiple words. The above-mentioned method can accurately match and extract multiple words from the word database based on user information, determine the practice mode adapted to the user information, and improve the adaptability of word display and practice mode to users by retrieving and playing the video animation corresponding to each word and retrieving target exercises from a preset exercise bank according to the playback progress and practice mode of each video animation. This enhances the diversity of word content presentation and practice mode, making the learning content and learning mode more targeted and interesting.

[0055] Figure 2 This is a flowchart of another word learning data generation method provided in the embodiments of this application, such as... Figure 2 As shown, it includes: S201. Obtain user information and identify multiple semantically related words and practice modes corresponding to the user information.

[0056] S202. Play the video animation corresponding to each word, determine the training level according to the playback progress of each video animation, match the exercises in the preset exercise bank with the words and training levels corresponding to the video animations, determine the target exercises according to the matching results, extract the target exercises from the preset exercise bank, and collect the answer operation data of each target exercise in real time. Among them, if the playback progress of any video animation is at the preset progress value, the training level is determined as the first training level.

[0057] S203. If the answer operation data corresponding to each first training level meets the preset challenge conditions, determine the training level as the third training level.

[0058] Among them, the preset challenge conditions can refer to the pre-set configurable pass thresholds, such as a first training level's answer accuracy rate ≥90%, no core wrong answers, and completion of all questions, used to determine whether one is qualified to advance. The third training level can refer to the challenge enhancement practice stage that is higher than the basic practice, corresponding to the challenge questions within the group that comprehensively improve vocabulary and increase difficulty.

[0059] In one embodiment, determining the training level as the third training level when the answer operation data corresponding to each first training level meets the preset challenge conditions can be achieved by: real-time collection of the user's answer operation data in the first training level, including the accuracy rate, completion status, and number of incorrect answers for all exercises; comparing this data with preset challenge conditions such as achieving the accuracy target, no critical errors, and completion of all questions; and when all answer data meets the challenge conditions, the user is considered to have mastered the basic learning content, and the current training level is officially upgraded from the first training level to the third training level for more challenging practice. For example, if a lower elementary school student completes a single word basic instant practice exercise with a 90% accuracy rate and no uncorrected errors, the challenge conditions are met, and the training level is switched to the third training level, which includes challenge exercises. In another possible embodiment, the third training level can be modified into a fun challenge training, using points and badges as incentives instead of a hard full-correction requirement, making it suitable for younger students with weaker self-control.

[0060] S204. Based on the practice mode, retrieve the challenge questions corresponding to the third training level from the preset question bank, and collect the answer operation data of the challenge questions in real time.

[0061] Among them, the challenge exercises can refer to the high-difficulty comprehensive exercises exclusive to the third training level.

[0062] In one embodiment, the method of "retrieving challenge questions corresponding to the third training level from a preset question bank according to the practice mode and collecting the answer operation data of the challenge questions in real time" can be as follows: Based on the user's grade and the appropriate practice mode, the system accurately selects and retrieves advanced challenge questions corresponding to the third training level from the preset question bank, and collects answer operation data such as correctness of answers, answer time, answer options, and submission status in real time during the user's answering process. For example, when a user in grades 4-6 of primary school is in self-practice mode, the system extracts word spelling and multi-word integrated application challenge questions from the question bank and pushes them to the user, while simultaneously recording the user's answer results and answer time for each question.

[0063] S205. Perform data analysis and processing on the answer operation data of the target exercises and the answer operation data of the challenge exercises to generate learning result data.

[0064] In one embodiment, the method for generating learning result data by analyzing and processing the answer operation data of the target exercises and the challenge exercises can be as follows: First, collect and organize the answer operation data generated by the user during the completion of the target exercises and challenge exercises, including correct and incorrect answers, accuracy rate, answer time, incorrect question numbers, and weak knowledge points. Then, statistically analyze the number of correct answers, the number of incorrect answers, and the total number of answers for each word across all exercises, based on the word dimension. Calculate the word accuracy rate as: word accuracy rate = number of correct answers for a single word / total number of answers for a single word. The user's mastery of a word is assessed based on its accuracy rate (e.g., ≥90% indicates proficiency, 60%-89% indicates average, and <60% indicates weakness). Similarly, the number and accuracy of correct answers are analyzed by question type, including phonetic-Chinese translation, listening comprehension, spelling, and comprehensive application, to identify the user's weak question types. Likewise, the accuracy and completion rates for each training level (basic, comprehensive, and challenging) are analyzed to compare mastery differences at different difficulty levels. Based on these statistical results, the words, question types, and training levels corresponding to frequently missed questions are determined, forming a list of weaknesses. The final output includes core indicator data such as mastery level of each word, overall accuracy rate, weak question types, and weak words. For example, taking a user learning the words apple, banana, and orange as an example, the system first performs statistics by word: apple 5 questions correct / out of 5 (100% accuracy, proficient), banana 4 questions correct / out of 5 (80% accuracy, average), orange 2 questions correct / out of 5 (40% accuracy, weak); by question type: listening comprehension accuracy is 95%, spelling accuracy is only 50%, spelling is identified as a weak question type; by training level: basic practice accuracy is 90%, challenge practice accuracy is only 55%, higher-level challenges are identified as a weak area. Finally, the overall accuracy is calculated to be 73%, and orange, spelling, and challenge practice are identified as the core weak areas.

[0065] As described above, after a user completes the basic exercises of the first training level and their answer data reaches the preset challenge standard, they are automatically upgraded to the third training level and pushed corresponding challenge exercises. Answer data from both types of exercises are collected simultaneously and analyzed in a unified manner. This not only dynamically increases the difficulty of exercises based on the user's actual learning performance, achieving progressive tiered training that balances basic consolidation and ability enhancement, but also fully integrates learning data from all stages, so that the analysis results can comprehensively reflect the user's true level from basic mastery to comprehensive application and accurately pinpoint knowledge gaps.

[0066] Optionally, after collecting the answer operation data of the challenge exercises in real time, the method further includes: identifying the incorrect exercises based on the answer operation data of the target exercises and the answer operation data of the challenge exercises, and counting the number of incorrect exercises; if there are multiple incorrect exercises, resetting the order of the multiple incorrect exercises, and retrieving each incorrect exercise in the reset order.

[0067] Among them, "incorrect exercises" can refer to target exercises and challenge exercises that are determined to be answered incorrectly after comparison with the standard answers by the system, that is, exercises that do not match the standard answers. "Order reset" can refer to randomly shuffling and reordering multiple incorrect exercises.

[0068] In one embodiment, the method for determining and counting incorrect exercises based on the answering data of the target exercises and the challenge exercises can be as follows: Retrieve the answering data generated by the user after completing the target and challenge exercises, including answer options and submission results; compare the user's answer to each question with the system's preset standard answer; mark questions with inconsistent answers or no answers as incorrect exercises; then count all marked incorrect exercises to obtain the total number of incorrect exercises in this learning session. For example, after a user completes 5 target exercises and 3 challenge exercises, if the system finds that the 2nd and 4th target exercises and the 1st challenge exercise were answered incorrectly, these 3 questions are marked as incorrect exercises, and the total number of incorrect exercises is counted as 3.

[0069] In one embodiment, the method of "resetting the order of multiple incorrect exercises and retrieving each incorrect exercise in the reset order" can be as follows: when the total number of incorrect exercises is multiple, the original question order of all incorrect exercises is reset by a random shuffling algorithm to prevent users from forming a memorized answering order. Then, the corresponding incorrect exercises are retrieved from the exercise bank or the incorrect question set in the reset order and pushed out for display.

[0070] In one embodiment, after entering the incorrect question practice stage, the system will by default repeat the incorrect questions. When the answer operation data corresponding to the incorrect questions meets the preset exit conditions, the repeated incorrect question practice mode can be exited. For example, if all incorrect questions appear and are answered correctly at least twice, the repeated incorrect question practice mode will be exited; or, if all incorrect questions appear and are answered correctly, the repeated incorrect question practice mode can also be exited.

[0071] The above-mentioned system accurately and automatically identifies and counts incorrect questions based on answer data, thereby objectively locating and quantifying the user's weak knowledge points. When there are multiple incorrect questions, the original question order is broken by resetting the order, which can effectively prevent users from relying on sequential memory to answer questions, improve the authenticity and validity of the review of incorrect questions, and retrieve incorrect questions in the reset order to achieve randomized and personalized push of incorrect questions, strengthen the consolidation effect of weak knowledge points, and make the practice process more intelligent and more in line with the actual mastery testing needs, thereby improving the accuracy and reliability of vocabulary learning data processing and question push.

[0072] Figure 3 This is a flowchart of a video animation generation method including pronunciation instruction provided in an embodiment of this application, such as... Figure 3 As shown, it includes: S301. Perform pronunciation decomposition on the pre-stored words to generate pronunciation teaching analysis data corresponding to each pre-stored word.

[0073] Among these, pre-stored words refer to the curriculum standard English vocabulary database that the system imports and saves in advance. Pronunciation breakdown processing refers to the standardized process of segmenting, splitting, and analyzing word pronunciation according to phonics rules or syllable division rules. Pronunciation teaching analysis data refers to the pronunciation explanation data generated after breakdown that can be directly used for teaching, including syllable segmentation, letter / letter combination pronunciation, spelling rhythm, mouth shape prompts, etc.

[0074] In one embodiment, the method for generating pronunciation teaching analysis data for each pre-stored word by decomposing its pronunciation can be as follows: For pre-stored English words, a decomposition strategy is differentiated based on word length. Short words are decomposed into independent pronunciation units using phonics rules, while long words are decomposed into multiple syllable segments according to syllable division rules. Simultaneously, standard pronunciation, spelling rhythm, and follow-along prompts are labeled for each segmented unit. The above data is then integrated to generate pronunciation teaching analysis data containing syllable decomposition, spelling logic, and pronunciation points. For example, the pre-stored word "information" is decomposed into four parts (in-for-ma-tion) according to syllables and each part is labeled with its pronunciation. The short word "cat" is decomposed into three pronunciation units (cat) according to phonics, ultimately forming complete analysis data that can be directly used for pronunciation teaching.

[0075] In one embodiment, a combination of phonetic symbol decomposition and mouth shape animation can be used to demonstrate the opening and closing angle of the mouth and the position of the tongue during pronunciation. This can optimize the standard of spoken pronunciation for primary school students with weak pronunciation and make up for the lack of mouth shape in simple spelling.

[0076] S302. Input the pronunciation teaching analysis data, spelling form, Chinese explanation and application scenario corresponding to each pre-stored word into the contextual animation resource model to obtain the video animation corresponding to each pre-stored word. The video animation includes multiple example sentences of different scenarios corresponding to the pre-stored word.

[0077] Among these, spelling refers to the complete alphabetical composition, capitalization, and writing structure of a word. Chinese definition refers to the standard Chinese definition and translation of the word, used to express its meaning. Application scenario refers to the applicable context of the word in daily life, learning, and social interactions. Contextual animation resource model refers to a pre-built 3D / dynamic contextual animation generation model that can automatically output teaching animations by integrating various information about the word. Multiple example sentences in different scenarios refer to multiple sets of example sentences around the same word, corresponding to different usage scenarios, for multi-contextual teaching.

[0078] In one embodiment, the method of inputting the pronunciation teaching analysis data, spelling, Chinese explanation, and application scenarios corresponding to each pre-stored word into a contextual animation resource model to obtain the corresponding video animation can be as follows: The pronunciation teaching analysis data, standard spelling, Chinese explanation, and application scenario information such as life / learning / social interaction for each pre-stored word are uniformly input into a pre-trained contextual animation resource model. The model automatically integrates and renders the pronunciation rhythm, word text display, explanation, and scene visuals, and arranges them in sequence to generate a video animation that integrates the word's sound, form, meaning, and context. Each video animation includes example sentences for life scenarios, learning scenarios, and social scenarios corresponding to the pre-stored word, and each example sentence includes a Chinese translation and grammatical analysis. This multi-contextual demonstration of word usage generates contextual teaching data, strengthening word application skills.

[0079] In one embodiment, 2D hand-drawn dynamic illustrations plus live-action voice-over short films can be used instead of 3D animation to reduce production costs and adapt to low-end learning terminals; or, AR real-scene interactive displays can be used to overlay words with real objects for recognition, enhancing real-scene memory, which is suitable for intelligent interactive learning devices.

[0080] As described above, by deconstructing pre-stored words into phonetics to generate standardized pronunciation teaching analysis data, abstract pronunciations can be transformed into structured pronunciation knowledge that can be learned and followed along with. Then, by inputting the pronunciation teaching analysis data, spelling, Chinese explanations, and application scenarios into a contextual animation resource model, video animations with example sentences from multiple scenarios can be generated. This enables visualized and dynamic teaching that integrates the sound, form, meaning, and context of words, effectively improving the intuitiveness and fun of word learning, helping students quickly understand the meaning of words, master pronunciation logic and usage, significantly reducing the difficulty of memorization and improving learning efficiency.

[0081] Figure 4 This is a schematic diagram of the structure of a word learning data generation device provided in an embodiment of this application, as shown below. Figure 4 As shown, it includes: Data acquisition module 41 is used to acquire user information; The word determination module 42 is used to determine multiple semantically related words corresponding to the user information; The practice mode determination module 43 is used to determine the practice mode corresponding to the user information; The learning data retrieval module 44 is used to play the video animation corresponding to each of the words, and retrieve the target exercises from the preset exercise bank according to the playback progress of each of the video animations and the exercise mode. The answer operation data acquisition module 45 is used to collect the answer operation data of each of the target exercises in real time; The learning result generation module 46 is used to perform data analysis and processing on the answer operation data to generate learning result data.

[0082] This application embodiment obtains user information and a word database, extracts multiple semantically related words from the word database based on the user information, and determines the practice mode corresponding to the user information; it retrieves and plays the video animation corresponding to each word, and retrieves target exercises from a preset exercise bank according to the playback progress and practice mode of each video animation, and collects the answer operation data of each target exercise in real time; it performs data aggregation and analysis on each answer operation data to generate learning report data for multiple words. The above-mentioned method can accurately match and extract multiple words from the word database based on user information, determine the practice mode adapted to the user information, and improve the adaptability of word display and practice mode to users by retrieving and playing the video animation corresponding to each word and retrieving target exercises from a preset exercise bank according to the playback progress and practice mode of each video animation. This enhances the diversity of word content presentation and practice mode, making the learning content and learning mode more targeted and interesting.

[0083] In one possible embodiment, the word determination module 42 is used for: Based on the user information, the number of target words is determined. Based on the semantic tags of each word in the word database, words with the same part-of-speech tag, the same logical tag, and the same scene tag are semantically aggregated and associated to obtain multiple word combinations with different semantic types. The semantic tags include part-of-speech tag, logical tag, and scene tag. Based on the target number of words, extract multiple semantically related words from the same word combination.

[0084] In one possible embodiment, the learning data retrieval module 44 is used for: The training level is determined based on the playback progress of each video animation. The exercises in the preset exercise library are matched with the words corresponding to the video animation and the training level. The target exercise is determined based on the matching result and extracted from the preset exercise library.

[0085] In one possible embodiment, the learning data retrieval module 44 is used for: If the playback progress of any of the video animations is at a preset progress value, the training level is determined to be the first training level; If the playback progress of each of the aforementioned video animations is a preset progress value, then the training level is determined to be the second training level.

[0086] In one possible embodiment, the learning data retrieval module 44 is further configured to: If the answer operation data corresponding to each of the first training levels meets the preset challenge conditions, the training level is determined to be the third training level, and the answer operation data is real-time collected data. According to the practice mode, the challenge exercises corresponding to the third training level are retrieved from the preset exercise bank; The answer operation data acquisition module 45 is also used for: Real-time collection of answer operation data for the challenge exercises; Accordingly, the learning outcome generation module 46 is used for: Data analysis and processing are performed on the answer operation data of the target exercises and the answer operation data of the challenge exercises.

[0087] In one possible embodiment, the learning data retrieval module 44 is further configured to: Based on the answer operation data of the target exercises and the answer operation data of the challenge exercises, identify the incorrect exercises and count the number of incorrect exercises; If there are multiple incorrect exercises, the multiple incorrect exercises are reset sequentially, and the incorrect exercises are retrieved in the reset order.

[0088] In one possible embodiment, a video animation generation module is further included, the video animation generation module being used for: The pre-stored words are processed to decompose their pronunciations, generating pronunciation teaching analysis data corresponding to each of the pre-stored words. The pronunciation teaching analysis data, spelling form, Chinese explanation and application scenario corresponding to each of the pre-stored words are input into the contextual animation resource model to obtain the video animation corresponding to each of the pre-stored words. The video animation includes multiple example sentences of different scenarios corresponding to the pre-stored words.

[0089] This application also provides an electronic device that can integrate a word learning data generation apparatus provided in this application. Figure 5 This is a schematic diagram of the structure of a word learning data generation device provided in an embodiment of this application, with reference to... Figure 5 The word learning data generation device includes: an input device 53, an output device 54, a memory 52, and one or more processors 51; the memory 52 is used to store one or more programs; when one or more programs are executed by one or more processors 51, the one or more processors 51 implement the word learning data generation method provided in the above embodiments. The input device 53, output device 54, memory 52, and processors 51 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.

[0090] The memory 52, as a computing device readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the word learning data generation method provided in any embodiment of this application. The memory 52 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the device, etc. Furthermore, the memory 52 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 52 may further include memory remotely located relative to the processor 51, and these remote memories can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0091] Input device 53 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the device. Output device 54 may include display devices such as a display screen.

[0092] The processor 51 executes various functional applications and data processing of the device by running software programs, instructions and modules stored in the memory 52, thereby realizing the above-mentioned word learning data generation method.

[0093] The word learning data generation apparatus, device, and computer provided above can be used to execute the word learning data generation method provided in any of the above embodiments, and have corresponding functions and beneficial effects.

[0094] This application embodiment also provides a storage medium for storing computer-executable instructions. When executed by a computer processor, the computer-executable instructions are used to execute the word learning data generation method provided in the above embodiment. The word learning data generation method includes: obtaining user information, determining multiple semantically related words corresponding to the user information and practice modes; playing video animations corresponding to each word, retrieving target exercises from a preset exercise bank according to the playback progress and practice mode of each video animation, and collecting answer operation data of each target exercise; performing data analysis and processing on each answer operation data to generate multiple learning result data.

[0095] Storage medium – any type of memory device or storage device. The term “storage medium” is intended to include: mounting media, such as CD-ROMs, floppy disks, or magnetic tape devices; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media (e.g., hard disks or optical storage); registers or other similar types of memory elements, etc. Storage media may also include other types of memory or combinations thereof. Furthermore, storage media may reside in a first computer system in which a program is executed, or may reside in a different second computer system connected to the first computer system via a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The term “storage medium” can include two or more storage media that may reside in different locations (e.g., in different computer systems connected via a network). Storage media may store program instructions (e.g., specifically implemented as a computer program) executable by one or more processors.

[0096] Of course, the computer-executable instructions provided in the embodiments of this application are not limited to the word learning data generation method described above, but can also perform related operations in the word learning data generation method provided in any embodiment of this application.

[0097] The word learning data generation apparatus, device, and storage medium provided in the above embodiments can execute the word learning data generation method provided in any embodiment of this application. For technical details not described in detail in the above embodiments, please refer to the word learning data generation method provided in any embodiment of this application.

[0098] The above description is merely a preferred embodiment and the technical principles employed in this application. This application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions that can be made by those skilled in the art will not depart from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of this application, the scope of which is determined by the scope of the claims.

Claims

1. A method for generating vocabulary learning data, characterized in that, include: Obtain user information and identify multiple semantically related words and practice patterns corresponding to the user information; Play the video animation corresponding to each of the aforementioned words, retrieve the target exercises from the preset exercise bank according to the playback progress of each of the aforementioned video animations and the exercise mode, and collect the answer operation data of each of the aforementioned target exercises; The data from each of the question-answering operations is analyzed and processed to generate learning result data.

2. The word learning data generation method according to claim 1, characterized in that, The step of determining multiple semantically associated words corresponding to the user information includes: Based on the user information, the number of target words is determined. According to the semantic tags of each word in the preset word database, words with the same part-of-speech tag, the same logical tag, and the same scene tag are semantically aggregated and associated to obtain multiple word combinations with different semantic types. The semantic tags include part-of-speech tag, logical tag, and scene tag. Based on the target number of words, extract multiple semantically related words from the same word combination.

3. The word learning data generation method according to claim 1, characterized in that, The step of retrieving target exercises from a preset exercise bank based on the playback progress of each video animation and the exercise mode includes: The training level is determined based on the playback progress of each video animation. The exercises in the preset exercise library are matched with the words corresponding to the video animation and the training level. The target exercise is determined based on the matching result and extracted from the preset exercise library.

4. The word learning data generation method according to claim 3, characterized in that, The step of determining the training level based on the playback progress of each video animation includes: If the playback progress of any of the video animations is at a preset progress value, the training level is determined to be the first training level. If the playback progress of each of the aforementioned video animations is a preset progress value, then the training level is determined to be the second training level.

5. The word learning data generation method according to claim 4, characterized in that, After collecting the answer operation data for each of the target exercises, the method further includes: If the answer operation data corresponding to each of the first training levels meets the preset challenge conditions, the training level is determined to be the third training level, and the answer operation data is real-time collected data. According to the practice mode, the challenge questions corresponding to the third training level are retrieved from the preset question bank, and the answer operation data of the challenge questions are collected in real time. Accordingly, the data analysis and processing of each of the answer operation data includes: Data analysis and processing are performed on the answer operation data of the target exercises and the answer operation data of the challenge exercises.

6. The word learning data generation method according to claim 5, characterized in that, After collecting the answer operation data of the challenge exercises in real time, the method further includes: Based on the answer operation data of the target exercises and the answer operation data of the challenge exercises, identify the incorrect exercises and count the number of incorrect exercises; If there are multiple incorrect exercises, the multiple incorrect exercises are reset sequentially, and the incorrect exercises are retrieved in the reset order.

7. The word learning data generation method according to claim 1, characterized in that, Before obtaining user information, the method also includes: The pre-stored words are processed to decompose their pronunciations, generating pronunciation teaching analysis data corresponding to each of the pre-stored words. The pronunciation teaching analysis data, spelling form, Chinese explanation and application scenario corresponding to each of the pre-stored words are input into the contextual animation resource model to obtain the video animation corresponding to each of the pre-stored words. The video animation includes multiple different scenario example sentences corresponding to the pre-stored words.

8. A vocabulary learning data generation device, characterized in that, include: The data acquisition module is used to acquire user information; A word determination module is used to determine multiple semantically related words corresponding to the user information; The practice mode determination module is used to determine the practice mode corresponding to the user information; The learning data retrieval module is used to play the video animation corresponding to each of the aforementioned words, and retrieve the target exercises from the preset exercise bank according to the playback progress of each of the aforementioned video animations and the exercise mode. The answer operation data acquisition module is used to collect answer operation data for each of the target exercises. The learning result generation module is used to perform data analysis and processing on the answer operation data to generate learning result data.

9. An electronic device, characterized in that, The device includes: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the word learning data generation method as described in any one of claims 1-7.

10. A storage medium for storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the word learning data generation method as described in any one of claims 1-7.