Word classification system and method
Through the combination of data sorting module, classification algorithm module, dynamic adjustment module and AI generation module, the problem of inefficiency in the existing word memory system is solved, automated classification and dynamic adjustment are realized, and system operation efficiency and memory consistency are improved.
Patent Information
- Application Number
- CN202510388526.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-11-25
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-22
AI Technical Summary
The existing word memory systems rely on manual classification and sorting, resulting in inefficiency, lack of scientific algorithm support, unable to dynamically adjust learning paths, unreasonable resource allocation, and lack of automatic root and affix recognition technology based on natural language processing, affecting memory consistency.
The data sorting module, classification algorithm module, dynamic adjustment module and AI generation module are used to analyze the number of syllables, spelling rules and root affix combinations of words through natural language processing technology, and use reinforcement learning algorithms to optimize the learning path, and generate standardized memory images and story chains based on the encoding table to reduce manual intervention and ensure the consistency of splitting rules.
It realizes automated classification and dynamic adjustment, improves system operation efficiency, reduces server computing burden, and improves memory consistency and learning efficiency.
Smart Images

Figure CN120354840A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data classification, and particularly to a word classification system and method. Background Art
[0002] Existing word memory systems mostly rely on manual classification and sorting, and have the following technical defects: traditional methods lack scientific algorithm support and it is difficult to dynamically generate classification criteria according to word frequency, spelling complexity, and root and affix combinations, resulting in strong subjectivity and time consumption in classification results; existing systems cannot monitor the progress and performance of learners in real time through algorithms and cannot dynamically adjust the learning path, resulting in unreasonable resource allocation; memory assistance content relies on manual creation, with low generation efficiency and difficulty in standardization, and cannot meet the needs of large-scale learning; the lack of automatic root and affix recognition technology based on natural language processing (NLP) leads to chaotic splitting rules and affects memory consistency. Summary of the Invention
[0003] The purpose of the present invention is to provide a word classification system and method, aiming to solve the problem of low efficiency in the existing word memory systems that mostly rely on manual classification and sorting.
[0004] To achieve the above purpose, in the first aspect, the present invention provides a word classification system, including a data sorting module, a classification algorithm module, a dynamic adjustment module, and an AI generation module, which are connected in sequence;
[0005] The data sorting module is used to extract word frequency data from a corpus and analyze the syllable count, spelling rules, and root and affix combinations of words through natural language processing technology;
[0006] The classification algorithm module is used to assign words to primary, intermediate, and advanced word libraries;
[0007] The dynamic adjustment module optimizes the learning path through a reinforcement learning algorithm and switches the word library level based on learner data;
[0008] The AI generation module generates standardized memory images and story chains based on an encoding table.
[0009] Among them, the data sorting module includes a data collection unit, an NLP analysis unit, a data cleaning unit, and a storage management unit, which are connected in sequence;
[0010] The data collection unit is used to batch grab words and their frequency data from a preset corpus;
[0011] The NLP analysis unit divides the word into syllables by a syllable counting algorithm;
[0012] The data cleaning unit is used to remove duplicate data, correct spelling errors, and fill in missing fields;
[0013] The storage management unit is used to store the cleaned structured data in a relational database.
[0014] Wherein, the classification algorithm module includes a frequency classification submodule, a complexity scoring submodule and a comprehensive classification submodule, and the frequency classification submodule and the complexity scoring submodule are respectively connected to the comprehensive classification submodule;
[0015] The frequency classification submodule divides words based on the frequency of occurrence of the words in the corpus;
[0016] The complexity scoring submodule generates a combination complexity score according to the word complexity score;
[0017] The comprehensive classification submodule distributes words to different levels of vocabulary according to the division frequency and combination complexity score.
[0018] Wherein, the dynamic adjustment module includes a data monitoring unit, an algorithm execution unit and a feedback optimization unit, and the data monitoring unit, the algorithm execution unit and the feedback optimization unit are connected in sequence;
[0019] The data monitoring unit collects data such as the learner's accuracy, response time, and review interval, and stores them in the learner's behavior log;
[0020] The algorithm execution unit outputs a vocabulary push priority based on the learner behavior log;
[0021] The feedback optimization unit adjusts the vocabulary push strategy according to the algorithm results and generates personalized learning suggestions.
[0022] The AI generation module includes an image generation unit, a story chain generation unit and a quality control unit, wherein the image generation unit and the story chain generation unit are respectively connected to the quality control unit;
[0023] The image generation unit matches the coding elements after word splitting according to the coding table, and uses a generative adversarial network to generate an image that matches the word semantics and splitting logic;
[0024] The story chain generation unit extracts the core semantics of words through dependency syntactic analysis and calls the pre-trained language model to generate a coherent story chain;
[0025] The quality control unit is used to automatically review the generated images and story chains.
[0026] In a second aspect, a word classification method for the word classification system described in the first aspect includes the following steps:
[0027] The data sorting module outputs structured data to the classification algorithm module for classification to obtain a classification result;
[0028] The classification result is input into the dynamic adjustment module to generate a push strategy in combination with learner data;
[0029] The AI generation module generates content based on the classification result and the coding table and finally pushes it to the user side;
[0030] The dynamic adjustment module feeds back data to the classification algorithm module in real time to trigger the redistribution of the word library.
[0031] A word classification system of the present invention includes a data sorting module, a classification algorithm module, a dynamic adjustment module, and an AI generation module. The data sorting module, the classification algorithm module, the dynamic adjustment module, and the AI generation module are connected in sequence; the data sorting module is used to extract word frequency data from a corpus and analyze the syllable count, spelling rules, and root and affix combinations of words through natural language processing technology; the classification algorithm module is used to assign words to primary, intermediate, and advanced word libraries; the dynamic adjustment module optimizes the learning path through a reinforcement learning algorithm and switches the word library level based on learner data; the AI generation module generates standardized memory images and story chains based on the coding table. The data sorting module outputs structured data to the classification algorithm module for classification to obtain a classification result; the classification result is input into the dynamic adjustment module to generate a push strategy in combination with learner data; the AI generation module generates content based on the classification result and the coding table and finally pushes it to the user side; the dynamic adjustment module feeds back data to the classification algorithm module in real time to trigger the redistribution of the word library. The present invention reduces manual intervention and improves the system operation efficiency through algorithmic automatic classification and dynamic adjustment; based on the NLP-based root and affix recognition technology, it ensures the consistency of the splitting rules and reduces the server computing burden. Thereby, it solves the problem that the existing word memory systems mostly rely on manual classification and sorting with low efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 It is a schematic diagram of a word classification system provided by the present invention.
[0034] Figure 2 It is a schematic diagram of the data sorting module.
[0035] Figure 3 It is a schematic diagram of the classification algorithm module.
[0036] Figure 4 It is a schematic diagram of the dynamic adjustment module.
[0037] Figure 5 It is a schematic diagram of the AI generation module
[0038] Figure 6 It is a schematic diagram of the frequency classification sub-module.
[0039] Figure 7 It is a schematic diagram of the complexity evaluation sub-module.
[0040] Figure 8 It is a schematic diagram of the comprehensive classification sub-module.
[0041] Figure 9 It is a flowchart of a word classification method provided by the present invention.
[0042] In the figure: 1 - data sorting module, 2 - classification algorithm module, 3 - dynamic adjustment module, 4 - AI generation module, 11 - data acquisition unit, 12 - NLP analysis unit, 13 - data cleaning unit, 14 - storage management unit, 21 - frequency classification sub-module, 22 - complexity evaluation sub-module, 23 - comprehensive classification sub-module, 31 - data monitoring unit, 32 - algorithm execution unit, 33 - feedback optimization unit, 41 - image generation unit, 42 - story chain generation unit, 43 - quality control unit, 211 - word frequency statistics unit, 212 - frequency division unit, 221 - syllable analysis unit, 222 - spelling rule analysis unit, 223 - root and affix analysis unit, 231 - score integration unit, 232 - thesaurus allocation unit. Detailed implementation manners
[0043] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.
[0044] Please refer to Figures 1 to 8 , in the first aspect, the present invention provides a word classification system, including a data sorting module 1, a classification algorithm module 2, a dynamic adjustment module 3, and an AI generation module 4. The data sorting module 1, the classification algorithm module 2, the dynamic adjustment module 3, and the AI generation module 4 are connected in sequence;
[0045] The data sorting module 1 is used to extract word frequency data from the corpus and analyze the number of syllables, spelling rules and root affix combinations of words through natural language processing technology;
[0046] The classification algorithm module 2 is used to allocate words to elementary, intermediate and advanced vocabulary;
[0047] The dynamic adjustment module 3 optimizes the learning path through a reinforcement learning algorithm and switches the vocabulary level based on the learner data;
[0048] The AI generation module 4 generates standardized memory images and story chains based on the coding table.
[0049] In this embodiment, the data sorting module 1 outputs structured data to the classification algorithm module 2 for classification to obtain the classification results; the classification results are input into the dynamic adjustment module 3, and the push strategy is generated in combination with the learner data; the AI generation module 4 generates content according to the classification results and the coding table, and finally pushes it to the user end; the dynamic adjustment module 3 feeds back data to the classification algorithm module 2 in real time, triggering the redistribution of the vocabulary. The present invention reduces manual intervention and improves the operating efficiency of the system through automatic classification and dynamic adjustment of algorithms; based on the root and affix recognition technology of NLP, it ensures the consistency of the splitting rules and reduces the computing burden of the server. This solves the problem that the existing word memory system relies heavily on manual classification and sorting, which is inefficient.
[0050] Further, the data sorting module 1 includes a data acquisition unit 11, an NLP analysis unit 12, a data cleaning unit 13 and a storage management unit 14, and the data acquisition unit 11, the NLP analysis unit 12, the data cleaning unit 13 and the storage management unit 14 are connected in sequence;
[0051] The data collection unit 11 is used to batch capture words and their frequency data from a preset corpus;
[0052] The NLP analysis unit 12 divides the word into syllables by a syllable counting algorithm;
[0053] The data cleaning unit 13 is used to remove duplicate data, correct spelling errors, and fill in missing fields;
[0054] The storage management unit 14 is used to store the cleaned structured data in a relational database.
[0055] In this embodiment, the data acquisition unit 11 scrapes words and their frequency data from the COCA (Corpus of Contemporary American English) corpus through the Scrapy framework of Python, and supports offline data import in JSON and CSV formats. The data scraping frequency is once a day to ensure the timeliness of the word frequency data. The NLP analysis unit 12 uses regular expressions based on phonetic rules (such as [aeiouy]+ to match syllables). For example, "unbelievable" is divided into un-be-liev-a-ble (5 syllables). The Porter Stemmer algorithm of the NLTK library is used to separate the root words. For example, "happiness" is split into "happy" (root word) and "-ness" (suffix). The data cleaning unit 13 deletes duplicate word entries through the drop_duplicates() function of Pandas. The SymSpell algorithm is integrated to automatically correct spelling mistakes (such as correcting "accomodate" to "accommodate"). For low-frequency words without word frequency data, the average frequency value of the corpus (such as 10 times per million words) is filled by default. The storage and management unit 14 stores the cleaned data in the MySQL database.
[0056] Further, the classification algorithm module 2 includes a frequency classification sub-module 21, a complexity evaluation sub-module 22, and a comprehensive classification sub-module 23. The frequency classification sub-module 21 and the complexity evaluation sub-module 22 are respectively connected to the comprehensive classification sub-module 23;
[0057] The frequency classification sub-module 21 divides words based on the frequency of occurrence of the words in the corpus;
[0058] The complexity evaluation sub-module 22 generates a combined complexity score according to the word complexity score;
[0059] The comprehensive classification sub-module 23 assigns words to different hierarchical thesauruses according to the divided frequency and the combined complexity score.
[0060] In this embodiment, the frequency classification sub-module 21 includes a word frequency statistics unit 211 and a frequency division unit 212;
[0061] The word frequency statistics unit 211 is used to calculate the frequency of occurrence of words in the corpus and generate a frequency distribution histogram;
[0062] The frequency division unit 212 divides words into three categories: high-frequency, medium-frequency, and low-frequency according to preset quantiles.
[0063] The complexity assessment sub-module 22 includes a syllable analysis unit 221, a spelling rule analysis unit 222, and a root and affix analysis unit 223;
[0064] The syllable analysis unit 221 generates a complexity score based on the number of syllables;
[0065] The spelling rule analysis unit 222 determines whether a word conforms to common spelling rules and outputs a rule matching degree score;
[0066] The root and affix analysis unit 223 is used to count the number of roots and affixes and generate a combined complexity score.
[0067] The comprehensive classification sub-module 23 includes a score integration unit 231 and a thesaurus assignment unit 232;
[0068] The score integration unit 231 is used to perform a weighted sum of the frequency and complexity scores to generate a comprehensive score (weight distribution: frequency 60%, complexity 40%);
[0069] The thesaurus assignment unit 232 assigns words to different levels of the thesaurus according to the comprehensive score (primary: high frequency + simple, intermediate: medium frequency + medium, advanced: low frequency + complex).
[0070] Further, the dynamic adjustment module 3 includes a data monitoring unit 31, an algorithm execution unit 32, and a feedback optimization unit 33, and the data monitoring unit 31, the algorithm execution unit 32, and the feedback optimization unit 33 are connected in sequence;
[0071] The data monitoring unit 31 collects data such as the learner's correct rate, response time, review interval, etc., and stores it in the learner behavior log;
[0072] The algorithm execution unit 32 outputs the push priority for the thesaurus based on the learner behavior log;
[0073] The feedback optimization unit 33 adjusts the thesaurus push strategy according to the algorithm result and generates personalized learning suggestions.
[0074] In this embodiment, the data monitoring unit 31 collects learner behavior data through front-end data embedding (JavaScript event monitoring) and stores it in the Elasticsearch log. The algorithm execution unit 32 uses the Q-learning algorithm. The state space is the thesaurus level (primary / intermediate / advanced), and the action space is the push priority (high / medium / low). When the learner's correct rate is lower than 70%, the feedback optimization unit 33 automatically downgrades the current thesaurus (e.g., advanced → intermediate), and pushes review suggestions by email (e.g., "Please review the root '-able'").
[0075] Further, the AI generation module 4 includes an image generation unit 41, a story chain generation unit 42, and a quality control unit 43. The image generation unit 41 and the story chain generation unit 42 are respectively connected to the quality control unit 43;
[0076] The image generation unit 41 matches the encoded elements after word splitting according to the encoding table, and uses a generative adversarial network to generate an image that matches the word semantics and splitting logic;
[0077] The story chain generation unit 42 extracts the core semantics of words through dependency syntactic analysis and calls a pre-trained language model to generate a coherent story chain;
[0078] The quality control unit 43 is used to automatically review the generated images and story chains.
[0079] In this embodiment, the image generation unit 41 uses the StyleGAN2 model. The input is the encoding table matching result (such as "b → uncle"), and the output is a 512×512 pixel PNG image. The story chain generation unit 42 calls the GPT-3 model. The input is the word semantic analysis result (such as "bake → roast") to generate a story chain; the quality control unit 43 uses OpenCV to detect whether the image contains encoded elements (such as the image of "uncle"); calculates the semantic similarity between the generated story and the encoding table through the BERT model (threshold ≥ 0.8)
[0080] Please refer to Figure 9 , on the second aspect, a word classification method for the word classification system described in the first aspect, includes the following steps:
[0081] S1 The data sorting module 1 outputs structured data to the classification algorithm module 2 for classification to obtain a classification result;
[0082] Specifically, the data sorting module 1 reads structured data (including word frequency, number of syllables, word roots and affixes) from the MySQL database. The classification algorithm module 2 calls the frequency classification sub-module 21 and the complexity evaluation sub-module 22 to generate primary, intermediate, and advanced word libraries, and stores them in the Redis cache to support fast query.
[0083] S2 Input the classification result into the dynamic adjustment module 3 to generate a push strategy in combination with the learner data;
[0084] Specifically, the dynamic adjustment module 3 calculates the word library push priority through the Q-learning algorithm according to the learner's historical behavior data (such as the correct rate of the last 10 exercises). For example, if the learner's correct rate is continuously higher than 80%, the words in the advanced word library will be pushed first.
[0085] The S3 AI generation module 4 generates content based on the classification result and the coding table, and finally pushes it to the user side;
[0086] Specifically, the AI generation module 4 matches and splits elements from the coding table (for example, the letter "b" corresponds to "uncle"), generates an image and a story chain, and then pushes them to the user-side APP through the RESTful API. The front end displays them in the form of cards (the image is on the left and the story text is on the right).
[0087] The S4 dynamic adjustment module 3 feeds back data to the classification algorithm module 2 in real time, triggering the redistribution of the thesaurus.
[0088] Specifically, when the dynamic adjustment module 3 detects that the learner's progress is abnormal (for example, the error rate > 50% for 3 consecutive times), it triggers the classification algorithm module 2 to redistribute the thesaurus. For example, some advanced vocabulary is temporarily demoted to the intermediate thesaurus, and the AI generation module 4 is notified through the message queue (Kafka) to regenerate the adapted content.
[0089] The above-disclosed is only a preferred embodiment of a word classification system and method of the present invention. Of course, the scope of the rights of the present invention cannot be limited by this. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. A word classification system, characterized in that: It includes a data sorting module, a classification algorithm module, a dynamic adjustment module and an AI generation module, wherein the data sorting module, the classification algorithm module, the dynamic adjustment module and the AI generation module are connected in sequence; The data sorting module is used to extract word frequency data from the corpus and analyze the number of syllables, spelling rules and root and affix combinations of words through natural language processing technology; The classification algorithm module is used to assign words to elementary, intermediate and advanced vocabulary; The dynamic adjustment module optimizes the learning path through a reinforcement learning algorithm and switches the vocabulary level based on the learner data; The AI generation module generates standardized memory images and story chains based on the coding table.
2. The word classification system according to claim 1, characterized in that: The data sorting module includes a data acquisition unit, an NLP analysis unit, a data cleaning unit and a storage management unit, and the data acquisition unit, the NLP analysis unit, the data cleaning unit and the storage management unit are connected in sequence; The data collection unit is used to batch capture words and their frequency data from a preset corpus; The NLP analysis unit divides the word into syllables by a syllable counting algorithm; The data cleaning unit is used to remove duplicate data, correct spelling errors, and fill in missing fields; The storage management unit is used to store the cleaned structured data in a relational database.
3. The word classification system according to claim 1, characterized in that: The classification algorithm module includes a frequency classification submodule, a complexity scoring submodule and a comprehensive classification submodule, and the frequency classification submodule and the complexity scoring submodule are respectively connected to the comprehensive classification submodule; The frequency classification submodule divides words based on the frequency of occurrence of the words in the corpus; The complexity scoring submodule generates a combination complexity score according to the word complexity score; The comprehensive classification submodule distributes words to different levels of vocabulary according to the division frequency and combination complexity score.
4. The word classification system according to claim 1, characterized in that: The dynamic adjustment module includes a data monitoring unit, an algorithm execution unit and a feedback optimization unit, and the data monitoring unit, the algorithm execution unit and the feedback optimization unit are connected in sequence; The data monitoring unit collects data such as the learner's accuracy, response time, and review interval, and stores them in the learner's behavior log; The algorithm execution unit outputs a vocabulary push priority based on the learner behavior log; The feedback optimization unit adjusts the vocabulary push strategy according to the algorithm results and generates personalized learning suggestions.
5. The word classification system according to claim 1, characterized in that: The AI generation module includes an image generation unit, a story chain generation unit and a quality control unit, wherein the image generation unit and the story chain generation unit are respectively connected to the quality control unit; The image generation unit matches the coding elements after word splitting according to the coding table, and uses a generative adversarial network to generate an image that matches the word semantics and splitting logic; The story chain generation unit extracts the core semantics of words through dependency syntactic analysis and calls a pre-trained language model to generate a coherent story chain; The quality control unit is used to automatically review the generated images and story chains.
6. A word classification method for the word classification system according to any one of claims 1-5, characterized in that, It includes the following steps: The data sorting module outputs structured data to the classification algorithm module for classification to obtain a classification result; The classification result is input into the dynamic adjustment module to generate a push strategy in combination with learner data; The AI generation module generates content according to the classification result and the coding table and finally pushes it to the user side; The dynamic adjustment module feeds back data to the classification algorithm module in real time to trigger the redistribution of the thesaurus.