Word classification method and system based on K-means clustering algorithm

The K-mean clustering algorithm classifies the overall structure of words, which solves the problem of confusion in English word learning and improves memory efficiency.

CN120296469APending Publication Date: 2025-07-11CHONGQING SITU YUANJING TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510356348.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art fails to effectively utilize the overall structure of words when categorizing English words, resulting in learners being prone to confuse similar words and affecting memory efficiency.

Method used

The K-mean clustering algorithm is used to extract the characteristic parameters of the word (word letter string) for clustering, calculate the similarity and class feature parameters are optimized to improve classification accuracy.

Benefits of technology

Classification is achieved according to the overall structure of words, improving the effect of learning and memorizing similar words.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses a word classification method and system based on a K-means clustering algorithm. Classification of words according to structures comprises the following steps: step 1, extracting feature parameters of new words; 2, clustering the feature parameters based on a preset K-means clustering algorithm; 3, calculating the similarity of the feature parameters of the new words relative to the class in which each feature parameter is temporarily included; and step 4, comparing the similarity of each class obtained in the step 3, incorporating the new words into the class with the corresponding maximum similarity, and removing the new words from other temporarily incorporated classes. And 5, the steps 1-4 are repeated, and the word alphabet strings selected every time are different until the word alphabet strings included in the new word are exhaustive. Words with similar overall structures can be classified, and classified learning of the words with similar overall structures can be provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text classification, and particularly relates to a word classification method and system based on the K-means clustering algorithm. Background Art

[0002] English is a spelling language. The word formation process is that letters form phonemes, then phonemes form morphemes, and finally morphemes form words. However, when learners face words with similar spellings, they are often visually confused, resulting in memory difficulties and affecting learning efficiency.

[0003] The patent with the publication number CN 110909161 B discloses an English word classification method based on density clustering and visual similarity. The steps are as follows: preprocessing English words; calculating the visual similarity and distance between the selected English word and each English word; if the number of words in the neighborhood set is greater than or equal to 2, then add the selected word to the empty cluster and select the unclassified English word; process the unvisited unclassified English words; for the visited unclassified English words, directly add them to the cluster; determine whether there are unclassified English words. If there are, select the unclassified English word, otherwise mark the cluster as a new class of words and mark them as visited; if all words have been visited, output each class.

[0004] This patent preprocesses the selected English word and each English word in the original set of words to be classified, calculates the visual similarity between the selected English word and each English word in the original set of words to be classified, and calculates the distance between the selected English word and each English word in the original set of words to be classified. It overcomes the problem that the character similarity between any two characters in the English words to be compared needs to be calculated before calculating the similarity of English words, which makes the calculation amount increase, resulting in the problems of slow calculation speed and unclear distinguishing features between words after classification. Therefore, it cannot better improve the efficiency of memorizing words. The present invention has the advantages of fast similarity calculation speed, and uses visual similarity to well measure the distinguishing features between English words for accurate classification, which is beneficial for memory.

[0005] This patent is based on individual letters and does not consider the overall structural similarity between words. However, learners are often confused by words that look similar as a whole during learning.

[0006] The k-means clustering algorithm is an iterative clustering analysis algorithm. Its steps are as follows: The data is divided into K groups, and K objects are randomly selected as the initial clustering centers. Then, the distance between each object and each seed clustering center is calculated, and each object is assigned to the clustering center closest to it. The clustering centers and the objects assigned to them represent a cluster. Each time a sample is assigned, the clustering centers of the cluster are recalculated based on the existing objects in the cluster. This process will be repeated continuously until a certain termination condition is met. The termination condition can be that no (or the minimum number of) objects are reassigned to different clusters, no (or the minimum number of) clustering centers change anymore, or the sum of squared errors is locally minimized. Summary of the Invention

[0007] To make up for the above deficiencies, the present invention aims to provide a word classification method and system based on the k-means clustering algorithm. By using the k-means clustering algorithm to classify words, it can classify words according to the similarity or proximity in the overall structure, and can group words with similar overall structures together, which is convenient for learning and memorization.

[0008] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0009] A word classification method based on the k-means clustering algorithm, which classifies words according to their structures. The step of classifying words according to their structures includes the following steps:

[0010] Step 1: Extract the characteristic parameters of the new word. The characteristic parameters are the word letter strings contained in the new word. The word letter strings are composed of adjacent letters contained in the new word, and the word letter strings are randomly selected.

[0011] Step 2: Cluster the characteristic parameters based on a preset k-means clustering algorithm, and temporarily incorporate the new word into the class containing class characteristic parameters similar to the characteristic parameters. The class characteristic parameters are letter strings, denoted as class letter strings.

[0012] Step 3: Calculate the similarity of the characteristic parameters of the new word with respect to each class into which the characteristic parameters are temporarily incorporated.

[0013] Step 4: Compare the similarities of each class obtained in Step 3, incorporate the new word into the class corresponding to the maximum similarity, and remove the new word from other classes into which it is temporarily incorporated.

[0014] Step 5: Repeat Steps 1-4, where each time the selected word letter strings are different until the word letter strings contained in the new word are exhausted.

[0015] Furthermore, after Step 5, it also includes:

[0016] Step 6: Redetermine the class feature parameters of the class incorporating the new word.

[0017] Preferably, in Step 4, the calculation method of the similarity is as follows:

[0018] Let the number of letters of the new word be A, the number of letters of the word letter string be n, and the number of letters of the class letter string be the same as that of the word letter string;

[0019] Compare the first to the nth letters of the word letter string and the class letter string one by one, count the number of identical letters as m, the number of i consecutive adjacent identical letters as pi, and the number of similar letters as h. The similarity D is determined by the following formula:

[0020] D = [m÷n + ∑(pi÷n×i) + 0.5×h÷n]÷A

[0021] where i = 2,..., m.

[0022] Preferably, the similar letters are "u" and "v", "c" and "o", "p" and "q", "b"

[0023] and "d".

[0024] Preferably, the method for clustering the feature parameters by the preset K-means clustering algorithm is as follows:

[0025] First step, initialization: Randomly select k words from the word library as the initial centroids to establish k classes;

[0026] Second step, assignment: Assign the remaining words to the class with the maximum similarity to the initial centroid;

[0027] Third step, update: Recalculate the centroid of each class;

[0028] Fourth step, repeat the second and third steps until the class assignment no longer changes or the maximum number of iterations is reached.

[0029] Preferably, in the second step, the similarity is calculated according to the method in Step 4, and the class letter string is the letter string randomly selected from the initial centroid.

[0030] Preferably, in the third step, the method for recalculating the centroid of each class is as follows:

[0031] Extract the feature parameters of all words in the class, and use the word with the maximum average similarity to all other words in the class as the centroid.

[0032] The present invention also discloses a system for implementing the above-mentioned word classification method based on the K-means clustering algorithm.

[0033] Preferably, the system includes:

[0034] An input module for reading in new words and extracting the characteristic parameters of the new words;

[0035] A first calculation module for implementing steps 2-5;

[0036] A storage module for classifying and storing words and being called by the calculation module.

[0037] Furthermore, the system further includes:

[0038] A second calculation module for implementing the first step, the second step, the third step, and the fourth step;

[0039] A database module for storing a word library and being called by the second calculation module.

[0040] The beneficial effects of the present invention are as follows:

[0041] The present invention classifies words according to the similarity or proximity in the overall structure of the words, can classify words with similar overall structures, and can provide classified learning of words with similar overall structures, improving the learning and memory effects. Specific embodiments

[0042] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below.

[0043] Embodiment 1

[0044] This embodiment discloses a word classification method based on the K-means clustering algorithm. This method classifies words according to their structures so as to utilize word categories with similar structures for learning.

[0045] Specifically, classifying words according to their structures includes the following steps:

[0046] Step 1: Extract the characteristic parameters of the new word. The characteristic parameters are the word letter strings included in the new word. The word letter string is composed of adjacent letters included in the new word in sequence, and the word letter string is randomly selected;

[0047] Step 2: Cluster the characteristic parameters based on the preset K-means clustering algorithm, and temporarily incorporate the new word into the class containing the class characteristic parameters similar to the characteristic parameters. The class characteristic parameters are letter strings, denoted as class letter strings;

[0048] Step 3: Calculate the similarity of the characteristic parameters of the new word with respect to each class into which the characteristic parameters are temporarily incorporated;

[0049] Step 4: Compare the similarities of each class obtained in Step 3, incorporate the new word into the class corresponding to the maximum similarity, and remove the new word from other classes into which it is temporarily incorporated.

[0050] Step 5: Repeat Steps 1-4, where each selected word letter string is different until the word letter strings included in the new word are exhausted. For example, for the first time, select the first to the third letters of the new word, for the second time, select the first to the fourth letters of the new word, for the third time, select the first to the fifth letters of the new word, and so on until the combination of letters is exhausted.

[0051] In Step 4, the calculation method of the similarity is as follows:

[0052] Let the number of letters in the new word be A, the number of letters in the word letter string be n, and the number of letters in the class letter string be the same as that in the word letter string;

[0053] Compare the first to the nth letters of the word letter string and the class letter string one by one. Count the number of identical letters as m, the number of i consecutive adjacent identical letters as pi, and the number of similar letters as h. The similarity D is determined by the following formula:

[0054] D = [m÷n + ∑(pi÷n×i) + 0.5×h÷n]÷A

[0055] where i = 2, …, m. The similar letters are "u" and "v", "c" and "o", "p" and "q", "b" and "d".

[0056] Calculated according to the above formula, it comprehensively considers the number of identical letters, and also considers whether the identical letters are consecutive, as well as the influence of similar letters on the similarity of words, and assigns different weights. It can fully reflect the similarity between words.

[0057] After Step 5, it further includes:

[0058] Step 6: Redetermine the class feature parameters of the class incorporated into the new word.

[0059] Step 6 can further optimize the class.

[0060] Embodiment 2

[0061] Based on Embodiment 1, the specific method for clustering the feature parameters by the preset K-means clustering algorithm in this embodiment is as follows:

[0062] The first step is initialization. Randomly select k words in the word library as the initial centroids to establish k classes;

[0063] The second step is assignment. Assign the remaining words to the class with the greatest similarity to the initial centroid;

[0064] The third step is update. Recalculate the centroid of each class;

[0065] Fourthly, repeat the second and third steps until the class assignment no longer changes or the maximum number of iterations is reached.

[0066] Among them, in the second step, the similarity is calculated according to the method in step 4, and the class letter string is the letter string randomly selected as the initial centroid.

[0067] In the third step, the method for recalculating the centroid of each class is as follows:

[0068] Extract the feature parameters of all words in the class, and use the word with the largest average similarity to all other words in the class as the centroid.

[0069] Embodiment 3

[0070] Based on Embodiments 1 and 2, this embodiment discloses a system for implementing the above-mentioned word classification method based on the K-means clustering algorithm.

[0071] Specifically, the system includes:

[0072] An input module for reading in new words and extracting the feature parameters of the new words.

[0073] A first calculation module for implementing steps 2-5.

[0074] A storage module for classifying and storing words and being called by the calculation module.

[0075] A second calculation module for implementing the first, second, third, and fourth steps.

[0076] A database module for storing a word library and being called by the second calculation module.

[0077] Of course, the present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.

Claims

1. A word classification method based on the K-means clustering algorithm, characterized in that, Classify words according to their structures, and the classification of words according to their structures includes the following steps: Step 1: Extract the characteristic parameters of a new word, where the characteristic parameters are the strings of word letters included in the new word. The string of word letters is composed of adjacent letters included in the new word, and the string of word letters is randomly selected; Step 2: Cluster the characteristic parameters based on the preset K-means clustering algorithm, and temporarily incorporate the new word into the class that includes the class characteristic parameters similar to the characteristic parameters. The class characteristic parameters are strings of letters, denoted as class strings of letters; Step 3: Calculate the similarity of the characteristic parameters of the new word with respect to each class into which the characteristic parameters are temporarily incorporated; Step 4: Compare the similarities of each class obtained in Step 3, incorporate the new word into the class corresponding to the maximum similarity, and remove the new word from other classes into which it is temporarily incorporated. Step 5: Repeat Steps 1-4, where the strings of word letters selected each time are different until the strings of word letters included in the new word are exhausted.

2. The word classification method based on the K-means clustering algorithm according to claim 1, characterized in that, After Step 5, it further includes: Step 6: Redetermine the class characteristic parameters of the class into which the new word is incorporated.

3. The word classification method based on the K-means clustering algorithm according to claim 1, wherein In Step 4, the calculation method of the similarity is as follows: Let the number of letters of the new word be A, the number of letters of the string of word letters be n, and the number of letters of the class string of letters is the same as the number of letters of the string of word letters; Compare the first to the nth letters of the string of word letters and the class string of letters one by one, count the number of identical letters as m, the number of consecutive adjacent identical letters as pi for i, and the number of similar letters as h. The similarity D is determined by the following formula: D = [m÷n + ∑(pi÷n×i) + 0.5×h÷n]÷A where i = 2, …, m.

4. The word classification method based on the K-means clustering algorithm according to claim 3, wherein The similar letters are "u” and "v”, "c” and "o”, "p” and "q”, "b” and "d”.

5. The word classification method based on the K-means clustering algorithm according to claim 3, characterized in that, In Step 2, the method for clustering the characteristic parameters by the preset K-means clustering algorithm is as follows: The first step: Initialization, randomly select k words from the word library as the initial centroids to establish k classes; The second step: Assignment, assign the remaining words to the class with the maximum similarity to the initial centroid; The third step: Update, recalculate the centroid of each class; The fourth step: Repeat the second step and the third step until the class assignment no longer changes or reaches the maximum number of iterations.

6. The word classification method based on the K-means clustering algorithm according to claim 5, characterized in that, In the second step, the similarity is calculated according to the method in Step 4, and the class string of letters is the string of letters randomly selected as the initial centroid.

7. The word classification method based on the K-means clustering algorithm according to claim 6, characterized in that, The third step, the method for recalculating the centroid of each class is as follows: Extract the characteristic parameters of all words in the class, and use the word with the maximum average similarity to all other words in the class as the centroid.

8. A system, characterized in that, The system is used to implement the word classification method based on the K-means clustering algorithm as described in any one of claims 1-7.

9. The system according to claim 8, wherein It includes: An input module, used to read in a new word and extract the characteristic parameters of the new word; A first calculation module, used to implement Steps 2-5; A storage module, used to store words classified and called by the calculation module.

10. The system according to claim 7, wherein When it refers to claims 5-7, it further includes: A second calculation module, used to implement the first step, the second step, the third step, and the fourth step; A database module, used to store the word library and called by the second calculation module.

Citation Information

Patent Citations

  • English word classification method based on density clustering and visual similarity

    CN110909161B