Large-scale online course group clustering method
By using graph convolution network and word embedding model to perform structural reorganization and denoising of online course catalogs, automated course classification without manual annotation is achieved, the problem of low classification accuracy in the existing technology is solved, and the accuracy and efficiency of classification are improved.
Patent Information
- Application Number
- CN202510216988.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
The existing online course classification methods mainly rely on manual annotation and existing standards, and the ability to represent the actual content of the course is insufficient, resulting in insufficient accuracy of classification results.
The large-scale online course clustering method is adopted to train the process of sample data generation, neural network training and automatic course classification, and use graph convolution network (GNN) and word embedding models to reorganize and denoise the course catalog to generate hierarchical classification of courses.
It realizes automated course classification without manual annotation, improves the accuracy and efficiency of classification, and can directly obtain hierarchical classification results.
Smart Images

Figure CN120067734A_ABST
Abstract
Claims
1. A method for clustering large-scale online courses, characterized in that: The specific process is: Training sample data generation: Capture the online course catalog, and use the prompt word to guide the model to reorganize the catalog structure and denoise it, build the catalog graph dgl_G, and then use the word embedding model to generate the initial embedding vector for each node. The initial embedding vectors of all nodes constitute the node embedding matrix; Neural network training: construct a GNN network based on the graph convolutional network, and use the directory graph and embedding matrix to train the neural network. During training, forward propagation is performed to obtain the updated node embedding of each layer, and then back propagation is performed to calculate the gradient and update the model parameters to ensure that the model fully learns the information in the directory structure and outputs the embedding of the root node; Automatic course clustering: Classify courses based on the root node embedding output by the neural network.
2. The method for clustering massive online courses according to claim 1, characterized in that: The specific process of automatic course classification is as follows: using the root node embedding at the end of training as the final overall representation of the course, performing dimensionality reduction and clustering processing on it, and averaging the embeddings of courses within a category as the vector representation of the category; The category names are generated by extracting multiple subject terms from the course catalog of each category.
3. The method for clustering massive online courses according to claim 2, characterized in that: The dimension reduction and clustering processing is as follows: using UMAP to reduce the dimension of course embedding and reduce it to 5 dimensions, using HDBSCAN to divide the several embeddings after UMAP dimension reduction into different clusters with a hierarchical structure, so as to realize the automatic classification of course groups; the category name generation is as follows: using the TF-IDF algorithm to extract the first 10 keywords for the course catalog of each category, and inputting them into LLM to generate category names that conform to human language habits and professional catalog style.
4. The method for clustering massive online courses according to claim 1, characterized in that: The specific process of automatic course classification is as follows: use K-means to perform preliminary clustering on the embedding of the course root nodes, and set the number of clusters to 12, corresponding to 12 subject categories; for each cluster, select several courses closest to the cluster center as representative courses to determine the subject category to which they belong.
5. The method for clustering massive online courses according to claim 2 or 4, characterized in that: When there is a new online course, the new course is assigned to the existing classification that is closest to the overall embedding representation of the course; the specific process is: (1) Generate a directory graph and node embedding matrix for the new course, and use a neural network to obtain the embedding of the root node of the new course; (2) Calculate the similarity between the root node embedding of the new course and the embedding of the lowest level category in the existing hierarchical classification, and include the course in the lowest level category with the highest similarity.
6. The method for clustering massive online courses according to claim 2 or 4, characterized in that: When the crawled online course catalog is: (1) Hierarchical structure: the hierarchical structure of chapters, sections, and subsections is extracted; (2) Ordered list: extract the hierarchical structure of chapters, sections, and subsections according to the list sequence number; (3) Unordered paragraphs: Based on the words between paragraphs that directly or indirectly indicate the chapter order and directory structure, the hierarchical structure of chapters, sections, and subsections is extracted; (4) Incomplete structure: The semantic relationship and order of appearance of each chapter in the course catalog are comprehensively considered to extract the hierarchical structure of chapters, sections, and subsections.
7. The method for clustering massive online courses according to claim 6, characterized in that: The prompt word guidance model includes "chapter", "section", and "subsection", corresponding to "chapter", "section", and "byte", wherein a chapter contains several sections stored in a sections list, and a section contains several subsections stored in a subsections list.
8. The method for clustering massive online courses according to claim 1, characterized in that: The denoising is to remove indicators of chapter structure and order, as well as common words commonly found in online course catalogs.
9. The method for clustering massive online courses according to claim 6, characterized in that: The directory graph dgl_G takes the course name as the root node and is generated in sequence according to the three-level node order of chapter, section, and subsection, and the inclusion relationship between nodes is used as the edge between every two levels of nodes, with a total of four layers of nodes; the row order of the node embedding matrix is the node order in the process of constructing dgl_G, the text of the node is input into the pre-trained embedding model to obtain the embedding vector, and the initial embedding list of all nodes constitutes the node embedding matrix.
10. The method for clustering massive online courses according to claim 1, characterized in that: During training, the mean square error loss function is used to calculate the difference between the model output and the original node embedding as the optimization target. The loss function formula is: Where N is the number of nodes, y i is the i-th element in the original node embedding, is the i-th element in the node embedding output by the model.