Large-scale online course group clustering method

By using graph convolution network and word embedding model to perform structural reorganization and denoising of online course catalogs, automated course classification without manual annotation is achieved, the problem of low classification accuracy in the existing technology is solved, and the accuracy and efficiency of classification are improved.

CN120067734APending Publication Date: 2025-05-30BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510216988.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing online course classification methods mainly rely on manual annotation and existing standards, and the ability to represent the actual content of the course is insufficient, resulting in insufficient accuracy of classification results.

Method used

The large-scale online course clustering method is adopted to train the process of sample data generation, neural network training and automatic course classification, and use graph convolution network (GNN) and word embedding models to reorganize and denoise the course catalog to generate hierarchical classification of courses.

Benefits of technology

It realizes automated course classification without manual annotation, improves the accuracy and efficiency of classification, and can directly obtain hierarchical classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067734A_ABST
    Figure CN120067734A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of online education, and particularly relates to a large-scale online course group clustering method, which specifically comprises the following steps of: generating training sample data: capturing an online course catalogue, recombining a catalogue structure by utilizing a cue word guide model, denoising, constructing a catalogue graph dglG, and constructing a database; generating an initial embedding vector for each node by using a word embedding model, wherein the initial embedding vectors of all the nodes form a node embedding matrix; and neural network training: constructing a GNN network based on a graph convolutional network, training the neural network by using the directory graph and the embedded matrix, executing forward propagation during training to obtain updated node embedding of each layer, then executing back propagation to calculate a gradient and update model parameters so as to ensure that the model fully learns information in the directory structure, and finally, performing neural network training. Embedding an output root node; and course automatic clustering: performing course category division based on root node embedding output by the neural network.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A method for clustering large-scale online courses, characterized in that: The specific process is: Training sample data generation: Capture the online course catalog, and use the prompt word to guide the model to reorganize the catalog structure and denoise it, build the catalog graph dgl_G, and then use the word embedding model to generate the initial embedding vector for each node. The initial embedding vectors of all nodes constitute the node embedding matrix; Neural network training: construct a GNN network based on the graph convolutional network, and use the directory graph and embedding matrix to train the neural network. During training, forward propagation is performed to obtain the updated node embedding of each layer, and then back propagation is performed to calculate the gradient and update the model parameters to ensure that the model fully learns the information in the directory structure and outputs the embedding of the root node; Automatic course clustering: Classify courses based on the root node embedding output by the neural network.

2. The method for clustering massive online courses according to claim 1, characterized in that: The specific process of automatic course classification is as follows: using the root node embedding at the end of training as the final overall representation of the course, performing dimensionality reduction and clustering processing on it, and averaging the embeddings of courses within a category as the vector representation of the category; The category names are generated by extracting multiple subject terms from the course catalog of each category.

3. The method for clustering massive online courses according to claim 2, characterized in that: The dimension reduction and clustering processing is as follows: using UMAP to reduce the dimension of course embedding and reduce it to 5 dimensions, using HDBSCAN to divide the several embeddings after UMAP dimension reduction into different clusters with a hierarchical structure, so as to realize the automatic classification of course groups; the category name generation is as follows: using the TF-IDF algorithm to extract the first 10 keywords for the course catalog of each category, and inputting them into LLM to generate category names that conform to human language habits and professional catalog style.

4. The method for clustering massive online courses according to claim 1, characterized in that: The specific process of automatic course classification is as follows: use K-means to perform preliminary clustering on the embedding of the course root nodes, and set the number of clusters to 12, corresponding to 12 subject categories; for each cluster, select several courses closest to the cluster center as representative courses to determine the subject category to which they belong.

5. The method for clustering massive online courses according to claim 2 or 4, characterized in that: When there is a new online course, the new course is assigned to the existing classification that is closest to the overall embedding representation of the course; the specific process is: (1) Generate a directory graph and node embedding matrix for the new course, and use a neural network to obtain the embedding of the root node of the new course; (2) Calculate the similarity between the root node embedding of the new course and the embedding of the lowest level category in the existing hierarchical classification, and include the course in the lowest level category with the highest similarity.

6. The method for clustering massive online courses according to claim 2 or 4, characterized in that: When the crawled online course catalog is: (1) Hierarchical structure: the hierarchical structure of chapters, sections, and subsections is extracted; (2) Ordered list: extract the hierarchical structure of chapters, sections, and subsections according to the list sequence number; (3) Unordered paragraphs: Based on the words between paragraphs that directly or indirectly indicate the chapter order and directory structure, the hierarchical structure of chapters, sections, and subsections is extracted; (4) Incomplete structure: The semantic relationship and order of appearance of each chapter in the course catalog are comprehensively considered to extract the hierarchical structure of chapters, sections, and subsections.

7. The method for clustering massive online courses according to claim 6, characterized in that: The prompt word guidance model includes "chapter", "section", and "subsection", corresponding to "chapter", "section", and "byte", wherein a chapter contains several sections stored in a sections list, and a section contains several subsections stored in a subsections list.

8. The method for clustering massive online courses according to claim 1, characterized in that: The denoising is to remove indicators of chapter structure and order, as well as common words commonly found in online course catalogs.

9. The method for clustering massive online courses according to claim 6, characterized in that: The directory graph dgl_G takes the course name as the root node and is generated in sequence according to the three-level node order of chapter, section, and subsection, and the inclusion relationship between nodes is used as the edge between every two levels of nodes, with a total of four layers of nodes; the row order of the node embedding matrix is ​​the node order in the process of constructing dgl_G, the text of the node is input into the pre-trained embedding model to obtain the embedding vector, and the initial embedding list of all nodes constitutes the node embedding matrix.

10. The method for clustering massive online courses according to claim 1, characterized in that: During training, the mean square error loss function is used to calculate the difference between the model output and the original node embedding as the optimization target. The loss function formula is: Where N is the number of nodes, y i is the i-th element in the original node embedding, is the i-th element in the node embedding output by the model.