An online learning behavior personalized recommendation system based on cluster analysis
The online learning behavior personalized recommendation system based on cluster analysis solves the problem of the lack of personalized recommendations on online learning platforms. By building student profiles and recommending similar learning friends, it improves learning effectiveness and enthusiasm, and forms an effective learning community.
Patent Information
- Application Number
- CN202311029204.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-16
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-08-16
AI Technical Summary
Existing online learning platforms lack in-depth analysis of students' learning processes and behaviors, and cannot provide personalized recommendation systems, resulting in students being unable to accurately select suitable course resources and achieving poor learning outcomes.
Design a personalized recommendation system for online learning behavior based on cluster analysis. By collecting and cleaning learning data, construct student profiles, and use cluster analysis to create personalized profiles for students from three dimensions: learning attitude style, learning interest preference, and learning level ability. Recommend similar learning friends to form a social learning ecosystem.
It enables in-depth analysis of students' learning behaviors, provides personalized course recommendations, improves learning outcomes, stimulates students' enthusiasm for learning, and forms an effective learning community.
Smart Images

Figure CN117056616B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of recommendation system and data mining technology, and particularly to an online learning behavior personalized recommendation system based on cluster analysis. BACKGROUND
[0002] With the development of artificial intelligence, big data and other technologies, the learning mode of students gradually changes from the traditional classroom mode to the online mode. At present, many online learning platforms have been explored in the field of distance education at home and abroad. For example, the online learning platform represented by MOOC realizes the fundamental change of knowledge dissemination mode by means of networking. Classroom education is no longer limited by time and place, and realizes ubiquitous learning. However, from the data analysis of students using the online learning platform, the following shortcomings can be seen. First, a large number of learning resources cause students to immerse in the environment of selecting resources, so that they cannot accurately select the course resources suitable for their own learning characteristics, resulting in problems such as inability to concentrate on learning and poor learning effect. Second, the online learning platform only provides a large number of course resources, lacks tracking and evaluation of students' learning process, learning behavior and learning effect, and cannot analyze the deep learning information of students, so it cannot truly realize closed-loop feedback and effectively assist students in learning. Therefore, on the basis of the existing online learning platform, it is an urgent task to realize individualized teaching for different types of students and provide personalized recommendation system.
[0003] At present, many researchers have conducted research on online learning behavior data of online learning platform and personalized recommendation system field. For example, Liu Feifan analyzed the time series characteristics of online learning behavior data and found that it has the general characteristics of time series. However, he only verified the feasibility of applying time series analysis method to process network learning behavior data, and did not give specific processing method model. For example, Gui Zhongyan et al. calculated the similarity of user learning behavior sequence, and adopted user-based collaborative filtering recommendation modeling. Wang Yonggu et al. proposed a theoretical model of learning resource personalized recommendation system based on collaborative filtering technology and discussed the key technologies in the model. However, they analyzed by using traditional recommendation systems such as collaborative filtering and content-based recommendation system, and failed to fully consider the personalized online learning behavior characteristics of students in the recommendation process, but recommended learning content according to the content rating.
[0004] In summary, the present application solves the existing problems by designing an online learning behavior personalized recommendation system based on cluster analysis. SUMMARY
[0005] The present application aims to provide an online learning behavior personalized recommendation system based on cluster analysis to solve the problems raised in the background art.
[0006] To achieve the above object, the present application provides the following technical solutions.
[0007] An online learning behavior personalized recommendation system based on cluster analysis is composed of three parts: the first part is to collect and clean the data obtained from the online learning platform to build an online learning student portrait label system;
[0008] The second part is to create an online student portrait for each online student based on cluster analysis from three dimensions of learning attitude style, learning interest preference and learning level ability;
[0009] The third part is to recommend similar classmates to each other based on the personalized recommendation model of the same type of learning friends, so as to realize the personalized recommendation of the same type of learning friends;
[0010] The online learning behavior personalized recommendation system based on cluster analysis is mainly divided into three functional modules, namely online learning behavior data information management module, data analysis modeling module and personalized recommendation module of the same type of learning friends, and the specific process of the online learning behavior personalized recommendation system is as follows:
[0011] S1, first, the process data of the student online learning can be analyzed to establish the student portrait model, therefore, the online learning behavior data information management module collects and obtains the learning behavior data of different dimensions collected from different functional modules of the platform;
[0012] S2, secondly, the data analysis modeling module extracts, cleans and classifies the data according to the demand label, and uses the cluster analysis algorithm to model and analyze the student online learning behavior data, constructs the online learning portrait of the student from three dimensions of learning attitude style, learning interest preference and learning level ability, and saves and displays the portrait data, which constitutes the recommendation data set of the online learning personalized recommendation service;
[0013] S3, finally, the personalized recommendation module of the same type of learning friends analyzes the data set by using a suitable recommendation algorithm to recommend the same type of learning friends to the student on the online learning platform, so as to provide personalized recommendation service and establish a solid foundation for the online social learning ecology.
[0014] As a preferred scheme of the present application, the online learning behavior data information management module includes a student information library and a data label definition library, the student information library mainly saves the learning behavior data of different dimensions collected from the online learning platform, and the online learning platform records the online learning behavior of the student from the registration of the platform account, and generates a large amount of continuous learning data;
[0015] The analysis of the students' online learning data can be divided into two dimensions:
[0016] (1) Basic information data: refers to the user profile information filled in manually by students when registering the platform, including but not limited to: name, gender, date of birth, political affiliation, education, major, college, and interest information;
[0017] (2) Online learning behavior data: refers to selecting 9 indicators recorded in the online learning platform to measure the parameters, and classifying these indicators according to the three dimensions of the online learning behavior model:
[0018] The first dimension is the learning operation behavior dimension, which refers to the operation behavior of students on the online learning platform. The specific data indicators to measure this dimension are the type and total number of access to course resources, the type and total duration of learning course resources, the type and completion rate of learning course resources, and the type and total number of downloading course documents;
[0019] Among them, the type and total number of access to course resources: the background classifies various course resource types on the platform, and directly counts the type and total number of clicks on the platform by different students according to the log record data of the online learning platform background;
[0020] The type and total duration of learning course resources: the log record data of the online learning platform background respectively counts the total duration of learning course resources by different students on the platform;
[0021] The type and completion rate of learning course resources: the completion rate of course resources on the online platform is divided into 0%-100% according to the progress bar, 0% means that the course has not been watched and completed, and 100% means that the course has been watched and completed. The completion rate of the same type of course resource is the average value of the completion rate of watching and learning course resources. According to the log record data of the background, the average value of the completion rate of learning courses by different students is calculated;
[0022] The type and total number of downloading course documents: the log record data of the online learning platform background directly counts the completion rate of browsing course materials by different students on the platform;
[0023] The second dimension is the learning problem solving behavior dimension, which refers to the problem solving ability of students on the online learning platform. The specific data indicators to measure this dimension are:
[0024] 1) Online exercise completion rate: most learning course resources are followed by corresponding course materials, and students further consolidate the knowledge points learned after watching the video course through course assignments. The background log records the completion rate of students submitting course assignments.
[0025] 2) the score of the online practice homework: after the student submits the practice homework, the background log records the score of the student's practice homework;
[0026] The third dimension is the collaborative learning behavior dimension, which refers to the learning behavior ability of students to cooperate and communicate with other students on the online learning platform. The specific data indicators to measure this dimension are the number of posts in the interaction area, the number of replies in the interaction area, and the number of learning notes published;
[0027] Among them, the number of posts in the interaction area: the background log records the learning behavior trajectory data of students in the interaction area, and counts the total number of questions or posts published by students;
[0028] The number of replies in the interaction area: the background log counts the total number of replies to comments in the interaction area to further measure the activity of students' collaborative learning behavior on the platform;
[0029] The number of learning notes published: the background log counts the number of learning notes published by students online;
[0030] The data tag definition library mainly saves the online learning student portrait tag system from the learning behavior data saved in the student information library, and constructs the online learning student portrait tag system from different data dimensions. Based on the interviews with experts in the education industry and a large number of literature materials, the construction of online learning student portrait is defined from the following three dimensions:
[0031] 1) Learning attitude style: from this dimension, we can define:
[0032] a. Active learning type: students have a long online learning time and a high learning frequency, submit homework frequently, and have a high activity level in the interaction area, such as publishing comments, submitting questions, and answering questions. The platform has high stickiness and dependence;
[0033] b. Independent learning type: students have a long online learning time and a high learning frequency, but have a low activity level in the interaction area, do not fully use the social functions of the platform, and have little interaction with teachers and other students;
[0034] c. Inert learning type: students basically have a low frequency of using the online learning platform after registration, and have a short duration of single course resource learning, and have no loyalty to the platform;
[0035] 2) Learning Interests and Hobbies: Based on the student's major and the type of learning resources accessed, the student's learning interest preferences are further determined. The course resource type data mainly refers to the classification of course resources on the online learning platform. Course resources are divided into 14 categories according to disciplines, namely philosophy, economics, law, education, literature, history, science, engineering, agriculture, medicine, military science, management, art, and physical education.
[0036] 3) Learning level ability: The student's learning level ability can be measured and analyzed from two aspects: learning problem-solving behavior and collaborative learning behavior. Specifically, it includes the total number and frequency of homework submissions and homework score data. The learning level ability is divided into four categories: excellent, good, average, and poor.
[0037] As a preferred embodiment of the present invention, the data analysis and modeling module mainly constructs an online learning profile tagging system to classify and statistically display the data from the online learning behavior data information management module according to demand tags. Specifically, it uses a clustering analysis algorithm to model and analyze the data generated by students, constructing an online learning student profile for each student from three dimensions: learning attitude style, learning interest preference, and learning level ability. These student profile data of all students on the platform constitute the recommendation dataset for the personalized recommendation service of online learning. The specific steps are as follows:
[0038] First, the acquired database data is cleaned and processed to determine the valid fields. The valid fields include serial numbers 1-16 and names. The names are arranged in order according to the serial numbers and include, but are not limited to, Good University ID, name, student ID, school, college, major, grade, first time using the platform, total duration of learning courses, course completion rate, courseware download rate, online exercise completion rate, number of times study notes are posted, number of interactive posts, number of interactive replies, and online exercise scores. The Good University ID serves as the primary key and, along with the student ID, represents the user's index and does not contribute to the calculation.
[0039] After identifying valid fields and discarding invalid records, the following processing is performed:
[0040] 1. Referring to the Ministry of Education's subject / major classification standards, all majors were replaced with standard major names, and first-level, second-level, and third-level classification codes and names were added;
[0041] 2. The school type code is classified and coded according to whether the school is a key university, i.e., a 985 / 211 "Double First-Class" university. If it is, it is coded as "1" and otherwise as "0".
[0042] 3. The college is coded by using a one-hot encoding, and the original data contains 68 different colleges, which are coded by using a fixed-length binary code, such as 0100101, and there are 68 different colleges, so the length of the code is 7;
[0043] 4. The X-level subject code is coded by using the subject classification standard of China;
[0044] 5. The first use of the platform time is converted into the first use of the platform time from now on (days);
[0045] 6. The total length of the learning course (hours, minutes and seconds) is converted into the total length of the learning course (seconds);
[0046] 7. The completion rate of the learning course, the completion rate of the downloaded courseware browsing, the completion rate of the online practice, and other percentage data are converted into 0-100 numerical values;
[0047] 8. The school type code, college, and X-level subject code are combined into a student information code;
[0048] 9. The first use of the platform time from now on (days), the total length of the learning course (seconds), the completion rate of the learning course, the completion rate of the online practice, and the completion rate of the downloaded courseware browsing are combined into a completion degree, and the average value is taken;
[0049] 10. The number of published learning notes, the number of posts in the interaction area, and the number of replies in the interaction area are combined into an activity degree, and the average value is taken. After processing and combining, the effective fields are retained to generate a recommended system combined effective field table. Based on the various fields mastered, the numerical error and sequence similarity are generated to form a recommended system index table. The three types of indexes are integrated into a comprehensive similarity index, as shown in the following formula:
[0050] Similarity=Cos SC -KL SC -SE SC -AE SC -SE Mode -AE Score .
[0051] As a preferred scheme of the present application, the same type learning friend personalized recommendation module refers to analyzing the recommendation data set by using a suitable recommendation algorithm to mine the same type learning friends on the online learning platform from the learning attitude style, learning interest and hobby, and learning level and ability, and at the same time, the course recommendation list of the same type learning friends who have learned or evaluated highly is recommended to the student, and the students can communicate and supervise each other, which lays a foundation for forming a social learning ecology in the future.
[0052] As a preferred scheme of the present application, the clustering algorithm refers to a K-Means model, wherein the K-Means divides samples into K clusters, and similarity within the cluster is as high as possible, while similarity between clusters is as low as possible. In machine learning, distance is usually used to measure similarity between samples, and the smaller the distance, the higher the similarity, and the larger the distance, the lower the similarity. By using this clustering algorithm to analyze online learning student data, the correlation pattern between the "activity-completion" of online students is mined from the data, so as to determine the optimal value of the cluster number K, and then the calinski_harabasz_score, the silhouette_score and the total_inertia are used as three indexes of unsupervised learning for evaluating the model.
[0053] Calinski-Harabasz score (calinski_harabasz_score):
[0054]
[0055] Wherein S B represents the between-class variance, S W represents the within-class variance, tr(·) represents the trace operation of the matrix, k represents the cluster number, and N represents the sample point number.
[0056] Average silhouette coefficient (silhouette_score):
[0057]
[0058] Wherein a n represents the average value of the dissimilarity of point n to other points in the same cluster, and b n represents the minimum value of the average dissimilarity of point n to other clusters.
[0059] Total inertia (total_inertia):
[0060]
[0061] Wherein x ij represents the i-th point in the j-th cluster, and u j represents the center point of the j-th cluster.
[0062] From the clustering algorithm, the higher the calinski_harabasz_score and the silhouette_score, and the smaller the total_inertia, the better. Therefore, the comprehensive index cmp_index is defined as:
[0063]
[0064] Although the value of cmp_index should be higher in theory, but considering the problem of "too many clusters lead to meaningless data patterns" in clustering algorithm, the optimal number of clusters is determined under the acceptable complexity and the acceptable number of clusters, and the learning attitude style characteristics of students are determined according to the completion degree and the activity degree:
[0065] 1) Inert learning type: the learning attitude style of the student belongs to the extreme low, which is represented by the extremely low participation in the video course;
[0066] 2) Independent learning type: the student has a long learning time, the discussion index is not high, and the independent learning ability is strong;
[0067] 3) Active learning type: the student has a long learning time, and the post discussion is very active, and the enthusiasm is high.
[0068] Compared with the prior art, the beneficial effects of the present application are:
[0069] The design innovation point of the present application is that the original traditional online learning platform based on artificial intelligence technology adds new function points: 1) the big data analysis technology is used to clearly process the data accumulated by the online learning platform, obtain the basic information data and online learning behavior data of students, and classify the data of students on the platform, and the online behavior data is classified from three dimensions of learning operation behavior, learning problem solving behavior and cooperative learning behavior, and then three dimensions of learning attitude style, learning interest and hobby and learning level and ability are selected to construct the online learning student portrait label system, based on the label system, the personalized online student portrait of each online student can be created. Effectively utilize the data assets deposited in the platform to better play the data value of the student portrait of thousands of people; 2) based on the online student portrait, the recommendation algorithm is used to find the same type of learning friends matched with each student, and the learning course list of the same type of learning friends is recommended to each other, and an online learning community is created, students can communicate and discuss with each other, and the learning enthusiasm of students is further stimulated. BRIEF DESCRIPTION OF DRAWINGS
[0070] Figure 1 The business logic diagram for personalized recommendation of the same type of learning friends of the present application;
[0071] Figure 2 The label system diagram of the online learning student portrait of the present application;
[0072] Figure 3 The flowchart of the construction process of the online learning student portrait of the present application;
[0073] Figure 4 The flowchart of the personalized recommendation algorithm based on the online learning student portrait of the present application. Detailed Implementation
[0074] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0075] For examples, please refer to Figures 1-4 The present invention provides a technical solution:
[0076] A personalized recommendation system for online learning behavior based on cluster analysis is mainly divided into three functional modules: online learning behavior data information management module, data analysis and modeling module, and personalized recommendation module for similar learning friends. The online learning behavior data information management module includes a student information database and a data tag definition database. The student information database mainly stores learning behavior data of different dimensions collected from the online learning platform. The online learning platform records students' online learning behavior from the moment they register their platform accounts, generating a large amount of continuous learning data. The workflow is as follows: (1) Basic information data: refers to the user information manually filled in by students when registering on the platform, mainly including: name, gender, date of birth, political affiliation, education, major, college, hobbies, etc.; (2) Online learning behavior data: refers to the nine indicators of student learning behavior recorded in the online learning platform as measurement parameters, and these indicators are classified according to the three dimensions of the online learning behavior model; the specifics are as follows:
[0077] The first dimension is the learning operation behavior dimension, which refers to the students' operation behavior on the online learning platform. The specific data indicators for measuring this dimension are the types of course resources accessed and their total number of times, the types of course resources learned and their total duration, the types of course resources learned and their completion rate, and the types of after-class document resources downloaded and their total number of times.
[0078] 1) Types of course resources accessed and their total number of visits
[0079] The backend classifies various types of course resources on the platform, and the log data recorded by the backend of the online learning platform directly shows the types of course resources accessed by different students and the total number of clicks.
[0080] 2) Types of learning course resources and their total duration
[0081] The total time spent by different students on the platform to learn course resources was calculated from the log data recorded in the backend of the online learning platform.
[0082] 3) Learning course resource type and its completion rate
[0083] The completion rate of course resources on the online platform is divided into 0%-100% according to the progress bar, 0% means that the course has not been watched and completed, and 100% means that the course has been watched and completed. The completion rate of the same type of course resource is the average value of the completion rate of watching and learning course resources, and the average value of the completion rate of different students learning courses is calculated according to the background log record data.
[0084] 4) Download after-class document resource type and its total number
[0085] The completion rate of different students browsing platform courseware materials is directly calculated from the log record data of the online learning platform background;
[0086] The second dimension is the learning problem solving behavior dimension, which refers to the students' ability to solve learning problems on the online learning platform. The data indicators for measuring this dimension are:
[0087] 1) Online exercise completion rate
[0088] Most of the learning course resources are followed by corresponding after-class learning materials. After learning the video course, students further consolidate the knowledge points through after-class homework. The background log records the completion rate of students submitting after-class homework.
[0089] 2) Online exercise score
[0090] After the students submit the exercise, the background log records the score of the student's exercise.
[0091] The third dimension is the collaborative learning behavior dimension, which refers to the students' ability to cooperate and communicate with other students on the online learning platform. The data indicators for measuring this dimension are the number of posts in the interactive area, the number of replies in the interactive area, and the number of learning notes published.
[0092] 1) Number of posts in the interactive area
[0093] The background log records the learning behavior trajectory data of students in the interactive area, and counts the total number of questions or comments posted by students.
[0094] 2) Number of replies in the interactive area
[0095] The background log counts the total number of replies to comments in the interactive area to further measure the activity of students' collaborative learning behavior on the platform.
[0096] 3) Number of learning notes published
[0097] The background log counts the number of learning notes published by students online.
[0098] The data tag definition library mainly saves the online learning student portrait tag system from the learning behavior data saved in the student information library. Based on the interviews with experts in the education industry and a large number of literature materials, the construction of the online learning student portrait in this paper defines the classification indicators from the following three dimensions:
[0099] 1) Learning attitude style: From the dimension of learning attitude style, the following can be defined: 1. Active learning type, this student has a long online learning time and a high learning frequency, submits a large number of homework and publishes comments, submits problems and answers questions in the mutual evaluation area, and has a high degree of stickiness and dependence on the platform. 2. Independent learning type, this student has a long online learning time and a high learning frequency, but the activity in the mutual evaluation area of the platform is very low, and does not fully use the social function of the platform, and the interaction with teachers and other students is very little. 3. Inert learning type, this student basically has a low use frequency after registering the online learning platform, and the duration of single learning of course resources is short, and has no loyalty to the platform.
[0100] 2) Learning interest: According to the type of student's major and the type of access to learning course resources, the student's learning interest preference type is further obtained; the course resource type data mainly refers to classifying the course resources on the online learning platform. Course resources are classified into 14 categories according to disciplines, namely philosophy, economics, law, education, literature, history, science, engineering, agriculture, medicine, military science, management, art, and sports.
[0101] 3) Learning level: From the dimensions of learning problem solving behavior and collaborative learning behavior, the learning level of the student can be measured and analyzed, especially the total number and frequency of submitting homework, the score of homework, etc. The learning level is divided into four categories: excellent, good, general, and poor.
[0102] The data analysis modeling module mainly includes classifying and statistically presenting the data of the online learning behavior data information management module according to the demand tags, characterized by using clustering analysis algorithm to model and analyze the data generated by the students, and constructing the online learning student portrait of the students from the dimensions of learning attitude style, learning interest preference, and learning level. The student portrait data of all students on the platform constitutes the recommendation data set of the online learning personalized recommendation service.
[0103] First, the obtained database data is cleaned and processed, and the effective fields are as shown in the following table:
[0104] Serial Number Name Serial Number Name 1 Good University ID 9 Total Length of Learning Course 2 Name 10 Course Completion Rate 3 Student ID 11 Courseware Download Rate 4 School 12 Online Exercise Completion Rate 5 College 13 Number of Times of Sending Learning Notes 6 Major 14 Number of Times of Sending Interactive Posts 7 Grade 15 Number of Times of Replying to Interactive Posts 8 First Time of Using Platform 16 Online Exercise Score
[0105] As shown in the table, where
Good University ID
Student ID
[0106] After determining the valid fields and discarding invalid records, the processing is as follows:
[0107] 1. Refer to the Ministry of Education's subject / professional classification standard, replace all
major
[0108] 2.
School Type Code
[0109] 3.
College
[0110] 4.
X-level subject code
[0111] 5. Convert
First use platform time
First use platform time from now (days)
[0112] 6. Convert
Total length of learning courses (hours, minutes, seconds)
Total length of learning courses (seconds)
[0113] 7. Convert percentage data such as
Learning course completion rate
Download courseware browsing completion rate
Online practice completion rate
[0114] 8. Merge
School Type Code
College
X-level subject code
Student Information Code
[0115] 9. Merge
First use platform time from now (days)
Total length of learning courses (seconds)
Learning course completion rate
Online practice homework completion rate
Download courseware browsing completion rate
Completion
[0116] 10. Merge
Number of learning notes posted
Number of posts in interactive area
Number of replies in interactive area
Activity
[0117] Serial Number Field Name Field Description 1 id Good University ID 2 stu Student Information Code 3 cpt Learning Completion Degree 4 act Student Activity Degree 5 score Student Score
[0118] According to the numerical error and sequence similarity of various fields, a recommended system index table is generated, as shown in the following table:
[0119]
[0120] Integrate the three indicators into a comprehensive similarity index, as shown in the following formula:
[0121] Similarity=Cos SC -KL SC -SE SC -AE SC -SE Mode -AE Score
[0122] The clustering algorithm refers to the K-Means model, which has a simple principle but ideal effect and strong interpretability. K-Means divides samples into K clusters, with high similarity within the cluster and low similarity between clusters. In machine learning, distance is usually used to measure the similarity between samples. The smaller the distance, the higher the similarity, and the larger the distance, the lower the similarity. This clustering algorithm is used to analyze online learning student data to mine the correlation pattern between "activity" and "completion" of online students from the data, so as to determine the optimal value of the number of clusters (K), and then use the Calinski-Harabasz index, the silhouette index and the within-group variance as three indicators for evaluating the model.
[0123] Calinski-Harabasz index (calinski_harabasz_score):
[0124]
[0125] Where S B represents the between-class variance, S W represents the within-class variance, tr(·) represents the trace operation of the matrix, k represents the number of clusters, and N represents the number of sample points.
[0126] Average silhouette coefficient (silhouette_score):
[0127]
[0128] Where a n represents the average value of the dissimilarity of point n to other points in the same cluster, and b n represents the minimum value of the average dissimilarity of point n to other clusters.
[0129] Total cluster inertia (total_inertia):
[0130]
[0131] where x ij represents the i-th point in the j-th cluster, u j represents the center point of the j-th cluster.
[0132] From the clustering algorithm, the higher the calinski_harabasz_score and silhouette_score, the better, and the smaller the total_inertia, the better, so consider the comprehensive index cmp_index defined as:
[0133]
[0134] Although the numerical value of cmp_index should be higher in theory, but considering the problem of "too large number of clusters leading to meaningless data patterns" in the clustering algorithm. Determine the optimal
cluster number
completion
activity
[0135] 1) Inert learning type: This type of student almost does not watch video courses, and the participation is extremely low, indicating that the learning attitude style belongs to extremely low;
[0136] 2) Independent learning type: This type of student has long learning time, and the discussion index is not high, and the independent learning ability is strong;
[0137] 3) Active learning type: This type of student has long learning time, and the post discussion is very active, and the enthusiasm is high.
[0138] The same type of learning friend personalized recommendation module refers to obtaining the top five students with the highest similarity to the student on the platform by appropriate recommendation algorithm according to
comprehensive similarity index
[0139] Specific implementation case:
[0140] Refer to the attached Figure 1The same type of learning friend personalized recommendation business logic diagram is shown, which mainly consists of three parts: the first part is to collect and clean the data obtained from the online learning platform to build an online learning student portrait label system; the second part is to build an online student portrait for each online student from three dimensions of learning attitude style, learning interest preference and learning level ability based on clustering analysis method; the third part is to recommend similar students to each other based on the same type of learning friend personalized recommendation model, so as to realize the same type of learning friend personalized recommendation; the system mainly consists of three functional modules, which are online learning behavior data information management module, data analysis modeling module and same type of learning friend personalized recommendation module. Data source is crucial in establishing a learning portrait model. First, the process data of students' online learning should be collected to analyze students' learning behavior and establish a student portrait model. Therefore, the online learning behavior data information management module collects different dimensions of learning behavior data collected from different functional modules of the platform. Secondly, in the data analysis modeling module, the data is extracted, cleaned, classified and counted according to the demand labels, and the clustering analysis algorithm is used to model and analyze the students' online learning behavior data, so as to build the online learning student portrait of students from three dimensions of learning attitude style, learning interest preference and learning level ability. These portrait data constitute the recommendation data set of online learning personalized recommendation service. Finally, the same type of learning friend personalized recommendation module analyzes these data sets by using appropriate recommendation algorithm to recommend the same type of learning friend to a student on the online learning platform, so as to provide personalized recommendation service and lay a solid foundation for the establishment of online social learning ecology.
[0141] Referring to the drawings Figure 2, which shows the label system diagram of the online learning student portrait. The student portrait is constructed by selecting learning attitude style, learning interest, and learning level and ability as three dimensions. (1) Learning attitude style: from the dimension of learning attitude style, it can be defined that: 1. Active learning type: this student has a long online learning time and a high learning frequency, submits a large number of homework, and is active in publishing comments, submitting questions, and answering questions in the mutual evaluation area, and has a high degree of stickiness and dependence on the platform. 2. Independent learning type: this student has a long online learning time and a high learning frequency, but the activity in the mutual evaluation area of the platform is very low, and the social function of the platform is not fully used. The interaction with teachers and other students is very little. 3. Inert learning type: this student basically has a low frequency of using the online learning platform after registration, and the single learning time of course resources is short, and there is no loyalty to the platform. (2) Learning interest: according to the professional type of the student and the type of accessing learning course resources, the learning interest preference type of the student is further obtained; (3) Learning level and ability: from the dimensions of learning problem solving behavior and collaborative learning behavior, the learning level and ability of the student can be measured and analyzed, especially the total number and frequency of submitting homework, the score of homework, etc. The learning level and ability are divided into four categories: excellent, good, general, and poor.
[0142] Referring to the accompanying drawings Figure 3 , which shows the online learning student portrait construction flowchart. The specific process is that the student registers and fills in the information on the online learning platform to obtain the basic information data of the student. A series of behavior data generated by the student after logging in to the learning platform, such as the type and total number of course resources accessed, the type and total time of learning course resources, etc. are also recorded and stored in the database by the background log of the platform. Through extraction and cleaning of these data, these data can become valuable data that can be further analyzed and counted. Through the online learning portrait label system established in the previous subsection, these data are classified and counted according to the demand labels, and the clustering analysis algorithm is used to model and analyze the data generated by the student. From the three dimensions of learning attitude style, learning interest preference, and learning level and ability, the online learning student portrait of the student is constructed.
[0143] Referring to the accompanying drawings Figure 4 , which shows the personalized recommendation algorithm flowchart based on the online learning student portrait. After data cleaning, online learning behavior data is obtained by K-Means model in clustering analysis, behavior classification label is obtained by K-Means model in clustering analysis, learning interest is obtained by serializing processing, and comprehensive sequence similarity is obtained by serializing processing. Learning achievement is processed by normalization, and finally the similarity index is obtained by weighting balance.
[0144] While embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary of the principles and application of the present application. Numerous modifications and adaptions can be effected without departing from the spirit and scope of the present application, which is not limited to the exact construction and arrangement described. It is intended, therefore, to cover all modifications and adaptions that fall within the scope of the claims and their equivalents.
Claims
1. A personalized recommendation system for online learning behavior based on cluster analysis, characterized in that, It consists of three parts: the first part is data collection and cleaning from online learning platforms to build a student profile and tagging system for online learning; The second part uses cluster analysis to create a personalized online student profile for each online student from three dimensions: learning attitude and style, learning interest and preference, and learning level and ability. The third part is based on a personalized recommendation model for similar learning friends, which recommends students with similar online learning behaviors to each other, thus realizing personalized recommendation of similar learning friends. The personalized recommendation system for online learning behavior based on clustering analysis is mainly divided into three functional modules: online learning behavior data information management module, data analysis and modeling module, and personalized recommendation module for similar learning friends. The specific process of the personalized recommendation system for online learning behavior is as follows: S1. First, we need to collect the process data of students' online learning in order to analyze students' learning behavior and build a student profile model. Therefore, the online learning behavior data information management module collects and obtains learning behavior data from different functional modules of the platform from different dimensions. S2, Secondly, in the data analysis and modeling module, the data is extracted and cleaned by constructing an online learning profile tagging system, and statistically displayed according to the needs of the tags. Clustering analysis algorithm is used to model and analyze the online learning behavior data of students. From the three dimensions of learning attitude style, learning interest preference, and learning level ability, the online learning student profile is constructed, saved and displayed. Thus, the profile data constitutes the recommendation dataset for the personalized recommendation service of online learning. The profile data constitutes the recommendation dataset for personalized online learning recommendation services. The specific steps include: cleaning and processing the acquired database data, determining valid fields, and then, after determining valid fields and discarding invalid records, processing as follows: The activity level is calculated by merging the number of times study notes are posted, the number of times posts are made in the interaction area, and the number of times replies are made in the interaction area. The average value is then used to generate a table of merged effective fields for the recommendation system after processing and merging. The field table includes a serial number, field name, and field description. The serial number is an Arabic numeral 1, 2, 3, 4, 5. The field names are id, stu, cpt, act, and score. The field descriptions are Good University ID, Student Information Code, Learning Completion Rate, Student Activity Rate, and Student Grade. Based on the various fields already identified, a recommendation system indicator table is generated using numerical error and sequence similarity. This indicator table includes user profile keywords, included fields, and key metrics. User profile keywords include learning attitude and style, learning interests and hobbies, and learning level and ability. Included fields include completion-activity type, student information code, and grades. Key metrics include squared error (SE). Mode Cosine similarity (Cos) SC KL divergence KL SC Square error SE SC Absolute error AE SC and absolute error AE Score ; The three types of indicators are integrated into a comprehensive similarity index, as shown in the following formula: Similarity=Cos SC -KL SC -SE SC -AE SC -SE Mode -AE Score ; S3. Finally, the personalized recommendation module for similar learning friends analyzes these datasets using appropriate recommendation algorithms to mine similar learning friends on the online learning platform and recommends them to a student, providing personalized recommendation services, thereby laying a solid foundation for building an online social learning ecosystem.
2. The personalized recommendation system for online learning behavior based on cluster analysis according to claim 1, characterized in that: The online learning behavior data information management module includes a student information database and a data tag definition database. The student information database mainly stores learning behavior data from different dimensions collected from the online learning platform. The online learning platform records students' online learning behavior from the moment they register their platform accounts, generating a large amount of continuous learning data. Analyzing students' online learning data can be mainly divided into two dimensions: (1) Basic information data: refers to the user information manually filled in by students when registering on the platform, including but not limited to: name, gender, date of birth, political affiliation, education, major, college, and hobbies; (2) Online learning behavior data: This refers to the selection of nine indicators that record students' learning behavior on online learning platforms as measurement parameters. These indicators are classified according to the three dimensions of the online learning behavior model: The first dimension is the learning operation behavior dimension, which refers to the students' operation behavior on the online learning platform. The specific data indicators for measuring this dimension are the types of course resources accessed and their total number of times, the types of course resources learned and their total duration, the types of course resources learned and their completion rate, and the types of after-class document resources downloaded and their total number of times. Among them, the types of course resources accessed and their total number of accesses: The backend classifies various types of course resources on the platform, and the log data recorded by the backend of the online learning platform directly calculates the types of course resources accessed by different students on the platform and their total number of clicks. The types of learning course resources and their total duration: The total duration of different students learning course resources on the platform is calculated from the log data recorded by the backend of the online learning platform. The types of learning course resources and their completion rates: The completion rate of course resources on the online platform is divided into 0%-100% according to the progress bar. 0% means that the course has not been watched and completed, and 100% means that the course has been watched and completed. The completion rate of the same type of course resources is the average of the completion rates of the watched and learned course resources. The average completion rate of different students' learning courses is calculated according to the log data recorded in the background. The types of course materials and their total number of downloads: The completion rate of different students browsing course materials on the platform is directly calculated from the log data recorded in the backend of the online learning platform. The second dimension is the learning problem-solving behavior dimension, which refers to students' ability to solve learning problems on online learning platforms. The specific data indicators for measuring this dimension are: 1) Completion rate of online practice assignments: Most learning course resources are accompanied by corresponding after-class learning materials. After completing the video courses, students can further consolidate the knowledge points they have learned through after-class assignments. The background log statistics show the completion rate of students' submitted after-class assignments. 2) Online practice assignment scores: After a student submits their practice assignment, the backend log records and calculates the student's practice assignment score. The third dimension is the collaborative learning behavior dimension, which refers to the learning behavior ability of students to cooperate and communicate with other students on the online learning platform. The specific data indicators for measuring this dimension are the number of posts in the interaction area, the number of replies in the interaction area, and the number of study notes published. Among them, the number of posts in the interaction area: The backend log records the learning behavior data of students in the peer review area, and counts the total number of questions or comments posted by students. Number of replies in the interactive area: The total number of times students replied to comments in the peer review area is counted in the backend logs to further measure the activity level of students' collaborative learning behavior on the platform; Number of times study notes are published: The backend logs show the number of times students publish study notes online; The data tag definition library primarily stores the tag system for online learning student profiles. It constructs this tag system from different data dimensions based on learning behavior data stored in the student information database. Based on interviews with education industry experts and extensive literature review, the construction of online learning student profiles defines classification indicators from the following three dimensions: 1) Learning attitude style: From the dimension of learning attitude style, we can define: a. Active learner type: Students spend a lot of time learning online and study frequently. They submit many homework assignments and are highly active in the peer review area, posting comments, submitting questions, and answering questions. They also have a high degree of stickiness and dependence on the platform. b. Independent learning type: Students spend a lot of time learning online and study frequently, but they are not very active in the peer review area of the platform, do not make full use of the social functions of the platform, and have little interaction with teachers and other students. c. Lazy learner: Students generally register for online learning platforms but use them very infrequently, with short sessions of learning course resources and no loyalty to the platform. 2) Learning Interests and Hobbies: Based on the student's major and the type of learning resources accessed, the student's learning interest preferences are further determined. The course resource type data mainly refers to the classification of course resources on the online learning platform. Course resources are divided into 14 categories according to disciplines, namely philosophy, economics, law, education, literature, history, science, engineering, agriculture, medicine, military science, management, art, and physical education. 3) Learning level ability: The student's learning level ability can be measured and analyzed from two aspects: learning problem-solving behavior and collaborative learning behavior. Specifically, it includes the total number and frequency of homework submissions and homework score data. The learning level ability is divided into four categories: excellent, good, average, and poor.
3. The personalized recommendation system for online learning behavior based on cluster analysis according to claim 1, characterized in that: The data analysis and modeling module mainly constructs an online learning profile tagging system to classify and statistically display the data from the online learning behavior data management module according to demand tags. Specifically, it uses clustering analysis algorithms to model and analyze the data generated by students, constructing an online learning student profile for each student from three dimensions: learning attitude style, learning interest preference, and learning level ability. These student profile data of all students on the platform constitute the recommendation dataset for the personalized recommendation service for online learning. The steps also include the following: First, the acquired database data is cleaned and processed to determine the valid fields. The valid fields include serial numbers 1-16 and names. The names are arranged in order according to the serial numbers and include, but are not limited to, Good University ID, name, student ID, school, college, major, grade, first time using the platform, total duration of learning courses, course completion rate, courseware download rate, online exercise completion rate, number of times study notes are posted, number of interactive posts, number of interactive replies, and online exercise scores. The Good University ID serves as the primary key and, along with the student ID, represents the user's index and does not contribute to the calculation. After identifying valid fields and discarding invalid records, the processing also includes the following:
1. Referring to the Ministry of Education's subject / major classification standards, all majors were replaced with standard major names, and first-level, second-level, and third-level classification codes and names were added; 2. The school type code is classified and coded according to whether the school is a key university, i.e., a 985 / 211 "Double First-Class" university. If it is, it is coded as "1" and otherwise as "0".
3. One-hot encoding is used for the colleges. The original data contains 68 different colleges, which are encoded with a fixed length binary code, such as 0100101. There are a total of 68 different colleges, so the encoding length is 7.
4. The X-level subject codes shall be coded according to my country's subject classification standards; 5. Convert the first time of platform use to the number of days since the first time of platform use; 6. Convert the total course duration (hours, minutes, seconds) to the total course duration (seconds); 7. Convert percentage data such as the completion rate of learning courses, the completion rate of downloading and viewing course materials, and the completion rate of online exercises into values from 0 to 100; 8. Merge the school type code, college, and X-level subject code into a student information code; 9. The completion rate is calculated by averaging the following factors: the time elapsed since the first use of the platform (in days), the total duration of the learning courses (in seconds), the completion rate of the learning courses, the completion rate of online exercises and assignments, and the completion rate of downloaded courseware browsing.
4. The personalized recommendation system for online learning behavior based on cluster analysis according to claim 1, characterized in that: The personalized recommendation module for similar learning friends refers to using appropriate recommendation algorithms to analyze the recommendation dataset and find similar or slightly better learning friends on the online learning platform in terms of learning attitude, style, interests, and ability. At the same time, it recommends a list of courses that similar learning friends have taken or highly rated to the student, so that they can communicate and supervise each other's learning, laying the foundation for the future development of a social learning ecosystem.
5. The personalized recommendation system for online learning behavior based on cluster analysis according to claim 1, characterized in that: The clustering analysis algorithm mentioned refers to the K-Means model, where K-Means divides samples into K clusters, with high similarity within each cluster and low similarity between clusters. In machine learning, distance is usually used to measure the similarity between samples; the smaller the distance, the higher the similarity, and the larger the distance, the lower the similarity. This clustering algorithm is used to analyze online learning student data to mine the correlation pattern between online students' "activity level and completion rate" to determine the optimal value of the number of clusters K. Then, the Kalinsky-Harabas index, silhouette index, and within-group variance are used as three unsupervised learning indicators to evaluate the model. The calinski-harabasz score: Where S B S represents the variance between classes. W tr(·) represents the within-class variance, tr(·) represents the trace operation of the matrix, k represents the number of clusters, and N represents the number of sample points; Average profile coefficient si lhouette_score: Where a n b represents the average dissimilarity of point n to other points within the same cluster. n The minimum average dissimilarity of point n to other clusters; Total inertia (inclusive) Where x ij Let u represent the i-th point in the j-th cluster. j Let M represent the center point of the j-th cluster and M represent the number of clusters. Based on clustering algorithms, higher calinski_harabasz_score and silhouette_score are better, while lower total_inertia is better. Therefore, considering a comprehensive index, cmp_index is defined as follows: While a higher cmp_index value theoretically leads to better results, considering the problem in clustering algorithms where "excessive cluster size leads to meaningless data patterns," the optimal number of clusters is determined within acceptable complexity and a reasonable number of clusters. Furthermore, students' learning attitudes and styles are analyzed based on completion and activity levels. 1) Lazy learner type: This type of student almost never watches video courses, and their extremely low participation indicates a very low learning attitude and style; 2) Independent learning type: This type of student spends a lot of time studying, has low participation rates, and has strong independent learning ability; 3) Active learner type: This type of student spends a lot of time studying and is very active in posting and discussing, showing high enthusiasm.
Citation Information
Patent Citations
A personalized recommendation method and system for online medical education resources
CN109582875A
Method and apparatus for training online prediction model, device and storage medium
US20210248513A1