Online education informatization system of big data cloud platform
Through diversified data collection, data preprocessing and intelligent analysis, combined with K-means and decision tree algorithms, a personalized recommendation system is built, which solves the problems of incomplete data collection and inaccurate analysis in the existing technology, and realizes the accuracy and real-time optimization of personalized recommendations.
Patent Information
- Application Number
- CN202510392113.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Due to incomplete data collection and inaccurate data analysis, the online education information system of the existing big data cloud platform has affected the personalized recommendation function and cannot provide customized learning paths and content based on students' actual situation and needs.
The data collection module is used to collect data through crawlers, API interfaces, log files and student feedback methods, and combined with data preprocessing and cleaning modules to clean invalid redundant data. The K-means algorithm is used to analyze students' learning rules and interests and preferences, and a prediction model is built through the decision tree algorithm. The personalized recommendation engine module combines real-time student data for personalized recommendations, and optimizes the recommendation strategy through real-time feedback and adjustment modules.
It realizes the comprehensiveness and richness of data collection, improves the accuracy of personalized recommendations, ensures that the recommended content is closely in line with students' current learning status and needs, and enhances learning interest and enthusiasm.
Smart Images

Figure CN120298175A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of online education technology, and particularly to an online education information system for a big data cloud platform. Background Art
[0002] Cloud computing is a computing method based on the Internet. Through the Internet, a huge computing program is automatically split into countless smaller sub-programs, and then the processing results are sent back to the user after being searched, calculated and analyzed by a huge system composed of multiple servers. In an education information system, cloud computing technology provides powerful data storage and processing capabilities, enabling educational institutions to efficiently manage, access and share educational resources. In the field of education, the applications of cloud computing include various forms such as cloud computing-assisted teaching and cloud computing-assisted education. Big data technology refers to the technology of quickly acquiring, processing, and analyzing a large amount of diverse data sets to extract valuable information and knowledge to support decision-making and business process optimization.
[0003] After retrieval, the invention patent with the Chinese patent number CN115442385B discloses an online education information teaching system for a big data cloud platform, belonging to the field of online education. An online education information teaching system for a big data cloud platform marks the APP student end / APP teacher end downloaded with a certain learning material as the basic storage end by the server, and uses the widely existing APP student end / APP teacher end as the unit of distributed storage. Without directly downloading from the database, relevant data can be downloaded from the nearby or unobstructed basic storage end according to the actual situation, effectively sharing the pressure on the database, ensuring the performance of the efficient operation of the system to a certain extent. At the same time, by defining the number of basic storage ends of the same learning material, when the number of basic storage ends exceeds the defined value, the corresponding learning material in the database can be actively deleted, which is beneficial to increasing the number of stored learning materials in the system, improving the usage performance of the system, and having low cost.
[0004] However, in the actual use process of the above system, due to incomplete data collection and inaccurate data analysis, the personalized recommendation function is affected. The system may not be able to provide a customized learning path and content according to the actual situation and needs of students. For example, for a certain student, he may be more interested in learning some content related to real life, but the system recommends a large amount of theoretical knowledge to him. Such inaccurate recommendations not only cannot meet the learning needs of students, but may also reduce the learning interest and enthusiasm of students. Therefore, an online education information system for a big data cloud platform is proposed. Summary of the Invention
[0005] The object of the present invention is to solve the drawback that in the prior art, due to incomplete data collection and inaccurate data analysis, the personalized recommendation function is affected, and the system may not be able to provide a customized learning path and content according to the actual situation and needs of students, and a kind of online education informatization system of a big data cloud platform is proposed.
[0006] In order to achieve the above object, the present invention adopts the following technical scheme:
[0007] An online education informatization system of a big data cloud platform, comprising:
[0008] A data collection module: responsible for collecting students' learning behaviors (such as clicks, viewing duration, answering questions), grades, interests, and course information data from various data sources through web crawlers, API interfaces, log files, and student feedback means, and the data collection module transmits the collected data to the data preprocessing and cleaning module;
[0009] A data preprocessing and cleaning module: responsible for preprocessing the collected raw data, cleaning invalid, redundant, and incorrect data to ensure the accuracy of subsequent analysis, and the data preprocessing and cleaning module transmits the processed data to the data analysis and modeling module;
[0010] A data analysis and modeling module: responsible for extracting features from the processed data, such as students' learning duration, course completion, learning progress, interest points, etc., constructing a feature set, using the K-means algorithm to analyze students' learning rules and interest preferences, and constructing a prediction model through the decision tree algorithm. The data analysis and modeling module transmits the recommendation results output after training the prediction model to the personalized recommendation engine module;
[0011] A personalized recommendation engine module: responsible for receiving the recommendation results and making personalized recommendations in combination with real-time student data, and the personalized recommendation engine module transmits the recommended content to the real-time feedback and adjustment module;
[0012] A real-time feedback and adjustment module: responsible for adjusting the recommendation strategy in real time according to students' feedback on the recommended content, such as clicks, learning duration, completion, etc., to improve the accuracy of the recommendation. The real-time feedback and adjustment module transmits the adjusted recommendation strategy to the data analysis and modeling module for long-term optimization and improvement.
[0013] The above technical scheme further includes:
[0014] Preferably, the specific steps of collecting students' learning behaviors, grades, interests, and course information data from various data sources through web crawlers, API interfaces, log files, and student feedback means are as follows:
[0015] Data collection by crawlers: Set the target URLs of the crawlers, determine the data fields to be scraped, such as student IDs, study durations, course information, student grades, etc., write crawler scripts for scraping web page content, parse the scraped web page content, extract structured data and store it in a database;
[0016] Data collection through API interfaces: Call the API interfaces provided by the education platform, pass authentication information such as API keys, tokens, etc., send requests to the API, obtain data on students' learning behaviors (such as course viewing, assignment grades), interest preferences, grade records, and course information, parse the data returned by the API, and store the data in the database;
[0017] Data collection from log files: Collect the log files of the platform, parse the logs, extract learning behavior data such as login time, study duration, learning content, grade submission status, etc., store the processed log data in the database for subsequent analysis and processing;
[0018] Data collection from student feedback: Create feedback channels where students can provide information through questionnaires, evaluation systems, or other feedback forms, collect feedback data, including students' evaluations of courses, learning experiences, interest preferences, etc., organize, classify, and label the feedback data to ensure its structuring and store it in the database.
[0019] Preferably, the data preprocessing and cleaning module includes a data formatting unit, a missing value processing unit, an outlier detection and processing unit, a data deduplication unit, a data normalization and standardization unit, and a data integration unit. The data formatting unit is responsible for converting data from different sources into a unified format. The missing value processing unit is responsible for identifying and processing missing values in the data. The outlier detection and processing unit is responsible for identifying and processing outliers or noisy data in the data. The data deduplication unit is responsible for removing duplicate data records. The data normalization and standardization unit is responsible for normalizing or standardizing numerical data. The data integration unit is responsible for integrating data from multiple data sources.
[0020] Preferably, the specific steps for extracting features from the processed data and constructing a feature set are as follows:
[0021] Feature selection: Calculate the correlation coefficient r between the feature and the target variable (such as grades), and select highly correlated features. The calculation formula for the correlation coefficient r is: where x i and y i are the observed values of the feature and the target variable respectively, and μ x and μ y are the corresponding means;
[0022] Feature construction: Extract time features from learning behavior data, such as the total learning time, average learning duration, etc., calculate learning frequencies, such as the number of learning times per week, the number of completed courses per month, etc., and combine multiple basic features to form new features, such as the combined feature of course completion rate and grades;
[0023] Feature encoding and transformation: Encode categorical features, such as subject categories (mathematics, physics, chemistry) into [1,0,0], [0,1,0], [0,0,1], and encode ordinal categorical features, such as grade levels (A, B, C, D) into [1,2,3,4];
[0024] Feature dimensionality reduction: Project high-dimensional data into a low-dimensional space through linear transformation, retaining the main information, and use the recursive feature elimination method to select features based on model performance (such as accuracy, AUC, etc.);
[0025] Feature set construction: Integrate the features processed through the above steps into a feature set, ensuring that each feature is valid and contributes to the model performance, and verify the constructed feature set.
[0026] Preferably, the specific steps for using the K-means algorithm to analyze students' learning patterns and interest preferences are as follows:
[0027] Initialization: The objective function of the K-means algorithm is to minimize the squared error of the student features within the cluster: where J is the objective function representing the sum of the squared distances from all students within the cluster to the cluster center, C k is the k-th cluster, x i is the feature vector of the i-th student, μ k is the center of the k-th cluster, K is the total number of clusters for student partitioning, and the centers μ of the k clusters k are randomly selected from K data points as the initial centers for initialization;
[0028] Clustering execution and iteration: Select the cluster center according to the Euclidean distance, and assign each student to the cluster center closest to it. The Euclidean distance calculation formula is: where d(x i ,μ k ) is the Euclidean distance from student i to the cluster center k, x i =(x i1 ,x i2 ,...,x im ) represents the feature vector of the i-th student, and μ k =(μ k1 ,μ k2 ,...,μ km) is the cluster center eigenvector, representing the mean of all students' features in the cluster. The cluster center is a virtual point representing the average of all students' features in the cluster. Calculate the cluster center μ k , making it the mean of all students' features within the current cluster: where |C k | is the number of students in the k-th cluster. Through repeated iteration until convergence, the cluster center no longer changes significantly, i.e., it converges;
[0029] Analysis of clustering results and extraction of interest preferences: Finally, obtain the cluster labels for each student. The cluster labels include Cluster 1 representing students with high learning enthusiasm (e.g., high video viewing duration, high participation, relatively high grades), Cluster 2 representing students with relatively low learning enthusiasm (e.g., low video viewing duration, low participation, relatively low grades), and Cluster 3 representing students with poor grades but interested in certain courses (e.g., relatively high discussion frequency and assignment submission frequency, but low grades).
[0030] Preferably, the specific steps for constructing the prediction model through the decision tree algorithm are as follows:
[0031] Let the selected features be X = {X1, X2,..., X n}, where X i represents the i-th feature;
[0032] The information gain of feature X i is: IG(D, X i ) = H(D) - H(D|X i ), where H(D) is the entropy of the dataset D, and H(D|X i ) is the conditional entropy of the dataset D given the feature X i . The calculation formula for H(D) is: where p k is the proportion of samples belonging to class s in the dataset, m is the total number of classes, and the conditional entropy H(D|X i of the dataset D given the feature X i is calculated as: where V(X i ) is all possible values of the feature X i , p(v) is the probability that the feature X i takes the value v, and H(D v ) is the entropy of the subset D i of the dataset D when the feature X v takes the value v. At each node, select the feature X i with the maximum information gain as the splitting feature for the current node, and repeat this process until the stopping condition is met, such as reaching the maximum depth or the information gain is less than a certain threshold;
[0033] Starting from the root node, select the splitting feature according to the data characteristics, divide the data set into multiple subsets, each subset corresponding to a child node, and recursively construct the tree until the purity of each leaf node meets the requirements, such as all samples belonging to the same category or reaching the preset tree depth. Input the training data (i.e., historical student learning behavior and performance data) into the decision tree model. Train the decision tree through the above steps, use cross-validation (such as K-fold cross-validation) to evaluate the performance of the model, and measure the prediction effect through indicators such as accuracy, recall rate, and F1 score.
[0034] Preferably, the specific steps for receiving the recommendation result and making personalized recommendations in combination with real-time student data are as follows:
[0035] First, receive the recommendation result output after the prediction model training from the data analysis and modeling module, and at the same time obtain real-time student data;
[0036] Based on the received recommendation result and real-time student data, the personalized recommendation engine module constructs or updates the user profile of the student;
[0037] The personalized recommendation engine module matches the recommendation result with the real-time student data, generates a personalized recommendation list, and sorts and displays the generated recommendation result.
[0038] Preferably, the specific process for adjusting the recommendation strategy in real time according to the feedback of the student on the recommended content is as follows:
[0039] Collect the feedback data of the student on the recommended content, analyze the collected feedback data. Based on the analysis result of the feedback data, the real-time feedback and adjustment module evaluates the current recommendation strategy. According to the evaluation result, the real-time feedback and adjustment module adjusts the recommendation strategy. After adjusting the recommendation strategy, verify the adjustment effect.
[0040] The present invention has the following beneficial effects:
[0041] 1. In the present invention, through the data collection module, using various means such as web crawler technology, API interfaces, log files, and student feedback, data such as students' learning behaviors, grades, interests, and course information are comprehensively collected from multiple data sources. This diversified data collection method ensures the comprehensiveness and richness of the data and reduces the problem of incomplete data collection.
[0042] 2. In the present invention, the data analysis and modeling module extracts features from the processed data, constructs a feature set, and uses the K-means algorithm to analyze the learning patterns and interest preferences of students. It can effectively group student data with similar features, thereby more accurately revealing the learning patterns and interest distributions of students. The personalized recommendation engine module receives the recommendation results from the data analysis and modeling module and combines real-time student data for personalized recommendations. This real-time updated recommendation mechanism ensures that the recommended content can closely match the current learning status and needs of students. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 FIG. is a system architecture diagram of an online education informatization system of a big data cloud platform proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] As Figure 1 shown, an online education informatization system of a big data cloud platform includes:
[0046] Data collection module: responsible for collecting students' learning behaviors, grades, interests, and course information data from various data sources through crawlers, API interfaces, log files, and student feedback means. The data collection module transmits the collected data to the data preprocessing and cleaning module;
[0047] Data preprocessing and cleaning module: responsible for preprocessing the collected raw data and cleaning invalid, redundant, and incorrect data. The data preprocessing and cleaning module transmits the processed data to the data analysis and modeling module;
[0048] Data analysis and modeling module: responsible for extracting features from the processed data, constructing a feature set, using the K-means algorithm to analyze students' learning patterns and interest preferences, and constructing a prediction model through the decision tree algorithm. The data analysis and modeling module transmits the recommendation results output after training the prediction model to the personalized recommendation engine module;
[0049] Personalized recommendation engine module: responsible for receiving the recommendation results and combining real-time student data for personalized recommendations. The personalized recommendation engine module transmits the recommended content to the real-time feedback and adjustment module;
[0050] Real-time Feedback and Adjustment Module: Responsible for adjusting the recommendation strategy in real time based on the students' feedback on the recommended content. The Real-time Feedback and Adjustment Module will transmit the adjusted recommendation strategy to the Data Analysis and Modeling Module.
[0051] In the embodiments of the present invention, the system, through the data collection module, uses various means such as web crawler technology, API interfaces, log files, and student feedback to widely collect data on students' learning behaviors, grades, interests, course information, etc. from various data sources. It also incorporates advanced artificial intelligence (AI) technologies such as natural language processing (NLP) and deep learning algorithms to more intelligently and efficiently mine and utilize student data. The NLP technology enables the system to understand and analyze students' unstructured data, such as study notes, forum discussions, etc., so as to more comprehensively understand students' learning behaviors and interest preferences. The deep learning algorithms further improve the accuracy and efficiency of data collection, enabling the system to automatically identify and extract key information, reducing manual intervention. This diversified data collection method ensures the comprehensiveness and richness of data, reducing the possibility of incomplete data collection. The collected raw data may contain invalid, redundant, or incorrect information, which, if directly used for analysis, will seriously affect the accuracy of the results. Therefore, the system is provided with a data preprocessing and cleaning module, which is responsible for cleaning and preprocessing this data to ensure that the data passed to subsequent modules is accurate and effective. The data analysis and modeling module uses the K-means algorithm to analyze students' learning patterns and interest preferences. This is a commonly used clustering algorithm that can accurately group students according to their learning patterns and interests. At the same time, by constructing a prediction model through the decision tree algorithm, the system can predict students' future learning needs and interest changes based on their learning history and behavior patterns. This data-based prediction ability provides strong support for personalized recommendations. The personalized recommendation engine module receives the recommendation results from the data analysis and modeling module and utilizes AI technology. Through deep learning algorithms, this module can analyze students' current learning status and needs in real time and, combined with historical recommendation data, continuously optimize the recommendation strategy to ensure that the recommended content closely fits the actual situation and needs of students. This AI-based personalized recommendation mechanism not only improves the accuracy of recommendations but also enhances students' learning experience and satisfaction. This real-time updated recommendation mechanism ensures that the recommended content closely fits students' current learning status and needs. The real-time feedback and adjustment module adjusts the recommendation strategy in real time according to students' feedback on the recommended content. This dynamic adjustment ability enables the system to continuously learn and optimize, ensuring that the recommended content always meets the actual situation and needs of students. At the same time, the system also has functions such as intelligent tutoring and answering questions, automated exams and evaluations, optimized allocation of educational resources, virtual learning assistants, and cross-platform learning experiences. Intelligent tutoring and answering questions means introducing an AI chatbot to provide students with instant learning tutoring and answering services. Using natural language processing technology, the robot can understand students' questions and provide accurate answers. Automated exams and evaluations means using AI technology to achieve automatic question generation, automatic grading, and performance analysis. Through big data analysis, it can discover students' learning weaknesses and improvement spaces and provide targeted feedback for teaching.Optimal allocation of educational resources means using AI algorithms to analyze the use of educational resources, optimize the allocation and distribution of educational resources, predict future demand trends for educational resources, and provide data support for educational planning. Virtual learning assistants mean introducing AI virtual assistants to provide students with round-the-clock learning support and services, and enhance the interactivity and fun of the learning environment through voice interaction and visual recognition technology. Cross-platform learning experience means using AI technology to synchronize and share learning data across platforms, and providing support for a variety of learning terminals, such as mobile phones, tablets, computers, etc., to meet students' diverse learning needs.
[0052] In one embodiment, the specific steps of collecting students' learning behavior, grades, interests, and course information data from various data sources through crawlers, API interfaces, log files, and student feedback are as follows:
[0053] Crawler collects data: set the crawler's target URL, determine the data fields that need to be crawled, write a crawler script to crawl web page content, parse the crawled web page content, extract structured data and store it in the database;
[0054] API interface to collect data: call the API interface provided by the education platform, pass authentication information, send requests to the API, obtain data on students' learning behavior, interest preferences, grade records, and course information, parse the data returned by the API, and store the data in the database;
[0055] Log file data collection: Collect the platform's log files, parse the logs, extract the learning behavior data, and store the processed log data in the database;
[0056] Student feedback data collection: Create feedback channels, collect feedback data, and organize, categorize and label the feedback data.
[0057] In the embodiment of the present invention, when using a crawler to collect data, it is necessary to follow the target website protocol to ensure that the data is captured legally and avoid excessive load on the website. When using an API interface to collect data, it is necessary to ensure that the API call frequency complies with platform restrictions to avoid exceeding the number of requests, and to reasonably handle API error returns (such as timeouts, no data, etc.). When collecting data through student feedback, it is necessary to note that student feedback data may contain subjective opinions, so sentiment analysis and other processing are required to extract valuable information.
[0058] In one embodiment, the data preprocessing and cleaning module includes a data formatting unit, a missing value processing unit, an outlier detection and processing unit, a data deduplication unit, a data normalization and standardization unit, and a data integration unit. The data formatting unit is responsible for converting data from different sources into a unified format. The missing value processing unit is responsible for identifying and processing missing values in the data. The outlier detection and processing unit is responsible for identifying and processing outliers or noisy data in the data. The data deduplication unit is responsible for removing duplicate data records. The data normalization and standardization unit is responsible for normalizing or standardizing numerical data. The data integration unit is responsible for integrating data from multiple data sources.
[0059] In the embodiments of the present invention, common methods for processing missing values include the deletion method, the filling method, and the interpolation method. In the deletion method, if a row has missing values, the entire row is deleted. If a column has more missing values than a set threshold, the entire column is deleted. In the filling method, for median filling, first find the median in the current feature column and use it to replace the missing values. For mode filling, use the most frequently occurring value in the column to replace the missing values. The outlier detection and processing unit identifies outliers by calculating the Z-Score of the data: where X is the data point, μ is the mean of the feature, and σ is the standard deviation of the feature. When the Z-Score exceeds a certain threshold, the data point can be considered an outlier.
[0060] In one embodiment, the specific steps for extracting features from the processed data and constructing a feature set are as follows:
[0061] Feature selection: Calculate the correlation coefficient r between the feature and the target variable, and select highly correlated features. The calculation formula for the correlation coefficient r is: where x i and y i are the observed values of the feature and the target variable respectively, and μ x and μ y are the corresponding means;
[0062] Feature construction: Extract time features from the learning behavior data, calculate the learning frequency, and combine multiple basic features to form new features;
[0063] Feature encoding and conversion: Encode categorical features and encode ordinal categorical features;
[0064] Feature dimensionality reduction: Project high-dimensional data into a low-dimensional space through linear transformation, retain the main information, and use the recursive feature elimination method to select features based on model performance;
[0065] Feature set construction: Integrate the features processed through the above steps into a feature set, and verify the constructed feature set.
[0066] In one embodiment, the specific steps for analyzing students' learning patterns and interest preferences using the K-means algorithm are as follows:
[0067] Initialization: The objective function of the K-means algorithm is to minimize the squared error of the student features within the clusters: where J is the objective function representing the sum of the squared distances from all students within the cluster to the cluster center, C k is the k-th cluster, x i is the feature vector of the i-th student, μ k is the center of the k-th cluster, K is the total number of clusters for student partitioning, and the centers μ of the k clusters k are randomly selected from K data points as the initial centers for initialization;
[0068] Clustering execution and iteration: Select the cluster center according to the Euclidean distance, and assign each student to the cluster center closest to it. The Euclidean distance calculation formula is: where d(x i , μ k ) is the Euclidean distance from student i to cluster center k, x i =(x i1 , x i2 ,..., x im ) represents the feature vector of the i-th student, and μ k =(μ k1 , μ k2 ,..., μ km ) is the feature vector of the cluster center. Calculate the center μ of each cluster k to make it the mean of all student features within the current cluster: where |C k | is the number of students in the k-th cluster, and iterate repeatedly until convergence;
[0069] Analysis of clustering results and extraction of interest preferences: Finally, obtain the cluster label for each student. The cluster labels include cluster 1 representing students with high learning enthusiasm, cluster 2 representing students with relatively low learning enthusiasm, and cluster 3 representing students with poor grades but interested in certain courses.
[0070] In the embodiment of the present invention, in K-means clustering, the goal is to partition the student data into K clusters such that the similarity of student features within each cluster is maximized and the similarity between clusters is minimized.
[0071] In one embodiment, the specific steps for constructing a prediction model using the decision tree algorithm are as follows:
[0072] Let the selected features be X = {X1, X2,..., X n}, where Xi represents the i-th feature;
[0073] Feature X i The information gain of is: IG(D, X i ) = H(D) - H(D|X i ), where H(D) is the entropy of the dataset D, and H(D|X i ) is the conditional entropy of the dataset D given the feature X i . The calculation formula of H(D) is: where p k is the proportion of samples belonging to class s in the dataset, m is the total number of classes, and the conditional entropy H(D|X i ) of the dataset D given the feature X i is calculated as: where V(X i ) are all possible values of the feature X i , p(v) is the probability that the feature X i takes the value v, and H(D v ) is the entropy of the subset D i when the feature X v takes the value v. At each node, the feature X with the maximum information gain is selected i as the splitting feature of the current node, and this process is repeated until the stopping condition is met;
[0074] Starting from the root node, the splitting feature is selected according to the data features, and the dataset is divided into multiple subsets, each subset corresponding to a child node. The tree is recursively constructed until the purity of each leaf node meets the requirements. The training data is input into the decision tree model, and the decision tree is trained through the above steps.
[0075] In the embodiments of the present invention, to avoid overfitting during the construction of the decision tree, pruning can be performed during the construction process of the decision tree to cut off unnecessary branches. The common methods are pre-pruning (stopping in advance during the construction process) and post-pruning (after the construction is completed, some unimportant nodes are deleted by evaluating the performance of the model). After the model training is completed, cross-validation (such as K-fold cross-validation) needs to be used to evaluate the performance of the model, and the prediction effect is measured by indicators such as accuracy, recall rate, and F1 score. The accuracy Accuracy represents the proportion of correctly classified samples in the total samples, where TP represents true positives, TN represents true negatives, FP represents false positives, and FN represents false negatives. The recall rate Recall represents the proportion of all positive class samples that are correctly classified as positive classes, The accuracy Precision represents the proportion of all samples predicted to be positive classes that are actually positive classes,
[0076] In one embodiment, the specific steps of receiving the recommendation results and making personalized recommendations by combining real-time student data are as follows:
[0077] First, receive the recommendation results output after the prediction model training from the data analysis and modeling module, and at the same time obtain the real-time student data;
[0078] Based on the received recommendation results and real-time student data, the personalized recommendation engine module constructs or updates the user profile of the student;
[0079] The personalized recommendation engine module matches the recommendation results with the real-time student data, generates a personalized recommendation list, and sorts and displays the generated recommendation results.
[0080] In the embodiment of the present invention, the recommendation results include a series of learning resources, courses, or learning paths, etc., which are obtained by analyzing data such as students' learning behaviors, grades, interests, and course information through clustering algorithms and decision tree algorithms. Student data includes students' current learning status, learning progress, online duration, interaction situation, etc. These data are collected in real time by the data collection module and transmitted to the personalized recommendation engine module for real-time personalized recommendation. The user profile is an embodiment of the digital image of the student, including multiple dimensions such as the student's basic information, learning style, interest preferences, and learning level. By constructing the user profile, the system can more deeply understand the needs and behavior patterns of the students. The sorting is based on multiple factors such as the popularity, novelty, and relevance of the recommended content. The display methods may include various forms such as lists, charts, and cards, so that students can intuitively see the recommendation results and make quick choices.
[0081] In one embodiment, the specific process of real-time adjustment of the recommendation strategy according to the feedback of students on the recommended content is as follows:
[0082] Collect the feedback data of students on the recommended content, analyze the collected feedback data, based on the analysis results of the feedback data, the real-time feedback and adjustment module evaluates the current recommendation strategy, and according to the evaluation results, the real-time feedback and adjustment module adjusts the recommendation strategy. After adjusting the recommendation strategy, verify the adjustment effect.
[0083] In the embodiment of the present invention, the feedback data includes the click-through rate, rating, and dwell time information of the students. The click-through rate CTR is used to measure the attractiveness of the recommended content, and its calculation formula is: Among them, the number of clicks refers to the number of times students click on the recommended content, and the number of displays refers to the number of times the recommended content is displayed. The average value μ of the user ratings score is used to measure the quality of the recommended content, and its calculation formula is: Among them, score iThe score given by each student, and n is the total number of score data. The learning average stay time Average Time is used to evaluate the students' interest in the recommended content, and its calculation formula is: where time i is the stay time of the i-th student on the recommended content, and n is the total number of students.
[0084] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An online education informatization system for a big data cloud platform, characterized in that, Including: Data collection module: Responsible for collecting students' learning behavior, grades, interests, and course information data from various data sources through crawlers, API interfaces, log files, and student feedback means. The data collection module transmits the collected data to the data preprocessing and cleaning module; Data preprocessing and cleaning module: Responsible for preprocessing the collected raw data and cleaning invalid, redundant, and incorrect data. The data preprocessing and cleaning module transmits the processed data to the data analysis and modeling module; Data analysis and modeling module: Responsible for extracting features from the processed data, constructing a feature set, using the K-means algorithm to analyze students' learning patterns and interest preferences, and constructing a prediction model through the decision tree algorithm. The data analysis and modeling module transmits the recommended results output after training the prediction model to the personalized recommendation engine module; Personalized recommendation engine module: Responsible for receiving the recommended results and making personalized recommendations in combination with real-time student data. The personalized recommendation engine module transmits the recommended content to the real-time feedback and adjustment module; Real-time feedback and adjustment module: Responsible for adjusting the recommendation strategy in real time according to students' feedback on the recommended content. The real-time feedback and adjustment module transmits the adjusted recommendation strategy to the data analysis and modeling module.
2. The online education informatization system of a big data cloud platform according to claim 1, characterized in that, The specific steps for collecting students' learning behavior, grades, interests, and course information data from various data sources through crawlers, API interfaces, log files, and student feedback means are as follows: Data collection by crawler: Set the target website of the crawler, determine the data fields to be crawled, write a crawler script for crawling web page content, parse the crawled web page content, extract structured data, and store it in the database; Data collection by API interface: Call the API interface provided by the education platform, transmit authentication information, send a request to the API, obtain data on students' learning behavior, interest preferences, grade records, and course information, parse the data returned by the API, and store the data in the database; Data collection from log files: Collect the log files of the platform, parse the logs, extract the learning behavior data, and store the processed log data in the database; Data collection from student feedback: Create a feedback channel, collect feedback data, and organize, classify, and label the feedback data.
3. An online education informatization system for a big data cloud platform according to claim 1, characterized in that, The data preprocessing and cleaning module includes a data formatting unit, a missing value processing unit, an outlier detection and processing unit, a data deduplication unit, a data normalization and standardization unit, and a data integration unit. The data formatting unit is responsible for converting data from different sources into a unified format. The missing value processing unit is responsible for identifying and processing missing values in the data. The outlier detection and processing unit is responsible for identifying and processing outliers or noisy data in the data. The data deduplication unit is responsible for removing duplicate data records. The data normalization and standardization unit is responsible for normalizing or standardizing numerical data. The data integration unit is responsible for integrating data from multiple data sources.
4. An online education informatization system for a big data cloud platform according to claim 1, characterized in that, The specific steps for extracting features from the processed data and constructing a feature set are as follows: Feature Selection: Calculate the correlation coefficient r between the feature and the target variable, and select highly correlated features. The calculation formula for the correlation coefficient r is: where x i and y i are the observed values of the feature and the target variable respectively, and μ x and μ y are the corresponding means; Feature construction: Extract time features from learning behavior data, calculate learning frequencies, and combine multiple basic features to form new features; Feature encoding and transformation: Encode categorical features and encode ordinal categorical features; Feature dimensionality reduction: Project high-dimensional data into a low-dimensional space through linear transformation, retain the main information, and use the recursive feature elimination method to select features based on model performance; Feature set construction: Integrate the features processed through the above steps into a feature set and verify the constructed feature set.
5. An online education informatization system for a big data cloud platform according to claim 1, characterized in that, The specific steps for analyzing students' learning patterns and interest preferences using the K-means algorithm are as follows: Initialization: The objective function of the K-means algorithm is to minimize the squared error of the student features within the clusters: where J is the objective function representing the sum of the squared distances from all students within the cluster to the cluster center, C k is the k-th cluster, x i is the feature vector of the i-th student, μ k is the center of the k-th cluster, K is the total number of clusters for student partitioning, and the centers μ of the k clusters k are randomly selected from K data points as the initial centers for initialization; Clustering execution and iteration: Select the cluster centers according to the Euclidean distance, and assign each student to the nearest cluster center. The Euclidean distance calculation formula is as follows: where d(x i , μ k ) is the Euclidean distance from student i to cluster center k, x i = (x i1 , x i2 ,..., x im ) represents the feature vector of the i-th student, μ k = (μ k1 , μ k2 ,..., μ km ) is the feature vector of the cluster center. Calculate each cluster center μ k to make it the mean of all students' features within the current cluster: where |C k | is the number of students in the k-th cluster, and iterate repeatedly until convergence; Cluster result analysis and interest preference extraction: The cluster labels for each student are finally obtained. The cluster labels include Cluster 1 representing students with high learning enthusiasm, Cluster 2 representing students with relatively low learning enthusiasm, and Cluster 3 representing students with poor academic performance but interested in certain courses.
6. The online education informatization system of a big data cloud platform according to claim 1, characterized in that, The specific steps for constructing a prediction model through the decision tree algorithm are as follows: Let the selected feature X = {X1, X2,..., X n}, where X i represents the i-th feature; Feature X i The information gain of i ) is: IG(D, X i ) = H(D) - H(D|X i ), where H(D) is the entropy of the dataset D, and H(D|X i ) is the conditional entropy of the dataset D given the feature X where p k is the proportion of samples belonging to class s in the dataset, m is the total number of classes, and the conditional entropy H(D|X i ) of the dataset D given the feature X i ) is calculated as: where V(X i ) is all possible values of the feature X i , p(v) is the probability that the feature X i takes the value v, and H(D v ) is the entropy of the subset D i when the feature X v takes the value v. At each node, select the feature X i with the maximum information gain as the splitting feature of the current node, and repeat this process until the stopping condition is met; Starting from the root node, select splitting features according to data features, divide the data set into multiple subsets, with each subset corresponding to a child node, recursively construct the tree until the purity of each leaf node meets the requirements, and input the training data into the decision tree model. Train the decision tree through the above steps.
7. An online education informatization system for a big data cloud platform according to claim 1, characterized in that, The specific steps for receiving the recommendation results and making personalized recommendations in combination with real-time student data are as follows: First, receive the recommendation results output after the prediction model training from the data analysis and modeling module, and at the same time obtain real-time student data; Based on the received recommendation results and real-time student data, the personalized recommendation engine module constructs or updates the student's user profile; The personalized recommendation engine module matches the recommendation results with real-time student data, generates a personalized recommendation list, and sorts and displays the generated recommendation results.
8. An online education informatization system for a big data cloud platform according to claim 1, characterized in that, The specific process for adjusting the recommendation strategy in real time according to the students' feedback on the recommended content is as follows: Collect the feedback data of students on the recommended content, analyze the collected feedback data, evaluate the current recommendation strategy based on the analysis results of the feedback data, adjust the recommendation strategy according to the evaluation results, and verify the adjustment effect after adjusting the recommendation strategy.
Citation Information
Patent Citations
An online education information teaching system based on big data cloud platform
CN115442385B