User data analysis method based on big data
By collecting and analyzing students' learning behavior, grades and online interactive data, a personalized learning model is constructed, and personalized teaching suggestions are provided to teachers, which solves the problem of inaccurately capturing students' differences in the existing technology, and improves teaching effectiveness and learning quality.
Patent Information
- Application Number
- CN202510450227.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, there are great differences in learning methods and habits between students, and algorithmic models cannot accurately capture and explain these differences, resulting in the inability to provide students with effective learning support and suggestions.
By collecting students' online learning platform data, grade management system data and online interactive data, perform data preprocessing and multi-dimensional analysis, build a personalized learning model, and provide personalized teaching suggestions.
It can more accurately identify students' individual differences and learning needs, provide targeted and personalized teaching suggestions, and improve teaching effectiveness and learning quality.
Smart Images

Figure CN120387908A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and in particular, to a method for analyzing user data based on big data. Background Art
[0002] Educational institutions can monitor teaching quality and evaluate teachers' teaching effects through big data analysis, so as to timely discover teaching problems and make improvements. By analyzing students' learning data, it is possible to predict students' learning outcomes and provide teaching suggestions for teachers. The application of big data can promote innovation in the education field, facilitate the transformation of education models and the improvement of teaching methods. Through big data analysis, it is possible to more accurately understand students' learning situations and provide more precise teaching support, thereby improving the quality of education.
[0003] In the prior art, there are significant differences in learning methods and habits among students. These differences may stem from various factors such as students' personal characteristics, family backgrounds, and learning environments. An educational institution may collect data such as students' historical academic achievements, learning durations, and homework completion situations, and attempt to use this data to predict students' future academic achievements. However, due to the large differences in learning methods and habits among students, the algorithm model may not be able to accurately capture and interpret the impact of these differences on academic achievements, resulting in the inability to provide effective learning support and suggestions for students. Therefore, a method for analyzing user data based on big data is proposed. Summary of the Invention
[0004] The purpose of the present invention is to solve the drawbacks in the prior art that there are significant differences in learning methods and habits among students, and the algorithm model may not be able to accurately capture and interpret the impact of these differences on academic achievements, resulting in the inability to provide effective learning support and suggestions for students, and to propose a method for analyzing user data based on big data.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] A method for analyzing user data based on big data includes the following steps:
[0007] S1: Data collection: Collect student data, where the student data includes online learning platform data, performance management system data, and online interaction data;
[0008] S2: Data preprocessing: Perform preprocessing operations on the data. The preprocessing operations include data cleaning, data standardization, and data integration. The data standardization uniformly processes data from different sources and in different formats, and the data integration integrates student data from multiple data sources to form a complete set of student learning data;
[0009] S3: Multi-dimensional data analysis: Use students' learning data for learning behavior analysis, performance analysis, and online interaction analysis. The learning behavior analysis includes evaluating students' learning engagement and learning habits, identifying students' learning peak and trough periods, and then determining whether students are more inclined to independent learning and identifying students' learning patterns. The performance analysis includes evaluating students' learning outcomes and identifying students' learning difficulties and weak links. The online interaction analysis includes evaluating students' social interaction skills and collaborative learning abilities and identifying students' learning communities and interaction patterns;
[0010] S4: Construction of personalized learning model: Construct a personalized learning model based on students' learning patterns, learning difficulties, and weak links;
[0011] S5: Intelligent teaching suggestions: Timely feedback students' learning data and analysis results to teachers. The personalized learning model provides personalized teaching suggestions for teachers according to students' learning behaviors, performance, and online interaction situations. The teaching suggestions include learning plans, teaching methods, and tutoring strategies for specific students. Evaluate the effectiveness of teaching suggestions by regularly comparing students' learning data and the implementation of teaching suggestions.
[0012] The above further includes:
[0013] Further, in S1, the online learning platform data includes learning time (cumulative / daily average), learning frequency (number of logins / week), learning path (course / chapter access order), online quiz scores (average score / highest score / lowest score), and homework completion status (completion rate / submission time). The performance management system data includes examination subjects, examination time, examination scores (raw score / percentile / grading system), and historical score records (progress / regression trend). The online interaction data includes the number of questions, answer quality (number of likes / comments), and discussion depth (reply level / keyword analysis).
[0014] Further, in S2, the data cleaning includes removing duplicate data, handling missing values, and detecting and handling outliers. The removal of duplicate data uses a unique identifier (such as student ID) to identify and delete duplicate records. For missing data in the handling of missing values, fill in (such as using mean, median, mode, etc.) or delete according to the specific situation. Outliers are detected through box plots and corrected or deleted according to business logic.
[0015] Further, in S3, the specific steps of the learning behavior analysis:
[0016] Learning time analysis:
[0017] Cumulative study time: Calculate the total study time of students within a certain period of time (such as one day, one week, or one month). The calculation formula is: Cumulative study time = Σ(daily study time);
[0018] Average daily study time: reflects the average study intensity of students. The calculation formula is: Average daily study time = cumulative study time / total number of days;
[0019] Learning frequency: records the number of times students log into the learning platform and the number of active days. The calculation formula is: learning frequency = number of logins / total days (or number of active days / total days);
[0020] Identify learning peaks and troughs:
[0021] Use exponential smoothing to identify the changing trends of students’ study time and determine their peak and trough periods based on the trend chart;
[0022] Judgment of autonomous learning tendency:
[0023] Analyze students' learning autonomy based on their study time, study frequency, and study path. If students show high learning engagement, stable study frequency, and autonomous study path, they may be more inclined to autonomous learning.
[0024] Learning pattern recognition:
[0025] Using random forest to analyze students' learning data, we can find students' learning patterns, including autonomous learning, dependent learning, and mixed learning.
[0026] By analyzing the split paths of the decision tree, we can explain how the model makes predictions based on the students’ learning data. For example, if a decision tree splits on the feature “study time”, and self-directed learners typically study longer, we can conclude that this decision tree plays an important role in identifying self-directed learners.
[0027] Furthermore, in S3, the specific steps of the performance analysis are:
[0028] Test score analysis: Calculate students' average score, highest score, lowest score, normative score, and contribution to the standard, and use trend charts to show changes in students' scores.
[0029] Identification of learning difficulties and weak links:
[0030] Through cluster analysis, we can identify which knowledge points or question types students perform poorly on, and further analyze students' learning difficulties based on their homework completion status.
[0031] Furthermore, in S3, the specific steps of the online interaction analysis are:
[0032] Social interaction ability assessment:
[0033] Using text mining and natural language processing techniques, analyze the interaction content of students on various platforms such as online classrooms and discussion forums, and calculate interaction metrics such as the number of questions asked, answers given, and likes of students;
[0034] Collaborative learning ability assessment:
[0035] Analyze the performance of students in group discussions and collaborative tasks, and calculate the collaborative contribution and efficiency of students;
[0036] Learning community and interaction pattern recognition:
[0037] Using network analysis techniques, identify the learning communities and interaction patterns of students, and analyze the association relationships between students, such as friendship and cooperation relationships.
[0038] Furthermore, in S4, the specific steps for constructing the personalized learning model are as follows: construct a personalized learning model based on the learning patterns, learning difficulties, and weak links of students
[0039] Data preparation:
[0040] Collect students' learning data, including learning behavior records (such as learning time, learning progress, number of clicks on learning resources, etc.), learning outcomes (such as homework grades, test scores, etc.), and students' basic information (such as age, gender, learning style, etc.), and divide the data into training set, validation set, and test set for model training, validation, and testing;
[0041] Feature extraction and selection:
[0042] Extract features related to learning patterns, learning difficulties, and weak links from the original data. For example, features such as the learning time distribution of students, preferences for learning resources, and types of wrong questions can be extracted.
[0043] Use chi-square test to screen the features that are most helpful for model construction, ensuring that the selected features can comprehensively reflect students' learning situations and avoiding introducing too much noise;
[0044] Model design:
[0045] Design the architecture of a recurrent neural network, including an input layer, a hidden layer (recurrent layer), and an output layer. The formula of the recurrent neural network is
[0046] h t = f(W hh h t-1 + W xh x t + b h )
[0047] Among them, h t represents the hidden state at the t-th time step, f is an activation function (such as tanh or ReLU), W hh is the weight matrix from the hidden state to the hidden state, W xh is the weight matrix from the input to the hidden state, x t is the input at the t-th time step, b h is the bias term of the hidden layer;
[0048] The formula for the output layer is
[0049] y t = g(W hy h t + b y )
[0050] Among them, y t represents the output at the t-th time step, g is the activation function of the output layer (such as softmax for classification tasks), W hy is the weight matrix from the hidden layer to the output layer, b y is the bias term of the output layer;
[0051] Model implementation:
[0052] Implement a recurrent neural network model using TensorFlow;
[0053] Model training:
[0054] Train the model using the training dataset, update the model's weights through the backpropagation algorithm, monitor the loss function value and evaluation metrics (such as accuracy, recall, etc.) during the training process, ensure that the model has good performance on the training set, and evaluate the model using the validation dataset. Adjust the model parameters according to the evaluation results until the model reaches satisfactory performance;
[0055] Overfitting handling:
[0056] Use L1 regularization to prevent the model from overfitting, monitor the performance on the validation set, and stop training in a timely manner to avoid overfitting;
[0057] Testing process:
[0058] Test the model using the test dataset, verify the generalization ability of the model, and adjust and optimize the model according to the test results.
[0059] Further, in S5, the learning plan includes goal setting, time planning, and content arrangement. The teaching methods include differentiated instruction, project-based learning, collaborative learning, and technology-assisted instruction. Differentiated instruction provides different teaching contents, speeds, and methods according to the abilities, interests, and needs of students. Project-based learning cultivates students' practical abilities, teamwork abilities, and problem-solving abilities by allowing them to participate in actual projects. Collaborative learning encourages students to have group discussions and cooperation to promote their communication and mutual assistance. Technology-assisted instruction uses modern technology tools, such as online learning platforms, virtual reality, etc., to provide students with richer and more diverse learning resources. The tutoring strategies include one-on-one tutoring, group tutoring, and online tutoring. One-on-one tutoring provides individual guidance for students' specific problems and offers targeted solutions. Group tutoring forms groups of students with similar learning needs for collective tutoring and discussion. Online tutoring uses online platforms to provide remote tutoring for students, facilitating their learning anytime and anywhere.
[0060] The present invention has the following beneficial effects:
[0061] 1. In the present invention, by collecting and analyzing students' learning behaviors, grades, and online interaction data, the individual differences and learning needs of students can be more accurately identified. At the same time, by constructing a personalized learning model, according to the actual learning needs and learning effects of students, more targeted and personalized teaching suggestions can be provided for teachers. This helps teachers better adapt to the learning styles and ability levels of students, improving teaching effects and learning quality.
[0062] 2. In the present invention, by using big data analysis and machine learning algorithms, in-depth mining and analysis of students' learning data are carried out to provide data-based scientific decision-making support for teachers. This helps teachers more accurately judge the learning status and learning effectiveness of students, thereby formulating more reasonable and effective teaching strategies and methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a step diagram of a user data analysis method based on big data proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0065] Please refer to Figure 1As shown in the figure, the present invention is a method for analyzing user data based on big data, comprising the following steps:
[0066] S1: Data collection: Collect student data, where the student data includes online learning platform data, performance management system data, and online interaction data;
[0067] S2: Data preprocessing: Perform preprocessing operations on the data. The preprocessing operations include data cleaning, data standardization, and data integration. The data standardization uniformly processes data from different sources and in different formats. The data integration integrates the student data from multiple data sources to form a complete set of student learning data;
[0068] S3: Multi-dimensional data analysis: Use the student learning data for learning behavior analysis, performance analysis, and online interaction analysis. The learning behavior analysis includes evaluating the student's learning engagement and learning habits, identifying the student's learning peak and trough periods, and then determining whether the student is more inclined to independent learning and identifying the student's learning pattern. The performance analysis includes evaluating the student's learning outcomes and identifying the student's learning difficulties and weak links. The online interaction analysis includes evaluating the student's social interaction ability and collaborative learning ability and identifying the student's learning community and interaction pattern;
[0069] S4: Construction of personalized learning model: Construct a personalized learning model based on the student's learning pattern, learning difficulties, and weak links;
[0070] S5: Intelligent teaching suggestions: Timely feedback the student's learning data and analysis results to the teacher. The personalized learning model provides personalized teaching suggestions for the teacher according to the student's learning behavior, performance, and online interaction. The teaching suggestions include learning plans, teaching methods, and tutoring strategies for specific students. By regularly comparing the student's learning data and the implementation of the teaching suggestions, evaluate the effectiveness of the teaching suggestions.
[0071] In one embodiment, for the above S1, in S1, the online learning platform data includes learning time (cumulative / daily average), learning frequency (number of logins / week), learning path (course / chapter access order), online quiz scores (average score / highest score / lowest score), and homework completion status (completion rate / submission time). The performance management system data includes examination subjects, examination time, examination scores (raw score / percentile / grading system), and historical performance records (progress / regression trend). The online interaction data includes the number of questions, answer quality (number of likes / number of comments), and discussion depth (reply level / keyword analysis).
[0072] In one embodiment, for the above-mentioned S2, in S2, the data cleaning includes removing duplicate data, handling missing values, and detecting and handling outliers. The duplicate data removal uses a unique identifier (such as student ID) to identify and delete duplicate records. For the missing data in the missing value handling, it is filled (such as using the mean, median, mode, etc.) or deleted according to the specific situation. The outlier detection and handling identify outliers through box plots and correct or delete them according to business logic.
[0073] In one embodiment, for the above-mentioned S3, in S3, the specific steps of the learning behavior analysis are as follows:
[0074] Learning time analysis:
[0075] Cumulative learning time: Calculate the total learning time of a student within a certain period (such as one day, one week, one month). The calculation formula is cumulative learning time = Σ(daily learning time);
[0076] Average daily learning time: Reflect the average learning intensity of a student. The calculation formula is average daily learning time = cumulative learning time / total number of days;
[0077] Learning frequency: Record the number of times a student logs in to the learning platform and the number of active days. The calculation formula is learning frequency = number of logins / total number of days (or number of active days / total number of days);
[0078] Student A's cumulative learning time within a week is 40 hours, the average daily learning time is 6 hours, logs in to the learning platform 5 times, and the number of active days is 5 days. This indicates that Student A has a relatively high learning engagement and has learning activities every day;
[0079] Identification of learning peak and trough periods:
[0080] Use the exponential smoothing method to identify the change trend of a student's learning time. According to the trend chart, determine the student's learning peak and trough periods;
[0081] By analyzing the learning time data of Student B, it is found that from 8 pm to 10 pm every day is the learning peak period, while from 12 noon to 2 pm is the learning trough period. Teachers can adjust the teaching time and homework assignment time based on this information;
[0082] Judgment of self-directed learning tendency:
[0083] Combining the student's learning time, learning frequency, and learning path, analyze the student's learning autonomy. If a student shows a relatively high learning engagement, stable learning frequency, and autonomous learning path, then they may be more inclined to self-directed learning;
[0084] Learning mode recognition:
[0085] Using random forest, by analyzing students' learning data, find students' learning patterns, which include self-directed learning type, dependent type, and mixed type;
[0086] By analyzing the splitting paths of decision trees, explain how the model makes predictions based on students' learning data. For example, if a certain decision tree splits on the feature of "learning time" and students with self-directed learning type usually have longer learning time, then we can consider that this decision tree plays an important role in identifying self-directed learning type students.
[0087] In one embodiment, for the above S3, in S3, the specific steps of the achievement analysis are as follows:
[0088] Exam score analysis: Calculate the average score, highest score, lowest score, norm-referenced score, and contribution to meeting the standard of students, and use a trend chart to show the changes in students' scores.
[0089] Student C's math score has increased from 75 points in the last semester to 85 points in this semester, indicating an improvement in their learning achievements;
[0090] Identification of learning difficulties and weak links:
[0091] Through cluster analysis, identify the knowledge points or question types on which students perform poorly, and combine the students' homework completion situations to further analyze the students' learning difficulties;
[0092] By analyzing Student D's exam scores and homework completion situations, it is found that they perform poorly in the geometry part of mathematics, which may be the learning difficulty.
[0093] In one embodiment, for the above S3, in S3, the specific steps of the online interaction analysis are as follows:
[0094] Assessment of social interaction ability:
[0095] Using text mining and natural language processing technologies, analyze the interaction content of students on various platforms such as online classrooms and discussion areas, and calculate interaction indicators such as the number of questions asked, answers given, and likes received by students;
[0096] Student E asked 10 questions, answered 20 questions, and received 50 likes within a month, indicating that their social interaction ability is relatively strong;
[0097] Assessment of collaborative learning ability:
[0098] Analyze the performance of students in group discussions and collaborative tasks, and calculate the collaborative contribution and collaborative efficiency of students;
[0099] In the group project, student F actively participated in the discussion, put forward multiple constructive opinions, and completed the tasks on time, demonstrating strong collaborative learning ability;
[0100] Learning community and interaction pattern recognition:
[0101] Use network analysis techniques to identify students' learning communities and interaction patterns, and analyze the association relationships between students, such as friendship, cooperation, etc.;
[0102] By analyzing the online interaction data of student G, it is found that G often discusses problems with classmates H and I, forming a small learning community. Teachers can use this information to promote mutual learning and help among students in this community.
[0103] In one embodiment, for S4 above, in S4, the specific steps for constructing the personalized learning model are as follows: construct a personalized learning model according to students' learning patterns, learning difficulties, and weak links
[0104] Data preparation:
[0105] Collect students' learning data, including learning behavior records (such as learning time, learning progress, number of clicks on learning resources, etc.), learning achievements (such as homework grades, test scores, etc.), and students' basic information (such as age, gender, learning style, etc.), and divide the data into training set, validation set, and test set for model training, validation, and testing;
[0106] Feature extraction and selection:
[0107] Extract features related to learning patterns, learning difficulties, and weak links from the original data. For example, features such as students' learning time distribution, learning resource preferences, and types of wrong questions can be extracted.
[0108] Use chi-square test to screen the features that are most helpful for model construction, ensuring that the selected features can comprehensively reflect students' learning situations while avoiding introducing too much noise;
[0109] Model design:
[0110] Design the architecture of the recurrent neural network, including the input layer, hidden layer (recurrent layer), and output layer. The formula of the recurrent neural network is
[0111] h t =f(W hh h t-1 +W xh x t +b h )
[0112] Where h trepresents the hidden state at the tth time step, f is the activation function (such as tanh or ReLU), W hh is the weight matrix from hidden state to hidden state, W xh is the weight matrix input to the hidden state, x t is the input of the tth time step, b h is the bias term of the hidden layer;
[0113] The output layer formula is
[0114] y t =g(W hy h t +b y )
[0115] Among them, y t represents the output of the t-th time step, g is the activation function of the output layer (such as softmax for classification tasks), W hy is the weight matrix from the hidden layer to the output layer, b y is the bias term of the output layer;
[0116] Model implementation:
[0117] Implement a recurrent neural network model using TensorFlow;
[0118] Model training:
[0119] Use the training dataset to train the model, update the model weights through the backpropagation algorithm, monitor the loss function value and evaluation indicators (such as accuracy and recall) during the training process to ensure that the model has good performance on the training set, and evaluate the model using the validation dataset. Adjust the model parameters based on the evaluation results until the model achieves satisfactory performance;
[0120] Overfitting processing:
[0121] Use L1 regularization to prevent the model from overfitting, monitor the performance on the validation set, and stop training in time to avoid overfitting;
[0122] Testing process:
[0123] Use the test data set to test the model, verify the generalization ability of the model, and adjust and optimize the model based on the test results.
[0124] In one embodiment, for the above-mentioned S5, in S5, the learning plan includes goal setting, time planning, and content arrangement. The teaching methods include differentiated instruction, project-based learning, collaborative learning, and technology-assisted instruction. The differentiated instruction provides different teaching contents, speeds, and methods according to the students' abilities, interests, and needs. The project-based learning cultivates their practical abilities, teamwork abilities, and problem-solving abilities by allowing students to participate in actual projects. The collaborative learning encourages group discussions and cooperation among students to promote their communication and mutual assistance. The technology-assisted instruction uses modern technology tools, such as online learning platforms, virtual reality, etc., to provide students with richer and more diverse learning resources. The tutoring strategies include one-on-one tutoring, group tutoring, and online tutoring. The one-on-one tutoring provides individual guidance for the specific problems of students and offers targeted solutions. The group tutoring forms groups of students with similar learning needs for collective tutoring and discussion. The online tutoring uses online platforms to provide remote tutoring for students, facilitating their learning anytime and anywhere:
[0125] Input layer:
[0126] Input features: Students' learning behavior data (such as online learning time, learning frequency, learning path, etc.), performance data (such as scores of previous exams, subject score distributions, etc.), and online interaction situations (such as the number of questions asked, the accuracy of answering questions, the frequency of communication with classmates and teachers, etc.);
[0127] Input shape: Determined according to the actual situation of the dataset. For example, the learning data of each student can be represented as a time series, and each time step contains multiple feature values;
[0128] Hidden layer (recurrent layer):
[0129] Number of hidden layers: Determined according to the complexity of the problem and the scale of the data. For a personalized learning model, multiple layers of RNNs may be required to capture learning patterns at different time scales.
[0130] Number of nodes: The number of nodes in each layer of RNN (i.e., the size of the hidden state) should be adjusted according to the size and complexity of the dataset. A larger number of nodes can improve the representational ability of the model, but may also lead to overfitting and an increase in computational costs.
[0131] Activation function: Select a suitable activation function (such as tanh, ReLU, etc.) to introduce non-linearity and help the model learn complex patterns.
[0132] Output layer:
[0133] Output features: Personalized teaching suggestions, such as recommended learning resources, areas that need to be focused on, suitable learning strategies, etc.
[0134] Output shape: Determined according to specific requirements. For example, it can be a classification label (such as "strengthening mathematical foundation", "improving reading comprehension ability", etc.), or a continuous value (such as the predicted exam score).
[0135] Suppose a student named Xiaoming encounters difficulties in mathematics learning. The intelligent teaching system collects his learning data and conducts an analysis. The analysis results show that Xiaoming has obstacles in understanding mathematical concepts and lacks effective learning methods. Based on these data, the system generates a personalized learning plan for Xiaoming, including special training for mathematical concepts, suggestions on learning methods suitable for him (such as using charts to assist understanding, etc.), and regular tutoring strategies. The teacher adjusts the teaching strategy according to these suggestions and provides targeted tutoring for Xiaoming. After a period of time, by comparing Xiaoming's learning data and the implementation of teaching suggestions, it is found that his mathematics score has been significantly improved, and his learning methods and habits have also been improved.
[0136] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for analyzing user data based on big data, characterized in that, It includes the following steps: S1: Data collection: Collect student data, which includes online learning platform data, performance management system data, and online interaction data; S2: Data preprocessing: Perform preprocessing operations on the data. The preprocessing operations include data cleaning, data standardization, and data integration. The data standardization uniformly processes data from different sources and in different formats, and the data integration integrates student data from multiple data sources to form a complete set of student learning data; S3: Multi-dimensional data analysis: Use student learning data for learning behavior analysis, performance analysis, and online interaction analysis. The learning behavior analysis includes evaluating students' learning engagement and learning habits, identifying students' learning peak and trough periods, and then determining whether students are more inclined to independent learning and identifying students' learning patterns. The performance analysis includes evaluating students' learning outcomes and identifying students' learning difficulties and weak links. The online interaction analysis includes evaluating students' social interaction and collaborative learning abilities and identifying students' learning communities and interaction patterns; S4: Construction of personalized learning model: Construct a personalized learning model based on students' learning patterns, learning difficulties, and weak links; S5: Intelligent teaching suggestions: Timely feedback students' learning data and analysis results to teachers. The personalized learning model provides personalized teaching suggestions for teachers based on students' learning behaviors, performances, and online interaction situations. The teaching suggestions include learning plans, teaching methods, and tutoring strategies for specific students. By regularly comparing students' learning data and the implementation of teaching suggestions, evaluate the effectiveness of teaching suggestions.
2. The user data analysis method based on big data according to claim 1, characterized in that, In S1, the online learning platform data includes learning time, learning frequency, learning path, online quiz scores, and homework completion status. The performance management system data includes exam subjects, exam times, exam scores, and historical score records. The online interaction data includes the number of questions, answer quality, and discussion depth.
3. A method for user data analysis based on big data according to claim 1, characterized in that In S2, the data cleaning includes removing duplicate data, handling missing values, and detecting and handling outliers. The removal of duplicate data uses unique identifiers to identify and delete duplicate records. For missing data, according to specific situations, fill in or delete them. The detection and handling of outliers identify outliers through box plots and correct or delete them according to business logic.
4. A method for user data analysis based on big data according to claim 1, characterized in that, In S3, the specific steps of the learning behavior analysis are as follows: Learning time analysis: Cumulative learning time: Calculate the total learning time of students within a certain period. The calculation formula is cumulative learning time = Σ(daily learning time); Average daily learning time: Reflect the average learning intensity of students. The calculation formula is average daily learning time = cumulative learning time / total number of days; Learning frequency: Record the number of times students log in to the learning platform and the number of active days. The calculation formula is learning frequency = number of logins / total number of days (or number of active days / total number of days); Identification of learning peak and trough periods: Use the exponential smoothing method to identify the change trend of students' learning time. According to the trend chart, determine students' learning peak and trough periods; Judgment of autonomous learning tendency: Analyze students' learning autonomy based on their learning time, learning frequency, and learning path; Learning pattern recognition: By analyzing students' learning data using random forest, students' learning patterns are found, including autonomous learning, dependent learning, and mixed learning.
5. A method for analyzing user data based on big data according to claim 1, characterized in that In S3, the specific steps of the performance analysis are: Test score analysis: calculate students' average score, highest score, lowest score, norm score, and contribution to the standard, and use trend charts to show changes in students' scores; Identification of learning difficulties and weak links: Through cluster analysis, we can identify which knowledge points or question types students perform poorly on, and further analyze students' learning difficulties based on their homework completion status.
6. The method for analyzing user data based on big data according to claim 1, wherein In S3, the specific steps of the online interaction analysis are: Social interaction ability assessment: Use text mining and natural language processing technologies to analyze students’ interactive content on various platforms and calculate student interaction indicators; Collaborative learning ability assessment: Analyze students' performance in group discussions and collaborative tasks, and calculate their collaborative contribution and efficiency; Learning communities and interaction pattern recognition: Use network analysis technology to identify students' learning communities and interaction patterns, and analyze the relationships between students.
7. A method for user data analysis based on big data according to claim 1, characterized in that, In S4, the specific steps of constructing the personalized learning model are: Data preparation: Collect students' learning data, including learning behavior records, learning outcomes, and basic student information, and divide the data into training sets, validation sets, and test sets for model training, validation, and testing; Feature extraction and selection: Extract features related to learning patterns, learning difficulties, and weak links from the original data, and use the chi-square test to select the features that are most helpful for model construction; Model design: Design the architecture of the recurrent neural network, including the input layer, hidden layer and output layer. The formula of the recurrent neural network is h t = f(W hh h t-1 + W xh x t + b h ) Among them, h t represents the hidden state at the t-th time step, f is the activation function, and W hh is the weight matrix from the hidden state to the hidden state, and W xh is the weight matrix from the input to the hidden state, x t is the input at the t-th time step, and b h is the bias term of the hidden layer; The output layer formula is y t = g(W hy h t + b y ) Among them, y t represents the output at the t-th time step, g is the activation function of the output layer, and W hy is the weight matrix from the hidden layer to the output layer, and b y is the bias term of the output layer; Model implementation: Implement a recurrent neural network model using TensorFlow; Model training: Use the training dataset to train the model, update the model weights through the backpropagation algorithm, monitor the loss function value and evaluation indicators during the training process, and use the validation dataset to evaluate the model. Adjust the model parameters based on the evaluation results until the model achieves satisfactory performance; Overfitting processing: Use L1 regularization to prevent the model from overfitting, monitor the performance on the validation set, and stop training in time to avoid overfitting; Testing process: Use the test data set to test the model, verify the generalization ability of the model, and adjust and optimize the model based on the test results.
8. A method for user data analysis based on big data according to claim 1, characterized in that, In S5, the learning plan includes goal setting, time planning, and content arrangement. The teaching methods include differentiated instruction, project-based learning, cooperative learning, and technology-assisted instruction. The differentiated instruction provides different teaching contents, speeds, and methods according to the students' abilities, interests, and needs. The project-based learning cultivates their practical abilities, teamwork abilities, and problem-solving abilities by having students participate in actual projects. The cooperative learning encourages group discussions and cooperation among students. The technology-assisted instruction uses modern technology tools to provide learning resources for students. The tutoring strategies include one-on-one tutoring, group tutoring, and online tutoring. The one-on-one tutoring provides individual guidance for the specific problems of students and offers solutions. The group tutoring forms groups of students with similar learning needs for collective tutoring and discussion. The online tutoring uses online platforms to provide remote tutoring for students.
Citation Information
Cited By
Student score analysis and feedback method and system based on natural language processing
CN122243694A