Student management system based on big data

Through data collection, analysis and mining of the big data student management system, a multi-dimensional student portrait is built, the data island problem in traditional student management is solved, personalized teaching and precise management are realized, and the quality and efficiency of educational management are improved.

CN120336403AInactive Publication Date: 2025-07-18ZHANGZHOU CITY UNIV

Patent Information

Application Number
CN202510757359.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional student management methods cannot integrate multi-source data, and it is difficult to provide a comprehensive and in-depth comprehensive student situation portrait. It is difficult for teachers to formulate personalized teaching plans, and it is impossible to grasp students' learning progress and grade trends in real time and accurately.

Method used

Adopt a student management system based on big data, through data collection, storage and preprocessing, data analysis and mining, student portrait construction and management application modules, the integration and accurate analysis of multiple data sources is realized, multi-dimensional student portraits are constructed, and personalized management strategies are provided.

Benefits of technology

It has achieved comprehensive integration and accurate analysis of student data, helping teachers and school administrators to formulate personalized management strategies, enhance the coordination of home-school co-education, and improve the quality and efficiency of education management.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention relates to the technical field of student management, in particular to a student management system based on big data. The system comprises a data acquisition module, a data storage and preprocessing module, a data analysis and mining module, a student portrait construction module, a management application module and the like. The data acquisition module acquires structured and unstructured data of students from multiple sources such as school educational administration and attendance; the storage and preprocessing module adopts a distributed storage architecture to guarantee the stability and reliability of the data, and performs preprocessing operations such as cleaning and standardization on the data; the data analysis and mining module deeply mines data values by using clustering analysis, association rule mining and time sequence analysis, and provides a basis for student portrait construction, and portraits cover dimensions such as basic information, learning and behavioral performance. According to the system, through comprehensive data integration, accurate analysis and multi-role collaborative application, the education management quality and efficiency are effectively improved, comprehensive development of students is assisted, and the system is suitable for student management scenes of various schools.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of student management, and specifically to a student management system based on big data. Background Art

[0002] With the vigorous development of the education cause, the scale of schools has been continuously expanding, and the number of students has been increasing day by day. Traditional student management methods are facing many severe challenges.

[0003] On the one hand, student-related information is scattered and stored in different systems. For example, the academic affairs system records course arrangements and exam scores, the attendance system separately stores attendance situations, and the campus card system masters students' consumption behaviors, etc. These data are isolated from each other, difficult to integrate and analyze, and unable to provide managers with a comprehensive and in-depth portrait of students' comprehensive situations.

[0004] On the other hand, teachers mainly rely on periodic exam scores and limited daily observations to understand students' learning situations. It is difficult to accurately grasp the learning progress, advantages and disadvantages of each student in real time, as well as the factors specifically affecting students' academic performance and attendance, etc. Thus, it is difficult to formulate personalized teaching plans that truly meet the individual needs of students, and it is also difficult to identify the trend of students' performance changes through existing data, and unable to intervene in advance to reduce possible problems such as performance decline. Summary of the Invention

[0005] A student management system based on big data, characterized by including: A data collection module: used to collect student-related data from multiple data sources, and the data sources include the school academic affairs system and the student daily attendance system; A data storage and preprocessing module: stores the data collected by the collection module, and preprocesses the relevant data for subsequent data analysis and mining operations; A data analysis and mining module: analyzes and mines the preprocessed data using clustering analysis, association rule mining, and time series analysis; A student portrait construction module: constructs a student portrait based on the results of the data analysis and mining, and the student portrait covers multiple dimensions such as students' basic information, learning situations, and behavioral performances; A management application module: the management application module includes a teacher-side application, a parent-side application, and a school management-level application, which are respectively used by teachers, parents, and school management personnel to assist in corresponding student management work.

[0006] Preferably, the data storage and preprocessing module adopts a distributed storage architecture. During deployment, the number of data nodes is reasonably planned according to the expected amount of data to be stored, the data block size is configured, and through network bandwidth monitoring and load balancing mechanisms, the efficient transmission of data between storage nodes and the load balancing of each node are ensured, preventing data storage hotspots and transmission bottleneck problems, and regularly performing integrity verification and redundant data repair on the stored data to maintain the stability and reliability of data storage.

[0007] Preferably, the specific method of using cluster analysis for data analysis in the data analysis and mining module is as follows: From the student dataset processed by the data storage and preprocessing module, relevant features for cluster analysis are screened out, including the average scores of each subject of students, learning behavior data, attendance, and participation in campus activities; Perform data standardization processing on the screened relevant features; According to the past classification of student groups and the conventional classification requirements of school management, the number of clusters K value is preset in advance; Take the standardized student data as input, and run the K-Means clustering algorithm according to the determined number of clusters to obtain K different student group clusters; Analyze and interpret the obtained clustering results to understand the common characteristics of students in each cluster.

[0008] Preferably, the specific method of using association rule mining for data analysis in the data analysis and mining module is as follows: From the student dataset processed by the data storage and preprocessing module, a transaction dataset for association rule mining is screened out. Specifically, a transaction includes data information such as the average scores of each subject of students, attendance, participation in campus activities, and the usage of online learning platforms. These data together form transaction records, forming a basic dataset for mining association rules; Perform data standardization processing on the screened relevant features; Set the support threshold and confidence threshold, and select the Apriori algorithm to perform association rule mining on the transaction dataset; Deeply analyze the mined association rules and screen out the factors that have a positive impact on academic performance, attendance, and participation in campus activities.

[0009] Preferably, use cluster analysis to determine the student group with relatively low academic performance, less than ideal attendance, and less participation in activities, and define it as the "to-be-improved" student group, and provide corresponding assistance and tutoring to the "to-be-improved" student group based on the factors that have a positive impact on academic performance, attendance, and participation in campus activities screened out by association rule mining.

[0010] Preferably, the specific method of using time series analysis to perform data analysis in the data analysis and mining module is: Filter out learning data and attendance data from the student data set processed by the data storage and preprocessing module, where the learning data includes the student's academic performance information in various subjects at different time points, and the attendance data includes the student's daily attendance records; Arrange the collected data in chronological order to form a time series data format; Perform a stationarity test on the time series data. If the score or attendance time series data fails the stationarity test, it will be stabilized. Use the maximum likelihood estimation method, etc., to use the existing time series data, perform parameter estimation through the model fitting function, and obtain the estimated values of the model parameters; Apply the model with estimated parameters to the time series data for residual analysis and fitting. If the residual sequence is approximately white noise, it means that the model fits the data well and can effectively capture the rules in the time series. If the residual does not meet the white noise characteristics, readjust the model parameters, fit and test again until a model with good fitting effect is obtained. Using the well-fitted and tested time series model, the future trend of students' grades and attendance performance is predicted.

[0011] Preferably, the student portrait specifically includes: Basic student information: including student name, gender, date of birth, grade, and class information, which is collected directly from the school's academic affairs system and kept updated in real time; Learning status: record the performance trends of students in various subjects, draw a line chart to intuitively display the performance changes, and clearly identify the subjects that students are good at and weak in; Behavioral performance: Integrate students' daily attendance records, participation in campus activities, and disciplinary records to form a comprehensive description of their behavioral performance. Through statistical analysis, we can determine the students' level of activity on campus and the degree of discipline they abide by.

[0012] Preferably, the system also includes a security and privacy protection mechanism, which uses encryption technology and access control mechanism to ensure the security and privacy of student data, including: The sensitive information of students is encrypted using a symmetric encryption algorithm. When storing data, the encrypted data is stored in a distributed storage system in ciphertext form. During data transmission, the HTTPS transmission protocol is used to ensure the confidentiality and integrity of data transmission. Decryption operations are performed only on legally authorized user ends, and the decryption key is kept and distributed through the key management system. The access control mechanism assigns different account permissions to users of different roles, sets the user login authentication process. Teachers can only access and operate on the relevant data of the students in the classes they are responsible for. Parents can only view the data of their own children. School administrators can obtain the student data of the whole school within the corresponding scope according to their management levels. Strict permission reviews are carried out for key operations such as data modification and deletion. At the same time, the operation logs of all users are recorded to facilitate tracing of abnormal operation behaviors and prevent data leakage and illegal access.

[0013] Compared with the existing technologies, the advantages of the present invention are as follows: The data collection module of the system can collect student data covering structured and unstructured data from multiple data sources, break data silos, and achieve comprehensive integration, laying a foundation for in-depth insight into student situations. At the same time, the data analysis and mining module uses technologies such as cluster analysis, association rule mining, and time series analysis to accurately mine the data value, find out the key factors affecting student development, help teachers and school administrators formulate personalized management strategies, achieve precise management, and meet the development needs of different students. The student portrait construction module presents student situations in multiple dimensions in an intuitive and visual way, facilitating teachers, parents, and administrators to quickly understand student characteristics and enhancing the synergy of home-school co-education. Specific implementation manners

[0014] A student management system based on big data includes: A data collection module: used to collect student-related data from multiple data sources, and the data sources include the school educational administration system and the student daily attendance system. A data storage and preprocessing module: stores the data collected by the collection module and preprocesses the relevant data for subsequent data analysis and mining operations. A data analysis and mining module: uses cluster analysis, association rule mining, and time series analysis to analyze and mine the preprocessed data. A student portrait construction module: constructs a student portrait based on the results of the data analysis and mining, and the student portrait covers multiple dimensions such as student basic information, learning situation, and behavior performance. A management application module: the management application module includes a teacher-side application, a parent-side application, and a school management-level application, which are used by teachers, parents, and school administrators respectively to assist in corresponding student management work.

[0015] In an embodiment, the data collection module is integrated with the existing school educational administration system, and through a Web service interface, structured data such as students' test scores, usual homework scores, and course selection situations in each subject are obtained once a week at a fixed time, and these data are sorted in a unified format and then transmitted to the data storage module.

[0016] The data acquisition module also uses the school's internal network to establish a real-time data connection with the students' daily attendance system. Whenever a student punches in for attendance, the attendance system immediately sends the attendance record to the data acquisition module. After the data acquisition module receives the record in real time and conducts a preliminary data format verification to ensure the accuracy and integrity of the data, it forwards the data to the data storage module.

[0017] In this embodiment, various types of student data collected are transmitted through the network and enter the distributed storage cluster based on HDFS, where they are classified and stored according to the data type and source. At the same time, according to the pre-configured redundant backup strategy, each data block is automatically backed up three times on different data nodes to ensure the high reliability of the data. Even if a certain node fails, it will not affect the normal use of the data.

[0018] After storage, preprocessing is performed on the data, which specifically includes: Data cleaning: Write Python scripts to comprehensively clean the stored data. For structured data, check the integrity of the data, such as deleting records with missing student IDs or empty key time fields; verify the accuracy of the data, correct or mark as abnormal data cases where the exam scores exceed the normal score range; remove duplicate data records to avoid interference with subsequent analysis; for unstructured data, remove useless characters, emojis and other interfering information to make the text content more standardized and facilitate subsequent processing. Data standardization: Unify the data format, convert the date format to the ISO8601 form uniformly, and record the time in the 24-hour system; perform normalization processing on numerical data, such as normalizing the exam scores to the interval [0,1], so that the score data of different subjects and different sources are in the same dimension, facilitating subsequent data analysis and mining operations.

[0019] In another embodiment, clustering analysis is performed on the data, which specifically includes: Feature selection and data preparation: Select the average scores of each subject of the students, the full-semester attendance situation, and the number of campus club activities participated as relevant features for clustering analysis from the preprocessed student dataset. Organize the data of these features according to individual students to form a two-dimensional data table, where each row represents a student and each column corresponds to the corresponding feature value, and then perform standardization processing on these feature data to ensure that each feature has the same importance in the clustering process.

[0020] Clustering operation: Based on the school's past general classification experience of student groups and daily management needs, the initial value of the clustering number K is set to 4. The K-Means clustering algorithm is used, with the standardized student data as the input. Four clustering centers are randomly initialized, and then through iterative calculations, the cluster to which each student belongs and the positions of the clustering centers are continuously adjusted until the change in the clustering centers is less than the preset threshold of 0.001, and finally four different clusters of student groups are obtained.

[0021] Result analysis and interpretation: A detailed analysis is carried out on the four clusters of student groups obtained by clustering. It is found that the students in one cluster have relatively high average academic scores, good attendance throughout the semester, and participate in more club activities. This cluster is defined as the "excellent and active" student group; the students in another cluster have medium average academic scores, are occasionally late for attendance, and participate in fewer club activities, and are classified as the "medium stable" student group; there is also a cluster of students with relatively low scores, more attendance problems, and low enthusiasm for participating in club activities, which is defined as the "to-be-improved" student group; finally, the students in the last cluster have good scores but average attendance and participate in fewer club activities, and can be called the "need-to-strengthen self-study" student group. Through such clustering analysis, teachers can clearly understand the characteristics of different types of students and provide a basis for subsequent personalized teaching and management.

[0022] In another embodiment, association rule mining is performed on the data, specifically including: Transaction dataset construction and data preparation: Select the transaction dataset for association rule mining from the preprocessed student data. Each transaction includes the average academic scores of each subject of the student, the attendance situation throughout the semester, and the number of campus club activities participated; these data are sorted according to individual students to form individual transaction records, thereby constructing a complete transaction dataset, and standardizing the relevant features in the dataset to make the data meet the requirements of the mining algorithm.

[0023] Parameter setting and algorithm execution: Set the support threshold to 0.2, which means that only the item sets that appear in the transaction dataset with a frequency of 20% or more are regarded as frequent item sets; the confidence threshold is set to 0.7, that is, the confidence of the association rule needs to reach 70% or more to be considered an effective and worthy-of-attention rule. Select the Apriori algorithm to perform association rule mining on the transaction dataset. This algorithm first scans the dataset to find frequent 1-item sets, then generates candidate 2-item sets based on the frequent 1-item sets, scans the dataset again to count their occurrences, and filters out the frequent 2-item sets. And so on, continuously iterate to generate higher-order candidate item sets and filter out the frequent item sets until no new frequent item sets can be generated. Finally, generate association rules that meet the confidence threshold requirements based on the frequent item sets.

[0024] Result Analysis and Application: Conduct in-depth analysis on the mined association rules and find association rules such as "Participating in math club activities → Good math grades (confidence 0.75)", "Online learning duration exceeding 6 hours per week → Improvement in average grades of all subjects (confidence 0.8)". This indicates that there is a strong correlation between participating in specific club activities and corresponding subject grades, and between online learning duration and overall academic performance. Teachers and school administrators can formulate corresponding strategies based on these association rules, such as encouraging students to actively participate in subject-related club activities and guiding students to reasonably increase their online learning duration, etc., to promote the improvement of students' academic performance and the development of their comprehensive qualities.

[0025] In another embodiment, cluster analysis is used to identify a group of students with relatively low academic performance, unsatisfactory attendance, and few participation in activities, and define them as the "students to be improved" group. Based on the factors with positive impacts on academic performance, attendance, and campus activity participation screened by association rule mining, corresponding assistance and tutoring are provided to the "students to be improved" group.

[0026] In another embodiment, time series analysis is performed on the data, specifically including: Data Preparation: Screen the final exam scores of each subject and attendance data of students for time series analysis. For the score data, with the semester as the time unit, collect the final exam scores of each subject of each student in the past 5 semesters; for the attendance data, with the week as the time unit, organize the records of the number of days of attendance per week and the number of late arrivals of each student in the past year. Arrange these data in chronological order to form a time series data format.

[0027] Stationarity Test and Treatment: Use the ADF test and KPSS test to conduct stationarity tests on the score and attendance time series data. For the non-stationary score time series, adopt the method of first-order difference for stationary processing, that is, calculate the difference between the scores of adjacent semesters to obtain a new time series, and conduct the stationarity test again until the series passes the stationarity test.

[0028] Model Selection, Fitting and Testing: Select the ARIMA model for fitting. Through the maximum likelihood estimation method, use the existing time series data to determine the parameter values of the autoregressive order p, the difference order d, and the moving average order q. For example, a set of parameter estimates is obtained as p = 2, d = 1, q = 1. Apply the model with estimated parameters to the time series data for fitting, and conduct residual analysis on the fitted model to check whether the residual series is approximately white noise; if the residual series does not satisfy the white noise characteristics, readjust the model parameters, such as changing the values of p and q, and conduct fitting and testing again until a well-fitted model is obtained.

[0029] Trend prediction and decision-making application: Use a well-fitted and tested time series model to predict the changing trends of students' future grades and attendance performance. For example, based on a student's Chinese language score ARIMA model, predict his Chinese language score and the confidence interval of the score in the next semester. If the prediction results show that the score may decline, the teacher can communicate with the student in advance to understand his learning situation and formulate a special tutoring plan, such as increasing Chinese language tutoring hours and recommending relevant learning materials, to help students avoid a decline in grades. At the same time, school administrators can also reasonably deploy Chinese language teaching staff according to the predicted trend of overall student grades to ensure teaching quality.

[0030] In another embodiment, the system also includes a security and privacy protection mechanism, which uses encryption technology and access control mechanism to ensure the security and privacy of student data, including: The sensitive information of students is encrypted using a symmetric encryption algorithm. When storing data, the encrypted data is stored in a distributed storage system in ciphertext form. During data transmission, the HTTPS transmission protocol is used to ensure the confidentiality and integrity of data transmission. Decryption operations are performed only on legally authorized user ends, and the decryption key is kept and distributed through the key management system. The access control mechanism assigns different account permissions to users with different roles and sets up a user login authentication process. Teachers can only access and operate relevant data of students in the class they are responsible for, parents can only view their own children's data, and school administrators can obtain the corresponding scope of student data for the entire school according to their management level. Strict permission review is performed for key operations such as data modification and deletion. At the same time, the operation logs of all users are recorded to facilitate tracing of abnormal operation behaviors and prevent data leakage and illegal access.

[0031] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0032] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A student management system based on big data, characterized in that, Including: Data acquisition module: used to collect student-related data from multiple data sources, and the data sources include the school academic affairs system and the student daily attendance system; Data storage and preprocessing module: stores the data collected by the acquisition module and preprocesses the relevant data for subsequent data analysis and mining operations; Data analysis and mining module: analyzes and mines the preprocessed data using clustering analysis, association rule mining, and time series analysis; Student portrait construction module: constructs a student portrait based on the results of the data analysis and mining, and the student portrait covers multiple dimensions such as student basic information, learning situation, and behavior performance; Management application module: The management application module includes teacher-side applications, parent-side applications, and school management-level applications, which are used by teachers, parents, and school administrators respectively to assist in corresponding student management work.

2. The specific method of using time series analysis for data analysis in the data analysis and mining module is as follows: Screen out learning data and attendance data from the student dataset processed by the data storage and preprocessing module, where the learning data includes the academic performance information of each subject of the student at different time points, and the attendance data includes the daily attendance records of the student; Arrange the collected data in chronological order to form a time series data format; Conduct a stationarity test on the time series data. If the grade or attendance time series data fails the stationarity test, perform stationarization processing on it; Use methods such as maximum likelihood estimation, utilize the existing time series data, and perform parameter estimation through the model fitting function to obtain the estimated values of the model parameters; Apply the model with estimated parameters to the time series data for residual analysis fitting. If the residual sequence is approximately white noise, it indicates that the model fits the data well and can effectively capture the patterns in the time series; if the residuals do not satisfy the white noise characteristics, readjust the model parameters and perform fitting and testing again until a well-fitting model is obtained; Use the fitted and tested time series model to predict the future change trends of student grades and attendance performance.

3. A student management system based on big data according to claim 1, characterized in that, The data storage and preprocessing module adopts a distributed storage architecture. During deployment, reasonably plan the number of data nodes according to the expected data volume to be stored, configure the data block size, and through network bandwidth monitoring and load balancing mechanisms, ensure the efficient transmission of data between storage nodes and the load balancing of each node, prevent the occurrence of data storage hotspots and transmission bottleneck problems, and regularly perform integrity verification and redundant data repair on the stored data to maintain the stability and reliability of data storage.

4. A student management system based on big data according to claim 1, characterized in that, The specific method of using clustering analysis for data analysis in the data analysis and mining module is as follows: Screen out relevant features for clustering analysis from the student dataset processed by the data storage and preprocessing module, including the average academic performance of each subject of the student, learning behavior data, attendance situation, and campus activity participation; Perform data standardization processing on the selected relevant features; Pre-set the clustering number K value according to the past classification of the student group and the conventional classification requirements of school management; The standardized student data is used as input, and the K-Means clustering algorithm is run according to the determined number of clusters to obtain K different student group clusters; The obtained clustering results were analyzed and interpreted to understand the common characteristics of students in each cluster.

5. The student management system based on big data according to claim 1, characterized in that, The specific method of using association rule mining to perform data analysis in the data analysis and mining module is: From the student data set processed by the data storage and preprocessing module, the transaction data set for association rule mining is screened out. Specifically, a transaction includes data information such as the student's average scores in each subject, attendance, participation in campus activities, and use of the online learning platform. These data together constitute transaction records and form the basic data set for mining association rules. Perform data standardization on the selected relevant features; Set the support threshold and confidence threshold, and select the Apriori algorithm to mine association rules on the transaction data set; The mined association rules are deeply analyzed to screen out factors that have a positive impact on academic performance, attendance, and participation in campus activities.

6. A student management system based on big data according to claim 3 or 4, characterized in that, Cluster analysis is used to identify the student group with relatively low grades, unsatisfactory attendance and less participation in activities. They are defined as the "student group to be improved". Based on association rule mining, factors that have a positive impact on academic performance, attendance and participation in campus activities are screened out to provide corresponding help and counseling to the "student group to be improved".

7. A student management system based on big data according to claim 1, characterized in that, The student portrait specifically includes: Basic student information: including student name, gender, date of birth, grade, and class information, which is collected directly from the school's academic affairs system and kept updated in real time; Learning status: record the performance trends of students in various subjects, draw a line graph to intuitively display the performance changes, and clearly identify the subjects that students are good at and weak in; Behavioral performance: Integrate students' daily attendance records, participation in campus activities, and disciplinary records to form a comprehensive description of their behavioral performance. Through statistical analysis, we can determine the students' level of activity on campus and the degree of discipline they abide by.

8. The student management system based on big data according to claim 1, characterized in that, The system also includes security and privacy protection mechanisms, using encryption technology and access control mechanisms to ensure the security and privacy of student data, including: The sensitive information of students is encrypted using a symmetric encryption algorithm. When storing data, the encrypted data is stored in a distributed storage system in ciphertext form. During data transmission, the HTTPS transmission protocol is used to ensure the confidentiality and integrity of data transmission. Decryption operations are performed only on legally authorized user ends, and the decryption key is kept and distributed through the key management system. The access control mechanism assigns different account permissions to users with different roles and sets up a user login authentication process. Teachers can only access and operate relevant data of students in the class they are responsible for, parents can only view their own children's data, and school administrators can obtain the corresponding scope of student data for the entire school according to their management level. Strict permission review is performed for key operations such as data modification and deletion. At the same time, the operation logs of all users are recorded to facilitate tracing of abnormal operation behaviors and prevent data leakage and illegal access.

Citation Information

Patent Citations

  • Student teaching management system based on campus data

    CN112465260A

  • Education quality evaluation method based on machine learning model

    CN119228599A

  • Student growth archive management method

    CN119621697A

  • Smart campus system based on SDN

    CN119887460A

  • System and method of education administration

    US8187004B1

Cited By

  • Method and system for evaluating digital literacy of teacher based on process data

    CN121329248A

  • A method and system for assessing teachers' digital literacy based on process data

    CN121329248B