Artificial Intelligence-Based Student Learning Status Assessment Method and System
By employing reverse neighborhood set calculations and adaptive clustering parameter adjustments, the method addresses non-linear data structures and individual student variations, enhancing the accuracy and coverage of learning state evaluations.
Patent Information
- Application Number
- CN202510201246.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-02-24
AI Technical Summary
In the existing student learning status assessment methods, student learning data has complex nonlinear structures, making it difficult to effectively mine data correlation and determine evaluation priorities, resulting in inaccurate evaluation; there are large differences between individual students, and evaluation is difficult to meet the needs of different students.
Using a method based on reverse neighborhood set, neighborhood density is calculated and center data and edge data are divided, clustering model is constructed, cluster allocation is performed through central neighborhood set and weighted undirected graph, combined with cluster parameter tuning, the optimal clustering parameters are found to adapt to the learning data characteristics of different students.
It improves the accuracy and coverage of learning status assessment, can better adapt to learning data of students at different levels and types, and achieve accurate assessment.
Smart Images

Figure CN119692630B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of learning assessment, and specifically refers to a method and system for assessing students' learning status based on artificial intelligence. Background Art
[0002] The method for assessing students' learning status utilizes artificial intelligence technology to analyze and process students' learning data, automatically assess the learning status of students, and can achieve intelligent monitoring and management of the entire process of students' learning, thereby improving teaching quality and learning efficiency. However, in the existing methods for assessing students' learning status, there are problems that students' learning data has a complex non-linear structure, and the data correlation cannot be effectively mined and the assessment focus cannot be determined, resulting in inaccurate assessment of students' learning status; in the existing methods for assessing students' learning status, there are large differences among individual students, and students' learning data presents complex and diverse characteristics, and the assessment is difficult to meet the needs of different students, resulting in the problem that the actual learning status of students cannot be accurately assessed. Summary of the Invention
[0003] In view of the above situation, to overcome the defects of the existing technologies, the present invention provides an artificial intelligence-based method and system for evaluating the learning status of students. Aiming at the problems existing in the existing methods for evaluating the learning status of students, that is, the learning data of students has a complex non-linear structure, the data associations cannot be effectively mined, and the evaluation focus cannot be determined, resulting in inaccurate evaluation of the learning status of students. This solution calculates the neighborhood density based on the reverse neighborhood set, divides the central data and the edge data, and better captures the complex structure of the learning data of students; combines the neighborhood density and the relative distance to obtain the concentration degree, selects the central data with the largest concentration degree as the centroid and constructs clusters, and more accurately determines the focus of evaluating the learning status of students; continuously updates the clusters by constructing the central neighborhood set to complete the distribution of the central data and ensure the construction quality of the clusters; constructs the edge neighborhood set, obtains the near-neighbor screening set by constructing a weighted undirected graph and calculating the cluster density, completes the distribution of the edge data, obtains the clustering model, improves the coverage and efficiency of the evaluation, and conducts a more accurate evaluation of the learning status of students that conforms to the actual situation; aiming at the problems existing in the existing methods for evaluating the learning status of students, that is, there are large differences among students, the learning data of students presents complex and diverse characteristics, and the evaluation is difficult to meet the needs of different students, resulting in the inability to accurately evaluate the actual learning status of students. This solution establishes a parameter search space for the clustering parameters, and the initial individual position is used as the representative of the clustering parameters, providing more possibilities for adapting to the characteristics of the learning data of different students; calculates the random offset probability, selects the individual position for offset operation based on the offset threshold, which helps to find clustering parameters more suitable for the characteristics of the learning data of different students; based on the global optimal position and the adjustment factor, different update strategies are adopted to update the individual position at different search stages to find the optimal clustering parameters, and according to the characteristics of the learning data of students and the search process, it better adapts to the learning data of different levels and types of students, and improves the accuracy and effectiveness of the learning status evaluation.
[0004] The technical solution adopted by the present invention is as follows: The artificial intelligence-based method for evaluating the learning status of students provided by the present invention includes the following steps:
[0005] Step S1: Data collection;
[0006] Step S2: Data preprocessing;
[0007] Step S3: Constructing a clustering model;
[0008] Step S4: Tuning the clustering parameters;
[0009] Step S5: Evaluating the learning status of students.
[0010] Further, in step S1, the data collection is to collect historical student learning data and real-time student learning data; both the historical student learning data and the real-time student learning data include learning platform data, homework data, exam data, and classroom interaction data, and the historical student learning data further includes learning status; the learning status is used as a data label, and the data label is not considered during clustering and is only used when selecting the cluster label.
[0011] Further, in step S2, the data preprocessing is to perform data cleaning, data transformation, data normalization, and construction of a learning dataset on the collected student learning data; data cleaning includes handling noise, outliers, duplicate values, and missing values; data transformation is to convert the data into vector form; data normalization is to unify the data range based on the maximum-minimum normalization method; constructing the learning dataset is to construct the learning dataset based on the student learning data after data cleaning, data transformation, and data normalization.
[0012] Further, in step S3, the construction of the clustering model specifically includes the following steps:
[0013] Step S31: Calculate the neighborhood density. Calculate the Euclidean distance between any two student learning data in the learning dataset to obtain the K nearest neighbors of each student learning data, and construct the forward neighborhood set , and then construct the reverse neighborhood set . Calculate the neighborhood density of each student learning data based on the reverse neighborhood set. The formula used is as follows:
[0014] ;
[0015] In the formula, x i and x j are the i-th and j-th student learning data in the learning dataset respectively, i and j are student learning data indices, ρ i is the neighborhood density of x i , is the Euclidean distance between x i and x j , and are the forward neighborhood set and the reverse neighborhood set of x i respectively, is the forward neighborhood set of x j , is the Euclidean distance between x i and its K-th nearest neighbor, exp(·) is the exponential function, and α is the neighborhood influence factor;
[0016] Step S32: Data partitioning. If the number of student learning data N1 ini If it is greater than or equal to βK, then x i is classified as central data; otherwise, x i is classified as marginal data; and a central data set and a marginal data set are constructed respectively; where β is a regulation factor;
[0017] Step S33: Calculate the concentration degree, and take the product of the neighborhood density and the relative distance of each student's learning data as the concentration degree ; where δ i and A i are the relative distance and the concentration degree of x i respectively;
[0018] Step S34: Central data allocation, including the following steps:
[0019] Step S341: Select the central data with the largest concentration degree in Q1 as the centroid Z v , construct the corresponding cluster C v based on Z v , and delete Z v from Q1; where Z v is the v-th centroid;
[0020] Step S342: For each central data in C v , construct a central neighborhood set respectively, and the formula used is as follows:
[0021] ;
[0022] In the formula, is the w-th central data in C v , is 's central neighborhood set, and are 's forward neighborhood set and backward neighborhood set respectively, is the intersection operator, x g is the g-th central data in Q1, ρ g is the neighborhood density of x g , γ is the central control factor, ρ max is the maximum neighborhood density among all students' learning data;
[0023] Step S343: For each central data in , add it to C v and delete it from Q1;
[0024] Step S344: Return to Step S342 for iteration until the central data in C v no longer increases, then the cluster C v is initially constructed, and go to Step S345;
[0025] Step S345: Return to Step S341 for iteration until the number of central data in Q1 is 0, then the central data allocation is completed, obtaining V1 centroids and clusters, and go to Step S35;
[0026] Step S35: Allocate edge data. For each edge data in Q2, perform the following steps in sequence:
[0027] Step S351: Construct the edge neighborhood set of the edge data x m using the following formula:
[0028] ;
[0029] In the formula, x m is the m-th edge data in Q2, P1 m is the edge neighborhood set of x m , and are the forward neighborhood set and the reverse neighborhood set of x m respectively, is the set of all student learning data in the learning dataset except Q2, x h is the h-th central data in it, ρ h is the neighborhood density of x h , is the edge control factor;
[0030] Step S352: For each student learning data m in P1 , calculate the Euclidean distance between and the V1 centroids respectively, and select the cluster constructed by the centroid with the minimum Euclidean distance as the cluster corresponding to . Take each student learning data in as a vertex, and the Euclidean distance between two student learning data as the weight of the edge corresponding to the two vertices, construct a weighted undirected graph , and obtain the minimum spanning tree of based on Prim's algorithm , and then calculate the cluster density of ; If the Euclidean distance between x and m is less than , then based on Construct the nearest neighbor screening set P2 m , and the formula used is as follows:
[0031] ;
[0032] In the formula, is the s-th student learning data in P1 m , is the corresponding cluster, is a weighted undirected graph constructed based on , is the minimum spanning tree of is the cluster density of and are respectively the total weight value and the number of edges of , P2 m is the nearest neighbor screening set of x m ;
[0033] Step S353: Find the student learning data with the smallest Euclidean distance between P2 m and x m , calculate the Euclidean distances between and the V1 centroids respectively, add x to the cluster constructed by the centroid with the smallest Euclidean distance, and remove x m from Q2; return to step S351, continue to process the next marginal data until the number of marginal data in Q2 is 0, then the marginal data allocation is completed; obtain the clustering model and output the clustering result; where m is the f-th student learning data in P2 m th .
[0034] Furthermore, in step S4, the clustering parameter tuning specifically includes the following steps:
[0035] Step S41: Initial individual positions, establish a parameter search space for the neighborhood influence factor α, adjustment factor β, center control factor γ, and marginal control factor ε in the clustering, randomly initialize U individual positions within the parameter search space, use the individual positions as representatives of the clustering parameters, and take the average value of the silhouette coefficients of all student learning data in the clustering result obtained based on the individual positions as the fitness value of the individual;
[0036] Step S42: Individual position offset, preset the offset threshold B th , calculate the random offset probability of each individual position , if , an offset operation is performed on the u-th individual position; otherwise, the u-th individual position remains unchanged; the formula used is as follows:
[0037] ;
[0038] ;
[0039] In the formula, is the u-th individual position at the t-th search, is random offset probability, u is the individual position index, t is the search number index, T is the maximum search number, and are two randomly selected individual positions at the t-th search, is the offset control factor, min(·) is the function to take the minimum value, is the individual position after offset, a max and a min are the maximum and minimum values of the scaling factor respectively;
[0040] Step S43: Update of individual positions. Different update strategies are used to update individual positions at different search stages; the formula used is as follows:
[0041] ;
[0042] In the formula, is the u-th individual position at the (t + 1)-th search, is the global optimal position at the t-th search, is the average value of all individual positions at the t-th search, c1, c2, c3 and c4 are non-interfering random numbers, UB and LB are the upper and lower limits of the parameter search space respectively, is the adjustment factor;
[0043] Step S44: Determination of the optimal clustering parameters. A fitness threshold is preset, the fitness value of the individual position is updated. When the fitness value of the global optimal position is higher than the fitness threshold, the global optimal position corresponds to the optimal clustering parameters, and the clustering parameter tuning is completed; otherwise, if the maximum search number is reached, return to Step S41 to re-initialize the individual positions; otherwise, increase the search number by 1 and return to Step S42 to continue the search.
[0044] Further, in step S5, the student learning status evaluation is to jointly input the pre - processed historical student learning data and real - time student learning data into a clustering model constructed based on optimal clustering parameters for clustering processing. Based on the output clustering results, select the label with the largest number of historical student learning data in the cluster as the cluster label, and use the cluster label to which the real - time student learning data belongs as the evaluation result to obtain the learning status of the real - time student learning data, thus completing the evaluation of the student learning status.
[0045] The student learning status evaluation system based on artificial intelligence provided by the present invention includes a data acquisition module, a data pre - processing module, a clustering model construction module, a clustering parameter tuning module, and a student learning status evaluation module;
[0046] The data acquisition module collects historical student learning data and real - time student learning data and sends the data to the data pre - processing module;
[0047] The data pre - processing module performs data cleaning, data conversion, data normalization, and learning dataset construction processing on the collected student learning data and sends the data to the clustering model construction module;
[0048] The clustering model construction module calculates the neighborhood density based on the reverse neighborhood set, divides the central data and the marginal data, combines the neighborhood density and the relative distance to obtain the concentration degree, selects the central data with the largest concentration degree as the centroid and constructs clusters, continuously updates the clusters by constructing the central neighborhood set to complete the central data allocation, constructs the marginal neighborhood set, and obtains the near - neighbor screening set by constructing a weighted undirected graph and calculating the cluster density to complete the marginal data allocation, thus obtaining a clustering model and sending the data to the clustering parameter tuning module;
[0049] The clustering parameter tuning module establishes a parameter search space for the clustering parameters, uses the initial individual position as the representative of the clustering parameters, calculates the random offset probability, selects the individual position for offset operation based on the offset threshold, and updates the individual position using different update strategies at different search stages based on the global optimal position and the adjustment factor to find the optimal clustering parameters and sends the data to the student learning status evaluation module;
[0050] The student learning status evaluation module inputs the student learning data into the clustering model with tuned parameters for processing to obtain the learning status of the real - time student learning data, thus completing the evaluation of the student learning status.
[0051] The beneficial effects achieved by the present invention using the above - mentioned solution are as follows:
[0052] (1) Aiming at the problem that in the existing student learning status evaluation methods, the student learning data has a complex non-linear structure, and the data associations cannot be effectively mined and the evaluation focus cannot be determined, resulting in inaccurate evaluation of the student learning status, this solution calculates the neighborhood density based on the reverse neighborhood set, divides the central data and the edge data, finds the relationship of the tightness around the data points in the complex data structure, and better captures the complex structure of the student learning data; combines the neighborhood density and the relative distance to obtain the concentration degree, selects the central data with the largest concentration degree as the centroid and constructs clusters, comprehensively considering the aggregation degree of the student learning data in the local area and the relative position relationship with other data, and more accurately determines the focus of the student learning status evaluation; continuously updates the clusters by constructing the central neighborhood set to complete the allocation of the central data, ensuring the construction quality of the clusters; constructs the edge neighborhood set, obtains the near-neighbor screening set by constructing a weighted undirected graph and calculating the cluster density, completes the allocation of the edge data, obtains the clustering model, more comprehensively evaluates the student learning status, improves the coverage and efficiency of the evaluation, and conducts a more accurate evaluation of the student learning status that conforms to the actual situation.
[0053] (2) Aiming at the problem that in the existing student learning status evaluation methods, there are large differences among students, the student learning data presents complex and diverse characteristics, and the evaluation is difficult to meet the needs of different students, resulting in the inability to accurately evaluate the actual learning status of students, this solution establishes a parameter search space for the clustering parameters. The initial individual position is used as the representative of the clustering parameters, which can explore various possible combinations of clustering parameters and provides more possibilities for adapting to the characteristics of different students' learning data; calculates the random offset probability, selects the individual position for offset operation based on the offset threshold to avoid the clustering parameters falling into local optima, and helps to find more suitable clustering parameters for the characteristics of different students' learning data; based on the global optimal position and the adjustment factor, adopts different update strategies to update the individual positions at different search stages to find the optimal clustering parameters, and more reasonably adjusts the clustering parameters according to the characteristics of the student learning data and the search process, which can better adapt to the learning data of different levels and types of students and improve the accuracy and effectiveness of the learning status evaluation. Brief Description of the Drawings
[0054] Figure 1 It is a schematic flowchart of the student learning status evaluation method based on artificial intelligence provided by the present invention;
[0055] Figure 2 It is a schematic diagram of the student learning status evaluation system based on artificial intelligence provided by the present invention;
[0056] Figure 3 It is a schematic flowchart of step S3;
[0057] Figure 4 It is a schematic flowchart of step S4.
[0058] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention. Detailed embodiments
[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0060] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. indicating the orientation or positional relationship are based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.
[0061] Embodiment 1. Refer to Figure 1 , the method for evaluating the learning state of students based on artificial intelligence provided by the present invention includes the following steps:
[0062] Step S1: Data collection, collecting historical student learning data and real-time student learning data;
[0063] Step S2: Data preprocessing, performing data cleaning, data conversion, data normalization and constructing a learning data set on the collected student learning data;
[0064] Step S3: Constructing a clustering model, calculating the neighborhood density based on the reverse neighborhood set, dividing the central data and the edge data, combining the neighborhood density and the relative distance to obtain the concentration degree, selecting the central data with the largest concentration degree as the centroid and constructing clusters, continuously updating the clusters by constructing the central neighborhood set to complete the distribution of the central data, constructing the edge neighborhood set, and obtaining the near-neighbor screening set by constructing a weighted undirected graph and calculating the cluster density to complete the distribution of the edge data, thereby obtaining the clustering model;
[0065] Step S4: Tuning the clustering parameters, establishing a parameter search space for the clustering parameters, using the initial individual position as the representative of the clustering parameters, calculating the random offset probability, selecting the individual position for offset operation based on the offset threshold, and updating the individual position using different update strategies at different search stages based on the global optimal position and the adjustment factor to find the optimal clustering parameters;
[0066] Step S5: Evaluation of students' learning status. Input the students' learning data into the clustering model with optimized parameters for processing to obtain the learning status of real-time students' learning data, thus completing the evaluation of students' learning status.
[0067] Embodiment 2. Refer to Figure 1 , this embodiment is based on the above embodiment. In step S1, both the historical students' learning data and the real-time students' learning data include learning platform data, homework data, exam data, and classroom interaction data. The historical students' learning data also includes the learning status; the learning platform data includes the time and frequency of students' logging in to the learning platform, the staying duration on each course page, and the progress of watching teaching videos; the homework data includes the scores of students' submitted homework, the duration of completing homework, and the number of times of submitting homework; the exam data includes the scores of the exam and the time used for the exam; the classroom interaction data includes the number of questions asked by students in class, the correct rate of answering questions, and the activity of participating in group discussions; the learning status includes excellent, good, average, and poor. The learning status is used as a data label, which is not considered during clustering and is only used when selecting the cluster label.
[0068] Embodiment 3. Refer to Figure 1 , this embodiment is based on the above embodiment. In step S2, data cleaning includes dealing with noise, outliers, duplicate values, and missing values; data transformation is to convert the data into vector form; data normalization is to unify the data range based on the maximum-minimum normalization method; constructing the learning dataset is to construct the learning dataset based on the students' learning data after data cleaning, data transformation, and data normalization.
[0069] Embodiment 4. Refer to Figure 1 and Figure 3 , this embodiment is based on the above embodiment. In step S3, constructing the clustering model specifically includes the following steps:
[0070] Step S31: Calculate the neighborhood density. Calculate the Euclidean distance between any two students' learning data in the learning dataset to obtain the K nearest neighbors of each student's learning data, and construct the forward neighborhood set , then construct the reverse neighborhood set . Calculate the neighborhood density of each student's learning data based on the reverse neighborhood set. The formula used is as follows:
[0071] ;
[0072] In the formula, x i and x j are respectively the i-th and j-th students' learning data in the learning dataset. i and j are the indexes of students' learning data. ρ i is the neighborhood density of x i , is xi The Euclidean distance between x j and and are the forward neighborhood set and the reverse neighborhood set of x i respectively, is the forward neighborhood set of x j , is the Euclidean distance between x i and its K-th nearest neighbor, exp(·) is the exponential function, and α is the neighborhood influence factor in the range of (0, 1);
[0073] Step S32: Data partitioning. If the number N1 of middle school students' learning data i is greater than or equal to βK, then x i is partitioned into central data; otherwise, x i is partitioned into marginal data; and the central data set and the marginal data set are constructed respectively; where β is the adjustment factor in the range of (0, 1);
[0074] Step S33: Calculate the concentration degree. Take the product of the neighborhood density and the relative distance of each student's learning data as the concentration degree ; where δ i and A i are the relative distance and the concentration degree of x i respectively, ρ j is the neighborhood density of x j , and ρ max is the maximum neighborhood density among all students' learning data;
[0075] Step S34: Central data allocation, including the following steps:
[0076] Step S341: Select the central data with the maximum concentration degree in Q1 as the centroid Z v , and construct the corresponding cluster C v based on Z v , and delete Z v from Q1; where Z v is the v-th centroid;
[0077] Step S342: For each central data in C v , construct the central neighborhood set respectively, and the formula used is as follows:
[0078] ;
[0079] In the formula, is C vThe w-th central data in is the central neighborhood set of and are respectively the forward neighborhood set and the reverse neighborhood set of is the intersection operator, x g is the g-th central data in Q1, ρ g is x g 's neighborhood density, γ is the central control factor within the range of (0, 1);
[0080] Step S343: For each central data in , add it to C v and delete it from Q1;
[0081] Step S344: Return to Step S342 for iteration until the central data in C v no longer increases, then the cluster C v is initially constructed, and go to Step S345;
[0082] Step S345: Return to Step S341 for iteration until the number of central data in Q1 is 0, then the central data allocation is completed, obtaining V1 centroids and clusters, and go to Step S35;
[0083] Step S35: Edge data allocation. For each edge data in Q2, perform the following steps in sequence:
[0084] Step S351: Construct the edge neighborhood set of the edge data x m , and the formula used is as follows:
[0085] ;
[0086] In the formula, x m is the m-th edge data in Q2, P1 m is the edge neighborhood set of x m , and are respectively m the forward neighborhood set and the reverse neighborhood set of x is the set of all student learning data in the learning dataset except Q2, x h is the h-th central data in h is x h 's neighborhood density, is the edge control factor within the range of (0, 1);
[0087] Step S352: For P1 mEach student learning data in , calculate respectively the Euclidean distance between and V1 centroids, and select the cluster constructed by the centroid with the smallest Euclidean distance as the corresponding cluster , take each student learning data in as vertices, and the Euclidean distance between two student learning data as the weight of the edge between the corresponding two vertices to construct a weighted undirected graph , obtain based on Prim's algorithm the minimum spanning tree , then calculate the cluster density of ; if x m and the Euclidean distance between is less than , then construct a nearest neighbor screening set P2 based on , and the formula used is as follows: m ;
[0088] ;
[0089] In the formula, is the s-th student learning data in P1 m , is the corresponding cluster, is the weighted undirected graph constructed based on , is the minimum spanning tree of is the cluster density of and are respectively the total weight and the number of edges of, P2 m is the nearest neighbor screening set of x m ;
[0090] Step S353: Find the student learning data with the smallest Euclidean distance between in P2 m and x m , calculate respectively the Euclidean distance between and V1 centroids, add x to the cluster constructed by the centroid with the smallest Euclidean distance, and delete x m from Q2; return to step S351, continue to process the next marginal data until the number of marginal data in Q2 is 0, then the marginal data allocation is completed; obtain the clustering model and output the clustering result; where m ; is the f-th student learning data in P2 m .
[0091] By performing the above operations, for the problem in the existing student learning status evaluation method that the student learning data has a complex non-linear structure, the data associations cannot be effectively mined, and the evaluation focus cannot be determined, resulting in inaccurate evaluation of the student learning status, this solution calculates the neighborhood density based on the reverse neighborhood set, divides the central data and the edge data, finds the tightness relationship around the data points in the complex data structure, and better captures the complex structure of the student learning data; combines the neighborhood density and the relative distance to obtain the concentration degree, selects the central data with the largest concentration degree as the centroid and constructs clusters, comprehensively considering the aggregation degree of the student learning data in the local area and the relative position relationship with other data, and more accurately determines the focus of the student learning status evaluation; continuously updates the clusters by constructing the central neighborhood set to complete the allocation of the central data and ensure the construction quality of the clusters; constructs the edge neighborhood set, obtains the near-neighbor screening set by constructing a weighted undirected graph and calculating the cluster density, completes the allocation of the edge data, obtains the clustering model, more comprehensively evaluates the student learning status, improves the coverage and efficiency of the evaluation, and conducts a more accurate evaluation of the student learning status that conforms to the actual situation.
[0092] Example 5, refer to Figure 1 and Figure 4 , based on the above example, in step S4, the clustering parameter tuning specifically includes the following steps:
[0093] Step S41: Initial individual positions. Establish a parameter search space for the neighborhood influence factor α, adjustment factor β, central control factor γ, and edge control factor ε in the clustering. Randomly initialize U individual positions within the parameter search space. Use the individual positions as representatives of the clustering parameters, and take the average of the silhouette coefficients of all student learning data in the clustering results obtained based on the individual positions as the fitness value of the individual;
[0094] Step S42: Individual position offset. Preset the offset threshold B th , calculate the random offset probability of each individual position , if , then perform an offset operation on the u-th individual position; otherwise, the u-th individual position remains unchanged; the formula used is as follows:
[0095] ;
[0096] ;
[0097] In the formula, is the u-th individual position at the t-th search, is 's random offset probability, u is the individual position index, t is the search number index, and T is the maximum search number. and are two individual positions randomly selected during the t-th search, , is an offset control factor within the range of (0, 0.1], and min(·) is a function to take the minimum value, is the offset individual position, a max and a min are respectively the maximum and minimum values of the scaling factor;
[0098] Step S43: Update the individual position. Different update strategies are adopted to update the individual position at different search stages; the used formula is as follows:
[0099] ;
[0100] In the formula, is the u-th individual position during the (t + 1)-th search, is the global optimal position during the t-th search. The global optimal position is the individual position with the largest fitness value, is the average value of all individual positions during the t-th search, c1, c2, c3, and c4 are random numbers within the range of (0, 1) that do not interfere with each other, UB and LB are respectively the upper and lower limits of the parameter search space, is an adjustment factor within the range of (0.1, 0.5);
[0101] Step S44: Determine the optimal clustering parameter. Preset a fitness threshold, update the fitness value of the individual position. When there is a global optimal position whose fitness value is higher than the fitness threshold, then the global optimal position corresponds to the optimal clustering parameter, and the clustering parameter tuning is completed; otherwise, if the maximum search number is reached, return to Step S41 to re-initialize the individual position; otherwise, increment the search number by 1 and return to Step S42 to continue the search.
[0102] By performing the above operations, in view of the problem that there are significant differences among individual students in the existing methods for evaluating students' learning status, the students' learning data presents complex and diverse characteristics, and the evaluation is difficult to meet the needs of different students, resulting in the inability to accurately evaluate the actual learning status of students, this solution establishes a parameter search space for the clustering parameters. The initial individual position serves as the representative of the clustering parameters, which can explore various possible combinations of clustering parameters and provides more possibilities for adapting to the characteristics of different students' learning data; calculates the random offset probability, selects the individual position for offset operation based on the offset threshold to avoid the clustering parameters falling into local optima, and helps to find more suitable clustering parameters for the characteristics of different students' learning data; based on the global optimal position and the adjustment factor, adopts different update strategies to update the individual position at different search stages to find the optimal clustering parameters, and more reasonably adjusts the clustering parameters according to the characteristics of students' learning data and the search process, which can better adapt to the learning data of different levels and types of students and improve the accuracy and effectiveness of learning status evaluation.
[0103] Embodiment Six. Refer to Figure 1 , based on the above embodiment, in step S5, the evaluation of the students' learning status is to jointly input the historical students' learning data and the real-time students' learning data after data preprocessing into the clustering model constructed based on the optimal clustering parameters for clustering processing. Based on the output clustering result, select the label with the largest number of historical students' learning data in the cluster as the cluster label, and use the cluster label to which the real-time students' learning data belongs as the evaluation result to obtain the learning status of the real-time students' learning data and complete the evaluation of the students' learning status.
[0104] Embodiment Seven. Refer to Figure 2 , based on the above embodiment, the student learning status evaluation system based on artificial intelligence provided by the present invention includes a data acquisition module, a data preprocessing module, a clustering model construction module, a clustering parameter tuning module, and a student learning status evaluation module;
[0105] The data acquisition module acquires historical students' learning data and real-time students' learning data and sends the data to the data preprocessing module;
[0106] The data preprocessing module performs data cleaning, data conversion, data normalization, and learning dataset construction processing on the acquired students' learning data and sends the data to the clustering model construction module;
[0107] The clustering model construction module calculates the neighborhood density based on the reverse neighborhood set, divides the central data and the edge data, combines the neighborhood density and the relative distance to obtain the concentration degree, selects the central data with the maximum concentration degree as the centroid and constructs clusters, continuously updates the clusters by constructing the central neighborhood set, completes the allocation of the central data, constructs the edge neighborhood set, and obtains the near-neighbor screening set by constructing a weighted undirected graph and calculating the cluster density, completes the allocation of the edge data, obtains the clustering model, and sends the data to the clustering parameter tuning module;
[0108] The clustering parameter tuning module establishes a parameter search space for the clustering parameters. The initial individual position is used as the representative of the clustering parameters, calculates the random offset probability, selects the individual position for offset operation based on the offset threshold, and updates the individual position using different update strategies at different search stages based on the global optimal position and the adjustment factor to find the optimal clustering parameters, and sends the data to the student learning status evaluation module;
[0109] The student learning status evaluation module inputs the student learning data into the clustering model after parameter tuning for processing, obtains the learning status of the real-time student learning data, and completes the evaluation of the student learning status.
[0110] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0111] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention.
[0112] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. All in all, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative efforts without departing from the purpose of the present invention, they should all fall within the protection scope of the present invention.
Claims
1. An artificial intelligence-based method for evaluating students' learning status, characterized in that: The method includes the following steps: Step S1: Data collection, collecting historical student learning data and real-time student learning data; Step S2: Data preprocessing; Step S3: Constructing a clustering model; Step S4: Tuning clustering parameters; Step S5: Evaluating the student learning state, inputting the student learning data into the clustering model with tuned parameters for processing, obtaining the learning state of the real-time student learning data, and completing the evaluation of the student learning state; In step S1, the data collection is to collect historical student learning data and real-time student learning data; both the historical student learning data and the real-time student learning data include learning platform data, homework data, exam data, and classroom interaction data, and the historical student learning data also includes the learning state; taking the learning state as a data label, the data label is not considered during clustering and is only used when selecting the cluster label; In step S3, the constructing of the clustering model specifically includes the following steps: Step S31: Calculate neighborhood density. Calculate the Euclidean distance between any two students' learning data in the learning dataset to obtain the K nearest neighbors of each student's learning data, and construct a forward neighborhood set , and then construct a reverse neighborhood set . Calculate the neighborhood density of each student's learning data based on the reverse neighborhood set. The formula used is as follows: ; where x i and x j are the learning data of the i-th and j-th students in the learning dataset respectively, i and j are the indices of the students' learning data, ρ i is the neighborhood density of x i , is the Euclidean distance between x i and x j , and are the forward neighborhood set and the reverse neighborhood set of x i respectively, is the forward neighborhood set of x j , is the Euclidean distance between x i and its K-th nearest neighbor, exp(·) is the exponential function, and α is the neighborhood influence factor; Step S32: Data partitioning. If the quantity N1 of the learning data of middle school students i is greater than or equal to βK, then partition x i as central data; otherwise, partition x i as marginal data; and construct a central data set and a marginal data set respectively; where β is an adjustment factor. Step S33: Calculate the concentration degree, and take the product of the neighborhood density and the relative distance of each student's learning data as the concentration degree ; where, the product of the neighborhood density and the relative distance of each student's learning data is taken as the concentration degree ; among them, δ i and A i are the relative distance and the concentration degree of x i respectively; Step S34: Central data allocation; Step S35: Marginal data allocation; In step S34, the central data allocation specifically includes the following steps: Step S341: Select the central data with the highest concentration in Q1 as the centroid Z v , based on Z v Construct the corresponding cluster C v , and remove Z v from Q1; where Z v is the v-th centroid; Step S342: For each central data in C v , construct a central neighborhood set respectively, and the formula used is as follows: ; wherein, is the w-th central data in C v , is the central neighborhood set of , and are respectively the forward neighborhood set and the backward neighborhood set of , is the intersection operator, x g is the g-th central data in Q1, ρ g is the neighborhood density of x g , γ is the central control factor, ρ max is the maximum neighborhood density among all students' learning data; Step S343: For each center data in , add it to C v and delete it from Q1; Step S344: Return to step S342 for iteration until the central data in C v no longer increases, then the cluster C v is initially constructed, and proceed to step S345; Step S345: Return to step S341 for iteration until the number of central data in Q1 is 0, then the central data allocation is completed, obtaining V1 centroids and clusters, and proceeding to step S35; In step S35, the marginal data allocation is to perform the following steps for each marginal data in Q2 in sequence: Step S351: Construct the edge data x m 's edge neighborhood set, and the formula used is as follows: ; where x m is the m-th edge data in Q2, and P1 m is the edge neighborhood set of x m . and are the forward neighborhood set and the backward neighborhood set of x m respectively. is the set of all student learning data in the learning dataset except Q2, and x h is the h-th central data in . ρ h is the neighborhood density of x h , and is the edge control factor. Step S352: For P1 m the s-th student learning data in , calculate respectively the Euclidean distances between it and the V1 centroids, and select the cluster constructed by the centroid with the smallest Euclidean distance as the corresponding cluster . Take each student learning data in as vertices, and the Euclidean distance between two student learning data as the weight of the edge corresponding to the two vertices, and construct a weighted undirected graph . Based on Prim's algorithm, obtain the minimum spanning tree of . Then calculate the cluster density of ; if the Euclidean distance between x m and is less than , then construct the nearest neighbor screening set P2 based on m . The formula used is as follows: ; In the formula, is the learning data of the s-th student in P1 m , is the corresponding cluster, is the weighted undirected graph constructed based on , is the minimum spanning tree, is the cluster density, and are respectively the total weight value and the number of edges of P2 m is the nearest neighbor screening set of x m ; Step S353: Find P2 m among the student learning data with the smallest Euclidean distance to x m ; calculate the Euclidean distances between each of them and the V1 centroids respectively, add x to the cluster constructed by the centroid with the smallest Euclidean distance, and delete x m from Q2; return to step S351 to continue processing the next marginal data until the number of marginal data in Q2 is 0, then the allocation of marginal data is completed; obtain the clustering model and output the clustering result; where m is the f-th student learning data in P2 ; m In step S4, the tuning of the clustering parameters specifically includes the following steps: Step S41: Initial individual positions, establishing a parameter search space for the neighborhood influence factor α, adjustment factor β, central control factor γ, and marginal control factor ε in the clustering, randomly initializing U individual positions within the parameter search space, using the individual positions as representatives of the clustering parameters, and taking the average of the silhouette coefficients of all student learning data in the clustering result obtained based on the individual positions as the fitness value of the individual; Step S42: Individual position offset, and preset an offset threshold B th , calculate the random offset probability of each individual position , if , then perform an offset operation on the u-th individual position; otherwise, the u-th individual position remains unchanged; the formula used is as follows: ; ; Wherein, is the position of the u-th individual at the t-th search, is random offset probability, u is the individual position index, t is the search number index, T is the maximum search number, and are two randomly selected individual positions at the t-th search, is the offset control factor, min(·) is the minimum value function, is the individual position after offset, a max and a min are the maximum and minimum values of the scaling factor respectively; Step S43: Updating individual positions, using different update strategies to update the individual positions at different search stages; the formulas used are as follows: ; In the formula, is the position of the u-th individual at the (t + 1)-th search, is the global optimal position at the t-th search, is the average value of all individual positions at the t-th search. c1, c2, c3, and c4 are non-interfering random numbers. UB and LB are the upper and lower limits of the parameter search space respectively, is the adjustment factor; Step S44: Determining the optimal clustering parameters, presetting a fitness threshold, updating the fitness values of the individual positions, when there is a fitness value of the global optimal position higher than the fitness threshold, then the global optimal position corresponds to the optimal clustering parameters, and the clustering parameter tuning is completed; otherwise, if the maximum search number is reached, return to step S41 to re-initialize the individual positions; otherwise, increment the search number by 1 and return to step S42 to continue the search; In step S5, the evaluation of the student learning state is to jointly input the preprocessed historical student learning data and real-time student learning data into the clustering model constructed based on the optimal clustering parameters for clustering processing, based on the output clustering result, selecting the label with the largest number of historical student learning data in the cluster as the cluster label, and taking the cluster label to which the real-time student learning data belongs as the evaluation result, obtaining the learning state of the real-time student learning data, and completing the evaluation of the student learning state.
2. The method for evaluating the learning status of students based on artificial intelligence according to claim 1, characterized in that: In step S2, the data preprocessing is to perform data cleaning, data transformation, data normalization, and construction of a learning dataset on the collected student learning data; data cleaning includes handling noise, outliers, duplicate values, and missing values; Data transformation is to convert the data into vector form; data normalization is to unify the data range based on the maximum - minimum normalization method; Constructing the learning dataset is to construct the learning dataset based on the student learning data after data cleaning, data transformation, and data normalization.
3. An artificial intelligence-based student learning status evaluation system for implementing the artificial intelligence-based student learning status evaluation method according to any one of claims 1-2, characterized in that: It includes a data acquisition module, a data preprocessing module, a clustering model construction module, a clustering parameter tuning module, and a student learning status evaluation module; The data acquisition module acquires historical student learning data and real - time student learning data and sends the data to the data preprocessing module; The data preprocessing module performs data cleaning, data transformation, data normalization, and construction of a learning dataset on the collected student learning data and sends the data to the clustering model construction module; The clustering model construction module calculates the neighborhood density based on the reverse neighborhood set, divides the central data and the edge data, combines the neighborhood density and the relative distance to obtain the concentration degree, selects the central data with the maximum concentration degree as the centroid and constructs clusters, continuously updates the clusters by constructing the central neighborhood set to complete the allocation of central data, constructs the edge neighborhood set, and obtains the near - neighbor screening set by constructing a weighted undirected graph and calculating the cluster density to complete the allocation of edge data, obtains the clustering model, and sends the data to the clustering parameter tuning module; The clustering parameter tuning module establishes a parameter search space for the clustering parameters, uses the initial individual position as the representative of the clustering parameters, calculates the random offset probability, selects the individual position for offset operation based on the offset threshold, and updates the individual position using different update strategies at different search stages based on the global optimal position and the adjustment factor to find the optimal clustering parameters and sends the data to the student learning status evaluation module; The student learning status evaluation module inputs the student learning data into the clustering model after parameter tuning for processing, obtains the learning status of the real - time student learning data, and completes the evaluation of the student learning status.
Citation Information
Patent Citations
Minimum spanning tree clustering algorithm and system based on density core
CN112364887A
Electric power construction safety early warning method based on heterogeneous data
CN118469303A
Abnormal behavior evaluation method and system based on data driving
CN119442122A