Learning effect analysis method for quantitatively judging clustering number and improving K-Means model

The number of clusters of the K-means model is quantitatively determined by the elbow algorithm and derivative analysis. Combined with the covariance matrix and principal component analysis, the problem of inaccurate specification of the number of clusters of the K-means algorithm is solved, the stability and accuracy of the clustering results are achieved, and the targeted nature of education and teaching is improved.

CN120744562APending Publication Date: 2025-10-03CHONGQING VOCATIONAL COLLEGE OF LIGHT IND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510908417.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The existing K-means algorithm requires manual specification of the number of clusters K, which leads to unstable and inaccurate clustering results. In practical applications, the inflection point of the elbow algorithm is not obvious or there are multiple candidate inflection points, which affects the reliability of the clustering results.

Method used

The elbow algorithm is used to calculate the total sum of squared errors and draw the SSE curve to quantitatively determine the inflection point. If the inflection point is obvious, it is directly determined. Otherwise, the first-order and second-order derivatives are calculated to find the optimal number of clusters K. Cluster analysis is performed in combination with the improved K-Means model, and the covariance matrix and principal component analysis are used to perform dimensionality reduction and visualize the clustering results.

Benefits of technology

The stability and accuracy of clustering results have been improved, and the characteristics of different student groups can be clearly identified, guiding teachers to propose targeted education and teaching improvement strategies to enhance teaching effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744562A_ABST
    Figure CN120744562A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of clustering analysis algorithms, in particular to a learning effect analysis method for quantitatively judging a clustering number to improve a K-Means model, and aims at finding out the optimal clustering number K of the K-Means model by quantitatively determining an inflection point in combination with an elbow algorithm, reducing the subjectivity of inflection point determination and improving the stability and accuracy of a clustering result. The method provided by the invention has a better classification effect, can recognize the characteristics of different student groups more clearly, guides teachers to propose targeted education and teaching improvement strategies according to different student groups, and improves the education and teaching level of the teachers. Therefore, the problem that an existing learning effect analysis method is low in accuracy is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cluster analysis algorithms, and in particular to a learning effectiveness analysis method for quantitatively judging the number of clusters and improving a K-Means model. Background Art

[0002] In the prior art, the K-means algorithm is a commonly used clustering algorithm that can be used in fields such as education and teaching to perform cluster analysis on student groups, so as to identify the characteristics of different student groups and guide teaching improvements.

[0003] However, the traditional K-means algorithm has some shortcomings: it requires manual specification of the number of clusters, K, which is typically tried from smallest to largest, and iteratively searches for new cluster centers. This also doesn't guarantee that the final clustering result will be the overall optimal solution, impacting the stability and accuracy of the clustering results. While the elbow algorithm can assist in finding the optimal number of clusters, K, in practice, the inflection point may be unclear or there may be multiple candidate inflection points. This requires subjective judgment and selection, which introduces a degree of subjectivity and reduces the reliability of the clustering results. Summary of the Invention

[0004] The purpose of the present invention is to provide a learning effectiveness analysis method for quantitatively judging the number of clusters and improving the K-Means model, aiming to solve the problem of low accuracy of existing learning effectiveness analysis methods.

[0005] To achieve the above object, the present invention provides a learning effectiveness analysis method for quantitatively judging the number of clusters and improving the K-Means model, comprising the following steps:

[0006] Collect data on multiple factors that affect students' learning outcomes to form a data set;

[0007] The elbow algorithm is used to calculate the total sum of squared errors corresponding to different cluster numbers K, and the SSE curve is plotted. If there is an inflection point in the SSE curve where the rate of decline slows down significantly, it is directly used as the optimal cluster number K. If the inflection point is not obvious or there are multiple candidate inflection points, the first-order and second-order derivatives of the SSE corresponding to the number of clusters are calculated, and the cluster number K corresponding to the maximum value of the second-order derivative is found as the optimal cluster number K.

[0008] According to the optimal number of clusters K, the improved K-Means model is used to perform cluster analysis on the data set to obtain clustering results of different student groups;

[0009] Based on the clustering results, cluster labels and indicators are established, and teachers are guided to propose targeted education and teaching improvement strategies according to different student groups.

[0010] In "Calculating the SSE for Different Numbers of Clusters K Using the Elbow Algorithm and Plotting the SSE Curve," the formula for calculating the SSE for different numbers of clusters K using the elbow algorithm is:

[0011]

[0012] Where k is the number of clusters, Ci is all the data samples in the i-th cluster, xj is the j-th data sample, μ i is the center point of the i-th cluster, ||x j -μ i || 2 Represents the sample xj to its cluster center μ i The square of the Euclidean distance.

[0013] Among them, in "If there is an inflection point in the SSE curve where the rate of decline slows down significantly, it is directly used as the optimal number of clusters K; if the inflection point is not obvious or there are multiple candidate inflection points, the first-order derivative and second-order derivative of the SSE corresponding to the number of clusters are calculated, and the number of clusters K corresponding to the maximum value of the second-order derivative is found as the optimal number of clusters K", the formula for calculating the first-order derivative of the SSE corresponding to the number of clusters is:

[0014]

[0015] Where (K)' represents the decrease in SSE when the number of clusters increases from K to K+1 (usually a negative value, the larger the absolute value, the faster the decrease); SSE(K+1)-SSE(K) is the difference in SSE corresponding to the number of adjacent clusters K.

[0016] Among them, in "If there is an inflection point in the SSE curve where the rate of decline slows down significantly, it is directly used as the optimal number of clusters K; if the inflection point is not obvious or there are multiple candidate inflection points, the first-order derivative and second-order derivative of the SSE corresponding to the number of clusters are calculated, and the number of clusters K corresponding to the maximum value of the second-order derivative is found as the optimal number of clusters K", the formula for calculating the second-order derivative of the SSE corresponding to the number of clusters is:

[0017] (K)”=SSE(K+1)-2×SSE(K)+SSE(K-1)=(K)’-(K-1)’

[0018] Where SSE(K) is the SSE value corresponding to the number of clusters K, SSE(K-1) is the SSE value corresponding to the number of clusters K-1, and SSE(K+1) is the SSE value corresponding to the number of clusters K+1.

[0019] The following steps are included in "Based on the optimal number of clusters K, using the improved K-Means model to perform cluster analysis on the data set and obtain clustering results for different student groups":

[0020] Calculate the covariance matrix of the original data and analyze the correlation between each factor;

[0021] The original data is projected into the two-dimensional space composed of these two principal components to achieve dimensionality reduction, and a scatter plot of the clustering results is drawn based on the data after dimensionality reduction to identify the characteristics of different student groups.

[0022] The present invention provides a learning effectiveness analysis method for improving the K-Means model by quantitatively determining the number of clusters. This method, combined with the elbow algorithm, quantitatively determines the inflection point to find the optimal number of clusters, K, for the K-means model. This reduces the subjectivity of inflection point determination and improves the stability and accuracy of clustering results. The proposed method achieves better classification results, can more clearly identify the characteristics of different student groups, and can guide teachers to propose targeted educational and teaching improvement strategies based on different student groups, thereby improving teachers' educational and teaching capabilities. This solves the problem of low accuracy in existing learning effectiveness analysis methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0024] Figure 1 and Figure 2 It is the elbow rule that determines the number of clusters.

[0025] Figure 3 is the visualization clustering result.

[0026] Figure 4 This is a flowchart of a learning effectiveness analysis method for quantitatively judging the number of clusters and improving the K-Means model provided by the present invention.

[0027] Figure 5 This is a flowchart that uses the improved K-Means model to perform cluster analysis on the data set based on the optimal number of clusters K and obtain the clustering results of different student groups. DETAILED DESCRIPTION

[0028] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0029] See also Figures 1 to 5 The present invention provides a learning effectiveness analysis method for quantitatively judging the number of clusters and improving the K-Means model, comprising the following steps:

[0030] S1 collects data on multiple factors that affect students’ learning outcomes to form a data set;

[0031] Specifically, factors influencing student learning outcomes are statistically analyzed to form a data set. Each factor needs to be presented in the form of data (with scores for each factor), such as, but not limited to, "student psychological status," "learning motivation," "attendance," "science performance," "liberal arts performance," "physical health," and "sleep duration."

[0032] S2 calculates the total sum of squared errors corresponding to different cluster numbers K using the elbow algorithm and draws the SSE curve;

[0033] The formula for calculating the total sum of squares of errors corresponding to different cluster numbers K using the elbow algorithm is:

[0034]

[0035] Where k is the number of clusters, Ci is all the data samples in the i-th cluster, xj is the j-th data sample, μ i is the center point of the i-th cluster, ||x j -μ i || 2 Represents the sample xj to its cluster center μ i The square of the Euclidean distance.

[0036] Specifically, the sum of squared errors (SSE) of each cluster number is calculated using the elbow algorithm.

[0037] S3, if there is an inflection point in the SSE curve where the rate of decline slows down significantly, it is directly used as the optimal number of clusters K; if the inflection point is not obvious or there are multiple candidate inflection points, the first-order derivative and second-order derivative of the SSE corresponding to the number of clusters are calculated, and the number of clusters K corresponding to the maximum value of the second-order derivative is found as the optimal number of clusters K;

[0038] The formula for calculating the first-order derivative of SSE corresponding to the number of clusters is:

[0039]

[0040] Where (K)' represents the decrease in SSE when the number of clusters increases from K to K+1 (usually a negative value, the larger the absolute value, the faster the decrease); SSE(K+1)-SSE(K) is the difference in SSE corresponding to the number of adjacent clusters K.

[0041] The formula for calculating the second-order derivative of SSE corresponding to the number of clusters is:

[0042] (K)”=SSE(K+1)-2×SSE(K)+SSE(K-1)=(K)’-(K-1)’

[0043] Where SSE(K) is the SSE value corresponding to the number of clusters K, SSE(K-1) is the SSE value corresponding to the number of clusters K-1, and SSE(K+1) is the SSE value corresponding to the number of clusters K+1.

[0044] Specifically, the elbow algorithm is used to draw the SSE curve to find the inflection point. The SSE curve will show two situations. One is that there is a point in the image where the SSE decrease rate slows down significantly, forming a shape similar to an "elbow". This point is the "inflection point". The inflection point can be directly used as the optimal number of clusters K, such as Figure 1 As shown in Figure 2, when the number of clusters K = 4, the total square error SSE decreases significantly slower, that is, K = 4 is the "inflection point". The second reason is that there are multiple points in the SSE image where the decrease rate slows down significantly or the curve is smooth and the "inflection point" cannot be directly determined, such as Figure 2 As shown in Figure 2, the optimal number of clusters K needs to be quantitatively determined.

[0045] Calculate the first and second derivatives of the SSE corresponding to the number of clusters. As the value of K increases, the SSE decreases monotonically, and the rate of decline gradually slows. The number of clusters represents the turning point from a "rapid decline" to a "slow decline." Mathematically, the inflection point is the point where the slope of the SSE curve changes at its maximum rate of change. This means that the point where the second derivative reaches its maximum value represents the optimal number of clusters, K.

[0046] The (K)' and (K)" corresponding to the number of clusters K are:

[0047]

[0048] When K=3, the second-order derivative is the largest, that is, when K=3, the rate of decrease of the SSE curve slows down significantly, so the optimal number of clusters K=3.

[0049] S4 uses the improved K-Means model to perform cluster analysis on the data set based on the optimal number of clusters K, and obtains clustering results for different student groups;

[0050] S41 calculates the covariance matrix of the original data and analyzes the correlation between each factor;

[0051] Specifically, we collected and organized a dataset containing multidimensional factors, presenting each factor as data. The raw data was then standardized to eliminate dimensional differences and ensure that all factors were equally important in the cluster analysis. We calculated the covariance matrix of the raw data, analyzed the correlations between factors, and performed eigenvalue decomposition on the covariance matrix, extracting the first two mutually orthogonal, uncorrelated principal components to maximize the preservation of the raw data variance.

[0052] S42 projects the original data into a two-dimensional space composed of these two principal components to achieve dimensionality reduction, and draws a scatter plot of the clustering results based on the data after dimensionality reduction to identify the characteristics of different student groups.

[0053] Specifically, through principal component analysis (PCA), the data set (students' psychological status, learning motivation, attendance rate, science scores, liberal arts scores, physical health, and sleep duration) is reduced to two dimensions, principal component 1 and principal component 2. These two principal components are the two linear combinations with the largest variance in the original data, and they are orthogonal to each other, that is, there is no correlation between them. After principal component analysis, the six groups of feature data samples are clustered into three label types. The clustering result scatter plot is as follows Figure 3 As shown in the figure, the scattered points of different colors represent different data clusters. There is no overlap between the clusters, which means that the selected number of clusters K can better distinguish different categories and effectively identify the characteristics of different student groups.

[0054] S5 establishes cluster labels and indicators based on clustering results, and guides teachers to propose targeted education and teaching improvement strategies based on different student groups.

[0055] Specifically, based on the clustering results, cluster labels and indicators are established to guide teachers to propose targeted education and teaching improvement strategies based on different student groups, thereby improving teachers' education and teaching level. Figure 3 The clustering results and the classification labels are constructed as follows:

[0056]

[0057] For students with high academic performance, we encourage them to maintain their learning advantages while strengthening their mental state and developing proactive learning motivation. We also help them realize the importance of other abilities besides academic performance, such as social skills, emotional management, and the ability to relieve psychological stress.

[0058] For students with balanced development, we encourage them to maintain their current good performance and further strengthen their learning ability and psychological quality. Teachers can design interdisciplinary research topics or practical projects in the teaching process to deepen their understanding of knowledge through practical operations.

[0059] For students with untapped potential, their learning potential should be stimulated while maintaining a positive mental state and learning motivation. Appropriate achievement rewards should be used to motivate them to pursue learning achievements.

[0060] Beneficial effects:

[0061] The present invention uses mathematical and intelligent algorithms to cluster and analyze the factors that affect students' learning outcomes, and classifies student groups according to their characteristics, so as to empower teachers with intelligence to propose targeted education and teaching improvement strategies based on different student groups, thereby improving teachers' education and teaching level. The present invention combines the mathematical essence of the elbow algorithm and the inflection point, and by calculating the second-order derivative of SSE corresponding to the number of clusters K, determines the inflection point in a quantitative manner, finds the optimal number of clusters K, and reduces the number of traversals to achieve the overall optimal solution. It reduces the subjectivity of using the elbow algorithm to determine the number of K-means clusters, and improves the stability and accuracy of the clustering results. The improved K-means clustering model can more clearly identify the characteristics of different student groups, and guide teachers to propose targeted education and teaching improvement strategies based on different student groups, thereby improving teachers' education and teaching level.

[0062] The above disclosure is only a preferred embodiment of the learning effectiveness analysis method of the improved K-Means model for quantitatively judging the number of clusters of the present invention. Of course, this cannot be used to limit the scope of rights of the present invention. Ordinary technicians in this field can understand that implementing all or part of the processes of the above embodiment and making equivalent changes in accordance with the claims of the present invention still fall within the scope of the invention.

Claims

1. A learning effectiveness analysis method for quantitatively judging the number of clusters and improving the K-Means model, characterized by: The following steps are involved: Collect data on multiple factors that affect students' learning outcomes to form a data set; The elbow algorithm is used to calculate the total sum of squared errors corresponding to different cluster numbers K, and the SSE curve is drawn; If there is an inflection point in the SSE curve where the rate of decline slows down significantly, it is directly used as the optimal number of clusters K. If the inflection point is not obvious or there are multiple candidate inflection points, the first-order derivative and second-order derivative of the SSE corresponding to the number of clusters are calculated, and the number of clusters K corresponding to the maximum value of the second-order derivative is found as the optimal number of clusters K. According to the optimal number of clusters K, the improved K-Means model is used to perform cluster analysis on the data set to obtain clustering results of different student groups; Based on the clustering results, cluster labels and indicators are established, and teachers are guided to propose targeted education and teaching improvement strategies according to different student groups.

2. The learning effectiveness analysis method of the improved K-Means model by quantitatively judging the number of clusters according to claim 1 is characterized in that: In "Calculating the SSE for Different Numbers of Clusters K Using the Elbow Algorithm and Plotting the SSE Curve," the formula for calculating the SSE for different numbers of clusters K using the elbow algorithm is: Where k is the number of clusters, Ci is all the data samples in the i-th cluster, xj is the j-th data sample, μ i is the center point of the i-th cluster, ||x j -μ i || 2 Represents the sample xj to its cluster center μ i The square of the Euclidean distance.

3. The learning effectiveness analysis method of the improved K-Means model by quantitatively judging the number of clusters according to claim 1, characterized in that: In the article "If there is an inflection point in the SSE curve where the rate of decline slows significantly, directly use it as the optimal number of clusters K; if the inflection point is not obvious or there are multiple candidate inflection points, calculate the first-order and second-order derivatives of the SSE corresponding to the number of clusters, and find the cluster number K corresponding to the maximum value of the second-order derivative as the optimal number of clusters K", the formula for calculating the first-order derivative of the SSE corresponding to the number of clusters is: Where (K)' represents the decrease in SSE when the number of clusters increases from K to K+1 (usually a negative value, the larger the absolute value, the faster the decrease); SSE(K+1)-SSE(K) is the difference in SSE corresponding to the number of adjacent clusters K.

4. The learning effectiveness analysis method of the improved K-Means model by quantitatively judging the number of clusters according to claim 3, characterized in that: In "If there is an inflection point in the SSE curve where the rate of decline slows significantly, directly use it as the optimal number of clusters K; if the inflection point is not obvious or there are multiple candidate inflection points, calculate the first-order and second-order derivatives of the SSE corresponding to the number of clusters, and find the number of clusters K corresponding to the maximum value of the second-order derivative as the optimal number of clusters K", the formula for calculating the second-order derivative of the SSE corresponding to the number of clusters is: (K)”=SSE(K+1)-2×SSE(K)+SSE(K-1)=(K)’-(K-1)’ Where SSE(K) is the SSE value corresponding to the number of clusters K, SSE(K-1) is the SSE value corresponding to the number of clusters K-1, and SSE(K+1) is the SSE value corresponding to the number of clusters K+1.

5. The learning effectiveness analysis method of the improved K-Means model by quantitatively judging the number of clusters according to claim 1, characterized in that: In the "According to the optimal number of clusters K, use the improved K-Means model to perform cluster analysis on the dataset and obtain clustering results for different student groups," the following steps are included: Calculate the covariance matrix of the original data and analyze the correlation between each factor; The original data is projected into the two-dimensional space composed of these two principal components to achieve dimensionality reduction, and a scatter plot of the clustering results is drawn based on the data after dimensionality reduction to identify the characteristics of different student groups.