Learning situation data analysis method based on semantic space geometric consistency verification
By constructing a semantic space and geometric consistency verification method, the problem that traditional learning analysis methods cannot understand the semantic context of student behavior is solved, realizing the interpretability and high accuracy of learning data analysis and improving the effectiveness of teaching guidance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUNAN UNIV OF SCI & TECH
- Filing Date
- 2026-06-04
- Publication Date
- 2026-07-03
Smart Images

Figure CN122334513A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of educational data mining, and in particular to a method for analyzing student learning data based on semantic space geometric consistency verification. Background Technology
[0002] Student learning analysis is a core task in the fields of smart education and learning data mining, and it is of great significance for realizing personalized teaching intervention. Traditional analysis methods mainly rely on discriminative models such as logistic regression, random forests, or graph neural networks to establish mapping relationships by mining structured data such as student clickstreams and assignment submissions. However, these traditional methods are essentially "black box" models. Although they can handle high-dimensional data, they can only capture students' quantitative behaviors and cannot deeply understand the "semantic context" behind the behaviors. For example, the same low interaction frequency may be due to students mastering the material quickly or it may be due to insufficient learning motivation. Traditional models cannot distinguish between these two very different psychological motivations, resulting in prediction results that lack interpretability from an educational perspective and are difficult to guide actual teaching. Summary of the Invention
[0003] To address the aforementioned technical problems, this invention provides a learning data analysis method based on semantic space geometric consistency verification, which features a simple algorithm and high prediction accuracy.
[0004] The technical solution of this invention to solve the above-mentioned technical problems is: a learning data analysis method based on semantic space geometric consistency verification, comprising the following steps:
[0005] Step 1: Semantically sequence the learning data to obtain natural language text;
[0006] Step 2: Construct a semantic space based on natural language text and self-regulating learning theory;
[0007] Step 3: Perform dynamic reasoning based on retrieval enhancement;
[0008] Step 4: Geometric consistency verification and semantic space self-evolution.
[0009] The specific process of step 1 in the above-mentioned learning data analysis method based on semantic space geometric consistency verification is as follows:
[0010] First, given a containing Learning log dataset of samples , , for The first in One sample, for The first in Each sample contains static demographic characteristics, interaction sequences, and historical scores.
[0011] Then, a mapping mechanism is established from heterogeneous tabular data to natural language text, decoupling the original data and generating textual archives using predefined templates.
[0012] The above-mentioned learning data analysis method based on semantic space geometric consistency verification, in step 1, the specific process of decoupling the original data and generating textual archives using a predefined template is as follows:
[0013] First, static profile description is performed, converting discrete variables, including gender, age group, and educational background, into natural language sentences.
[0014] Then, the original clickstream is aggregated and statistically analyzed to extract key statistical indicators, including active days and total clicks, and historical performance is transformed into trend descriptions.
[0015] Finally, semantic text generation is performed, which combines natural language sentences, key statistical indicators, and trend descriptions of historical performance to form a complete natural language text, providing a semantic foundation for deep reasoning in subsequent large language model LLM.
[0016] The specific process of step 2 in the above-mentioned learning data analysis method based on semantic space geometric consistency verification is as follows:
[0017] Step 21, Construct a cue word template: Based on self-regulation learning theory, construct a cue word template that includes a structured thought chain. :
[0018] ;
[0019] in, For prompt word templates; Set instructions for the role of the large language model; Task instructions to guide the learning behavior features of the model for distribution inference; The semantic archive text of known labeled samples in the training set; An example of dynamic context introduced to enhance retrieval; This indicates a text concatenation operation;
[0020] Step 22, Generate Inference Text: Input each sample from the training set into the large language model. Generate text containing the first explanatory text The inference text of the first prediction result:
[0021] ;
[0022] in, The inference text output by the large language model; For large language models;
[0023] Step 23, Generate the first logical vector: Extract the first explanatory text using a pre-trained embedding model. Based on the semantic features, the first logical vector is generated:
[0024] ;
[0025] in, For the first logical vector, , Represents the real number field. For vector dimensions; For pre-trained embedding models;
[0026] Step 24, Semantic baseline center calculation: Based on the real labels of the training set, the generated first logical vector is divided into groups using the set... and the set that did not pass and calculate through the set and the set that did not pass The geometric center serves as the semantic reference center:
[0027] ;
[0028] ;
[0029] in, For the semantic benchmark center of the state; The semantic baseline center for the failed state; For through a set; The number of the first logical vectors contained in the set; The set that did not pass; The number of first logical vectors not included in the set.
[0030] The specific process of step 3 in the above-mentioned learning data analysis method based on semantic space geometric consistency verification is as follows:
[0031] Step 31: For the samples to be predicted in the test set ,calculate With the training set Cosine similarity of samples to retrieve similar cases;
[0032] Step 32: Construct dynamic prompt words ;
[0033] Step 33: Guide the large language model through retrieval-enhanced generation. Generate text containing a second explanation by referencing similar cases. and prediction labels The first result.
[0034] In the above-mentioned learning data analysis method based on semantic space geometric consistency verification, the dynamic prompt words constructed in step 32 are:
[0035] ;
[0036] in, These are dynamic prompt words; To take the front The most similar sample; For the training set One sample; for The corresponding historical reasoning text; for The true prediction label; The semantic archive text of known labeled samples in the training set; Let be the magnitude of the vector; These are the samples to be predicted in the test set; This is a pre-trained embedding model.
[0037] The specific process of step 4 in the above-mentioned learning data analysis method based on semantic space geometric consistency verification is as follows:
[0038] Step 41, Geometric distance calculation: Generate the logical vector of the first result. , and calculate Euclidean distance to the two semantic reference centers:
[0039] ;
[0040] ;
[0041] in, The logical vector for the first result; For pre-trained embedding models; This is the second explanatory text; For the semantic benchmark center of the state; The semantic baseline center for the failed state; It is an L2 norm; for arrive The Euclidean distance; for arrive The Euclidean distance;
[0042] Step 42, Consistency Decision Logic: Deriving Implicit Semantic Tags ,like Then implicit semantic tags For passing, otherwise implicit semantic tags are used. It failed; then the implicit semantic tag was added. With predictive labels By comparing the results, the final prediction is obtained. .
[0043] In the above-mentioned learning data analysis method based on semantic space geometric consistency verification, step 42 involves implicit semantic tags. With predictive labels The comparison method is as follows:
[0044] like If the prediction confidence is high, then the output is considered high. , This is the final prediction result; at this point, the logical vector of the first result is... Dynamically add to the corresponding set based on the classification results. Or not passed the set In China, and updated in real time. and This enables the self-evolution of the reasoning feature space of a large language model;
[0045] like This indicates that the large language model is experiencing illusions, triggering a self-correction mechanism that requires the large language model to generate again.
[0046] The beneficial effects of this invention are as follows: By introducing prompting engineering guided by self-regulating learning theory and a geometric consistency verification mechanism, this invention proposes a method for analyzing student learning data. This method differs from simply directly calling a large language model. Instead, it first uses self-regulating learning theory to design specific prompt templates, driving the large language model to infer deep semantic information such as students' self-efficacy and time management strategies from the original logs. Subsequently, the designed geometric consistency verification mechanism calculates the consistency between the semantic information and the predicted classification, automatically filtering out "illusions" and low-quality inferences generated by the large language model, guiding it to produce interpretable and highly accurate student learning data analysis results. This method utilizes the powerful semantic understanding and reasoning capabilities of the large language model while ensuring the authenticity and reliability of the generated content through a rigorous verification mechanism. Attached Figure Description
[0047] Figure 1 This is the overall flowchart of the present invention.
[0048] Figure 2The above are two student visualization radar charts for learning analysis in this embodiment of the invention, where (a) is the student visualization radar chart for learning analysis of student number #420087 and (b) is the student visualization radar chart for learning analysis of student number #350188. Detailed Implementation
[0049] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0050] like Figure 1 As shown, the learning data analysis method based on semantic space geometric consistency verification includes the following steps:
[0051] Step 1: Semantically sequence the learning data to obtain natural language text.
[0052] The specific process of step 1 is as follows:
[0053] First, given a containing Learning log dataset of samples , , for The first in One sample, for The first in Each sample contains static demographic characteristics, interaction sequences, and historical scores.
[0054] Then, a mapping mechanism is established from heterogeneous tabular data to natural language text, decoupling the original data and generating textual archives using predefined templates.
[0055] The specific process of decoupling the original data and generating textual archives using predefined templates is as follows:
[0056] First, static profile description is performed, converting discrete variables, including gender, age group, and educational background, into natural language sentences.
[0057] Then, the original clickstream is aggregated and statistically analyzed to extract key statistical indicators, including the number of active days and total clicks, and historical performance (such as clickstream and performance report) is transformed into trend descriptions (such as steady increase or fluctuating decrease).
[0058] Finally, semantic text generation is performed, which combines natural language sentences, key statistical indicators, and trend descriptions of historical performance to form a complete natural language text, providing a semantic foundation for deep reasoning in subsequent large language model LLM.
[0059] Step 2: Based on natural language text, construct a semantic space that can measure the logic of learning behavior according to self-regulating learning theory.
[0060] The specific process of step 2 is as follows:
[0061] Step 21, Construct a cue word template: Based on self-regulation learning theory, construct a cue word template that includes a structured thought chain. :
[0062] ;
[0063] in, For prompt word templates; Set instructions for the role of the large language model; Task instructions to guide the learning behavior features of the model for distribution inference; The semantic archive text of known labeled samples in the training set; An example of dynamic context introduced to enhance retrieval; This indicates a text concatenation operation;
[0064] Step 22, Generate Inference Text: Input each sample from the training set into the large language model. Generate text containing the first explanatory text The inference text of the first prediction result:
[0065] ;
[0066] in, The inference text output by the large language model; For large language models;
[0067] Step 23, Generate the first logical vector: Extract the first explanatory text using a pre-trained embedding model. Based on the semantic features, the first logical vector is generated:
[0068] ;
[0069] in, For the first logical vector, , Represents the real number field. For vector dimensions; For pre-trained embedding models;
[0070] Step 24, Semantic baseline center calculation: Based on the real labels of the training set, the generated first logical vector is divided into groups using the set... and the set that did not pass and calculate through the set and the set that did not pass The geometric center serves as the semantic reference center:
[0071] ;
[0072] ;
[0073] in, For the semantic benchmark center of the state; The semantic baseline center for the failed state; For through a set; The number of the first logical vectors contained in the set; The set that did not pass; The number of first logical vectors not included in the set.
[0074] Step 3: Perform dynamic reasoning based on retrieval enhancement.
[0075] The specific process of step 3 is as follows:
[0076] Step 31: For the samples to be predicted in the test set ,calculate With the training set Cosine similarity of samples to retrieve similar cases;
[0077] Step 32: Construct dynamic prompt words The constructed dynamic prompt words are:
[0078] ;
[0079] in, These are dynamic prompt words; To take the front The most similar sample; For the training set One sample; for The corresponding historical reasoning text; for The true prediction label; The semantic archive text of known labeled samples in the training set; Let be the magnitude of the vector; These are the samples to be predicted in the test set; For pre-trained embedding models;
[0080] Step 33: Guide the large language model through retrieval-enhanced generation. Generate text containing a second explanation by referencing similar cases. and prediction labels The first result.
[0081] Step 4: Geometric consistency verification and semantic space self-evolution.
[0082] The specific process of step 4 is as follows:
[0083] Step 41, Geometric distance calculation: Generate the logical vector of the first result. , and calculate Euclidean distance to the two semantic reference centers:
[0084] ;
[0085] ;
[0086] in, The logical vector for the first result; For pre-trained embedding models; This is the second explanatory text; For the semantic benchmark center of the state; The semantic baseline center for the failed state; It is an L2 norm; for arrive The Euclidean distance; for arrive The Euclidean distance;
[0087] Step 42, Consistency Decision Logic: Deriving Implicit Semantic Tags ,like Then implicit semantic tags For passing, otherwise implicit semantic tags are used. It failed; then the implicit semantic tag was added. With predictive labels By comparing the results, the final prediction is obtained. .
[0088] Implicit semantic tags With predictive labels The comparison method is as follows:
[0089] like If the prediction confidence is high, then the output is considered high. , This is the final prediction result; at this point, the logical vector of the first result is... Dynamically add to the corresponding set based on the classification results. Or not passed the set In China, and updated in real time. and This enables the self-evolution of the reasoning feature space of a large language model;
[0090] like This indicates that the large language model is experiencing illusions, triggering a self-correction mechanism that requires the large language model to generate again.
[0091] The following experiments and evaluations will be conducted:
[0092] Experimental Environment and Implementation: The entire experiment was conducted on the OULAD dataset, covering 7 courses (denoted as AAA to GGG) and 22 semesters. The OULAD dataset exhibits significant class imbalance (pass rates ranging from only 34% to 72%) and a long-tailed distribution, severely testing the model's ability to identify students on the margins of academic difficulty. The experimental environment consisted of Python 3.10, a 12th Gen Intel(R) Core(TM) i7-12700 CPU, and an NVIDIA GTX 4060Ti GPU. The proposed method (denoted as SE-SPP) was implemented using the PyTorch framework. The Qwen model was selected as the inference core. To investigate the impact of model size on prediction performance, experiments covered four versions with different parameter sizes: lightweight research-scale Qwen-0.5B (Qwen model with 500 million parameters), Qwen-1.5B (Qwen model with 1.5 billion parameters), and mainstream open-source deployment-scale Qwen-7B (Qwen model with 7 billion parameters) and Qwen-14B (Qwen model with 14 billion parameters). The embedding model used mxbai-embed-large-v1 for text vectorization, with an output dimension of 1024.
[0093] Comparative analysis of experimental results: In order to fully verify the performance of SE-SPP, a detailed analysis was conducted from four dimensions: scale effect and overall accuracy, model adaptability, model advancement and interpretability.
[0094] (1) Scale effect and overall accuracy:
[0095] To verify the advantages of this invention, experiments were conducted on the OULAD dataset to compare SE-SPP with Qwen models of different parameter scales and with an early learning achievement prediction model that only uses a large language model for direct prediction (denoted as LLM-EPSP). The comparison results are shown in Table 1, which shows the F1 score and accuracy under different parameter scales.
[0096] Table 1 Performance impact of SE-SPP under different parameter scales
[0097]
[0098] As shown in Table 1, the performance of the proposed SE-SPP exhibits diminishing marginal returns with increasing parameter count in large language models. Although it achieves the best performance on Qwen-14B, it only shows a slight improvement over Qwen-7B, while computational resource consumption increases exponentially. By comparing it with LLM-EPSP, it can be concluded that SE-SPP, by introducing dynamic prompt word cases and geometric consistency checks on top of directly using static prompt word templates from large models, can significantly reduce the prediction accuracy of large language models. Even when using Qwen-0.5B and Qwen-1.5B with lower parameter counts, it can still achieve high prediction accuracy.
[0099] (2) Model fit analysis:
[0100] This invention uses accuracy, AUC (area under the ROC curve, i.e., the two-dimensional area enclosed by the ROC curve and the horizontal axis), and F1 score as indicators to evaluate the classification performance of the model. Table 2 shows the performance changes of various mainstream classifiers after introducing SE-SPP semantic enhancement and feeding the enhancement results as features. The mainstream classifiers selected include: Logistic Regression, Random Forest, Extreme Gradient Boosting Tree (XGBoost), Multilayer Perceptron (MLP), and CatBoost.
[0101] The experimental results are shown in Table 2. Table 2 shows that SE-SPP brings significant performance improvements to all baseline models.
[0102] Table 2 Performance improvement of traditional baseline models by SE-SPP enhancement results.
[0103]
[0104] As can be seen from Table 2, SE-SPP+CatBoost achieved the best performance with an F1 score of 0.9739, which is about 16.89% higher than the strongest traditional baseline model (XGBoost, with an F1 score of 0.8332). This demonstrates that the method of this invention has extremely high adaptability as a general feature engineering framework.
[0105] (3) Model advancement analysis:
[0106] To verify the advancement of the method of this invention, based on the results of the model adaptability analysis above, the best-performing combination (SE-SPP+CatBoost) was selected to represent the method of this invention, and compared with the latest learning performance prediction methods. The comparison methods include:
[0107] Study-GNN: A graph neural network-based approach that focuses on mining the topological relationships between students, courses, and resources.
[0108] SAPP: A method based on sequence modeling of long short-term memory networks to capture temporal dynamics.
[0109] PBC: A Bayesian classification-based method that focuses on probabilistic modeling of behavioral features.
[0110] As shown in Table 3, SE-SPP outperforms the existing Study-GNN, SAPP and PBC in all indicators.
[0111] Table 3 Performance comparison with traditional classification methods
[0112]
[0113] Compared to Study-GNN, which relies on graph topological features, SE-SPP achieves a significantly higher F1 score. While graph neural networks (GNNs) effectively capture explicit entity relationships, they struggle to reveal the implicit psychological motivations behind student interactions, leading to poor performance in identifying complex risk patterns. SE-SPP still maintains its leading position compared to strong sequence modeling baselines such as SAPP and PBC, which utilize complex ensemble strategies and granular behavioral process data. These methods excel at capturing temporal dynamics, but SE-SPP's strength lies in its semantic abductive ability. By transforming sparse logs into features of explicit self-regulating learning theory, it can better generalize to students with atypical behavioral patterns, thus achieving the highest F1 score among all compared methods.
[0114] (4) Case analysis and interpretability verification:
[0115] To visually demonstrate the interpretability of the model, two representative students (student IDs #420087 and #350188) were selected from course AAA for visualization analysis, resulting in the following radar chart: Figure 2 As shown, in order to highlight individual differences, Figure 2 The paper introduces the average value calculated from all students as the gray baseline.
[0116] Output of student #420087:
[0117] Student #420087 belongs to the typical "high potential, low self-discipline" group. As can be seen from the radar chart (left), their score on the self-efficacy dimension is significantly higher than the average passer benchmark, stemming from their relatively good educational background and youth. However, their graph exhibits an extreme asymmetric collapse: despite having a moderate level of virtual learning environment (VLE) engagement, their time management and metacognitive levels are far below the baseline. Traditional models often mislead students with their high potential, while SE-SPP accurately points out that the root cause of their failure lies in a lack of regulatory ability, rather than a lack of competence.
[0118] Output of student #350188:
[0119] Student #350188 is a typical "older, high-engagement" student. Compared to the average baseline, this student showed a significant weakness in self-efficacy, reflecting a common lack of confidence among older learners. However, they exhibited a marked expansion in time management and metacognitive abilities, even exceeding the average passer level. This pattern reveals a compensatory learning strategy: compensating for insufficient self-efficacy through extreme time management and high-frequency interaction. SE-SPP identified this positive compensatory pattern, correctly classifying it as pass, avoiding misjudgment caused by low-efficacy characteristics.
Claims
1. A learning situation data analysis method based on semantic space geometric consistency verification, characterized in that, Includes the following steps: Step 1: Semantically sequence the learning data to obtain natural language text; Step 2: Construct a semantic space based on natural language text and self-regulating learning theory; Step 3: Perform dynamic reasoning based on retrieval enhancement; Step 4: Geometric consistency verification and semantic space self-evolution. 2.The semantic space geometry consistency checking based learning situation data analysis method according to claim 1, characterized in that, The specific process of step 1 is as follows: First, given a containing Learning log dataset of samples , , for The first in One sample, for The first in Each sample contains static demographic characteristics, interaction sequences, and historical scores. Then, a mapping mechanism is established from heterogeneous tabular data to natural language text, decoupling the original data and generating textual archives using predefined templates.
3. The learning data analysis method based on semantic space geometric consistency verification according to claim 2, characterized in that, In step 1, the specific process of decoupling the original data and generating a textual archive using a predefined template is as follows: First, static profile description is performed, converting discrete variables, including gender, age group, and educational background, into natural language sentences. Then, the original clickstream is aggregated and statistically analyzed to extract key statistical indicators, including active days and total clicks, and historical performance is transformed into trend descriptions. Finally, semantic text generation is performed, which combines natural language sentences, key statistical indicators, and trend descriptions of historical performance to form a complete natural language text, providing a semantic foundation for deep reasoning in subsequent large language model LLM.
4. The learning data analysis method based on semantic space geometric consistency verification according to claim 1, characterized in that, The specific process of step 2 is as follows: Step 21, Construct a cue word template: Based on self-regulation learning theory, construct a cue word template that includes a structured thought chain. : ; in, For prompt word templates; Set instructions for the role of the large language model; Task instructions to guide the learning behavior features of the model for distribution inference; The semantic archive text of known labeled samples in the training set; An example of dynamic context introduced to enhance retrieval; This indicates a text concatenation operation; Step 22, Generate Inference Text: Input each sample from the training set into the large language model. Generate text containing the first explanatory text The inference text of the first prediction result: ; in, The inference text output by the large language model; For large language models; Step 23, Generate the first logical vector: Extract the first explanatory text using a pre-trained embedding model. Based on the semantic features, the first logical vector is generated: ; in, For the first logical vector, , Represents the real number field. For vector dimensions; For pre-trained embedding models; Step 24, Semantic baseline center calculation: Based on the real labels of the training set, the generated first logical vector is divided into groups using the set... and the set that did not pass and calculate through the set and the set that did not pass The geometric center serves as the semantic reference center: ; ; in, For the semantic benchmark center of the state; The semantic baseline center for the failed state; For through a set; The number of the first logical vectors contained in the set; The set that did not pass; The number of first logical vectors not included in the set.
5. The learning data analysis method based on semantic space geometric consistency verification according to claim 1, characterized in that, The specific process of step 3 is as follows: Step 31: For the samples to be predicted in the test set ,calculate With the training set Cosine similarity of samples to retrieve similar cases; Step 32: Construct dynamic prompt words ; Step 33: Guide the large language model through retrieval-enhanced generation. Generate text containing a second explanation by referencing similar cases. and prediction labels The first result.
6. The learning data analysis method based on semantic space geometric consistency verification according to claim 5, characterized in that, In step 32, the constructed dynamic prompt words are: ; in, These are dynamic prompt words; To take the front The most similar sample; For the training set One sample; for The corresponding historical reasoning text; for The true prediction label; The semantic archive text of known labeled samples in the training set; Let be the magnitude of the vector; These are the samples to be predicted in the test set; This is a pre-trained embedding model.
7. The learning data analysis method based on semantic space geometric consistency verification according to claim 5, characterized in that, The specific process of step 4 is as follows: Step 41, Geometric distance calculation: Generate the logical vector of the first result. , and calculate Euclidean distance to the two semantic reference centers: ; ; in, The logical vector for the first result; For pre-trained embedding models; This is the second explanatory text; For the semantic benchmark center of the state; The semantic baseline center for the failed state; It is an L2 norm; for arrive The Euclidean distance; for arrive The Euclidean distance; Step 42, Consistency Decision Logic: Deriving Implicit Semantic Tags ,like Then implicit semantic tags For passing, otherwise implicit semantic tags are used. It failed; then the implicit semantic tag was added. With predictive labels By comparing the results, the final prediction is obtained. .
8. The learning data analysis method based on semantic space geometric consistency verification according to claim 7, characterized in that, In step 42, implicit semantic tags With predictive labels The comparison method is as follows: like If the prediction confidence is high, then the output is considered high. , This is the final prediction result; at this point, the logical vector of the first result is... Dynamically add to the corresponding set based on the classification results. Or not passed the set In China, and updated in real time. and This enables the self-evolution of the reasoning feature space of a large language model; like This indicates that the large language model is experiencing illusions, triggering a self-correction mechanism that requires the large language model to generate again.