Personalized learning long-tail data processing method fusing nerve collapse detection and regulation
By integrating the methods of nerve collapse detection and regulation in personalized learning, using LoRA fine-tuning and text modal collapse degree calculation, the problem of low classification accuracy caused by long-tail data distribution is solved, and the classification efficiency of personalized learning data is improved.
Patent Information
- Application Number
- CN202510726538.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-06-03
AI Technical Summary
In the existing personalized learning methods, long-tail data distribution leads to low classification accuracy, and it is difficult to collect high-quality educational data, and the problem of data imbalance has not been effectively solved.
A personalized learning long-tail data processing method that integrates fusion nerve collapse detection and regulation is adopted. The model weight is adjusted to improve classification efficiency through LoRA fine-tuning of large language models and the calculation of text modal collapse degree.
It effectively improves the ability of large models to classify personalized learning data, solves the problem of low classification accuracy of long-tail data, and this method is model-independent and is suitable for different personalized learning tasks.
Smart Images

Figure CN120234657A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of machine learning and intelligent education, and specifically relates to large language model text processing technology, especially a personalized learning long-tail data processing method that combines neural collapse detection and regulation for the long-tail data distribution problem in personalized learning methods. Background Art
[0002] Data-based personalized learning is widely applied in important educational scenarios such as knowledge tracing, cognitive diagnosis, and computerized adaptive testing. In recent years, researchers have continuously made progress in in-depth exploration. They not only enhance the expressive ability of the model through more complex architectures but also focus on constructing larger-scale and more diverse data sets to improve the generalization ability of the model in different scenarios. For example, in cognitive diagnosis tasks, by deeply analyzing students' learning behaviors and cognitive characteristics, personalized diagnostic reports can be generated to help teachers monitor students' progress in real time and identify learning bottlenecks, thus providing technical support for a student-centered teaching model. However, these methods generally assume that the data used for data-driven methods is of high quality and accurately labeled.
[0003] Due to the protection of minors' personal privacy and the difficulty of relying on educational experts for annotation, collecting high-quality and standardized educational data is both challenging and unrealistic. Therefore, data-based personalized learning benchmarks usually exhibit an imbalanced or long-tail distribution. However, existing research has largely ignored this problem. Experimental results show that the performance of large language models varies with the change of data balance. When the data imbalance is weak, the model performs well; conversely, when the data imbalance intensifies, the model performance significantly decreases. Summary of the Invention
[0004] The purpose of the present invention is to solve the above problems existing in the prior art and provide a personalized learning long-tail data processing method that combines neural collapse detection and regulation, to solve the problem of low classification accuracy in long-tail data, thereby improving the classification efficiency of large models for personalized learning data.
[0005] The specific technical solution adopted by the present invention is as follows:
[0006] A personalized learning long-tail data processing method that combines neural collapse detection and regulation, comprising the following steps:
[0007] S1. Obtain a personalized learning text data set containing a long-tail distribution, and use the method of stratified sampling to construct a test set, and the remaining data is used as a training set;
[0008] S2. Calculate the multi - head attention through the parallel attention heads of each layer of the large - language model. Based on the hidden state of the previous layer and the multi - head attention, obtain the hidden state of each transformer layer. Based on the difference between the inner product of the hidden states of each two different sample inputs in the transformer layer and the fixed - length constraint of the text representation, obtain the text modality collapse degree TCD;
[0009] S3. According to the text modality collapse TC loss and the task - specific loss Define the comprehensive loss function in the LoRA fine - tuning process. Use the binary function to calculate the text modality collapse degree between different category samples and accumulate it as the text modality collapse TC loss, and calculate the task - specific loss through intra - class cohesion and inter - class repulsion. The comprehensive loss function is composed of the linear combination of the text modality collapse TC loss and the task - specific loss , and the calculation expression is as follows
[0010]
[0011] where, λ is a constant parameter that controls the influence of the text modality collapse TC loss;
[0012] S4. Perform LoRA fine - tuning on the pre - trained large - language model (such as GPT, BERT, etc.) and the training set, and modify the weight matrix of the pre - trained large - language model by adding low - rank matrices A and B on the basis of the weight matrix . The formula is:
[0013]
[0014] where, is the weight update matrix, is the input feature vector, is the hidden - layer output of the model after low - rank adjustment, is a constant parameter for scaling, is the rank of the low - rank matrices A and B.
[0015] The hidden - layer output m of the model passes through the fully - connected layer to generate the prediction probability of each category, and calculate the loss for m according to the comprehensive loss function. The text modality collapse degree TCD and the comprehensive loss function update the parameters of the low - rank matrices A and B through backpropagation;
[0016] S5. Test and evaluate the fine - tuned large - language model according to the content of the test set to complete the classification of the test set.
[0017] Furthermore, the calculation of the text modality collapse degree TCD specifically includes:
[0018]
[0019] Among them, Avg represents taking the arithmetic mean of all elements that meet the conditions. and is the sample input, N is the number of categories. is the fixed-length constraint of the text expression. represents the hidden state of the l-th layer. .
[0020] Furthermore, the calculation of the text modality collapse TC loss, its calculation expression includes:
[0021]
[0022] Among them, ∑ is the summation. is the binary function, B is the batch size. is the category of the i-th sample; the text modality collapse TC loss is constrained by the following conditions:
[0023]
[0024] Among them, represents the Euclidean norm. represents any sample i from the 1st to the B-th sample in a batch.
[0025] Furthermore, the test evaluation on the fine-tuned large language model uses the reported classification accuracy and recall rate as indicators.
[0026] Compared with the prior art, the beneficial effects achieved by the present invention:
[0027] The present invention introduces neural collapse into the text embedding and LoRA fine-tuning of the large language model, improving the model's classification ability for personalized learning data. It effectively ensures the personalization of model training and solves the problem of low classification accuracy for long-tail data. The method of the present invention has the characteristic of being model-independent, can be compatible with various architectures and methods, and can be widely applied to different personalized learning tasks, especially tasks such as educational data, mathematical reasoning, and dialogue classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is the effect diagram of the personalized learning long-tail data processing method that fuses neural collapse detection and regulation under different data ratios provided by the embodiment of the present invention;
[0029] Figure 2 is the flow framework diagram of a personalized learning long-tail data processing method that fuses neural collapse detection and regulation provided by the embodiment of the present invention. Detailed Implementation Modes
[0030] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe in detail the specific implementation modes of the present invention with reference to the accompanying drawings. Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. The technical features in each embodiment of the present invention can be combined correspondingly without conflict.
[0031] Embodiment
[0032] The corpus datasets mainly used in this embodiment are TMWPL and PMTD. The overall process is as follows:
[0033] 1. Dataset Selection
[0034] To evaluate the performance, this embodiment considers two different downstream tasks of personalized learning: math problem classification and teaching strategy classification. The following is a description of the downstream datasets used for each task:
[0035] S1. Math Problem Classification: The math problem classification task is evaluated using TMWPL (TIMSS Math Word Problem Labeled Dataset). This dataset contains math word problems designed for grades 3 to 6 and is suitable for analyzing students' performance in different math cognitive dimensions. Each math problem is annotated by experts and covers multiple cognitive dimensions such as recall, reasoning, and analysis. Among the cognitive dimensions, the recall has the highest sample frequency, with a total of 5045 samples (accounting for 63.85% of the entire TMWPL dataset), while the analysis has the lowest sample frequency, with only 216 samples (accounting for 2.73% of the entire TMWPL dataset). The test set contains 50 samples for each cognitive dimension, and the remaining samples are used for training.
[0036] S2. Teaching Strategy Classification: The teaching strategy classification task uses PMTD (Personalized Math Tutoring Dialogue Dataset), which aims to record the one-on-one teaching interactions between teachers and students, with a particular focus on the guiding teaching strategies adopted when solving challenging math problems. The PMTD dataset is constructed based on the IRF framework (Initiation-Response-Feedback) and the Scaffolding Theory. Among all dimensions, R-FR (students' factual responses) contains 1014 samples, accounting for 24.06% of the total interaction samples; while R-RR (students' rejection responses) only contains 62 samples, accounting for 1.47%. The test set consists of 40 samples for each category, and the remaining samples are used for training.
[0037] 2. Implementation Details
[0038] Refer to Figure 2 , this embodiment introduces a personalized learning long-tail data processing method that integrates neural collapse detection and regulation, specifically including the following steps:
[0039] S1. Obtain the personalized learning text datasets TMWPL and PMTD containing long-tail distributions, and use the method of stratified sampling to construct a test set, with the remaining data as the training set;
[0040] S2. Calculate the multi-head attention through the parallel attention heads of each layer of the large language model. Based on the hidden state of the previous layer and the multi-head attention, obtain the hidden state of each transformer layer. Based on the difference between the inner product of the hidden states of two different sample inputs in the transformer layer and the fixed-length constraint of the text representation, obtain the text modality collapse degree TCD;
[0041] S3. Define the comprehensive loss function in the LoRA fine-tuning process according to the text modality collapse TC loss and the task-specific loss Calculate the text modality collapse degree between different category samples using a binary function and accumulate it as the text modality collapse TC loss, and calculate the task-specific loss through intra-class cohesion and inter-class repulsion. The comprehensive loss function is composed of a linear combination of the text modality collapse TC loss and the task-specific loss , and the calculation expression is as follows
[0042]
[0043] where λ is a constant parameter that controls the influence of the text modality collapse TC loss;
[0044] S4. Perform LoRA fine-tuning on the pre-trained large language model Qwen2.5-instruct and the training set, and modify the weight matrix of the pre-trained large language model, which is achieved by adding low-rank matrices A and B to the weight matrix . The formula is:
[0045]
[0046] where, is the weight update matrix, is the input feature vector, is the hidden layer output of the model after low-rank adjustment, is a constant parameter for scaling, is the rank of the low-rank matrices A and B.
[0047] The output m of the hidden layer of the model passes through a fully connected layer to generate the prediction probability for each category, and the loss is calculated for m according to the comprehensive loss function. The text modality collapse degree TCD and the comprehensive loss function Update the parameters of the low-rank matrices A and B through backpropagation;
[0048] S5. Test and evaluate the fine-tuned large language model according to the content of the test set, and complete the classification of the test set.
[0049] 3. Evaluate the effect of handling long-tail data
[0050] As shown in Table 1, the 7B parameter model adjusted by the present invention has achieved remarkable results, and its performance is better than that of the 7B to 14B parameter models, including dense and expert mixed (MoE) architectures. This remarkable achievement in all evaluation benchmarks indicates that the present invention can not only provide state-of-the-art performance but also maintain computational efficiency with a smaller number of parameters. Especially in the TMWPL task, the classification accuracy has achieved the largest improvement, which is 13.72% higher than the strongest baseline model, as can be proven from the results of the NCAL-Qwen2.5-Instruct model.
[0051] Table 1 Comparison of the performance of different models in educational tasks
[0052]
[0053] 4. Evaluate the performance of the present invention under different data ratios
[0054] The performance of the representative large language model Owen2.5 changes with the change of data balance. Figure 1 (a-d) The experimental results shown indicate that when testing different minimum and maximum class number ratios Among them, when The higher the value, indicating the weaker the data imbalance, the better the performance of the model. On the contrary, as The value decreases, that is, the deterioration of data imbalance, the performance of the model will decline, and the accuracy drops from 71.71 ( Figure 1 (a)) to 61.14 ( Figure 1 (c)). As the data distribution becomes more balanced, the accuracy of the model also improves. However, due to insufficient data volume, effective fine-tuning cannot be carried out, and the performance drops at the 0.25 ratio. Our analysis shows that there is a strong correlation between the TCD value and the classification accuracy in all model variants. Specifically, the larger the TCD value, The lower the value, which indicates that the degree of deviation from the ideal ETF structure directly affects the performance of the model.
[0055] 5 Evaluate the performance of the present invention with and without prompt words
[0056] In the context of an instruction tuning model, the design of prompts can significantly affect the zero-shot ability of the model, especially in differentiating between directly generated categories and categories generated after applying Chain of Thought (COT) reasoning. To address this issue, the present invention involves the impact of prompt pairs specified for classification tasks on model performance. Specifically, the optimal model input combines the task description and dataset features, while the scenario without prompts only includes dataset features. The results in Table 2 show that in the scenario where only dataset features are input without any additional prompts for LoRA fine-tuning, all models experience a significant performance decline.
[0057] Table 2 Impact of Prompts on the Present Invention
[0058] The embodiments described above are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical fields can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A personalized learning long-tail data processing method integrating neural collapse detection and regulation, characterized in that, Including the following steps: S1. Obtain a personalized learning text dataset containing a long-tail distribution, and use the method of stratified sampling to construct a test set, with the remaining data as the training set; S2. Calculate the multi-head attention through the parallel attention heads of each layer of the large language model. Based on the hidden state of the previous layer and the multi-head attention, obtain the hidden state of each transformer layer. Based on the difference between the inner product of the hidden states of each two different sample inputs in the transformer layer and the fixed-length constraint of the text expression, obtain the text modality collapse degree TCD; S3. Determine the TC loss according to the text modality collapse and the task-specific loss Define the comprehensive loss function in the LoRA fine-tuning process. Use a binary function to calculate the degree of text modality collapse between samples of different categories and accumulate it as the text modality collapse TC loss. Calculate the task-specific loss through intra-class cohesion and inter-class repulsion. The comprehensive loss function is composed of a linear combination of the text modality collapse TC loss and the task-specific loss , and the calculation expression is as follows: ; where λ is a constant parameter that controls the influence of the text modality collapse TC loss; S4. Perform LoRA fine-tuning on the pre-trained large language model and the training set, and modify the weight matrix of the pre-trained large language model by adding low-rank matrices A and B based on the weight matrix ; the formula is: ; Among them, is the weight update matrix, is the input feature vector, is the hidden layer output of the model after low-rank adjustment, is a constant parameter for scaling, is the rank of the low-rank matrices A and B; The output m of the hidden layer of the model passes through the fully connected layer to generate the prediction probability for each category, and the loss is calculated for m according to the comprehensive loss function; the text modality collapse degree TCD and the comprehensive loss function Update the parameters of the low-rank matrices A and B through backpropagation; S5. Test and evaluate the fine-tuned large language model according to the content of the test set, and complete the classification of the test set.
2. The personalized learning long-tail data processing method integrating neural collapse detection and regulation according to claim 1, wherein The calculation of the text modality collapse degree TCD specifically includes: ; Among them, Avg represents taking the arithmetic mean of all elements that meet the conditions, and are sample inputs, N is the number of categories, is the fixed-length constraint of the text expression, represents the hidden state of the l-th layer, .
3. The personalized learning long-tail data processing method integrating neural collapse detection and regulation according to claim 1, characterized in that, The calculation of the text modality collapse TC loss, and its calculation expression includes: ; where ∑ is the summation, is a binary function, B is the batch size, is the category of the i-th sample; the text modality collapse TC loss is constrained by the following conditions: ; Among them, represents the Euclidean norm, represents any sample i among the first to the Bth samples in a batch.
4. The personalized learning long-tail data processing method integrating neural collapse detection and regulation according to claim 1, characterized in that, For the test and evaluation of the large language model described in step 5, the classification accuracy and recall rate are used as indicators.
Citation Information
Patent Citations
Random group division channel whitening method for self-supervised learning
CN118246512A
Slope detection, evaluation, classification and prediction model based on neural network model
CN118395281A
Neural collapse theory-based pre-training model class incremental learning identification method
CN119760495A