Student mental health prediction method and system based on multi-view semi-supervision

Through the multi-view semi-supervised method, reliable negative samples and collaborative training are screened, combined with the cost-sensitive learning mechanism, the problem of difficult access to negative samples in the mental health prediction of college students is solved, and more efficient prediction and comment generation is achieved, supporting college mental health intervention.

CN120260932APending Publication Date: 2025-07-04SHANDONG JIANZHU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510732706.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The mental health prediction of college students faces the problem of difficulty in obtaining and labeling of negative samples. The existing technology fails to effectively utilize PU learning, resulting in poor model performance and inability to accurately predict mental health status.

Method used

Using a multi-view semi-supervised method, reliable negative samples are selected through a reliable negative sample selection module, combined with a multi-view collaborative training module and a cost-sensitive learning mechanism, a mental health prediction model is constructed, and a generative language model is used to generate comments.

Benefits of technology

It improves the accuracy and interpretability of mental health predictions, especially in terms of comprehensive check rate and F1 score performance over traditional methods, providing professional and targeted decision-making support to facilitate mental health interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260932A_ABST
    Figure CN120260932A_ABST
Patent Text Reader

Abstract

The invention provides a student mental health prediction method and system based on multi-view semi-supervision, and relates to the field of mental health prediction. A positive sample set containing student education feature vectors and psychological health labels and a to-be-predicted unmarked sample set are obtained, the positive sample set and the to-be-predicted unmarked sample set are input into the psychological health prediction model, and a reliable negative sample selection module screens out a reliable negative sample set according to the distance between the positive sample set and the to-be-predicted unmarked sample set; the multi-view cooperative training module comprises classifiers corresponding to multiple education data, and is used for extracting education data and labels in the positive sample set and the reliable negative sample set, training the classifiers, performing psychological health prediction on unmarked sample sets to be predicted, and synthesizing false labels of the classifiers based on a cost-sensitive mechanism to obtain a psychological health prediction result; and obtaining a psychological health prediction result of the to-be-predicted unmarked sample and generating a comment. The prediction performance of the model is improved by introducing a PU learning normal form, combining multi-view student related data and classifier cooperative training and fusing complementary information among different view data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of mental health prediction, and particularly to a method and system for predicting students' mental health based on multi-view semi-supervision. Background Art

[0002] Due to factors such as learning pressure and interpersonal relationships, the mental health of college students faces severe challenges. Traditional methods such as manual surveys and conversations have achieved remarkable results in the psychological assessment and screening of college students, but they consume a huge amount of manpower and material resources. For the massive data generated by the informatization of colleges and universities, the research on predicting students' mental health using machine learning technology has received increasing attention from educators and researchers. However, the inner true mental state of students is extremely hidden, making it very difficult to label samples, especially negative samples (i.e., students without mental problems) are difficult to define. Previous studies used the evaluation results from psychological assessment scales as sample labels for modeling, but the scales are highly subjective, and the authenticity and reliability of the obtained conclusions are difficult to guarantee, which seriously affects the performance of the model.

[0003] PU learning (Positive Unlabeled Learning) is a semi-supervised learning method aimed at training a binary classification model using a small number of positive samples and a large number of unlabeled samples to solve the problem of difficult acquisition or labeling of negative samples. It is widely used in the field of machine learning, such as tasks like malicious URL identification and gene detection.

[0004] Although PU learning has shown advantages in many fields, it has not been fully explored and applied in the prediction of students' mental health. Past research in this field has mostly focused on the traditional full-sample labeling mode, ignoring this innovative method of PU learning. The mental health data of college students has uniqueness, and the definition of negative samples becomes extremely complex due to the concealment of students' mental states, which is essentially different from scenarios such as malicious URL identification and gene detection. In the scenario of mental health prediction, the models and methods of PU learning in other fields cannot be simply applied. Currently, the existing technology fails to effectively utilize PU learning to solve the problem of difficult acquisition and labeling of negative samples when facing the task of predicting college students' mental health, and no research has successfully constructed a student mental health prediction model based on PU learning, resulting in the difficulty of fully exploring the potential value of massive educational data in this field and the inability to achieve accurate prediction and effective intervention of students' mental health status. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a method and system for predicting students' mental health based on multi-view semi-supervised learning. By introducing the PU learning paradigm and combining multi-view student-related data, a collaborative training framework of multiple integrated classifiers is constructed to improve the model performance by fusing the complementary information between different view data. In addition, the latest generative language model technology is used to interpret and explain the model prediction results, generating targeted mental health status comments, providing more valuable interpretable decision support for psychological counselors and student workers, and helping to carry out students' mental health intervention work more efficiently.

[0006] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a method for predicting students' mental health based on multi-view semi-supervised learning, including: Obtaining a basic sample set; the basic sample set is a positive sample set containing students' educational feature vectors and mental health labels and an unlabeled sample set to be predicted, wherein the educational feature vectors contain multiple items of educational data; Inputting the basic sample set into a mental health prediction model, the mental health prediction model includes a reliable negative sample selection module and a multi-view collaborative training module; the reliable negative sample selection module is used to screen out a reliable negative sample set according to the maximum mean distance between the positive sample set and the unlabeled sample set to be predicted; the multi-view collaborative training module includes classifiers corresponding to multiple items of educational data respectively, which are used to extract the educational data and labels in the positive sample set and the reliable negative sample set, train the classifiers, predict the mental health of the unlabeled sample set to be predicted, and obtain the mental health prediction result of the unlabeled sample set to be predicted by synthesizing the pseudo-labels of each classifier based on the cost-sensitive learning mechanism; Using an improved generative language model to generate comments for the mental health prediction results.

[0007] In a second aspect, the present invention provides a system for predicting students' mental health based on multi-view semi-supervised learning, including: A data acquisition unit configured to obtain a basic sample set; the basic sample set is a positive sample set containing students' educational feature vectors and mental health labels and an unlabeled sample set to be predicted, wherein the educational feature vectors contain multiple items of educational data; A prediction unit, configured to input a basic sample set into a mental health prediction model, where the mental health prediction model includes a reliable negative sample selection module and a multi-view collaborative training module; the reliable negative sample selection module is used to screen out a reliable negative sample set according to the maximum mean distance between a positive sample set and an unlabeled sample set to be predicted; the multi-view collaborative training module includes classifiers corresponding to multiple education data respectively, and is used to extract the education data and labels in the positive sample set and the reliable negative sample set, train the classifiers, perform mental health prediction on the unlabeled sample set to be predicted, and comprehensively integrate the pseudo-labels of each classifier based on a cost-sensitive learning mechanism to obtain the mental health prediction result of the unlabeled sample set to be predicted; A comment generation unit, configured to generate a comment for the mental health prediction result by using an improved generative language model.

[0008] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in a method for predicting students' mental health based on multi-view semi-supervised learning described in the first aspect are implemented.

[0009] In a fourth aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in a method for predicting students' mental health based on multi-view semi-supervised learning described in the first aspect are implemented.

[0010] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention first obtains a positive sample set and an unlabeled sample set containing educational feature vectors and mental health labels, then screens out a reliable negative sample set through a reliable negative sample selection module, uses a multi-view collaborative training module to train classifiers for prediction, and comprehensively integrates the results based on a cost-sensitive learning mechanism. Then, an improved generative language model is used to extract key features from the prediction results and convert them into natural language descriptions. Semantic vectors are extracted through local semantic and global relationship encoding, and a three-level prompt template is designed to construct a prompt word template, thereby generating a comment on the mental health status. The present invention fully explores the potential value of a large amount of educational data in colleges and universities, breaks through the limitations of the traditional full-sample annotation mode, avoids simply applying PU learning methods in other fields, can more accurately predict students' mental health status, provides support for effective intervention, and fills the gap in constructing a prediction model based on PU learning in this field. The method is superior to the traditional method in four indicators: accuracy, precision, recall, and F1. Especially in the recall indicator, it is 20.73 percentage points higher than the sub-optimal method.

[0011] (2) The present invention innovatively introduces the PU learning paradigm to the task of predicting students' mental health, and proposes a new mental health prediction method, which alleviates to a certain extent the challenging problem that it is extremely difficult to define negative samples due to the extreme concealment of students' true mental states.

[0012] (3) The present invention integrates multi-source heterogeneous data such as the teaching affairs system, psychological assessment system, campus behavior logs, and social attribute information (cadre appointments, competition awards) of a certain university to form a dataset containing multi-views, multi-view attributes, and student mental state labels, which can provide data support for educators and researchers to carry out mental health prediction research.

[0013] (4) The present invention designs a novel multi-view collaborative training framework. Through the collaborative learning strategy between views, classifiers of different views can share information by exchanging pseudo-labels, and by virtue of the complementarity between views, improve the utilization efficiency of the mental health prediction model for unlabeled data.

[0014] (5) The cost-sensitive learning mechanism proposed by the present invention, when working in collaboration with the classifier, can assign sample labels that are more in line with the actual situation of students' mental health to the classifier according to the misclassification cost, enabling the classifier to more accurately identify the characteristics of various samples. During the model training process, the training focus can be flexibly adjusted according to the cost difference, the model parameters can be optimized, training biases can be avoided, and the model training can be made more scientific and efficient. Finally, in the prediction link, the misjudgment probability is greatly reduced, the prediction accuracy is improved, which helps the school accurately grasp the mental health situation of students, lays a foundation for timely and effective psychological intervention, and effectively excavates the potential value of educational data.

[0015] (6) By extracting key features from the mental health prediction results and converting them into natural language descriptions, the present invention can accurately capture important information. Based on local semantics and global relationship encoding to extract semantic vectors, the text semantics can be comprehensively captured. Designing a three-level prompt template to construct a prompt word template provides a clear direction for comment generation. The finally generated mental health status comments are closely related to the semantics of the input features, providing professional, accurate, and targeted decision support for psychological counselors and student workers, and helping to effectively carry out students' mental health work.

[0016] The advantages of the additional aspects of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute a limitation to the present invention.

[0018] Figure 1The main flowchart of a method for predicting students' mental health based on multi-view semi-supervised learning provided by an embodiment of the present invention; Figure 2 The overall architecture diagram of a method for predicting students' mental health based on multi-view semi-supervised learning provided by an embodiment of the present invention. Detailed implementation manners

[0019] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0020] In view of the challenging problem that the true mental state of students is extremely hidden, making it difficult to define negative samples, the present invention proposes a method for predicting mental health based on PU learning (Positive-Unlabeled Learning).

[0021] First, by given a small amount of high-quality positive sample data (i.e., students with mental problems), and a large number of unlabeled samples, several samples with the largest average distance from the positive samples are selected from the unlabeled samples as "negative samples" to construct an initial labeled dataset.

[0022] It should be understood that the mental health prediction described in the present invention focuses on judging whether a student belongs to the category labels of "mentally healthy" or "mentally unhealthy" through a model, and does not limit and study the specific determination criteria for "mentally unhealthy". Those skilled in the art can define the connotations of the two category labels in the form of standardized scales, interview records, etc. according to the existing psychological assessment system or business requirements. The core of the present invention is to provide a method framework for screening reliable negative samples based on positive samples and then constructing a prediction model, rather than solving the original determination problem of mental states.

[0023] Next, the present invention introduces a multi-view collaborative training mechanism to construct a mental health prediction model. Specifically, the labeled data of each view are used to train the model respectively, and the error rate of each model is measured using the test set. Then, the classifiers with error rates exceeding a given threshold are excluded and do not participate in the current round of collaborative training, while the remaining classifiers are used as assisting classifiers to predict the unlabeled samples to obtain pseudo-labels.

[0024] Considering the differences in misclassification costs in the task of students' mental health, the present invention introduces a cost-sensitive mechanism to set different thresholds for decision-making when predicting unlabeled samples, rather than simple majority voting. In this way, pseudo-labeled samples with higher confidence are screened out and added to the labeled dataset of the previous round. Subsequently, the classifier is retrained using the expanded data. The above process is iterated continuously until all models do not update, and a trained model is obtained for accurately predicting the mental health status of students, and the prediction results are divided into mentally healthy and mentally unhealthy.

[0025] Embodiment 1 As Figure 1 shown, this embodiment discloses a method for predicting students' mental health based on multi-view semi-supervised learning, which includes the following steps: S1: Obtain a basic sample set; the basic sample set is a positive sample set containing students' educational feature vectors and mental health labels and an unlabeled sample set to be predicted, where the educational feature vectors include multiple educational data from different data sources; S2: Input the basic sample set into a mental health prediction model, which includes a reliable negative sample selection module and a multi-view collaborative training module; the reliable negative sample selection module is used to screen out a reliable negative sample set according to the maximum mean distance between the positive sample set and the unlabeled sample set to be predicted; the multi-view collaborative training module includes classifiers corresponding to multiple educational data respectively, which are used to extract the educational data and labels in the positive sample set and the reliable negative sample set, train the classifiers, predict the mental health of the unlabeled sample set to be predicted, and comprehensively obtain the mental health prediction results of the unlabeled sample set to be predicted based on the cost-sensitive learning mechanism by combining the pseudo-labels of each classifier; S3: Use an improved generative language model to generate comments for the mental health prediction results.

[0026] Next, combined with Figure 1 this, a method for predicting students' mental health based on multi-view semi-supervised learning disclosed in this embodiment will be described in detail.

[0027] (I) Problem formalization This embodiment formalizes the problem of students' mental health modeling and prediction as a PU learning (Positive and Unlabeled Learning) problem. Unless otherwise specified, hereinafter, this embodiment uses lowercase letters to represent scalars, lowercase bold letters to represent vectors, and uppercase letters to represent matrices or sets.

[0028] 1. Problem definition Given a data set D = , where ∈ Rd is the d-dimensional real vector representation of the i-th student, which is the vector representation of the student's features formed after integrating each view; is its label, and |D| represents the modulus of the set D, that is: the number of student samples.

[0029] In the PU learning scenario, the label has two values, including: = 1 (positive sample) and = "unknown", which respectively represent that the student has mental health problems and the mental health status of the student is unknown.

[0030] The goal of mental health modeling and prediction for students is to build a model f that can accurately predict the mental health status of students and identify potential students with mental health problems from unlabeled samples.

[0031] 2. Formal representation Positive sample set: P = { | = 1}; Unlabeled sample set: U = { | = "unknown"}; Reliable negative sample set: ; Final training set: L = P ∪ RN.

[0032] The task of the mental health prediction problem based on PU learning is to screen out reliable negative samples RN from the unlabeled student samples U and merge them with the positive class sample set P to obtain the labeled mental health data set L. Combining L and the remaining large amount of unlabeled data set (U - RN), establish a mapping f from the student representation vector to of ) → (i = 1,..., |D|), so as to predict the mental health status of newly arrived students.

[0033] (2) Algorithm framework As Figure 2 shown, the model designed in this embodiment uses a two-stage strategy for modeling, that is: screen out the reliable negative sample set RN, and then model and predict. Correspondingly, the model consists of two parts: a reliable negative sample selection module and a multi-view collaborative training module.

[0034] 1. Reliable negative sample selection module The smoothness hypothesis believes that similar data points should have similar labels in the feature space. In other words, the unlabeled samples that are farthest from the positive samples are more likely to belong to the negative class.

[0035] Therefore, in this embodiment, several unlabeled samples with the largest average distance from the positive samples are selected as reliable negative samples to obtain the pseudo-negative sample set RN.

[0036] The specific steps are as follows: (1) For each unlabeled sample ( , " = unknown") ∈ U, use formula (1) to calculate its Euclidean distance dis( , = 1") ∈ P: , ):

[0037] Among them, d is the feature dimension, and respectively represent the student sample vectors and values on the k-th feature.

[0038] (2) For each unlabeled sample , use formula (2) to calculate its average Euclidean distance from all positive samples,

[0039] where is the number of samples in the positive sample set P.

[0040] (3) According to the average distance sort all unlabeled samples (j = 1, …, |U|) in descending order, and select several samples with the top rankings as reliable negative samples to form the data set RN.

[0041] This embodiment is based on the smooth assumption and uses the Euclidean distance to screen samples, which can effectively find samples with large differences from positive samples as reliable negative samples from a large number of unlabeled samples. This process is accurate and efficient, provides a high-quality negative sample set for subsequent model training, reduces noise interference, and optimizes the training data structure. It enables the model to better distinguish the features of positive and negative samples during training, enhances the discriminant ability of the model for students' mental health status, and improves the accuracy and reliability of prediction.

[0042] 2. Multi-view collaborative training module Merge the selected reliable negative samples RN and the existing positive sample set P to obtain the labeled data set L, and train the mental health prediction model on this basis.

[0043] Educational data often presents characteristics such as multi-source and multi-dimensionality. For example: academic performance, competition awards, cadre positions, family situations, accommodation behaviors, etc., which reflect all aspects of students' study and life. Therefore, this embodiment regards data from different aspects as different views and introduces multi-view learning to build the model.

[0044] This embodiment gives the architecture diagram of the multi-view collaborative training module by constructing multiple views and training corresponding classifiers, as shown in Figure 2 . The specific steps are as follows: (1) For the reliable negative samples and the original positive samples in the preprocessed student data set composed of n view attributes, denoted as the labeled set , where represents the i-th labeled sample, and the remaining large number of unlabeled samples are denoted as the set U.

[0045] (2) Based on the labeled sample sets of each view , respectively train and initialize n base classifiers , and each classifier corresponds to one view.

[0046] (3) Each view corresponds to a classifier , extract the data related to this view and the corresponding labels from the labeled training data set as the training samples of the view classifier for training this classifier.

[0047] (4) For each classifier , calculate the current error rate according to the labeled samples under the corresponding view. If i ≤ n and the error rate of the current classifier is less than the error rate of the previous round, then execute (5)-(7). If i > n then execute (8).

[0048] (5) Considering that the losses caused by different classification errors in the student mental health prediction task are different. That is to say, if a mentally healthy student (negative sample) is predicted as having mental abnormalities (positive sample), it will lead to a slight increase in the workload of college mental health service staff; however, for a student with potential mental problems (positive sample) predicted as a negative sample, there may be a lack of attention due to missed detection, and there is a possibility of the mental problems getting worse.

[0049] Therefore, given an unlabeled sample u, introduce a cost-sensitive learning mechanism.

[0050] Solution 1: Adopt a cost-sensitive threshold decision to screen pseudo-labels: Use formula (3) to adopt a cost-sensitive loss strategy to screen high-confidence pseudo-labels:

[0051] where n is the total number of classifiers, and are the numbers of classifiers that predict u as a positive sample and a negative sample among other classifiers respectively.

[0052] In addition to the above method, this embodiment also provides a cost-sensitive mechanism from the perspective of information entropy calculation, further considering the impact of misclassification cost on model training.

[0053] The first solution determines the pseudo-labels of samples based on the opinions of multiple classifiers through a simple and clear judgment formula. It can quickly process a large number of unlabeled samples, greatly improving the classification efficiency, and has obvious advantages when the amount of student mental health data is huge. It is easy to operate and has low computational overhead. It can quickly and preliminarily locate the category of students' mental states in scenarios such as daily monitoring, facilitating school psychological service personnel to promptly grasp the overall situation and prioritize attention to student groups that may have psychological problems.

[0054] The second solution uses cost-sensitive weighted information entropy to screen pseudo-labels: When calculating the information entropy, introduce the class weight (cost ratio) and adjust the entropy calculation formula:

[0055] Among them, is the weight of class i, related to the misclassification cost: represents the probability that the sample belongs to class i.

[0056]

[0057] Among them, is the cost matrix. Since we want the weight of positive class samples to be higher, so .

[0058] The specific steps are as follows: ① For each class i, calculate its total misclassification cost:

[0059] ② Normalize the weights:

[0060] Among them, is the total misclassification cost of the class.

[0061] ③ Calculate the weighted entropy:

[0062] ④ When splitting nodes, use the weighted entropy to calculate the information gain.

[0063] For example, if it is found that the features of a certain unlabeled sample have a higher weighted information gain in the positive class after calculating the weighted entropy, then there is more reason to assign it the pseudo-label of the positive class, and then add it to the labeled sample set of the corresponding view to participate in subsequent model training.

[0064] The second solution introduces class weights to comprehensively consider the misclassification cost, making the judgment of sample classes more refined. It can deeply explore the relationship between sample features and misclassification cost, especially suitable for scenarios where students' mental states are complex and hidden. It can accurately adjust the model training direction according to different misclassification costs, making the model more accurate in identifying various samples, effectively avoiding misjudging students with mental health as having mental abnormalities or missing students with potential psychological problems, and improving the accuracy of the model's prediction of students' mental health.

[0065] As an implementation method, the cost-sensitive learning mechanism of Solution 1 or Solution 2 can be selected according to needs. For example, in the middle stage of the semester, students' mental states are relatively stable but still need continuous attention. At this time, if you pursue a fast and simple way to screen pseudo-labels, you can use the cost-sensitive threshold determination to screen pseudo-labels (Solution 1). This solution determines pseudo-labels based on the judgment results of the majority classifier through a simple formula, with a small amount of calculation. It can quickly give a preliminary prediction when processing a large number of unlabeled samples in daily batches, improving the monitoring efficiency. For example, in the daily collection and preliminary evaluation of mental health data, student samples can be quickly classified to facilitate the timely discovery of student groups that may have problems.

[0066] In the early and late stages of the semester, freshmen have just entered the campus and are facing adaptation problems in many aspects such as the environment and interpersonal relationships. During the exam week and graduation season, students are facing pressures such as assessments, further studies, and employment, resulting in unstable mental states. At this time, it is more necessary to comprehensively consider the misclassification cost, and it is better to use the cost-sensitive weighted information entropy to screen pseudo-labels (Solution 2). By introducing class weights, this solution comprehensively considers factors such as the total misclassification cost and can more accurately evaluate sample classes. In the case of the complex and diverse mental health conditions of freshmen, it can more carefully explore students with potential psychological problems and avoid missed detections or over-interventions caused by misjudgments.

[0067] (6) Add the sample u that meets the conditions and its pseudo-label to the labeled sample set of the corresponding view of the current classifier to expand the training data.

[0068] (7) Decide whether to update the current classifier according to the error rate and the number of expanded samples.

[0069] (8) For the updated classifier, use the labeled sample set under the corresponding view after amplification to retrain and record the current error rate and the number of amplified samples.

[0070] (9) If any classifier is updated, repeat steps (4)-(8); if all classifiers are not updated, terminate the training and output the final classifier set .

[0071] In this embodiment, by introducing a cost-sensitive learning mechanism and combining it with a classifier, more reasonable sample labels can be provided for the classifier according to different cost strategies, enhancing the classifier's adaptability to the complex situations of students' mental states. For model training, it can adjust the training direction based on the misclassification cost, optimize the model parameters, avoid the model from overly biasing towards certain types of samples, and make the training more efficient and reasonable. In terms of the prediction results, the misjudgment situation is reduced, and the prediction accuracy is improved, enabling the school to more accurately grasp the mental health status of students, providing a reliable basis for timely and effective psychological intervention, fully exploring the value of educational data, and assisting in the work of ensuring students' mental health.

[0072] (III) Experimental Verification To verify the effectiveness of this embodiment, it will be compared with traditional supervised learning methods and ensemble methods to verify the effectiveness of the method of this embodiment. Next, the experimental dataset, experimental settings will be introduced first, and then the comparison and analysis of experimental results.

[0073] 1. Dataset The data of this embodiment comes from the information of graduates of an engineering college in a certain university in the past 3 years. After preprocessing such as data alignment, data cleaning, null value deletion / filling, anonymization, and normalization, the mental health dataset MentalData is obtained. This dataset has 1088 student samples, including: demographic information (gender, political status, ethnicity, household registration, etc.), cadre positions, competition awards, academic performance (foreign language scores, major scores, etc.), comprehensive evaluation score rankings, PU Pocket, etc. information from 6 views, with a total of 71-dimensional attributes. Among them, PU Pocket refers to the credits obtained by students in ideological and political and moral cultivation, social practice, cultural and artistic, and physical and mental development, academic technology and innovation and entrepreneurship, social work and skill development, volunteer service, etc. The comprehensive evaluation score includes the comprehensive evaluation of the first year, the second year... based on time series.

[0074] The 71-dimensional original features of the 6 views are directly concatenated into a 71-dimensional vector. After data alignment, data cleaning, null value deletion / filling, and anonymization, the data is preprocessed again. The categorical variables (such as ethnicity, household registration) are encoded, and the numerical variables (such as scores) are normalized. Finally, multi-view feature data is obtained.

[0075] 2. Experimental Settings The experimental environment is the Windows 10 Professional Edition 64-bit (Build 19045) operating system, with the hardware configuration of an Intel Xeon W-2255 CPU (20 cores, base frequency 3.7GHz) and 64GB of memory. The graphics card is an NVIDIA GeForce RTX 3090. All experiments are carried out in the PyCharm 2021.3 development environment.

[0076] The value of the number of views \(n\) is set to 6. For the selection of the number of reliable negative samples, in this embodiment, the top 20% of the unlabeled samples with the lowest similarity are selected and regarded as reliable negative samples. Referring to the Tri-training algorithm, in this embodiment, a random forest is used as the base classifier learning algorithm A, which is implemented by calling the RandomForestClassifier library function in scikit-learn. When conducting the comparative experiment, relevant program modules in the basic library of the machine learning library Scikit-learn are called in this embodiment. All algorithms are evaluated using 5-fold cross-validation, and the optimal hyperparameters are selected by means of grid search. In this embodiment, four indicators, namely accuracy, precision, recall, and F1-score, are used to evaluate the performance of the model.

[0077] 3. Comparative Experiment Table 1 presents the comparison results between the method described in this embodiment and common supervised learning methods, including: Bayesian model, multi-layer perceptron, k-nearest neighbor, support vector machine, decision tree, logistic regression, gradient boosting tree GBDT, extreme gradient boosting XGBoost, random forest, and Tri-training, etc. In this embodiment, the optimal and sub-optimal results are represented by bold and underline respectively. As shown in Table 1, the proposed method has achieved good performance in all four indicators, especially in recall and F1-score, far exceeding other comparative algorithms. In the mental health prediction scenario, recall directly reflects the effective detection ability of the model for high-risk individuals, while traditional models generally perform poorly in this indicator. Although the Bayesian model (0.6484) has a relatively high recall, its precision (0.2459) is extremely low, indicating a large number of false positives (wrong positive classes); the recall of the remaining algorithms is all lower than 0.35, and the logistic regression method (0.0836) can hardly identify high-risk samples, resulting in serious missed detections. In contrast, through the cost-sensitive learning mechanism, the proposed method, while ensuring a relatively high recall (0.8557), improves the precision to 0.7765 (a 156% increase compared to the sub-optimal model), and the F1-score reaches 0.8123 (a 129% increase compared to the sub-optimal model), indicating its effectiveness in balancing the costs of missed detections and false positives.

[0078] Table 1 Experimental Results of the Comparison between this Embodiment and Traditional Methods

[0079] Even compared with the integrated methods, the proposed method shows significant advantages in four metrics. In particular, it far outperforms the comparison algorithms in recall and F1-score. Specifically, the proposed method improves the recall metric by 36.44 percentage points compared to the second-best model XGBoost (0.4913), significantly reducing the risk of missed detections. In terms of the F1-score (0.8123), it improves by 129% compared to the second-best method (Random Forest 0.3541), indicating a better balance between high recall and high precision (0.7765). In terms of accuracy, although the proposed method only improves by 2.0% compared to the second-best method (GBDT 0.7546), the coordinated optimization of its recall and precision verifies the robustness of the model in data imbalance scenarios. The recall rates of Gradient Boosting Tree XGBoost and Random Forest (0.4913, 0.4738) are relatively high, but the precision rates (0.2742, 0.2858) are less than 30%, indicating that the model's identification of high-risk individuals is accompanied by a large number of false positives. As a semi-supervised learning method, Tri-Training achieved sub-optimal performance in precision (0.7692), but its recall rate (0.2083) is only 24.3% of the proposed method, indicating that its high precision comes at the cost of sacrificing the missed detection rate. In contrast, the coordinated optimization of the high recall rate (0.8557) and high precision rate (0.7765) of the proposed method makes it have better prospects for practical application.

[0080] (IV) Improved Generative Language Model Taking the prediction results output by the above algorithm framework and the multi-view feature data as inputs, it aims to automatically generate detailed and personalized mental health status comments, providing valuable interpretable decision-making support for psychological counselors and student workers.

[0081] 1. Semantic Vector Conversion of Feature Vectors and Output Results (1) Semantic Representation of Feature Vectors and Prediction Results 1) Key Feature Extraction Technology The Gradient-weighted Feature Importance (GFI) method is used to determine the key features in multi-view features. By calculating the partial derivative matrix of the prediction results with respect to the input features, the contribution degree of each feature is quantified. The specific calculation formula is:

[0082] Where, represents the global importance score of the -th feature, N represents the total number of samples, represents the predicted value of the -th sample with respect to the -th feature The absolute value of the gradient, representing the feature standard deviation.

[0083] Select features with values higher than the threshold τ as key features. Usually, τ is set to the 85th percentile of the GFI distribution of all features. This can screen out features with high influence on the prediction result and provide key information for subsequent semantic representation and comment generation.

[0084] 2) Structured data to text method Design a hierarchical templated description framework to convert different types of structured key feature data into natural language text descriptions for better interaction with language models. Specifically as follows: Numerical features: Use interval mapping language description. For example, map the academic performance GPA of students into corresponding natural language descriptions according to a certain interval range. "Academic performance is in the top 15%" corresponds to GPA ≥ 3.7. In this way, continuous numerical features are converted into easily understandable text expressions.

[0085] Categorical features: Construct a semantic mapping dictionary for conversion. For example, for the categorical feature of social practice, encode it as {"high": "actively participate", "low": "less participate"}. In this way, when encountering the value of a categorical feature, it can be directly converted into the corresponding natural language description according to the dictionary.

[0086] Temporal features: Apply trend description operators. Taking the difference value Δt in the time series as an example, when Δt > 0, generate the description "increased by {abs(Δt)}% compared with the previous year"; when Δt < 0, generate the description "decreased by {abs(Δt)}% compared with the previous year". In this way, the change trend of temporal features can be clearly expressed.

[0087] (2) Semantic vector generation technology Based on the converted natural language text description, adopt a dual-channel semantic encoding architecture, and combine local semantic encoding and global relationship encoding to generate semantic vectors to more comprehensively capture the semantic information of the text. The specific steps are as follows: Local semantic encoding: Use the [CLS] token vector of BERT for the natural language text description to capture sentence-level semantics. When the BERT model processes text, it encodes the entire sentence into a fixed-length vector, and the [CLS] token vector is the representative of this encoded vector, which contains the semantic information of the entire sentence.

[0088] Global relationship encoding: Construct the relationship matrix A between the features of the natural language text description through the Graph Attention Network (GAT), and output the graph embedding vector . GAT can automatically learn the mutual relationships between features, and assign different weights to different features through the attention mechanism, so as to better capture the global relationships between features.

[0089] The final semantic vector is the linear projection of the two, and the calculation formula is as follows:

[0090] Among them, , ∈ are trainable parameters, d is the target dimension, and b is the bias term. In this way, the local semantic information and the global relationship information are fused together to generate a vector that can comprehensively represent the text semantics , providing richer semantic features for the subsequent GPT-based comment generation.

[0091] 2. GPT-based comment generation technology (1)Conditional prompt construction Design a three-level prompt template system to provide clear generation guidance for the GPT model, enabling it to generate comments that meet the requirements according to the given features.

[0092] The system instruction clarifies the role and task requirements for generating comments, that is, as a mental health analyst, generate comments not exceeding 300 words.

[0093] The feature summary part provides a natural language description of the key features, which are obtained through the previous feature extraction and transformation steps and can summarize the important information of the input data.

[0094] The prediction context explains the risk level detected by the model and the main influencing factors, enabling the GPT model to consider these background information when generating comments, making the comments more targeted and reasonable.

[0095] (2)Dynamic context injection During the Transformer decoding process, the semantic vector is used as the continuous memory vector and is incorporated into the model's generation process through the K-V cache injection mechanism to ensure that the generation process is always constrained by the feature semantics. The specific injection method is implemented through the following formula:

[0096] Among them, || represents tensor concatenation, Q is the query vector, is the key vector dimension. This formula combines the semantic vector Concatenate with the original key vector K (the text content in the three - level prompt template system) and value vector V (the feature representation of the input token), and then use the concatenated vector in the attention calculation. In this way, when generating each output token, the model will consider the information carried by the semantic vector, so as to generate comments related to the semantics of the input features, and avoid generating content that is irrelevant to the features or semantically inconsistent.

[0097] This embodiment extracts key features based on the mental health prediction results and converts them into natural language, realizing the effective conversion of data into readable information. Subsequently, semantic vectors are obtained through local semantic encoding and global relationship encoding to comprehensively mine the semantic connotations of the text. Then, a three - level prompt template system is designed to construct prompt templates, providing clear guidance for comment generation. In this process, the semantic vectors are incorporated into the generation link to constrain the generated content to be semantically related to the features. The finally generated comments are highly targeted, providing valuable interpretable decision - making support for psychological counselors and student workers, and strongly promoting the efficient progress of students' mental health intervention and guidance work.

[0098] This embodiment addresses the challenging problem of difficult negative sample definition in the student mental health prediction task, and proposes a mental health modeling method based on multi - view semi - supervised learning. Given a limited number of positive samples and a large number of unlabeled samples, an algorithm framework combining a two - stage strategy and multi - view collaborative training is designed based on PU learning. In the first stage, reliable negative samples are selected, and based on the smoothness assumption, several negative samples with the largest difference from the positive samples are screened out from the unlabeled samples to construct a labeled data set containing positive and negative samples. In the second stage, a multi - view collaborative training framework is designed, and the models obtained from 6 different view data are jointly trained in the same framework. By leveraging the information complementarity of different view data, the unlabeled samples are used to iteratively update the models of different views. Experimental results on real - world data sets show that the proposed method outperforms 6 traditional supervised learning methods and 4 commonly used ensemble methods. This embodiment not only provides a relatively effective modeling solution for student mental health prediction, but also offers new ideas for dealing with similar small - sample and weakly - labeled problems.

[0099] Embodiment 2 This embodiment provides a student mental health prediction system based on multi - view semi - supervised learning, including: A data acquisition unit, configured to obtain a basic sample set; the basic sample set is a positive sample set containing students' educational feature vectors and mental health labels and an unlabeled sample set to be predicted, where the educational feature vectors include multiple items of educational data; A prediction unit, configured to input a basic sample set into a mental health prediction model, where the mental health prediction model includes a reliable negative sample selection module and a multi-view collaborative training module; the reliable negative sample selection module is used to screen out a reliable negative sample set according to the maximum mean distance between a positive sample set and an unlabeled sample set to be predicted; the multi-view collaborative training module includes classifiers corresponding to multiple educational data respectively, and is used to extract the educational data and labels in the positive sample set and the reliable negative sample set, train the classifiers, perform mental health prediction on the unlabeled sample set to be predicted, and comprehensively combine the pseudo-labels of each classifier based on a cost-sensitive learning mechanism to obtain the mental health prediction result of the unlabeled sample to be predicted; A comment generation unit, configured to generate a comment for the mental health prediction result by using an improved generative language model.

[0100] Embodiment III This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a method for predicting students' mental health based on multi-view semi-supervision as described in Embodiment I above.

[0101] Embodiment IV This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a method for predicting students' mental health based on multi-view semi-supervision as described in Embodiment I above.

[0102] The steps or modules involved in Embodiments II to IV above correspond to those in Embodiment I, and the specific implementation manners can refer to the relevant description part of Embodiment I. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0103] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for predicting students' mental health based on multi-view semi-supervision, characterized in that, Including: Obtain a basic sample set; the basic sample set is a positive sample set containing student education feature vectors and mental health labels and an unlabeled sample set to be predicted, where the education feature vectors contain multiple education data; Input the basic sample set into a mental health prediction model, the mental health prediction model includes a reliable negative sample selection module and a multi-view collaborative training module; the reliable negative sample selection module is used to screen out a reliable negative sample set according to the maximum mean distance between the positive sample set and the unlabeled sample set to be predicted; the multi-view collaborative training module includes classifiers corresponding to multiple education data respectively, which are used to extract the education data and labels in the positive sample set and the reliable negative sample set, train the classifiers, perform mental health prediction on the unlabeled sample set to be predicted, and comprehensively integrate the pseudo-labels of each classifier based on the cost-sensitive learning mechanism to obtain the mental health prediction result of the unlabeled sample to be predicted; Use an improved generative language model to generate comments for the mental health prediction result.

2. The method for predicting the mental health of students based on multi-view semi-supervised learning according to claim 1, wherein The mental health label of the positive sample set is that there are mental health problems.

3. A method for predicting students' mental health based on multi-view semi-supervised learning as claimed in claim 1, wherein, The multiple education data include demographic information, cadre appointment information, competition award information, academic performance information, comprehensive evaluation score ranking, and extracurricular activity credits.

4. The method for predicting students' mental health based on multi-view semi-supervised learning according to claim 1, characterized in that, The screening out of a reliable negative sample set according to the maximum mean distance between the positive sample set and the unlabeled sample set to be predicted specifically includes: Respectively extract the education feature vectors of the positive sample set and the unlabeled sample set to be predicted, calculate the Euclidean distance between the two education feature vectors and find the average value; Sort all the unlabeled samples to be predicted in descending order according to the average value, and select the first positive integer samples as reliable negative samples to form a data set.

5. The method for predicting students' mental health based on multi-view semi-supervised learning according to claim 1, wherein The extracting of the education data and labels in the positive sample set and the reliable negative sample set, training the classifiers, and performing mental health prediction on the unlabeled sample set to be predicted specifically includes: Combine the positive sample set and the reliable negative sample set into a labeled training data set; For the classifier corresponding to each item of education data, extract the education data and labels of this item from the labeled training data set, and train each classifier; Use the classifier with the current error rate lower than the error rate of the previous round as an auxiliary classifier to predict the unlabeled sample set to be predicted and generate pseudo-labels; When the error rates of all classifiers no longer decrease in the current iteration, the training is completed.

6. The method for predicting students' mental health based on multi-view semi-supervised learning according to claim 1, wherein, The cost-sensitive learning mechanism includes using a cost-sensitive threshold determination to screen pseudo-labels, specifically: wherein, is the prediction result, n is the total number of classifiers, and are respectively the numbers of classifiers that predict the unlabeled samples to be predicted as positive samples and negative samples in other classifiers; The cost-sensitive learning mechanism also includes using a cost-sensitive weighted information entropy to screen pseudo-labels, specifically: For the positive sample and negative sample categories, calculate the total misclassification cost respectively; Normalize the total misclassification cost of each category to obtain a normalized weight; Calculate the weighted entropy using the normalized weight; Calculate the information gain according to the weighted entropy, assign pseudo-labels to the unlabeled samples according to the information gain situation, and add them to the labeled sample set of the corresponding view for subsequent model training.

7. The method for predicting students' mental health based on multi-view semi-supervised learning according to claim 1, wherein The using of an improved generative language model to generate comments for the mental health prediction result specifically includes: Based on the mental health prediction result, extract the key features in the input student education feature vector and convert them into natural language descriptions; Extract semantic vectors for natural language descriptions based on local semantic encoding and global relationship encoding; Design a three-level prompt template system according to the natural language description and construct a prompt word template; Generate mental health status comments based on the prompt word template and the semantic vectors.

8. A student mental health prediction system based on multi-view semi-supervised learning, characterized in that, Including: A data acquisition unit configured to acquire a basic sample set; the basic sample set is a positive sample set containing student education feature vectors and mental health labels and an unlabeled sample set to be predicted, where the education feature vectors include multiple items of education data; A prediction unit configured to input the basic sample set into a mental health prediction model, the mental health prediction model including a reliable negative sample selection module and a multi-view collaborative training module; the reliable negative sample selection module is used to screen out a reliable negative sample set according to the maximum average distance between the positive sample set and the unlabeled sample set to be predicted; the multi-view collaborative training module includes classifiers corresponding to multiple items of education data respectively, and is used to extract the education data and labels in the positive sample set and the reliable negative sample set, train the classifiers, perform mental health prediction on the unlabeled sample set to be predicted, and comprehensively obtain the mental health prediction results of the unlabeled sample to be predicted based on the cost-sensitive learning mechanism and the pseudo-labels of each classifier; A comment generation unit configured to generate comments for the mental health prediction results by using an improved generative language model.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in a method for predicting students' mental health based on multi-view semi-supervised learning as described in any one of claims 1-7.

10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a method for predicting students' mental health based on multi-view semi-supervised learning as described in any one of claims 1-7.