Interpretable depression tendency recognition system based on large language model and first-order logic

By combining large language models and first-order logical rules, high accuracy and interpretability of depression tendency recognition in the field of mental health is achieved, and the transparency of deep learning models in this field is solved, which improves the transparency of the system and user trust.

CN120376095APending Publication Date: 2025-07-25ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510261745.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing deep learning models lack interpretability in the field of mental health, especially in the recognition of depression tendencies, and it is difficult to transparentize their decision-making processes. There are limitations and high computational complexity in existing interpretable methods such as LIME and SHAP.

Method used

Combining large language models and first-order logical rules, through input layer data preprocessing and encoding, model layer depth representation and symbolic representation are integrated, decision-making layer dynamic alignment and logical reasoning, and output layer transparent prediction results, traceable decision-making reasoning paths are generated to realize depression tendency recognition.

Benefits of technology

It significantly improves the accuracy and interpretability of depression tendency recognition, reaching an accuracy rate of 98.86%, reducing the cost of rule base updates, and enhancing the transparency of the system and user trust.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120376095A_ABST
    Figure CN120376095A_ABST
Patent Text Reader

Abstract

The invention discloses an interpretable depression tendency recognition system based on a large language model and first-order logic, which comprises an input layer, a model layer, a decision-making layer and an output layer, wherein the input layer is used for acquiring a depression tendency data set and psychological health symptom description data generated by the large language model, and performing format conversion; the model layer is used for representing data based on a deep learning model and a symbol model of first-order logic; the decision-making layer is used for aligning results of the deep learning model and the symbol model through an alignment operator, representing the alignment operator, reasoning and mapping the alignment operator to a symbol space to realize representation of logic antecedents, combining the logic antecedents by using a clustering operator, and accurately estimating logic consequent according to information of the logic antecedents to obtain a prediction result; and the output layer is used for presenting the prediction result of the decision-making layer and the symptom-based first-order logic rule to a user. According to the method, a traceable decision reasoning path is generated by utilizing a first-order logic rule, and the interpretability requirement of a user is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models, and in particular to an interpretable depression tendency recognition system based on a large language model and first-order logic. Background Art

[0002] In today's society, college students are facing unprecedented academic pressure, employment competition, interpersonal relationship challenges, and a rapidly changing social environment, all of which affect their mental health. Therefore, it is particularly necessary to use deep learning models to study college students' mental health. With its powerful data processing capabilities and pattern recognition accuracy, deep learning models can mine complex features and patterns associated with mental health from massive amounts of psychological assessment data, social media behavior records, physiological indicators and other multi-dimensional information. This not only helps to detect potential psychological problems early and provide a scientific basis for timely intervention and support, but also promotes a deep understanding of the dynamic changes in college students' mental health, laying the foundation for the development of more effective prevention and intervention strategies.

[0003] There is an urgent need to study an interpretable intelligent system for college students' mental health. With the expansion of college enrollment and the increase in academic pressure, college students' mental health problems have become increasingly prominent and have become an important factor affecting students' growth and social stability. Traditional mental health assessment and intervention methods often rely on manual work, which is not only limited in resources, but also difficult to be timely, comprehensive and personalized. Therefore, it is particularly important to develop an intelligent system that can automatically, efficiently and accurately identify and intervene in college students' mental health problems. More importantly, such a system must be highly interpretable. Mental health assessment and intervention is a highly sensitive and complex field, and the system's decisions and recommendations must be understood and accepted by professionals, students and their parents. An interpretable intelligent system can clearly show the logic and basis behind its decision-making, enhance users' trust in the system, and thus improve the actual application effect and acceptance of the system. In addition, interpretability can also help reveal the deep-seated causes and laws of mental health problems, provide valuable research materials for scientific researchers, and promote theoretical development and practical innovation in the field of mental health. Therefore, studying an interpretable intelligent system for college students' mental health is not only the key to improving the quality and efficiency of college students' mental health services, but also an important way to promote the progress of mental health scientific research.

[0004] Disease identification is an important task in the field of healthcare, and its accuracy is directly related to the health of patients and the effectiveness of treatment plans. Among the many disease identification methods, linear regression, decision tree, logistic regression and KNN (K-NearestNeighbors) are widely used due to their interpretability and practicality.

[0005] Deep learning has achieved remarkable results in the field of artificial intelligence, especially in image recognition, natural language processing, computer vision and other aspects. However, the black-box nature of deep learning models has gradually emerged, making model interpretation and interpretability become increasingly important. According to the relationship between interpretability methods and decision-making models, the existing interpretable deep learning algorithms can be divided into two categories: model-agnostic explanations and model-intrinsic explanations. (1) Model-agnostic explanations are also known as post-hoc explanations, mainly including interpretable methods such as LIME (Local Interpretable Model-agnostic Explanations) and SHAP (SHapley Additive exPlanations). Although LIME can provide local explanations, the explanations of LIME are based on local regions and may not represent the behavior of the global model. The calculation of SHAP values assumes that features are independent, but in practical applications, there are often correlations between features, which may lead to inaccurate explanations. In addition, the calculation complexity of SHAP values is relatively high, especially for large datasets and complex models. (2) Model-intrinsic explanations are also known as self-explanations, mainly including interpretable methods such as the Attention mechanism and neuro-symbolic models. The interpretability of the Attention mechanism is limited. Although the Attention mechanism can show the parts that the model focuses on, it does not directly provide a complete explanation of the model's decision-making process. Neuro-symbolic methods also face the following challenges: Fusion difficulty: How to effectively fuse symbolic knowledge with neural networks and maintain the advantages of both is a difficult problem. Interpretability enhancement: Although neuro-symbolic methods combine the interpretability of symbolic reasoning, how to further improve the overall interpretability of the model is still a challenge, especially in high-risk decision-making fields.

[0006] By providing specific input prompts for pre-trained models, the models can be guided to perform better on specific tasks. These input prompts are usually a piece of text used to guide the model to generate the required output. Among them, Zero-shot Learning can achieve guiding the model to complete tasks only through natural language instructions without giving the model any examples. Recently, a series of studies have emerged to provide external knowledge through the prompt engineering of large models, thereby enhancing the interpretability of the models. At present, how to effectively integrate the external knowledge generated by large models with the internal mechanism of decision-making models is still an unsolved problem. More fusion methods need to be explored. In addition, with the wide application of Prompt technology, the interpretability of the models becomes more and more important. How to make the decision-making models maintain good interpretability while introducing external knowledge is very necessary in high-risk decision-making fields.

[0007] In interpretable artificial intelligence, first-order logic rules are used to build knowledge bases for knowledge representation and reasoning. Through first-order logic rules, an artificial intelligence system can clearly show its decision-making process, enabling users to understand the system's behavior and supervise and correct it. As the model parameters increase and the model becomes more complex, first-order logic rules play an important role in interpretable artificial intelligence. It not only provides a clear and precise way of knowledge representation for artificial intelligence systems but also realizes a traceable reasoning process, enhancing the interpretability of the system.

[0008] Since deep learning models (especially pre-trained language models) are complex non-linear structures with millions of parameters, it is difficult to explain their decisions. Neuro-symbolic models based on rule learning are an interpretive technique tailored for deep neural networks, which learn logical rules to explain the internal decision-making logic of the model. Currently, most applications of deep learning algorithms in the field of mental health are mainly black-box algorithms. In the field of mental health, especially in the field of depression tendency recognition, it is still unknown how to fuse symbolic models based on first-order logic language with pre-trained models to establish an interpretable depression recognition system. Summary of the Invention

[0009] The purpose of the present invention is to propose an interpretable depression tendency recognition system based on large language models and first-order logic, which deeply integrates the advantages of large language models and first-order logic, aiming to judge whether a sample has a tendency to suffer from depression by learning first-order logic rules based on depressive symptoms.

[0010] The purpose of the present invention is achieved through the following technical solutions: An interpretable depression tendency recognition system based on large language models and first-order logic, the system includes:

[0011] An input layer, which is used for data preprocessing and encoding, obtaining a depression tendency dataset and mental health symptom description data generated by a large language model, and performing format conversion;

[0012] A model layer, which is used for the fusion of deep representation and symbolic representation, extracting features from the data of the input layer based on a deep learning model, and based on a symbolic model of first-order logic, symbolically representing the input data using first-order logic predicates based on symptoms to construct a symbolic space;

[0013] A decision layer, which is used for dynamic alignment and logical reasoning, aligning the results of the deep learning model and the symbolic model through an alignment operator, performing characterization and reasoning mapping on the alignment operator to the symbolic space to realize the representation of the logical antecedent, using an aggregation operator to combine the logical antecedents, and accurately estimating the logical consequent according to the information of the logical antecedents to obtain a prediction result;

[0014] The output layer is used for predicting results and rule explanations, presenting the prediction results of the decision-making layer and the first-order logic rules based on symptoms to the user.

[0015] Further, in the input layer, the depression tendency dataset contains actual case sample information, which is divided into a training set and a test set. And a symptom set containing instance descriptions is constructed based on a large language model, and both the samples and the symptom set are tokenized.

[0016] Further, the deep learning model converts the data after format conversion into dense vectors of a fixed size, and inputs the vectors obtained from the embedding layer into the language model Bert trained based on the transformer architecture for characterization, extracting feature information in the high-dimensional space.

[0017] Further, the symbolic model based on first-order logic uses first-order logic predicates to symbolically represent the input data and the set of symptom instance descriptions, and uses predicates to represent the positive and negative correlation relationships between samples and symptoms respectively.

[0018] Further, the specific process of aligning the results of the deep learning model and the symbolic model through the alignment operator is as follows: align the high-dimensional feature representation of the deep learning model with the low-dimensional predicate representation of the first-order logic symbolic model through the alignment operator; use logical inference rules to map the results of the alignment operator to the predicate set in the symbolic space to obtain the symptom predicate list after reasoning.

[0019] Further, the alignment operator is used to represent the dot product of two feature vectors, which is related to their lengths and angles in the high-dimensional space. By the alignment operator, the angle between the feature vectors in the high-dimensional space is inferred to evaluate their semantic similarity relationship.

[0020] Further, the specific process of combining the logical antecedents using the aggregation operator is as follows: for each input sample, linearly aggregate the relevant symptom predicates and truth values. The disjunctive form of the symptom predicates of the input sample is the antecedent part of the logical rule; obtain the feature representation of each sample, and calculate the weighted sum of the symptom relationships of all samples in the feature space to obtain the weighted sum vector of symptom relationships.

[0021] Further, the feature representation of each sample is obtained using the conjunctive normal form CNF, and linear aggregation is performed based on the aggregation operator.

[0022] Further, the specific process of accurately estimating the logical consequent based on the information of the logical antecedent is as follows: calculate the combined result of the logical antecedent using the successor operator based on the sigmoid function, and based on the calculation result, combine with the logical antecedent corresponding to the input sample for reasoning to obtain the final prediction result and the corresponding logical consequent.

[0023] Furthermore, the output layer presents the prediction result and the reasoning process to the user through the calculation process in the feature space and the reasoning process in the symbol space in the decision-making layer, and presents the first-order logic rules in the form of a visualized graph.

[0024] Advantages of the present invention:

[0025] (1) In terms of interpretability, the first-order logic rules are used to generate a traceable decision-making and reasoning path, making the internal decision-making process of the system transparent and meeting the interpretability requirements of users.

[0026] (2) In terms of performance, by combining large language model technology and first-order logic rule learning, compared with traditional interpretable technologies (such as interpretable enhanced machine models, KNN, post-hoc interpretable methods, etc.), the accuracy of depression tendency recognition is significantly improved (experimental data shows that the accuracy rate reaches 98.86%).

[0027] (3) By constructing a symptom set with the help of large language model technology, the present invention can automatically learn first-order logic rules during the model decision-making process, without the need to frequently update the rule base manually, which reduces costs and improves flexibility at the same time. Description of the drawings

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0029] Figure 1 It is a schematic diagram of an interpretable depression tendency recognition system based on a large language model and first-order logic rules of the present invention.

[0030] Figure 2 It is a visualization graph of the symptoms of test sample X of the present invention 15 showing a positive correlation, where the yellow circles represent symptoms of normal or positive mental states, the blue circles represent symptoms of negative mental states, and the red arrows represent positive correlation relationships.

[0031] Figure 3 It is a visualization graph of the symptoms of test sample X of the present invention 15 showing a negative correlation, where the yellow circles represent symptoms of normal or positive mental states, the blue circles represent symptoms of negative mental states, and the green arrows represent negative correlation relationships.

[0032] Figure 4 It is a visualization graph of the symptoms of test sample X of the present invention 61Visualization diagram of symptoms showing a positive correlation, where yellow circles represent symptoms with normal or positive mental states, blue circles represent symptoms with negative mental states, and red arrows represent positive correlation relationships.

[0033] Figure 5 is the test sample X of the present invention 61 Visualization diagram of symptoms showing a negative correlation, where yellow circles represent symptoms with normal or positive mental states, blue circles represent symptoms with negative mental states, and green arrows represent negative correlation relationships. Detailed implementation manners

[0034] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described below with reference to the accompanying drawings and implementation examples. It should be understood that the specific implementation examples described herein are only used to explain the present invention and are not used to limit the present invention.

[0035] An interpretable depression tendency recognition system based on large language model and first-order logic provided by the present invention (Interpretable Depression Tendency Recognition Framework Based on Large Language Model and First Order Logic Rule, LF-IDTR) deeply integrates the advantages of large language model and first-order logic, aiming to accurately judge the potential depression tendency of input samples by learning first-order logic rules based on depression symptoms. Among them, the data describing depression tendency symptoms is generated by a large model (Baidu Wenxin Yiyan). As Figure 1 shown, LF-IDTR mainly consists of an input layer, a model layer, a decision layer and an output layer.

[0036] I. Input layer: Data preprocessing and encoding

[0037] The input layer mainly includes two types of key data as inputs: one type is from the real-world college student depression tendency dataset, which contains rich actual case information; the other type is the mental health symptom descriptions generated by an advanced large language model. These descriptions, with their high generalization and accuracy, provide rich semantic backgrounds for the model. Through the preprocessing process, the present invention converts these data into a format that can be used by deep learning models, laying a solid foundation for subsequent analysis and recognition.

[0038] Specifically, the present invention first obtains the publicly available college student depression tendency dataset on the Kaggle platform and randomly divides it into a training set and a test set in a ratio of 8:2 to ensure the rigor and effectiveness of model training and evaluation. Let the dataset be D = {X, Y}, where X represents the set of input samples, and Y corresponds to the labels of the samples, that is, the judgment results of depression tendency.

[0039] To more deeply understand and process these data, the present invention uses the Tokenizer of the Bert model to tokenize each input sample, converting X into the form of tokens, i.e., X'.

[0040] In addition, referring to the reference (Uddin, Md Zia, et al. "Deep learning for prediction of depressive symptoms in a large textual dataset." Neural Computing and Applications 34.1 (2022): 721 - 744.), and combining with the prompt technology of the Baidu Wenxin Yiyan large model, the present invention constructs a symptom set S for mental health. This symptom set contains 27 symptom instance descriptions, i.e., S = [S1, S2, S3,... S 26 , S 27 . Each instance description precisely depicts different aspects of the depressive tendency. Similarly, the present invention also uses the Tokenizer of the Bert model to encode these symptom instance descriptions, obtaining the tokenized S', enabling it to participate in the subsequent processing of the model in the form of tokens.

[0041]

[0042]

[0043]

[0044]

[0045]

[0046] Table 1 Symptom set and its description in mental health

[0047] Specifically, as a key component of the input layer, the symptom set is crucial for the overall performance of the framework. However, traditional symptom sets are often based on extensive data collection and statistics, and may not accurately reflect the symptom manifestations of specific domain populations. Therefore, the present invention utilizes domain knowledge in the field of psychology (i.e., (Uddin, Md Zia, et al. "Deep learning for prediction of depressive symptoms in a large textual dataset." Neural Computing and Applications 34.1 (2022): 721-744.) and the massive knowledge retrieval ability of large language models to generate a more general symptom set, thereby reducing the costs of expert retrieval and annotation. Through these human - understandable mental health symptoms, users can be helped to understand the basis and logic of model decisions, thus enhancing the credibility and acceptability of the model.

[0048] II. Model Layer: Fusion of Deep Representation and Symbolic Representation

[0049] The model layer is the core part of the framework, combining two technologies: deep learning and first - order logic. In the deep learning model part, through a complex neural network structure, the input data is deeply mined and represented, extracting feature information in a high - dimensional space. And in the symbolic model part based on first - order logic, the input data is symbolically represented using first - order logic predicates to construct a clear and interpretable symbolic space. The fusion of these two parts not only improves the representation ability of the model but also provides diversified information support for subsequent decision - making.

[0050] The model layer mainly includes two parts; the first part is the deep learning model part. Through a complex neural network structure, the tokenized input data X’ and S’ are respectively represented to extract feature information in a high - dimensional space. First, the present invention uses an embedding layer to convert the discrete token sequences X’ and S’ into dense vectors of a fixed size, namely and Then, the vectors obtained from the embedding layer are input into the language model Bert trained based on the transformer architecture for representation, obtaining the final feature representations of the input data, namely and

[0051] In the second part of the model layer, the present invention innovatively introduces a symbolic model to achieve the representation and analysis of input data in the symbolic space. Specifically, in the symbolic model part, the input sample data and the symptom set are symbolically represented by first-order logic predicates. This method not only enhances the interpretability of the model but also enables the model to more intuitively understand the semantic relationship between the sample and the symptom. Specifically, the present invention defines predicates in the symbolic model part, and symbolically represents the input sample data and the symptom set through first-order logic predicates. Using the predicate Re_to_Sym(X i ,S j ) to represent the positive correlation between sample X i and symptom S j . Correspondingly, is used to represent the negative correlation between sample X i and symptom S j . Combining with the meanings of the symptom sets detailed in Table 1, the semantic connection between the sample and the specific mental health symptoms can be clearly interpreted. For example, Re_to_Sym(X1,S6) indicates that the input sample X1 is positively correlated with symptom 6 (sadness); indicates that the input sample X1 is negatively correlated with symptom 15 (positive attitude).

[0052] III. Decision-making layer: Dynamic alignment and logical reasoning

[0053] The decision-making layer is the key link for the framework to achieve intelligent judgment. First, through the dynamic alignment module, the present invention precisely aligns the high-dimensional representation of the deep learning model with the low-dimensional representation of the first-order logic symbolic model through the alignment operator. Subsequently, inference rules are used to represent and reason about the results of the alignment operator and map them to the symbolic space to achieve the representation of the logical antecedent. On this basis, the CNF (conjunctive form) aggregation operator is further used to combine the logical antecedents. Finally, through the logical consequent estimation module, the present invention accurately estimates the logical consequent based on the information of the logical antecedent, thereby obtaining the prediction result of the model. This process not only reflects the rigor of logical reasoning but also demonstrates the flexibility of deep learning in solving complex problems.

[0054] The first module of the decision-making layer is the dynamic alignment module: First, through the alignment operator match(·), the high-dimensional feature representation of the deep learning model (i.e., and ) is aligned with the low-dimensional predicate representation Re_to_Sym(X i ,S j ) of the first-order logic symbolic model. Specifically, the alignment operator match(·) is defined as follows:

[0055] M = Match(X,S)

[0056] In the feature space, the Match(·) operator is used to represent the dot product of two feature vectors, i.e., M i,j ∈M represents the feature vector of X i and the corresponding element product sum of the feature vector of S j Since the dot product of feature vectors is related to their lengths and angles in the high-dimensional space, the angle between feature vectors in the high-dimensional space can be inferred through the Match(·) operator, so as to evaluate their semantic similarity relationship. Specifically, the following logical inference rules are adopted to map the result of the alignment operator Match(·) to the predicate set Re_to_Sym(·,·) in the symbol space, and the specific operations are as follows:

[0057]

[0058] where the predicate Posit(M i,j ) indicates that the element M i,j is a positive integer, indicates that the element M i,j is not a positive integer. In the decision layer processing, the symptom predicate list after reasoning is represented by Re_P as follows:

[0059] Re_P = [[Re_to_Sym(X1,S1),…,Re_to_Sym(X1,S 27 )],

[0060] [Re_to_Sym(X2,S1),…,Re_to_Sym(X2,S 27 )],

[0061] …………

[0062] [Re_to_Sym(X n ,S1),…,Re_to_Sym(X n ,S 27 )]]

[0063] where Re_P[i] represents the symptom subset related to the i-th sample in the list, i.e., Re_to_Sym(X i ,·).

[0064] The second module in the decision layer is the aggregation module based on CNF: for the symptom predicates and truth values related to each input sample, a linear aggregation is performed using an aggregation operator based on CNF (conjunctive normal form) as follows:

[0065]

[0066] where A idenotes the disjunctive formula of symptoms predicates for the i-th input sample, that is, the antecedent part of the logical rule. In the present invention, the symbol is used to represent the logical antecedents of all input samples. Similarly, in the feature space, the aggregation module based on CNF can obtain the vector representation of each sample as:

[0067]

[0068] Through the above feature representation, the symptom relationship values of all samples can be weighted and calculated in the feature space. In the present invention, the symbol is used to represent the weighted sum vector of symptom relationships of all input samples.

[0069] The third module in the decision layer is the logical consequent estimation module: in the logical consequent estimation module, through the successor operator Estim(·) based on the sigmoid function, the symptom relationship weighted sum vector A T is calculated to obtain the final prediction result and the corresponding logical consequent. The specific calculation process is as follows:

[0070]

[0071] where w and b are two parameters in the training process, the AMax function represents the candidate label that makes the sigmoid function reach the maximum value, and the calculation of the sigmoid function is Through the calculation result of the above Estim(·), combined with the logical antecedent corresponding to the input sample, the following reasoning process can be obtained:

[0072]

[0073] where Y e represents the set of prediction labels for all input samples. During the training process, the prediction probability and the given label Y are calculated through the cross-entropy loss function L:

[0074]

[0075] where Y i ∈Y represents the true label of the i-th input sample, represents the probability that the model predicts the i-th input sample as the positive class (that is, having a depressive tendency); is the output data calculated through in the successor operator Estim(·).

[0076] IV. Output layer: Prediction result and rule explanation

[0077] The output layer is responsible for presenting the prediction results of the decision-making layer and the first-order logic rules based on symptoms to the user in an intuitive and understandable manner. Through the learned first-order logic rules, not only can accurate depression tendency prediction results be provided to the user, but also the logical rules on which the model makes judgments can be shown. This transparent output method not only enhances the interpretability of the model but also improves the user's trust in the model's decision-making.

[0078] Through the calculation process in the feature space and the reasoning process in the symbol space in the decision-making layer, the present invention can not only obtain the prediction results of the model but also the reasoning process of the model. Therefore, in the output layer, not only the prediction results corresponding to the input samples are output, but also the first-order logic rules of the reasoning process are output. To provide a more intuitive and understandable explanation to the end user, the first-order logic rules are presented in the form of a visualization graph in combination with Table 1. Specific cases can be referred to in the experimental results section.

[0079] The present invention evaluates the performance of LF-IDTR on the publicly available college student depression tendency dataset. The results show that compared with other interpretable methods, the LF-IDTR proposed by the present invention can achieve excellent performance while achieving high interpretability. In addition, the present invention conducts two case studies to illustrate how HDRL provides reliable logical explanations for its prediction results.

[0080] The present invention uses the publicly available college student anxiety and depression dataset on Kaggle (https: / / www.kaggle.com / datasets / sahasourav17 / students-anxiety-and-depression-dataset / data). This dataset is in Excel format and includes approximately 6,500 data from social media, Facebook comments, posts, etc. The data is annotated by fou. All those selected for data annotation are undergraduate students with good English. The dataset is processed to remove blank line data. The data is randomly divided, and 5,574 training samples and 1,395 test samples are extracted.

[0081] To evaluate the performance of the interpretable depression tendency recognition framework proposed in the present invention on real data, the present invention mainly makes comparisons with other interpretable methods, including intrinsically interpretable methods and post-hoc interpretable methods. In the experiments of the present invention, consistent settings were established on all datasets: the batch size was 128, and the size of the word embedding (embedding_dim) was 300. For the Bert model used in the model layer, the present invention adopted the pre-trained model proposed in the literature (Devlin J. Bert: Pre-training of deep bidirectional transformers for language understanding[J]. arXiv preprint arXiv:1810.04805, 2018.), where the size of the hidden vector (feature vector) was 128. The framework of the entire model was trained using the default learning rate of the Adam optimizer in Keras (i.e., learning_rate = 0.001). To ensure a fair comparison of model performance, both the Bert+LIME and Bert+SHAP baseline models and the system of the present invention were trained using the cross-entropy loss function and the Adam optimizer. In addition, during the training process of the model, the present invention set the maximum number of iterations to 200 and used early stopping to manage the training process, where the hyperparameters were set to min_delta = 1e-4 and patient = 6. All experiments were conducted on a Linux (Ubuntu)-based workstation with GTX 2080Ti and RTX 3090 GPUs, and all implementations were in the Python3 programming language. In addition, all experiments were run five times, and the present invention reported the average performance. For the large model experiments in the input layer, the present invention used the domestic large model Baidu Wenxin Yiyan to generate some symptoms and instance-based symptom descriptions.

[0082] The experimental results of the present invention mainly include two parts: model performance results and interpretable results, where the interpretable results are demonstrated through two cases. To ensure the robustness and reliability of the system of the present invention and the baseline in noisy high-risk scenarios, the present invention uses weighted precision, weighted recall, and weighted F1 score as evaluation metrics to evaluate model performance. The evaluation results of the present invention and other interpretable methods on the depression tendency dataset are shown in Table 2:

[0083] Table 2

[0084]

[0085]

[0086] As shown in Table 2, the method of the present invention (i.e., LF-IDTR) can show significant competitiveness compared with all benchmark methods in the interpretable depression tendency recognition task. Generally speaking, the method of the present invention has achieved the best performance in all evaluation metrics. When focusing on the interpretable benchmark models, the interpretable boosted machine (EBM) model has shown better performance than other benchmark models in three evaluation metrics. However, it is obvious that the performance of all interpretable benchmark models is significantly lower than the method proposed by the present invention, highlighting the advantages of the method of the present invention in providing high performance and high interpretability. Further, when the present invention conducts a comparative analysis of the LDF method with post-hoc interpretable methods (such as LIME and SHAP), an interesting phenomenon is observed: the performance of the LIME and SHAP methods is consistent with the performance of the original model (Bert). This finding is attributed to the inherent characteristics of post-hoc interpretable methods, that is, as model-agnostic techniques, they do not make any changes to the original structure of the model.

[0087] The interpretable results are as follows:

[0088] Case 1: Test sample X 15 The interpretable results are shown in Table 3:

[0089] Table 3

[0090]

[0091]

[0092] In this case study, the present invention uses the real sample X in the test data of the college student depression tendency dataset 15 to demonstrate the interpretable results of the LF-IDTR framework. As shown in Table 3, LF-IDTR allows exploring the internal principle of the model framework through the neuro-symbolic basis. According to the learned logical rules, users can easily understand the model decision-making process, that is, the test sample X 15 is recognized as having no depression tendency by the framework because the semantic meaning of this sample is positively correlated with symptom 2, symptom 7, symptom 13, symptom 14, symptom 15, symptom 16, symptom 21, symptom 24, symptom 25, and negatively correlated with other symptoms.

[0093] To enable end-users to have a more intuitive understanding of the interpreted symptoms, the present invention combines the symptom set generated by the large language model in the input layer to show the symptom visualization diagram positively correlated with the test sample X 15 (as shown in Figure 2 ) and the symptom visualization diagram negatively correlated with the test sample X 15 (as shown in Figure 3 ). According to these two visualization result diagrams, it can be seen that the test sample X 15It is mainly positively correlated with symptoms indicating normal or positive mental states such as "optimistic, positive attitude, emotional stability, contentment, mindfulness", etc.; and negatively correlated with symptoms indicating negative mental states such as "not caring about anything, always being very tired, sad, crying", etc. In addition, it is observed that in this example, there are 2 symptoms indicating negative mental states that are positively correlated with test sample X 15 and there is a part of the symptoms indicating normal or positive mental states that are negatively correlated with test sample X 15 . This phenomenon may be due to the unclear relationship between the semantic features of the test sample and these symptoms, resulting in errors in the dynamic alignment module. To deeply study this reason, in Case 2, a test sample identified by the framework as having a depressive tendency was selected for comparative analysis.

[0094] Case 2: Test sample X 61 The interpretable results are shown in Table 4 as follows:

[0095] Table 4

[0096]

[0097]

[0098] In this case study, the present invention uses the real sample X in the test data of the college student depressive tendency dataset 61 to demonstrate the interpretable results of the LF-IDTR framework. As shown in Table 4, according to the logical rules learned by the LF-IDTR framework, users can easily understand the model decision-making process, that is, test sample X 61 is identified by the framework as having a depressive tendency because the semantic meanings of this sample are positively correlated with symptoms 1, 3, 4, 5, 8, 9, 10, 11, 12, 17, 18, 19, 20, 22, 23, 26, 27, and negatively correlated with other symptoms. To enable end-users to have a more intuitive understanding of the interpreted symptoms, the present invention combines the symptoms generated by the large language model in the input layer to display the visualization graph of the symptoms positively correlated with test sample X 61 (as shown in Figure 4 ) and the visualization graph of the symptoms negatively correlated with test sample X 61 (as shown in Figure 5 ). According to these two visualization result graphs, it can be seen that test sample X 61 is mainly positively correlated with symptoms indicating negative mental states such as "always being very tired, negative thinking, sad, crying", etc.; and negatively correlated with symptoms indicating normal or positive mental states such as "optimistic, positive attitude, emotional stability, contentment, mindfulness", etc. In addition, it is observed that in this example, there are 2 symptoms indicating negative mental states that are positively correlated with test sample X 61is negatively correlated, and there is a part of the symptoms indicating normal or positive mental state that are positively correlated with test sample X 61 Combined with the visualization results of Case Study 1, it is found that test sample X 15 is positively correlated with symptom 2, symptom 7, symptom 13, symptom 14, symptom 15, symptom 16, symptom 21, symptom 24, symptom 25; and these symptoms happen to be negatively correlated with test sample X 15 This shows that LF-IDTR makes decisions mainly based on the semantic correlation between the test sample and the given symptoms during the process of identifying depressive tendencies.

[0099] The above embodiments are used to explain the present invention rather than limit the present invention. Any modification and change made to the present invention within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.

Claims

1. An interpretable depression tendency recognition system based on large language models and first-order logic, characterized in that The system includes: An input layer for data preprocessing and encoding, obtaining a dataset of depression tendencies and data on mental health symptom descriptions generated by a large language model, and performing format conversion; A model layer for the fusion of deep representation and symbolic representation, extracting features from the data of the input layer based on a deep learning model, and based on a symbolic model of first-order logic, using first-order logic predicates based on symptoms to symbolically represent the input data and constructing a symbolic space; A decision layer for dynamic alignment and logical reasoning, aligning the results of the deep learning model and the symbolic model through an alignment operator, performing representation and reasoning on the alignment operator and mapping it to the symbolic space to achieve the representation of the logical antecedent, using an aggregation operator to combine the logical antecedents, and accurately estimating the logical consequent based on the information of the logical antecedent to obtain a prediction result; An output layer for prediction results and rule interpretation, presenting the prediction results of the decision layer and the first-order logic rules based on symptoms to the user.

2. The interpretable depression tendency recognition system based on large language models and first-order logic according to claim 1, wherein In the input layer, the dataset of depression tendencies contains actual case sample information, which is divided into a training set and a test set, and a symptom set containing instance descriptions is constructed based on a large language model, and both the samples and the symptom set are tokenized.

3. The interpretable depression tendency recognition system based on large language models and first-order logic according to claim 1, characterized in that, The deep learning model converts the data after format conversion into a dense vector of a fixed size, and inputs the vector obtained from the embedding layer into the language model Bert trained based on the transformer architecture for representation, extracting feature information in a high-dimensional space.

4. The interpretable depression tendency recognition system based on a large language model and first-order logic according to claim 1, wherein The symbolic model based on first-order logic uses first-order logic predicates to symbolically represent the input data and the set of symptom instance descriptions, and uses predicates to represent the positive and negative correlation relationships between samples and symptoms respectively.

5. The interpretable depression tendency recognition system based on a large language model and first-order logic according to claim 1, characterized in that, The specific process of aligning the results of the deep learning model and the symbolic model through the alignment operator is as follows: aligning the high-dimensional feature representation of the deep learning model with the low-dimensional predicate representation of the first-order logic symbolic model through the alignment operator; using logical inference rules to map the result of the alignment operator to the predicate set in the symbolic space to obtain a list of symptom predicates after reasoning.

6. The interpretable depression tendency recognition system based on a large language model and first-order logic according to claim 5, wherein, The alignment operator is used to represent the dot product of two feature vectors, which is related to their lengths and angles in the high-dimensional space. The alignment operator is used to infer the angle between the feature vectors in the high-dimensional space, thereby evaluating their semantic similarity relationship.

7. An interpretable depression tendency recognition system based on a large language model and first-order logic according to claim 1, characterized in that, The specific process of combining the logical antecedents using the aggregation operator is as follows: linearly aggregating the symptom predicates and truth values related to each input sample, and the disjunctive form of the symptom predicates of the input sample is the antecedent part of the logical rule; obtaining the feature representation of each sample, and calculating the weighted sum of the symptom relationship values of all samples in the feature space to obtain a weighted sum vector of symptom relationships.

8. The interpretable depression tendency recognition system based on a large language model and first-order logic according to claim 1, characterized in that, The feature representation of each sample is obtained using the conjunctive normal form CNF, and linear aggregation is performed based on the aggregation operator.

9. The interpretable depression tendency recognition system based on a large language model and first-order logic according to claim 1, characterized in that, The specific process of accurately estimating the logical consequent based on the information of the logical antecedent is as follows: calculating the combined result of the logical antecedent using the successor operator based on the sigmoid function, and based on the calculation result, reasoning in combination with the logical antecedent corresponding to the input sample to obtain the final prediction result and the corresponding logical consequent.

10. The interpretable depression tendency recognition system based on a large language model and first-order logic according to claim 1, characterized in that, The output layer presents the prediction result and the reasoning process to the user together through the calculation process in the feature space and the reasoning process in the symbol space in the decision layer, and presents the first-order logic rules in the form of a visualized graph result.