Tree structure-based interpretable hierarchical text classification method

By adopting a deep learning model based on tree structure in the hierarchical text classification task, combining feature extraction module and tree classification module, the problem of the lack of interpretability in hierarchical text classification is solved, and the interpretability and efficient classification of the model are realized.

CN120216698APending Publication Date: 2025-06-27BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510338213.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Deep learning models lack interpretability in hierarchical text classification tasks and are difficult to explain their decision-making process.

Method used

A interpretable deep learning model based on tree structure is adopted, and text features are extracted and hierarchical classification module is performed through a combination of feature extraction module and tree classification module. The feature extraction module uses a pre-trained language model and interprets it through visual attention weights; the tree classification module uses a tree structure to make decisions, and the prediction results of the parent node and the child node jointly determine the final classification results.

Benefits of technology

A deep learning model that provides interpretability in the hierarchical text classification task is realized. By visualizing attention weights and decision-making paths, the model's attention to text and the basis for classification decisions is clearly displayed, and the "black box characteristics" of the deep learning model are overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216698A_ABST
    Figure CN120216698A_ABST
Patent Text Reader

Abstract

The invention discloses an interpretable hierarchical text classification method based on a tree structure, and belongs to the field of hierarchical text classification of natural language processing. The method is composed of a feature extraction module and a tree classification module, and each module comprises a corresponding function part and an interpretability part. Firstly, a feature extraction module selects a proper pre-training model and is responsible for processing an input text and extracting feature codes; this stage may visualize the attention weight of the feature encoder to assess and show the degree of attention to each word. And then, the tree classification module performs final classification on the texts according to the text features of the feature extraction module. The final decision of each node is jointly decided by a father node and a child node of the tree structure. In the stage, the prediction probability of each decision node can be output, and the basis and process of final classification of the method are displayed. According to the method, the accuracy and interpretability of hierarchical text classification are improved by utilizing the pre-training model and the tree structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Based on deep learning technology, the present invention studies an interpretable hierarchical text classification method based on a tree structure. First, a suitable pre-trained language model is selected as the backbone feature encoder to process the input text and extract corresponding text features for integration with the tree-shaped classification module. At this stage, the attention weights of the feature encoder can be visualized to evaluate and display the degree of attention to each word. Secondly, at the parent node of the tree-shaped classification module, the features to be passed to the child nodes are processed and predictions are made simultaneously; the prediction results of the child nodes are jointly determined according to the results of the child nodes and the parent nodes. This hierarchical structure allows the model to use top-down decision dependencies to optimize predictions while maintaining the flexibility of classification. The present invention belongs to the field of natural language processing, and specifically designs deep learning and hierarchical text classification technologies. Background Art

[0002] With the continuous development of AI technology, deep learning-related technologies have been widely applied in the field of natural language processing. The rapid development of natural language processing technology has further promoted the progress of the field of hierarchical text classification. Currently, many tasks can be defined as hierarchical text classification tasks. For example, some works model the text classification problem in the news field as a hierarchical classification task, which can quickly determine the detailed field content of a news report. At the same time, in the field of psychology, the work of extracting cognitive distortion paths, which is crucial for the treatment of depression, can perform hierarchical text classification on the patient's speech, extract what type of cognitive distortion the patient has, and finally prescribe the right medicine according to the specific distortion type. This has immeasurable effects on the treatment in the field of mental health.

[0003] Currently, the methods of hierarchical text classification can be divided into three categories: flat, integrated, and structured methods. Among them, the flat method treats all labels as the same level, ignores the hierarchical relationship between them, and classifies the text from a flat perspective. This method ignores the hierarchical relationship between labels and is prone to errors. However, its advantage is that it is simple and convenient to implement. The integrated method slightly involves training multiple models, and each model focuses on different parts or levels of the hierarchical structure. This method benefits from the diversity of models, can capture all aspects of the hierarchical structure, and thus improves the performance of complex classification tasks. However, it also increases the computational complexity. The structured method considers the structural characteristics of the labels of hierarchical text classification at the overall level and makes decisions from top to bottom with a tree-shaped classification module. The computational complexity of this method is better than that of the integrated method and is the main research direction of current hierarchical text classification.

[0004] There have been some studies on structured methods for hierarchical text classification. Given the current black-box nature of deep learning itself, although some structured methods for hierarchical text classification can consider the hierarchical characteristics of labels at the overall level, they do not have the interpretability that is crucial in many fields. The hierarchical text classification method of the present invention can not only consider the hierarchy of labels from top to bottom at the structured level, but also explain the hierarchical classification results at the word level and classification path. Summary of the Invention

[0005] Although some large-scale pre-trained language models have been proposed in the field of hierarchical text classification, there is still a major challenge: the inherent nature of deep learning models as "black boxes" makes them difficult to interpret. The present invention proposes an interpretable deep learning model with a tree structure for hierarchical text classification tasks. The present invention is divided into two modules: a feature extraction module and a tree-shaped classification module. Both the feature encoder in the feature extraction module and the decision path in the tree-shaped classification module can be visualized for interpretation. First, the feature extraction module uses a pre-trained language model to process the input text and generate corresponding feature encodings. During this process, the attention weights can be visualized to evaluate and display the degree of attention of the model to each word. Then, the tree-shaped classification module uses the text features generated by the feature extraction module for the final classification. The decision-making process of this module is based on a tree structure, where the decision of each classification node is jointly determined by its parent node and child nodes, and the parent node and child nodes correspond to the parent class and subclass in the hierarchical classification task respectively. At this stage, the system can output the prediction probability of each decision node, thus clearly showing the basis and process of the classification decision. The specific technical details of each module will be shown below:

[0006] Step 1: Construct an interpretable hierarchical text classification model.

[0007] (1) The text for hierarchical classification first passes through the feature extraction module.

[0008] The feature extraction model mainly includes two parts: namely, the feature encoder part and the encoder interpretation part. Among them, the feature encoder can also be called the backbone function encoder, which is a pre-trained language model that can process the input text and convert it into text embeddings. Then the extracted features are integrated with the tree-structured classification module. The method of the module of the present invention is not limited to a specific encoder, it is flexible and can use various model architectures. The final predicted result is determined by the classification module, and the backbone encoder is selected according to its performance in downstream tasks and is used as a form of pre-training. By integrating the proposed high-level module, this step can also enhance the final performance.

[0009] Model interpretation refers to a specially designed technique for revealing the working principle of a black-box model, while "interpretability" refers to the property that the model itself is designed to be easy to understand and interpret. In the present invention, the encoder interpretation part in the feature extraction module aims to improve the interpretability of the Transformer-based model. This method evaluates whether the model can effectively capture the relationships between keywords crucial for downstream tasks by visualizing the attention weights of the feature encoder. Specifically, the attention mechanism of the encoder can be understood from two aspects: firstly, the word-level attention weights, which determine the degree of attention the model pays to other words when encoding a specific word; secondly, the role of attention heads, where each attention head captures different aspects of the relationships between words by assigning different attention weights. To achieve this goal, the present invention uses a toolkit called BertViz (Vig, 2019) to visualize the weights of each attention head, thus providing effective interpretation and analysis for the Transformer-based model.

[0010] (2) Tree-structured classification module

[0011] The tree-structured classification module consists of a parent node and child nodes. During implementation, a file containing the relationships between all nodes is input and automatically converted into a model structure. The parent node and child nodes have a similar architecture: a dropout layer, followed by a fully connected (FC) layer and an activation function. The difference between the parent node and child nodes is that the parent node processes the features to be passed to the child nodes and also makes predictions. Therefore, it includes a layer with ReLU activation for processing and an FC layer with a sigmoid function for decision-making. While the child nodes only need to make predictions, so their activation function is only sigmoid. The final decision of each child node is determined by combining the probabilities of the parent and child through a multiplication operation. The model supervises all nodes, so the results of the child nodes are jointly determined by the sub-decision nodes and the parent decision nodes. The choice of the loss function of the model is also flexible, and the loss functions at the parent and child levels are not necessarily the same. Appropriate loss functions can be selected respectively.

[0012] From the perspective of decision-making, the method of the present invention is also interpretable. In the present invention, the decision path interpretation part in the tree-structured classification module depends on the parent decision node and the child decision node from top to bottom. The predicted probability values of the text data at each node (including the parent node and child nodes) are shown in the form of a tree diagram, thus clearly explaining the reason and basis for the model to obtain this prediction result for the data, thereby overcoming the "black-box characteristic" of the deep learning model itself.

[0013] Step 2: Select the hierarchical text classification dataset for model application and construct the specific form of the tree - structured classification module according to the dataset. This method is applicable to hierarchical text classification tasks of any language type, any topic, and any domain. In the process of designing and implementing hierarchical text classification tasks, in order to systematically organize the classification system and facilitate subsequent model training and evaluation, this method constructs a classification rule containing the names of parent classes and sub - classes and stores it in the CSV file format. Suppose there are m parent classes and n sub - classes in a specific hierarchical text classification task, and the content writing rule for each row is shown in formula (1):

[0014] Row r =(P k ,C k,j ) #(1)

[0015] Where:

[0016] Row r is the content writing rule for each row, r is the row number, 1 ≤ r ≤ n;

[0017] Each row of the CSV file has two columns, namely the parent class P k and the sub - class C k,j ;

[0018] In the parent class P k , k is the parent - class index, 1 ≤ k ≤ m;

[0019] In the sub - class C k,j , j is the sub - class index corresponding to the k - th parent class, 1 ≤ j ≤ n k ;

[0020] n k is the number of sub - classes of the k - th parent class,

[0021] The following introduces two different hierarchical text classification tasks for method verification:

[0022] (1) Cognitive distortion path extraction task

[0023] Extracting the distortion types of cognitive paths contained in user statements belongs to the category of the psychology field. According to Beck's cognitive - behavioral therapy theory, psychological problems can be extended to a structured ABCD framework. Specifically, cognitive paths can be divided into four main categories (parent classes) and nineteen sub - categories (sub - classes). The constructed dataset is a Chinese dataset, which contains a total of 4742 sentences, divided into 2835 training - set sentences, 932 validation - set sentences, and 975 test - set sentences.

[0024] (2) Task of classifying papers into fields

[0025] The task of classifying known papers into fields also belongs to a type of hierarchical text classification. The publicly available dataset is called WOS-46985, which is an English dataset. The field scope includes 7 first-level labels (parent categories) and 134 second-level labels (sub-categories). This provides a classification challenge for the model from coarse-grained to fine-grained. The publicly available dataset contains 46,985 papers, that is, 46,985 pieces of data.

[0026] Step 3: The model is fine-tuned on the dataset to obtain a classification model more suitable for a specific field.

[0027] Taking the cognitive distortion path extraction task as an example, relevant information about the training stage of the interpretable hierarchical text classification model is introduced. The selection of the backbone network feature extractor can be flexibly chosen according to the field and language features of the dataset. Since the cognitive path extraction dataset is a Chinese psychology field dataset, the model Chinese MentalBERT is selected for the feature extraction part. This model is fine-tuned and trained on the basis of the English psychology field model MentalBERT using a Chinese psychology field corpus. The model has 110 million parameters, has 12 Transformer encoder blocks, each layer has 12 attention heads, and the total number of attention heads is 12×12 = 144. The hidden layer dimension is 768, and each input token is represented as a 768-dimensional vector. The dimension of each head is the hidden layer dimension divided by the number of attention heads, that is, 768 / 12 = 64 dimensions.

[0028] This method is implemented using an NVIDIA GeForce RTX 4090 24GB GPU combined with the PyTorch framework during the validation process, and 100 epochs are set for training. An early stopping strategy is adopted to prevent the model from overfitting. The patience of the early stopping strategy is set to 20, and the best validation model is selected for test evaluation. For the interpretable tree structure model, we set the learning rate to 3e-5, the batch size to 16, and use MSE (mean squared error) to calculate the loss. The random seed is set to 42. The model parameters can be flexibly adjusted according to different datasets. Description of the Drawings

[0029] Figure 1 Flowchart of the interpretable hierarchical text classification method based on a tree structure Detailed Implementation Manner

[0030] According to the above description, the following is a specific implementation process, but the scope protected by this patent is not limited to this implementation process.

[0031] Step 1: Construct an interpretable hierarchical text classification model.

[0032] Step 1.1: Feature Extraction Module of the Interpretable Hierarchical Text Classification Method Based on Tree Structure

[0033] Step 1.1.1: Encoder Part of the Feature Extraction Model

[0034] In this part, the hierarchical classification text is first processed by the feature extraction module. The core of the feature extraction module is the feature encoder, which is a key component of the entire module. The feature encoder can adopt various pre-trained language model architectures, such as BERT, GPT, RoBERTa. These pre-trained models can effectively extract high-dimensional semantic information from the input text and convert the original text into text feature embeddings of a fixed dimension, as shown in formula (1):

[0035] f x =M e (x)#(1)

[0036] where f x is the extracted text feature embedding, Me is the pre-trained language model, and x is the text data.

[0037] The feature encoder is responsible for processing the input text and converting it into a latent semantic representation. This part adopts the Transformer architecture, which contains multiple layers of self-attention mechanisms and feed-forward neural networks to capture long-range dependencies in the text. In specific implementation, different pre-trained models can be selected according to the text language and type. For example, for the English open-source global citation dataset WOS, BERT, RoBERTa, XLNet can be selected; for the Chinese sentiment analysis dataset, the Chinese MentalBERT model can be selected.

[0038] Step 1.1.2: Interpretability Part of the Feature Extraction Module

[0039] The core goal of the encoder interpretation part of this step is to improve the interpretability of the Transformer-based model, making its decision-making process more transparent and easy to understand. By analyzing the weights and attention mechanisms of the model, the keywords that the model focuses on when processing the input text and the relationships between them can be revealed, thus helping to understand the decision-making basis of the model.

[0040] (1) Word-level Attention Weights

[0041] By analyzing the attention mechanism of the pre-trained language model based on Transformer, the present invention reveals the focus of the model when processing input text, thereby enhancing the interpretability of the model. Specifically, this method first loads the pre-trained model and its tokenizer, and preprocesses and tokenizes the input text. Subsequently, the attention weights of each layer are extracted through model inference, and an average operation is performed on all attention heads and layers to obtain the attention weights at the token level. To focus on the actual semantic information, special tokens (such as [CLS] and [SEP]) are excluded, and each token is mapped to its corresponding attention weight. Finally, by analyzing and sorting the attention weights, the tokens that the model pays the most attention to when processing text are extracted. Experimental results show that this method can effectively reveal the key basis for the model's decision-making, provide intuitive visual support for understanding the model's behavior, and thus enhance the interpretability of the model.

[0042] (2) Attention head weights

[0043] In the multi-head self-attention mechanism, the model calculates different attention weights in parallel through multiple attention heads. Each attention head focuses on capturing different relationships between words or different semantic levels in the text. To display the weights of different attention heads, the present invention uses the BertViz toolkit to effectively display and interpret the weights of different attention heads when processing the pre-trained model based on Transformer, visually visualize the role of each attention head, and help understand how each word affects the encoding process of the entire sentence and how the model processes the input text at different levels.

[0044] Step 1.2: Tree-structured interpretable hierarchical text classification method - tree structure classification module

[0045] The present invention uses the text features extracted by the pre-trained model in Step 1 to make top-down predictions through the parent decision node and the child decision node.

[0046] Step 1.2.1: Tree-shaped classification part

[0047] (1) Parent decision node

[0048] The parent decision node processes the text features extracted by the pre-trained model and passes them to the child nodes, and also makes predictions. It includes an activation layer with a ReLU function and a fully connected layer with a Sigmoid function. The activation layer with the ReLU function is responsible for processing the feature embedding fx obtained by the feature extraction module to obtain the feature fp processed by the parent decision node, as shown in formula (2):

[0049]

[0050] Where The features extracted for the parent node, is the parent node feature conversion function, f x is the feature of the feature extraction module in step 1. Feature conversion function is the ReLu function, as shown in formula (3):

[0051]

[0052] The features after processing the parent node Inputting the fully connected layer with the Sigmoid function can get the final decision of the parent decision node As shown in formula (4):

[0053]

[0054] in is the final decision of the parent decision node, Decision transition function for the parent node, Features extracted for the parent node.

[0055] (2) Sub-decision nodes

[0056] The child decision node receives the processed features from its parent node and makes predictions based on them. The child node does not need to perform feature processing and directly applies a dropout layer and then a fully connected layer with a Sigmoid activation function to generate its prediction results. The Sigmoid activation function outputs its predicted probability value, which indicates the likelihood that the input belongs to a given category. The dropout layer acts as a feature conversion function and processes the feature fp passed from the parent decision node, as shown in formula (5):

[0057]

[0058] in is the intermediate feature obtained by the child node, is the sub-node feature conversion function, Features passed to the parent node.

[0059] The final decision of each child node is determined by combining the probability results of the parent node and the child node. This mechanism ensures that a child category can only be activated when its parent category is also active, thereby enforcing hierarchical consistency in classification. As shown in formula (6):

[0060]

[0061] in For the final prediction result, is the prediction result of the parent decision node, is the decision conversion function for child nodes, is the intermediate feature of the child node.

[0062] The decision conversion functions in formulas (3) and (5) and are both Sigmoid functions, as shown in formula (7):

[0063]

[0064] (3) Final decision

[0065] The final classification probability set is composed of the decisions of the parent node and the child node, as shown in formula (8):

[0066]

[0067] where is the prediction result of the parent node, is the prediction result of the child node. Each of them is constrained by its respective parent node . This hierarchical structure allows the model to use top-down decision dependencies to refine predictions while maintaining classification flexibility.

[0068] The choice of the loss function is also flexible. The loss functions at the parent and child levels do not require to be the same. The final loss calculation method is as shown in formula (9):

[0069]

[0070] where, w c + w p = 1. w c is the weight of the loss of the sub-category, and w p is the weight of the loss of the parent category. L final is the final loss during the model training process, is the loss of the child node part, is the loss of the parent node part. is the logits of the sub-category prediction, y c is the true label of the sub-category, is the logits of the parent-category prediction, y p is the true label of the parent category.

[0071] By combining the losses of the child node and the parent node through weighting, the importance of different levels in the hierarchical classification task is balanced. During its implementation, the parent node and the child node each support the use of multiple different loss functions, and the model is optimized through backpropagation.

[0072] Step 1.2.2: Decision path visualization part

[0073] According to the description of the tree - shaped classification module of the present invention, the model can output the prediction probabilities of each node, including the parent node and the child nodes, rather than just simple classification results. Therefore, from the perspective of decision - making, the model of the present invention is essentially interpretable, and its decision - making process is driven by a top - down hierarchical structure, relying on the weights and probability distributions of the parent nodes. This hierarchical decision - making mechanism not only improves the classification efficiency but also enhances the interpretability of the model through explicit probability distributions, enabling developers to easily understand the decision - making basis of the model and providing important references for subsequent model optimization.

[0074] Step 2: Select the hierarchical text classification dataset for model application and construct the specific form of the tree - shaped structure classification module according to the dataset.

[0075] This method is applicable to hierarchical text classification tasks of any language type, any theme, and any domain. In the process of designing and implementing hierarchical text classification tasks, in order to systematically organize the classification system and facilitate subsequent model training and evaluation, this method constructs a classification rule containing the names of parent classes and child classes and stores it in the CSV file format. Suppose there are m parent classes and n child classes in a specific hierarchical text classification task, and the writing rule for the content of each row is shown in formula (10):

[0076] Row r =(P k ,C k,j )#(10)

[0077] Where:

[0078] Row r is the writing rule for the content of each row, r is the row number, 1 ≤ r ≤ n;

[0079] Each row of the CSV file has two columns, namely the parent class P k and the child class C k,j ;

[0080] In the parent class P k , k is the parent - class index, 1 ≤ k ≤ m;

[0081] In the child class C k,j , j is the child - class index corresponding to the k - th parent class, 1 ≤ j ≤ n k ;

[0082] n k is the number of child classes of the k - th parent class,

[0083] Taking the cognitive distortion path extraction task as an example, introduce the specific form of the tree - structured classification module of the interpretable hierarchical text classification model. The dataset has 4 parent classes: Activating, Belief, Consequence, and Disputation. Each parent class has several sub - classes. The 4 parent classes have a total of 19 sub - classes, so the file has 19 lines:

[0084] Activating_Event,Disease_symptom

[0085] Activating_Event,Social_relation

[0086] Activating_Event,Life

[0087] Activating_Event,Study_and_work

[0088] Activating_Event,Emotional

[0089] Belief,All_or_nothing_thinking

[0090] Belief,Over_generalization

[0091] Belief,Mental_filter

[0092] Belief,Disqualifying_the_positive

[0093] Belief,Jumping_to_conclusions

[0094] Belief,Magnification_and_minimization

[0095] Belief,Emotional_reasoning

[0096] Belief,Should_statements

[0097] Belief,Labeling_and_mislabeling

[0098] Belief,Blaming_oneself_others

[0099] Consequence,Emotional_effect

[0100] Consequence,Behavioral_effect

[0101] Disputation,Habitual_disputation

[0102] Disputation,Effective_disputation

[0103] Step 3: The model is fine-tuned on the dataset to obtain a classification model more suitable for the specific domain.

[0104] Taking the cognitive distortion path extraction task as an example, relevant information about the training stage of the interpretable hierarchical text classification model is introduced. The choice of the backbone network feature extractor can be flexibly selected according to the domain and language features of the dataset. Since the cognitive path extraction dataset is a Chinese psychology domain dataset, the model Chinese MentalBERT is selected for the feature extraction part. This model is fine-tuned and trained on the basis of the English psychology domain model MentalBERT using a Chinese psychology domain corpus. The model has 110 million parameters, has 12 Transformer encoder blocks, each layer has 12 attention heads, and the total number of attention heads is 12×12 = 144. The hidden layer dimension is 768, and each input token is represented as a 768-dimensional vector. The dimension of each head is the hidden layer dimension divided by the number of attention heads, that is, 768 / 12 = 64 dimensions.

[0105] This method is implemented using an NVIDIA GeForce RTX 4090 24GB GPU combined with the PyTorch framework during the verification process, and 100 epochs are set for training. An early stopping strategy is adopted to prevent the model from overfitting. The patience of the early stopping strategy is set to 20, and the best validation model is selected for test evaluation. For the interpretable tree structure model, we set the learning rate to 3e-5, the batch size to 16, and use MSE (mean squared error) to calculate the loss. The random seed is set to 42. The model parameters can be flexibly adjusted according to different datasets.

Claims

1. An interpretable hierarchical text classification method based on a tree structure, characterized in that: The following steps are involved: Step 1: Build an interpretable hierarchical text classification model; (1) The hierarchically classified text first passes through the feature extraction module; The feature extraction model consists of two parts: the feature encoder part and the encoder interpretation part; the extracted features are then integrated with the tree-structured classification module; The attention mechanism of the encoder can be understood from two aspects: first, the word-level attention weight, which determines the degree of attention the model pays to other words when encoding a specific word; second, the role of attention heads, each of which captures different levels of the relationship between words by assigning different attention weights; (2) Tree structure classification module The tree-structured classification module consists of parent nodes and child nodes. During the implementation process, a file containing the relationships between all nodes is input and automatically converted into a model structure; the parent node and the child node have similar architectures: a dropout layer, followed by a fully connected FC layer and an activation function; The difference between a parent node and a child node is that the parent node processes the features to be passed to the child node and also makes predictions; therefore, the parent node includes a layer with ReLU activation for processing and an FC layer with sigmoid function for decision making; while the child node only needs to make predictions, so its activation function is only sigmoid; the final decision of each child node is determined by combining the probabilities of the parent and the child through a multiplication operation; The decision path explanation part in the tree structure classification module depends on the parent decision node and the child decision node from top to bottom; the predicted probability value of the text data at each node including the parent node and the child node is displayed in the form of a tree diagram; Step 2: Construct a classification rule containing the names of parent and child categories and store it in CSV file format; suppose the specific hierarchical text classification task has m parent categories and n child categories, and the content writing rule for each line is shown in formula (1): Row r =(P k ,C k,j )#(1) in: Row r Write rules for the content of each line, r is the line number, 1≤r≤n; Each row of the CSV file has two columns, namely the parent class P k and subclass C k,j ; In the parent class P k In, k is the parent class index, 1≤k≤m; In subclass C k,j In the above formula, j is the subclass index corresponding to the k-th parent class, 1≤j≤n k ; n k is the number of subclasses of the kth parent class, The two different hierarchical text classification tasks tested are as follows: (1) Cognitive distortion path extraction task (2) The task of classifying papers into different fields Step 3: The model is fine-tuned on the data set to obtain a more suitable classification model; The Chinese MentalBERT model is used for feature extraction; 100 epochs are set for training; an early stopping strategy is used to prevent model overfitting, and the patience of the early stopping strategy is set to 20. The best validation model is selected for testing; for the interpretable tree structure model, the learning rate is set to 3e-5, the batch size is set to 16, and the mean square error MSE is used to calculate the loss; the random seed is set to 42.