Method and device for solving imbalance of Chinese anesthesia data set
By constructing a special anesthesia data set and combining the FDCL loss function and BertGCN model, the problem of imbalance in Chinese anesthesia data set is solved, which significantly improves the model's performance in ASA grading and anesthesia risk prediction, especially in the rare ASA rating categories, providing more accurate and effective anesthesia risk assessment support.
Patent Information
- Application Number
- CN202510122254.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-23
AI Technical Summary
The Chinese anesthesia dataset has an imbalance problem, resulting in poor performance in ASA grading and anesthesia risk prediction, especially in the rare ASA grading categories.
A dedicated anesthesia dataset was constructed and combined with the focus dual contrast learning (FDCL) loss function and the BertGCN model were used to improve the model's performance in ASA grading and anesthesia risk prediction.
Through this method, the robustness and accuracy of the model are significantly improved, especially in the prediction of the less common ASA grade categories, providing more accurate and effective support for anesthesia risk assessment.
Smart Images

Figure CN120032908A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical data processing, and in particular to a solution to an imbalanced Chinese anesthesia data set, and a solution device to an imbalanced Chinese anesthesia data set. Background Art
[0002] ASA (American Society of Anesthesiologists) classification is the traditional preoperative evaluation standard, which is usually divided into grades I, II, III, IV, and V. Grade I and II patients have good tolerance to anesthesia and surgery, and the anesthesia process is smooth. Grade III patients have certain risks during anesthesia, and they need to be fully prepared before anesthesia and take effective measures to prevent possible complications. Grade IV patients have extremely high anesthesia risks, and even if they are well prepared before surgery, the perioperative mortality rate is still very high. For grade V dying patients, both anesthesia and surgery are extremely dangerous. The use of a complete ASA classification system can effectively predict the frequency and severity of adverse events, thereby improving the patient's treatment outcomes. Therefore, it is very necessary to understand the patient's basic preoperative information and divide the ASA grade and risk level accordingly.
[0003] In China, many attempts have been made to combine artificial intelligence with anesthesiology to solve various problems that may be encountered during the perioperative period. However, methods specifically for ASA grading and risk assessment before anesthesia are still lacking. The main reason is that there is no Chinese data set specifically for risk assessment and ASA grading before anesthesia, and it is impossible to optimize and verify the existing models in a targeted manner, which limits the research on ASA and anesthesia risk grading methods in Chinese working environments.
[0004] Due to the requirements of patient privacy and medical experts, it is difficult to construct training data sets in the medical field, including anesthesia data sets. There are problems such as difficulty in obtaining data and uneven distribution of data categories, which increases the difficulty of learning various artificial intelligence models. Especially in the field of anesthesia, patients' ASA grades are generally concentrated in level 2, followed by level 3 and level 1, while level 4 and level 5 are less common in practice, and their importance is even higher than other risk levels. This category imbalance problem seriously affects the robustness of the model, making the model tend to predict categories with higher frequency, while performing poorly on categories with lower frequency, limiting the practicality of the model. Summary of the invention
[0005] In order to overcome the defects of the prior art, the technical problem to be solved by the present invention is to provide a solution to the imbalance of Chinese anesthesia dataset. It constructs a special anesthesia dataset and combines it with the innovative FDCL loss function to improve the performance of the BertGCN model in ASA grading and anesthesia risk prediction, thereby providing more accurate and effective technical support for anesthesia risk assessment in clinical medicine.
[0006] The technical solution of the present invention is: this solution to the imbalance of Chinese anesthesia data set includes the following steps:
[0007] (1) Constructing an anesthesia dataset: Obtaining patient information required for ASA classification and anesthesia risk assessment. The dataset contains options for judging anesthesia risk and ASA grade, and separates each piece of information to prevent ambiguity caused by the combination of different information.
[0008] (2) The basic pre-trained model is BertGCN, which constructs a heterogeneous graph on the dataset, uses words and documents as nodes, initializes the node embedding using pre-trained Bert, and then trains BerGCN by combining the Bert and GCN models and taking advantage of the two models. The large-scale pre-trained model learns the implicit and rich semantic information in the language by training on a large-scale unlabeled corpus. The graph neural network uses the co-occurrence relationship between words to learn the representation of nodes and edges, while processing complex structures and maintaining global information in text classification. BerGCN is used to extract text features from the anesthesia dataset, and the trained model is used to predict anesthesia risk.
[0009] (3) For the constructed anesthesia dataset, the focal dual contrastive learning (FDCL) model is used to enhance the classification ability. FDCL combines supervised contrastive learning with the focal loss function. It uses supervised contrastive learning to enhance the model's ability to distinguish between different categories, and uses the focal loss function to focus on a few categories to enhance the model's classification ability for these categories.
[0010] The present invention can help researchers better understand the risk assessment process of anesthetized patients, better deal with the data skew problem in anesthesia datasets, improve the robustness and accuracy of the model, and achieve significant performance improvements in ASA grading and anesthesia risk prediction tasks.
[0011] A device for solving the imbalance of Chinese anesthesia data set is also provided, the device comprising:
[0012] A dataset construction module, which constructs an anesthesia dataset: obtains patient information required for ASA classification and anesthesia risk assessment. The dataset contains options for anesthesia risk and ASA grade judgment, and separates each piece of information to prevent ambiguity caused by different combinations of information;
[0013] The training prediction module uses the BertGCN model as its basic pre-training model. This model constructs a heterogeneous graph on the dataset, takes words and documents as nodes, initializes the node embedding using the pre-trained Bert, and then trains BerGCN by combining the Bert and GCN models and taking advantage of the two models. The large-scale pre-training model learns the implicit and rich semantic information in the language by training on a large-scale unlabeled corpus. The graph neural network uses the co-occurrence relationship between words to learn the representation of nodes and edges, while processing complex structures and maintaining global information in text classification. It uses BerGCN to extract text features from the anesthesia dataset and uses the trained model to predict anesthesia risk.
[0014] The loss function module uses the focal dual contrast learning (FDCL) model to enhance the classification ability for the constructed anesthesia dataset. FDCL combines supervised contrast learning with the focal loss function, using supervised contrast learning to enhance the model's ability to distinguish between different categories, and using the focal loss function to focus on a few categories to enhance the model's classification ability for these categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 ASA classification and anesthesia risk data distribution diagram are shown.
[0016] Figure 2 A diagram of the model structure is shown.
[0017] Figure 3 A tSNE plot showing the learned representation of ASA grade on the anesthesia dataset.
[0018] Figure 4 The average L2 norm of the last classifier layer weight gradients is shown.
[0019] Figure 5 The confusion matrix of the model predictions is shown.
[0020] Figure 6 A flow chart of a solution to the imbalance of Chinese anesthesia dataset according to the present invention is shown. DETAILED DESCRIPTION
[0021] like Figure 6 As shown in Figure 1, this solution to the imbalance of the Chinese anesthesia dataset includes the following steps:
[0022] (1) Constructing an anesthesia dataset: Obtaining patient information required for ASA classification and anesthesia risk assessment. The dataset contains options for judging anesthesia risk and ASA grade, and separates each piece of information to prevent ambiguity caused by the combination of different information.
[0023] (2) The basic pre-trained model is BertGCN, which constructs a heterogeneous graph on the dataset, uses words and documents as nodes, initializes the node embedding using pre-trained Bert, and then trains BerGCN by combining the Bert and GCN models and taking advantage of the two models. The large-scale pre-trained model learns the implicit and rich semantic information in the language by training on a large-scale unlabeled corpus. The graph neural network uses the co-occurrence relationship between words to learn the representation of nodes and edges, while processing complex structures and maintaining global information in text classification. BerGCN is used to extract text features from the anesthesia dataset, and the trained model is used to predict anesthesia risk.
[0024] (3) For the constructed anesthesia dataset, the focal dual contrastive learning (FDCL) model is used to enhance the classification ability. FDCL combines supervised contrastive learning with the focal loss function. It uses supervised contrastive learning to enhance the model's ability to distinguish between different categories, and uses the focal loss function to focus on a few categories to enhance the model's classification ability for these categories.
[0025] The present invention can help researchers better understand the risk assessment process of anesthetized patients, better deal with the data skew problem in anesthesia datasets, improve the robustness and accuracy of the model, and achieve significant performance improvements in ASA grading and anesthesia risk prediction tasks.
[0026] Preferably, in step (1), the data set includes 31 options for females and 29 options for males for judging anesthesia risk and ASA grade; when collecting ASA grades, ASA grade predictions for emergency surgery are excluded, and data items with unclear classification are removed; for data with the same information but different ASA grades and anesthesia risk assessments, only one of the ASA grades or anesthesia risk assessments is removed, and other information is retained; for data with the same patient information but different results, the majority of options are retained or the entire ambiguous data is deleted.
[0027] Preferably, in step (2), when constructing the heterogeneous graph, each item is taken as a node, and edges are constructed between nodes according to the occurrence of the item in a specific level and the co-occurrence of the item in the entire corpus. The edges include term frequency-inverse document frequency TF-IDF and positive pointwise mutual information PPMI. The weight of the edge between two nodes i and j is defined as:
[0028]
[0029] Among them, A ij The elements in the matrix represent the relationship measure between nodes i and j; PPMI(i,j)
[0030] The positive mutual information between nodes i and j is used to measure their relevance; TF-IDF(ij)
[0031] The term frequency of node i in document j - inverse document frequency, which is used to evaluate the importance of a word in a document.
[0032] Preferably, in step (2), Bert is used to obtain feature embeddings of the input data, and these embeddings are used as the node feature matrix X of GCN. The output of GCN is fed into the softmax layer for classification:
[0033] Z GCN =softmax(g(X,A)) (2)
[0034] Where g represents the GCN model and Z_GCN is the prediction result of the model.
[0035] Preferably, in step (2), the prediction of GCN and the prediction of BERT are combined:
[0036] Z=λZ GCN +(1-λ)Z Bert (3)
[0037] Where Z_Bert is the prediction result of Bert, and λ controls the trade-off between the two prediction results of GCN and Bert.
[0038] Preferably, in step (3), the total loss function is:
[0039]
[0040] Among them, y k Indicates the degree of match between the predicted probability and the true label, γ is an adjustable focusing parameter,
[0041] α is a hyperparameter that controls the impact of the loss term;
[0042]
[0043] Where N is the number of categories, y i is the actual label, is the predicted probability, θ is the classifier label feature, z is the input data feature, τ is the temperature factor, P i is the positive sample set, A i is a set of comparison samples.
[0044] Preferably, the method also includes step (4), using the anesthesia data set to predict the risk level and ASA grade of anesthesia respectively; drawing a t-SNE graph on the test set of ASA grade; selecting 300 samples from the ASA classification test set and calculating the average L2 norm of the weight gradient of the last classifier layer on BertGCN; calculating their normalized confusion matrix on the test set of anesthesia ASA grade prediction.
[0045] A device for solving the imbalance of Chinese anesthesia data set is also provided, the device comprising:
[0046] A dataset construction module, which constructs an anesthesia dataset: obtains patient information required for ASA classification and anesthesia risk assessment. The dataset contains options for anesthesia risk and ASA grade judgment, and separates each piece of information to prevent ambiguity caused by different combinations of information;
[0047] The training prediction module uses the BertGCN model as its basic pre-training model. This model constructs a heterogeneous graph on the dataset, takes words and documents as nodes, initializes the node embedding using the pre-trained Bert, and then trains BerGCN by combining the Bert and GCN models and taking advantage of the two models. The large-scale pre-training model learns the implicit and rich semantic information in the language by training on a large-scale unlabeled corpus. The graph neural network uses the co-occurrence relationship between words to learn the representation of nodes and edges, while processing complex structures and maintaining global information in text classification. It uses BerGCN to extract text features from the anesthesia dataset and uses the trained model to predict anesthesia risk.
[0048] The loss function module uses the focal double contrast learning FDCL model to enhance the classification ability for the constructed anesthesia dataset. FDCL combines supervised contrast learning with the focal loss function and uses supervised contrast learning to enhance the model's ability to distinguish between different categories.
[0049] At the same time, the focal loss function is used to focus on a few categories to enhance the model's classification ability for these categories.
[0050] Preferably, in the training prediction module, when constructing the heterogeneous graph, each item is taken as a node, and edges are constructed between the nodes according to the occurrence of the item in a specific level and the co-occurrence of the item in the entire corpus. The edges include term frequency-inverse document frequency TF-IDF and positive pointwise mutual information PPMI. The weight of the edge between two nodes i and j is defined as:
[0051]
[0052] Among them, A ijThe elements in the matrix represent the relationship measure between nodes i and j; PPMI(i,j)
[0053] The positive mutual information between nodes i and j is used to measure their relevance; TF-IDF(ij)
[0054] The word frequency of node i in document j - inverse document frequency, used to evaluate the importance of words in the document;
[0055] Use Bert to get the feature embedding of the input data, and use these embeddings as the node feature matrix X of GCN. The output of GCN is fed into the softmax layer for classification:
[0056] Z GCN =softmax(g(X,A)) (2)
[0057] Where g represents the GCN model, and Z_GCN is the prediction result of the model;
[0058] Combine the predictions of GCN and Bert:
[0059] Z=λZ GCN +(1-λ)Z Bert (3)
[0060] Where Z_Bert is the prediction result of Bert, and λ controls the trade-off between the two prediction results of GCN and Bert.
[0061] Preferably, in the loss function module, the total loss function is:
[0062]
[0063] Among them, y k Indicates the degree of match between the predicted probability and the true label, γ is an adjustable focusing parameter,
[0064] α is a hyperparameter that controls the impact of the loss term;
[0065]
[0066] Where N is the number of categories, y i is the actual label, is the predicted probability, θ is the classifier label feature, z is the input data feature, τ is the temperature factor, P i is the positive sample set, A i is a set of comparison samples.
[0067] In order to verify the effectiveness of the proposed scheme, the accuracy of this method in anesthesia prediction in Chinese was first demonstrated, and compared with traditional machine learning methods and deep learning methods to verify the effectiveness of the technical route of the present invention.
[0068] The anesthesia data set was used to predict the risk level of anesthesia and the ASA grade respectively. The results of the FDCL and baseline methods are shown in Table 1. It can be observed that in the prediction of anesthesia risk, the best traditional machine learning method is the naive Bayes classifier, with an accuracy of 63.44%. In the deep learning method, the accuracy is generally higher than that of the traditional machine learning method, and the convolutional neural network (CNN) reaches the highest accuracy of 65.53%. When fine-tuning the pre-trained model, the present invention adopts the MACBert pre-trained model, which is more friendly to Chinese processing, as the benchmark model, and the accuracy after fine-tuning is 65.13%. After combining MACBert with the graph neural network (GCN), the accuracy is further improved to 66.38%. We replaced the cross entropy loss function in BertGCN with the FDCL method of the present invention, and the accuracy reached 67.38%, which is better than all the above methods.
[0069] In the ASA classification prediction, the best performing traditional machine learning method is the random forest, with an accuracy of 75.63%. Among the deep learning methods, the best performing one is also the convolutional neural network (CNN), reaching 76.71%. When using the pre-trained model MacBert for text classification, the accuracy is 76.21%, while the accuracy of BertGCN is 76.05%. Using this method, the accuracy is improved by 1.08%, which is better than other methods tested.
[0070] Table 1 Accuracy of different models in predicting anesthesia risk and ASA grade
[0071]
[0072] To show how this method improves representation on the anesthesia dataset, a t-SNE plot is plotted on the ASA graded test set. MacBert was used as the encoder for fine-tuning, as shown in Figure 3 The results of CE (cross entropy), Dual (dual contrastive learning) and FDCL are shown. It can be seen from the figure that compared with the cross entropy loss and the dual contrastive learning loss, the results of the present invention have a better effect on learning text representation.
[0073] 300 samples were selected from the ASA classification test set, and the average L2 norm of the weight gradient of the last classifier layer on BertGCN was calculated. Figure 4As shown in the figure, the number of each category in the ASA classification is sorted from large to small (Ⅱ, Ⅲ, Ⅰ, Ⅳ, Ⅴ). It can be seen that compared with the cross entropy loss and supervised contrastive learning (Dual), this method can better adjust the gradient norm, making the change from the most frequent category to the least frequent category more balanced.
[0074] In order to specifically demonstrate the improvement of this method compared with the basic method, their normalized confusion matrix was calculated on the test set of anesthesia ASA grade prediction, as shown in Figure 5 As shown. It can be observed that the accuracy is mainly concentrated in the category of ASA grade II. This is because most of the samples in the data set belong to ASA grade II, while the recognition accuracy of data of other ASA grades is very low. After using supervised contrastive learning Dual to modify the loss function, it can be seen that the recognition accuracy of ASA grade III, which is the second largest, is improved, but the recognition accuracy of other categories is limited. Compared with the above two methods, this method improves the accuracy of ASA grade III from 0.21 to 0.65, and ASA IV from 0 to 0.56. These results show that this method can improve the recognition accuracy of the model when dealing with unbalanced data and few sample categories, especially better performance on less common ASA grades.
[0075] Figure 1 The final processing results showed that there were 9643 valid data for ASA classification and 9309 valid data for anesthesia risk assessment.
[0076] Figure 2 The input of the model is data consisting of labels and text. Bert is used to extract text features of anesthesia data, and labels and patient information are used to construct a heterogeneous graph. The GCN module extracts features by using the heterogeneous graph and text features extracted by the Bert model. Combining the predictions of the graph neural network and the Bert model, the final ASA grade and anesthesia risk level are obtained.
[0077] Figure 3 t-SNE can intuitively see the distribution between different categories in text data and reveal the similarities or differences between categories. The distribution of different text data points (such as sentences, documents, etc.) in a two-dimensional plane can be viewed in the form of a scatter plot. If samples of certain categories are clustered together in the t-SNE graph, it means that they are relatively similar in high-dimensional space; on the contrary, if the sample distribution is more dispersed, it means that they are quite different in high-dimensional space.
[0078] Figure 4The average L2 norm of the weight gradient of the last classifier layer. If there are fewer samples in some categories, the model may have weaker learning signals (gradients) for these categories during training, which will result in smaller gradients for the minority class. Using a weighted loss function improves the learning signal of minority class samples, which in turn affects the size of the gradient. In this way, the minority class will receive more attention during training, resulting in a more balanced gradient.
[0079] Figure 5 The confusion matrix reflects the model's predictions for each category, through which we can intuitively understand which categories are predicted better and which categories are predicted worse.
[0080] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Any simple modification, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention are still within the protection scope of the technical solution of the present invention.
Claims
1. A solution to the imbalance of Chinese anesthesia dataset, characterized by: The method comprises the following steps: (1) Constructing an anesthesia dataset: Obtaining patient information required for ASA classification and anesthesia risk assessment. The dataset contains options for judging anesthesia risk and ASA grade, and separates each piece of information to prevent ambiguity caused by the combination of different information. (2) The basic pre-trained model is BertGCN, which constructs a heterogeneous graph on the dataset, uses words and documents as nodes, initializes the node embedding using pre-trained Bert, and then trains BerGCN by combining the Bert and GCN models and taking advantage of the two models. The large-scale pre-trained model learns the implicit and rich semantic information in the language by training on a large-scale unlabeled corpus. The graph neural network uses the co-occurrence relationship between words to learn the representation of nodes and edges, while processing complex structures and maintaining global information in text classification. BerGCN is used to extract text features from the anesthesia dataset, and the trained model is used to predict anesthesia risk. (3) For the constructed anesthesia dataset, the focal dual contrastive learning (FDCL) model is used to enhance the classification ability. FDCL combines supervised contrastive learning with the focal loss function. It uses supervised contrastive learning to enhance the model's ability to distinguish between different categories, and uses the focal loss function to focus on a few categories to enhance the model's classification ability for these categories.
2. The solution to the imbalance of Chinese anesthesia data set according to claim 1 is characterized by: In step (1), the data set includes 31 options for females and 29 options for males for judging anesthesia risk and ASA grade; when collecting ASA grades, ASA grade predictions for emergency surgery are excluded, and data items with unclear classification are removed; for data with the same information but different ASA grades and anesthesia risk assessments, only one of the ASA grades or anesthesia risk assessments is removed, and the other information is retained; for data with the same patient information but different results, the majority of options are retained or the entire ambiguous data is deleted.
3. The solution to the imbalance of Chinese anesthesia data set according to claim 2 is characterized by: In step (2), when constructing the heterogeneous graph, each item is taken as a node, and edges are constructed between nodes according to the occurrence of the item in a specific level and the co-occurrence of the item in the entire corpus. The edges include term frequency-inverse document frequency TF-IDF and positive pointwise mutual information PPMI. The weight of the edge between two nodes i and j is defined as: Among them, A ij The elements in the matrix represent the relationship measure between nodes i and j; PPMI(i,j) is the positive point-wise mutual information between nodes i and j, which is used to measure their relevance; TF-IDF(ij) is the term frequency-inverse document frequency of node i in document j, which is used to evaluate the importance of words in a document.
4. The solution to the imbalance of Chinese anesthesia data set according to claim 3 is characterized by: In step (2), Bert is used to obtain the feature embedding of the input data, and these embeddings are used as the node feature matrix X of GCN. The output of GCN is sent to the softmax layer for classification: FROM GCN =softmac(g(X,A)) (2) Where g represents the GCN model and Z_GCN is the prediction result of the model.
5. The solution to the imbalance of Chinese anesthesia data set according to claim 4 is characterized by: In step (2), the prediction of GCN and the prediction of Bert are combined: Z=λZ GCN +(1-λ)Z Bert (3) Where Z_Bert is the prediction result of Bert, and λ controls the trade-off between the two prediction results of GCN and Bert.
6. The solution to the imbalance of Chinese anesthesia data set according to claim 5 is characterized by: In step (3), the total loss function is: Among them, y k Indicates the degree of match between the predicted probability and the true label, γ is an adjustable focus parameter, and α is a hyperparameter that controls the impact of the loss term; Where N is the number of categories, y i is the actual label, is the predicted probability, θ is the classifier label feature, z is the input data feature, τ is the temperature factor, P i is the positive sample set, A i is a set of comparison samples.
7. The solution to the imbalance of Chinese anesthesia data set according to claim 6 is characterized by: The method also includes step (4), using the anesthesia dataset to predict the risk level and ASA grade of anesthesia respectively; drawing a t-SNE graph on the test set of ASA grade; selecting 300 samples from the ASA classification test set and calculating the average L2 norm of the weight gradient of the last classifier layer on BertGCN; and calculating their normalized confusion matrix on the test set of anesthesia ASA grade prediction.
8. A device for solving the imbalance of Chinese anesthesia data set, characterized by: The device includes: a data set construction module, which constructs an anesthesia data set: obtains patient information required for ASA classification and anesthesia risk assessment, the data set contains options for anesthesia risk and ASA grade judgment, and separates each information to prevent ambiguity caused by the combination of different information; The training prediction module uses the BertGCN model as its basic pre-training model. This model constructs a heterogeneous graph on the dataset, takes words and documents as nodes, initializes the node embedding using the pre-trained Bert, and then trains BerGCN by combining the Bert and GCN models and taking advantage of the two models. The large-scale pre-training model learns the implicit and rich semantic information in the language by training on a large-scale unlabeled corpus. The graph neural network uses the co-occurrence relationship between words to learn the representation of nodes and edges, while processing complex structures and maintaining global information in text classification. It uses BerGCN to extract text features from the anesthesia dataset and uses the trained model to predict anesthesia risk. The loss function module uses the focal dual contrast learning (FDCL) model to enhance the classification ability for the constructed anesthesia dataset. FDCL combines supervised contrast learning with the focal loss function, using supervised contrast learning to enhance the model's ability to distinguish between different categories, and using the focal loss function to focus on a few categories to enhance the model's classification ability for these categories.
9. The device for solving the imbalance of Chinese anesthesia data set according to claim 8, characterized in that: In the training prediction module, when constructing the heterogeneous graph, each item is taken as a node, and edges are constructed between nodes according to the occurrence of the item in a specific level and the co-occurrence of the item in the entire corpus. The edges include term frequency-inverse document frequency TF-IDF and positive point-by-point mutual information PPMI. The weight of the edge between two nodes i and j is defined as: Among them, A ij The elements in the matrix represent the relationship measure between nodes i and j; PPMI(i,j) is the positive point mutual information between nodes i and j, which is used to measure their relevance; TF-IDF(ij) is the term frequency-inverse document frequency of node i in document j, which is used to evaluate the importance of words in documents; Use Bert to get the feature embedding of the input data, and use these embeddings as the node feature matrix X of GCN. The output of GCN is fed into the softmax layer for classification: Z GCN =softmax(g(X,A)) (2) Where g represents the GCN model, and Z_GCN is the prediction result of the model; Combine the predictions of GCN and Bert: Z=λZ GCN +(1-λ)Z Bert (3) Where Z_Bert is the prediction result of Bert, and λ controls the trade-off between the two prediction results of GCN and Bert.
10. The device for solving the imbalance of Chinese anesthesia data set according to claim 9, characterized in that: In the loss function module, the total loss function is: Among them, y k Indicates the degree of match between the predicted probability and the true label, γ is an adjustable focus parameter, and α is a hyperparameter that controls the impact of the loss term; Where N is the number of categories, y i is the actual label, is the predicted probability, θ is the classifier label feature, z is the input data feature, τ is the temperature factor, P i is the positive sample set, A i is a set of comparison samples.