Medical text classification method and device

By combining the integrated encoder with BERT and ERNIE encoder, the problem of low accuracy in medical text classification is solved by using the enhanced medical text data set and large language model to solve the problem of low accuracy in medical text classification and achieve efficient and accurate medical text classification.

CN120162437APending Publication Date: 2025-06-17UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510566190.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art has problems in medical text classification with low accuracy, sensitive parameters, and the need to manually extract features, especially when large-scale medical text data processing.

Method used

Using an integrated encoder, the BERT encoder and its fully connected layer and the ERNIE encoder and its target fully connected layer are combined. It uses the enhanced medical text data set for training, retrieves professional term interpretations from external knowledge databases through a large language model for data augmentation, and the classified medical text is embedded and voting processed to obtain the final classification result.

Benefits of technology

It improves the accuracy and efficiency of medical text classification, can effectively process large-scale medical text data, and improves the reliability of classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162437A_ABST
    Figure CN120162437A_ABST
Patent Text Reader

Abstract

The invention discloses a medical text classification method and device. The method comprises the steps of obtaining a to-be-classified medical text; performing embedding processing on the to-be-classified medical text to obtain embedding of the to-be-classified medical text; embedding and inputting the to-be-classified medical text into a pre-trained integrated encoder to obtain a first classification result and a second classification result; wherein the integrated encoder is formed by combining a BERT encoder, a full connection layer corresponding to the BERT encoder, an ERNIE encoder and a target full connection layer corresponding to the ERNIE encoder; the integrated encoder is obtained by training an enhanced medical text data set in advance; and voting the first classification result and the second classification result to obtain a final classification result. Therefore, the multiple encoders are encoded by utilizing the enhanced medical text data and the corresponding classification results are obtained, and all the classification results are voted, so that large-scale efficient classification of the medical texts can be realized, and the classification accuracy is further effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer application technologies, and in particular, to a method and device for classifying medical texts. Background Art

[0002] With the continuous development of the medical field and the rapid expansion of medical information systems, the quantity of medical text data has increased sharply, including medical research papers, textbooks, clinical records, and medical test reports, etc. Therefore, classifying and organizing these medical texts can effectively promote the management and application of medical information, thereby improving its utilization efficiency.

[0003] In the existing technologies, medical text classification mainly goes through three stages, namely rule-based classification methods, traditional machine learning-based classification methods, and deep learning-based classification methods. Specifically, the rule-based classification method uses medical domain knowledge and expert experience to formulate a series of classification rules, and then classifies medical texts through the formulated series of classification rules. The traditional machine learning-based classification method refers to using traditional machine learning algorithms such as random forest, support vector machine, and naive Bayes, etc. to classify medical texts. The deep learning-based classification method is the current mainstream and has better effects in medical text classification, and classifies medical texts by constructing deep neural networks such as convolutional neural network (CNN) and recurrent neural network (RNN), etc. However, the exploration of large language models in the field of medical text classification is still insufficient.

[0004] Although the accuracy of classifying medical texts using the rule classification method is relatively high, its portability is poor, and it also requires a certain amount of manpower. Moreover, when classifying large-scale medical text data, the classification results will not be accurate enough. And the traditional machine learning-based classification method is sensitive to parameters, requires manual feature extraction, and has average accuracy. Also, the exploration of the deep learning-based classification method in the field of medical text classification is still insufficient, so the classification results are not accurate enough. Summary of the Invention

[0005] Based on the deficiencies of the above-mentioned existing technologies, the present application provides a method and device for classifying medical texts to solve the problem of inaccurate medical text classification brought by the existing technologies.

[0006] To achieve the above object, the present application provides the following technical solutions:

[0007] The first aspect of the present application provides a method for classifying medical texts, including:

[0008] Obtaining the medical text to be classified;

[0009] Perform embedding processing on the medical text to be classified to obtain the embedding of the medical text to be classified;

[0010] Input the embedding of the medical text to be classified into a pre-trained integrated encoder to obtain a first classification result and a second classification result; wherein, the integrated encoder is composed of a BERT encoder and its corresponding fully connected layer and an ERNIE encoder and its corresponding target fully connected layer; the integrated encoder is pre-trained using an enhanced medical text dataset; the enhanced medical text dataset is pre-obtained by using a large language model to retrieve professional terms from an external knowledge database for the medical text dataset, obtaining explanations of the professional terms, and writing the explanations of the professional terms into the medical text dataset for data enhancement;

[0011] Perform voting processing on the first classification result and the second classification result to obtain the final classification result corresponding to the medical text to be classified.

[0012] Optionally, in the above medical text classification method, the performing embedding processing on the medical text to be classified to obtain the embedding of the medical text to be classified includes:

[0013] Perform token embedding processing on the medical text to be classified to obtain the token embedding of the medical text to be classified;

[0014] Perform segment embedding processing on the medical text to be classified to obtain the segment embedding of the medical text to be classified;

[0015] Perform position embedding processing on the medical text to be classified to obtain the position embedding of the medical text to be classified;

[0016] Perform summation calculation on the token embedding of the medical text to be classified, the segment embedding of the medical text to be classified, and the position embedding of the medical text to be classified to obtain the embedding of the medical text to be classified.

[0017] Optionally, in the above medical text classification method, the training method of the integrated encoder includes:

[0018] Obtain a medical text dataset;

[0019] Use a large language model to retrieve professional terms from an external knowledge database for the medical text dataset to obtain explanations of multiple professional terms;

[0020] Add the explanation of each professional term to the medical text dataset for data enhancement processing to obtain an enhanced medical text dataset;

[0021] Perform embedding processing on the enhanced medical text dataset to obtain an embedded medical text dataset;

[0022] Input the embedded medical text dataset into the BERT encoder included in the integrated encoder to obtain a feature vector, and input the embedded medical text dataset into the ERNIE encoder included in the integrated encoder to obtain a target feature vector;

[0023] Input the feature vector into the fully connected layer included in the integrated encoder to obtain a first sample classification result, and input the target feature vector into the target fully connected layer included in the integrated encoder to obtain a second sample classification result;

[0024] Calculate a first loss function between the first sample classification result and the actual classification result corresponding to the medical text dataset, and calculate a second loss function between the second sample classification result and the actual classification result;

[0025] Determine whether the first loss function converges and whether the second loss function converges;

[0026] If the first loss function does not converge and the second loss function does not converge, use the Adam optimizer to adjust the parameters of the BERT encoder and the parameters of the ERNIE encoder, and return to execute the step of inputting the embedded medical text dataset into the BERT encoder included in the integrated encoder to obtain a feature vector, and inputting the embedded medical text dataset into the ERNIE encoder included in the integrated encoder to obtain a target feature vector;

[0027] If the first loss function converges and the second loss function converges, determine that the integrated encoder is a trained integrated encoder.

[0028] Optionally, in the above classification method for medical texts, the step of inputting the embedded medical text dataset into the BERT encoder included in the integrated encoder to obtain a feature vector includes:

[0029] After inputting the embedded medical text dataset into the BERT encoder included in the integrated encoder, sequentially perform encoding processing on the embedded medical text dataset through multiple Transformer encoders to obtain a feature vector.

[0030] Optionally, in the above classification method for medical texts, the step of sequentially performing encoding processing on the embedded medical text dataset through multiple Transformer encoders after inputting the embedded medical text dataset into the BERT encoder included in the integrated encoder to obtain a feature vector includes:

[0031] For each of the Transformer encoders, the input data is calculated through the multi-head self-attention mechanism module in the Transformer encoder to obtain a multi-head self-attention vector;

[0032] The multi-head self-attention vector and the input data are subjected to residual connection through the residual connection module in the Transformer encoder to obtain a first connection vector;

[0033] The first connection vector is subjected to normalization calculation through the first layer normalization module in the Transformer encoder to obtain a normalized vector;

[0034] The normalized vector is calculated through the feed-forward network in the Transformer encoder to obtain a feed-forward network vector;

[0035] The normalized vector and the feed-forward network vector are subjected to residual connection through the residual connection module to obtain a second connection vector;

[0036] The second connection vector is subjected to normalization calculation through the second layer normalization module in the Transformer encoder to obtain a feature vector.

[0037] Optionally, in the above medical text classification method, the inputting the embedded medical text dataset into the ERNIE encoder included in the integrated encoder to obtain a target feature vector includes:

[0038] After the embedded medical text dataset is input into the ERNIE encoder included in the integrated encoder, the embedded medical text dataset is encoded through a plurality of Transformer encoders in sequence to obtain a feature vector;

[0039] The feature vector is encoded with knowledge through a plurality of knowledge encoders in sequence to obtain a target feature vector.

[0040] Optionally, in the above medical text classification method, the encoding the feature vector with knowledge through a plurality of knowledge encoders in sequence to obtain a target feature vector includes:

[0041] For each of the knowledge encoders, the feature vectors are calculated through the first multi-head self-attention mechanism module in the knowledge encoder to obtain token multi-head self-attention vectors, and the entities are calculated through the second multi-head self-attention mechanism module in the knowledge encoder to obtain entity multi-head self-attention vectors; wherein, the entities are pre-recognized from the medical text to be classified by using the ERNIE encoder, encoded by the TransE model, and then input into the knowledge encoder;

[0042] The heterogeneous information fusion layer in the knowledge encoder is used to determine whether there is an entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector;

[0043] If there is an entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector, the token multi-head self-attention vector and the entity vector are calculated by using an entity fusion algorithm through the heterogeneous information fusion layer to obtain a target feature vector;

[0044] If there is no entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector, the token multi-head self-attention vector is calculated by using a fusion algorithm through the heterogeneous information fusion layer to obtain a target feature vector.

[0045] The second aspect of the present application provides a classification device for medical texts, including:

[0046] A text acquisition unit for acquiring a medical text to be classified;

[0047] An embedding processing unit for performing embedding processing on the medical text to be classified to obtain an embedding of the medical text to be classified;

[0048] A text input unit for inputting the embedding of the medical text to be classified into a pre-trained integrated encoder to obtain a first classification result and a second classification result; wherein, the integrated encoder is composed of a BERT encoder and its corresponding fully connected layer and an ERNIE encoder and its corresponding target fully connected layer; the integrated encoder is pre-trained by using an enhanced medical text dataset; the enhanced medical text dataset is pre-obtained by using a large language model to retrieve professional terms from an external knowledge database, obtaining explanations of the professional terms, and writing the explanations of the professional terms into the medical text dataset for data enhancement;

[0049] A voting unit for performing voting processing on the first classification result and the second classification result to obtain a final classification result corresponding to the medical text to be classified.

[0050] Optionally, in the above-mentioned medical text classification device, the embedding processing unit includes:

[0051] A token embedding processing unit for performing token embedding processing on the medical text to be classified to obtain the token embedding of the medical text to be classified;

[0052] A segment embedding processing unit for performing segment embedding processing on the medical text to be classified to obtain the segment embedding of the medical text to be classified;

[0053] A position embedding processing unit for performing position embedding processing on the medical text to be classified to obtain the position embedding of the medical text to be classified;

[0054] An adding unit for performing summation calculation on the token embedding of the medical text to be classified, the segment embedding of the medical text to be classified, and the position embedding of the medical text to be classified to obtain the embedding of the medical text to be classified.

[0055] Optionally, in the above-mentioned medical text classification device, it further includes:

[0056] An acquisition unit for acquiring a medical text data set;

[0057] A retrieval unit for using a large language model to retrieve professional terms from an external knowledge database for the medical text data set to obtain explanations of multiple professional terms;

[0058] An enhancement processing unit for adding the explanation of each professional term to the medical text data set for data enhancement processing to obtain an enhanced medical text data set;

[0059] A processing unit for performing embedding processing on the enhanced medical text data set to obtain an embedded medical text data set;

[0060] A first input unit for inputting the embedded medical text data set into the BERT encoder included in the integrated encoder to obtain a feature vector, and inputting the embedded medical text data set into the ERNIE encoder included in the integrated encoder to obtain a target feature vector;

[0061] A second input unit for inputting the feature vector into the fully connected layer included in the integrated encoder to obtain a first sample classification result, and inputting the target feature vector into the target fully connected layer included in the integrated encoder to obtain a second sample classification result;

[0062] A loss function calculation unit for calculating a first loss function between the first sample classification result and the actual classification result corresponding to the medical text data set, and calculating a second loss function between the second sample classification result and the actual classification result;

[0063] A judgment unit, configured to judge whether the first loss function converges and whether the second loss function converges;

[0064] An adjustment unit, configured to, if the first loss function does not converge and the second loss function does not converge, use an Adam optimizer to adjust the parameters of the BERT encoder and the parameters of the ERNIE encoder, and return to execute inputting the embedded medical text dataset into the BERT encoder included in the integrated encoder to obtain a feature vector, and inputting the embedded medical text dataset into the ERNIE encoder included in the integrated encoder to obtain a target feature vector;

[0065] A determination unit, configured to, if the first loss function converges and the second loss function converges, determine that the integrated encoder is a trained integrated encoder.

[0066] Optionally, in the above medical text classification device, the first input unit includes:

[0067] A first encoding processing unit, configured to, after inputting the embedded medical text dataset into the BERT encoder included in the integrated encoder, sequentially perform encoding processing on the embedded medical text dataset through a plurality of Transformer encoders to obtain a feature vector.

[0068] Optionally, in the above medical text classification device, the first encoding processing unit includes:

[0069] A first calculation unit, configured to respectively calculate, for each of the Transformer encoders, the input data through the multi-head self-attention mechanism module in the Transformer encoder to obtain a multi-head self-attention vector;

[0070] A first residual connection unit, configured to perform a residual connection on the multi-head self-attention vector and the input data through the residual connection module in the Transformer encoder to obtain a first connection vector;

[0071] A first normalization calculation unit, configured to perform a normalization calculation on the first connection vector through the first layer normalization module in the Transformer encoder to obtain a normalized vector;

[0072] A second calculation unit, configured to calculate the normalized vector through the feed-forward network in the Transformer encoder to obtain a feed-forward network vector;

[0073] A second residual connection unit, configured to perform a residual connection on the normalized vector and the feed-forward network vector through the residual connection module to obtain a second connection vector;

[0074] A second normalization calculation unit, configured to perform a normalization calculation on the second connection vector through the second layer normalization module in the Transformer encoder to obtain a feature vector.

[0075] Optionally, in the above medical text classification device, the first input unit includes:

[0076] A second encoding processing unit, configured to input the embedded medical text dataset into the ERNIE encoder included in the integrated encoder, and then perform encoding processing on the embedded medical text dataset through multiple Transformer encoders in sequence to obtain a feature vector;

[0077] A knowledge encoding unit, configured to perform knowledge encoding on the feature vector through multiple knowledge encoders in sequence to obtain a target feature vector.

[0078] Optionally, in the above medical text classification device, the knowledge encoding unit includes:

[0079] A third calculation unit, configured to, for each of the knowledge encoders, calculate the feature vector through the first multi-head self-attention mechanism module in the knowledge encoder to obtain a token multi-head self-attention vector, and calculate the entity through the second multi-head self-attention mechanism module in the knowledge encoder to obtain an entity multi-head self-attention vector; wherein, the entity is pre-trained by using a knowledge graph on the TransE model, and the trained knowledge graph is embedded into the knowledge encoder;

[0080] An entity judgment unit, configured to judge whether there is an entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector through the heterogeneous information fusion layer in the knowledge encoder;

[0081] A fourth calculation unit, configured to, if there is an entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector, calculate the token multi-head self-attention vector and the entity vector through the heterogeneous information fusion layer by using an entity fusion algorithm to obtain a target feature vector;

[0082] A fifth calculation unit, configured to, if there is no entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector, calculate the token multi-head self-attention vector through the heterogeneous information fusion layer by using a fusion algorithm to obtain a target feature vector.

[0083] A classification method for medical texts provided by this application. By obtaining the medical text to be classified, secondly performing embedding processing on the medical text to be classified to obtain the embedding of the medical text to be classified, and then inputting the embedding of the medical text to be classified into a pre-trained integrated encoder to obtain a first classification result and a second classification result. Among them, the integrated encoder is composed of a BERT encoder and its corresponding fully connected layer and an ERNIE encoder and its corresponding target fully connected layer. The integrated encoder is pre-trained using an enhanced medical text dataset, and the enhanced medical text dataset is pre-obtained by using a large language model to retrieve professional terms from an external knowledge database for the medical text dataset, obtaining the explanations of the professional terms, and writing the explanations of the professional terms into the medical text dataset for data enhancement. Finally, voting processing is performed on the first classification result and the second classification result to obtain the final classification result corresponding to the medical text to be classified. Thus, using multiple encoders to encode the enhanced medical text data and obtaining corresponding classification results, and voting on all classification results can achieve large-scale and efficient classification of medical texts, and further effectively improve the accuracy of classification. Brief Description of the Drawings

[0084] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0085] Figure 1 It is a schematic flowchart of a classification method for medical texts provided by an embodiment of the present application;

[0086] Figure 2 It is a schematic flowchart of an embedding processing of medical texts provided by an embodiment of the present application;

[0087] Figure 3 It is a schematic flowchart of a training method for an integrated encoder provided by an embodiment of the present application;

[0088] Figure 4 It is a schematic flowchart of a coding processing method for a medical text dataset provided by an embodiment of the present application;

[0089] Figure 5 It is a schematic flowchart of a method for obtaining a target feature vector provided by an embodiment of the present application;

[0090] Figure 6 It is a schematic flowchart of a knowledge coding method for a medical text dataset provided by an embodiment of the present application;

[0091] Figure 7 Schematic structural diagram of a medical text classification method based on large language model data augmentation provided by an embodiment of the present application;

[0092] Figure 8 Schematic structural diagram of a classification device for medical texts provided by another embodiment of the present application. Detailed implementation manners

[0093] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0094] In the present application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0095] The embodiments of the present application provide a classification method for medical texts, as Figure 1 shown, which specifically includes the following steps:

[0096] S101. Obtain the medical text to be classified.

[0097] Optionally, the medical text to be classified can be sent to the system through the receiving interface provided by the system, so that the system can obtain the medical text to be classified and perform subsequent classification processing.

[0098] S102. Perform embedding processing on the medical text to be classified to obtain the embedding of the medical text to be classified.

[0099] It can be understood that in order to make full use of the semantic information, structural information and sequential information of the medical text, so as to improve the understanding and reasoning ability of the subsequent integrated encoder for the medical text, the medical text to be classified can be pre-embedded.

[0100] Optionally, in another embodiment of the present application, a specific implementation manner of step S102 is as follows Figure 2 shown, including the following steps:

[0101] S201. Perform token embedding processing on the medical text to be classified to obtain the token embedding of the medical text to be classified.

[0102] It can be understood that the main goal of token embedding processing is to map each word or sub-word (Token) in the medical text to be classified into a low-dimensional vector space, so that the model can use the semantic information of the vocabulary for classification tasks.

[0103] Specifically, a token model (such as Word2Vec, GloVe, etc.) can be selected, and the medical text to be classified is input into the selected token model to obtain the embedding vector of each medical vocabulary, that is, the token embedding (token embedding(X) ) of the medical text to be classified.

[0104] S202. Perform segment embedding processing on the medical text to be classified to obtain the segment embedding of the medical text to be classified.

[0105] It can be understood that segment embedding processing is to convert the entire paragraph (rather than a single token) into a low-dimensional vector representation, retaining the overall semantic information in the text.

[0106] Specifically, a segment embedding model, such as Doc2Vec, Transformer-based Models, etc., is selected, and then the medical text to be classified is input into the selected segment embedding model to obtain a low-dimensional vector representation, that is, the segment embedding (segment embedding(X) ) of the medical text to be classified.

[0107] S203. Perform position embedding processing on the medical text to be classified to obtain the position embedding of the medical text to be classified.

[0108] It can be understood that the goal of position embedding processing is to add position information to each token (or sub-word) in the medical text to be classified, so that the model can recognize the relative position of the words in the sentence, thereby improving the model's understanding ability of the text order.

[0109] Specifically, the medical text to be classified is input into the position embedding model for token position learning to obtain the position embedding representation, that is, the position embedding (position embedding(X) ) of the medical text to be classified.

[0110] S204. Calculate the sum of the token embeddings, segment embeddings, and position embeddings of the medical text to be classified to obtain the embedding of the medical text to be classified.

[0111] Specifically, after obtaining the token embedding(X) ), (segment embedding(X) ), and (position embedding(X) ) of the medical text to be classified, they need to be added together before being used as the input to the integrated encoder, so that the integrated encoder can better capture the sentence structure and syntactic relationships, thereby improving the classification accuracy.

[0112] That is, the calculation expression for the embedding E(X) of the medical text to be classified is:

[0113] E(X) = token embedding(X) + segment embedding(X) + position embedding(X) .

[0114] S103. Input the embedding of the medical text to be classified into a pre-trained integrated encoder to obtain a first classification result and a second classification result.

[0115] Among them, the integrated encoder is composed of a BERT encoder and its corresponding fully connected layer and an ERNIE encoder and its corresponding target fully connected layer, and the integrated encoder is pre-trained using an enhanced medical text dataset. The enhanced medical text dataset is pre-obtained by using a large language model to retrieve professional terms from an external knowledge database for the medical text dataset, obtaining explanations of the professional terms, and writing the explanations of the professional terms into the medical text dataset for data enhancement.

[0116] It should be noted that in the embodiments of the present application, an integrated encoder is constructed by combining two encoders, namely a BERT encoder and an ERNIE encoder, and the encoding results of the two encoders obtain corresponding prediction results through different fully connected layers (linear classifiers). Therefore, when the embedding of the medical text to be classified is input into the pre-trained integrated encoder, prediction results of the BERT encoder and the ERNIE encoder will be obtained respectively, that is, the first classification result and the second classification result.

[0117] Optionally, the embodiments of the present application provide a training method for an integrated encoder, as Figure 3 shown, including the following steps:

[0118] S301. Obtain a medical text dataset.

[0119] Optionally, each medical text can be retrieved from a database, and then each retrieved medical text can be integrated to obtain a medical text dataset.

[0120] S302. Use a large language model to retrieve professional terms from an external knowledge database for the medical text dataset, and obtain explanations of multiple professional terms.

[0121] It should be noted that in order to effectively reduce the hallucination problem brought by directly applying the subsequent integrated encoder to the field of medical text classification, and at the same time solve the problems of many professional terms and difficulty in understanding in medical texts, in the embodiments of the present application, the API of the large language model is called to identify the medical professional terms in the medical text dataset, and the external knowledge database is used to retrieve the medical text dataset to obtain the noun explanations and descriptions of the corresponding medical professional terms.

[0122] For example, for the medical professional term "myocardial infarction" in "The patient has a history of myocardial infarction", then retrieve the web knowledge database to find the noun explanation and description of "myocardial infarction": "Myocardial infarction is the damage to the myocardium when coronary blood flow decreases or stops..."

[0123] S303. Add the noun explanation of each professional term to the medical text dataset for data augmentation processing to obtain an enhanced medical text dataset.

[0124] Specifically, after obtaining multiple professional terms, the noun explanation of each professional term can be added to the medical text dataset for data augmentation processing, which helps the integrated encoder better understand the semantics of medical texts.

[0125] For example, the enhanced medical text is: "The patient has a history of myocardial infarction. Myocardial infarction is the damage to the myocardium when coronary blood flow decreases or stops..."

[0126] S304. Perform embedding processing on the enhanced medical text dataset to obtain an embedded medical text dataset.

[0127] It should be noted that for the specific implementation manner of step S304, reference can be made correspondingly to step S102 in the above method embodiment, which will not be elaborated here.

[0128] S305. Input the embedded medical text dataset into the BERT encoder included in the integrated encoder to obtain feature vectors, and input the embedded medical text dataset into the ERNIE encoder included in the integrated encoder to obtain target feature vectors.

[0129] Specifically, in terms of the model architecture, BERT is a bidirectional encoding representation based on the Transformer encoder architecture. Based on the architecture of the multi-layer Transformer encoder of BERT, ERNIE adds multiple layers of knowledge encoders to achieve the effective fusion of text semantic information and knowledge entity information. ERNIE is equivalent to introducing a knowledge graph on the basis of BERT.

[0130] In addition, in terms of the pre-training task, BERT is pre-trained through the masked language model and the next sentence prediction task. Based on these two pre-training tasks of BERT, ERNIE designs a new pre-training task, namely the denoising entity auto-encoder task, to train the model's ability to combine text semantic information and knowledge entity information.

[0131] Therefore, by inputting the embedded medical text dataset into the BERT encoder and the ERNIE encoder for training and fine-tuning, it can be adapted to specific medical text classification tasks.

[0132] Optionally, in another embodiment of the present application, a specific implementation manner of inputting the embedded medical text dataset into the BERT encoder included in the integrated encoder in step S305 to obtain a feature vector includes the following steps:

[0133] After inputting the embedded medical text dataset into the BERT encoder included in the integrated encoder, the embedded medical text dataset is encoded through multiple Transformer encoders in sequence to obtain a feature vector.

[0134] It should be noted that the BERT encoder encodes the embedded medical text dataset through multiple layers of Transformer encoders to obtain a feature vector. Among them, the input of each layer of the Transformer encoder is the output of the previous layer of the Transformer encoder.

[0135] Optionally, in another embodiment of the present application, a specific implementation manner of encoding the embedded medical text dataset through multiple Transformer encoders in sequence after inputting it into the BERT encoder included in the integrated encoder to obtain a feature vector is as Figure 4 shown and includes the following steps:

[0136] S401. For each Transformer encoder respectively, calculate the input data through the multi-head self-attention mechanism module in the Transformer encoder to obtain a multi-head self-attention vector.

[0137] It should be noted that each layer of the Transformer encoder consists of a multi-head self-attention mechanism module, a feed-forward network, and two layer normalization modules, and residual connections are used.

[0138] It should also be noted that when the Transformer encoder is the j-th layer Transformer encoder, the calculation expression of this layer is:

[0139]

[0140] where X j is the input of the j-th layer Transformer encoder, and X j+1 is also the output of the j-th layer Transformer encoder, that is, the input of the (j + 1)-th layer Transformer encoder.

[0141] Then, each layer of the multi-head self-attention mechanism module will perform multi-head attention calculation on the input data, so as to embed the dependencies of the medical text dataset and enhance the representation ability of the model.

[0142] For the j-th layer Transformer encoder, the calculation formula of the multi-head self-attention mechanism is as follows:

[0143]

[0144] where head i is the attention head, are the weight matrices corresponding to the query, key, and value respectively, and Q, K, and V are the query matrix, key matrix, and value matrix respectively. is the scaling factor to prevent the result of the attention mechanism from being too large. Since it is a self-attention mechanism, X j serves as the query, key, and value at the same time. is the multi-head self-attention vector.

[0145] S402. Perform a residual connection on the multi-head self-attention vector and the input data through the residual connection module in the Transformer encoder to obtain the first connection vector.

[0146] It can be understood that the residual connection refers to adding the multi-head self-attention vector and the input data.

[0147] Therefore, the calculation expression of the first connection vector is:

[0148] X j + MultiHeadAttention(X j )

[0149] where Xj is the input to the j-th layer Transformer encoder, which is the multi-head self-attention vector.

[0150] S403. Calculate the normalization of the first connection vector through the first normalization module in the Transformer encoder to obtain the normalized vector.

[0151] Specifically, the calculation formula for the normalized vector (F j ) is:

[0152] F j =LayerNormalization(X j +MultiHeadAttention(X j ))

[0153] where X j is the input to the j-th layer Transformer encoder, is the multi-head self-attention vector, and LayerNormalization is the normalization function.

[0154] S404. Calculate the normalized vector through the feed-forward network in the Transformer encoder to obtain the feed-forward network vector.

[0155] It should be noted that the feed-forward network is a multi-layer perceptron using the ReLU function as the activation function. Therefore, the calculation formula for the normalized vector FFN(F j ) is:

[0156]

[0157] where is the weight matrix of the hidden layer and output layer of the position-based feed-forward network, is the bias of the hidden layer and output layer of the position-based feed-forward network, and both are learnable parameters.

[0158] S405. Perform a residual connection on the normalized vector and the feed-forward network vector through the residual connection module to obtain the second connection vector.

[0159] Specifically, the expression for the second connection vector is: F j +FFN(F j )

[0160] where F j is the normalized vector, and FFN(F j ) is the feed-forward network vector.

[0161] S406. Perform normalization calculation on the second connection vector through the second-layer normalization module in the Transformer encoder to obtain a feature vector.

[0162] Specifically, the feature vector X j+1 is calculated as follows:

[0163] X j+1 = LayerNormalization(F j + FFN(F j ))

[0164] where F j is the normalization vector, FFN(F j ) is the feed-forward network vector, and LayerNormalization is the normalization function of the second-layer normalization module.

[0165] It should be noted that for the BERT encoder, the feature vector output by the last layer of the Transformer encoder is the final vector representation of the medical text, that is, the output of the BERT encoder.

[0166] Optionally, in another embodiment of the present application, a specific implementation manner of inputting the embedded medical text dataset into the ERNIE encoder included in the integrated encoder in step S305 to obtain the target feature vector is as Figure 5 shown, including the following steps:

[0167] S501. After inputting the embedded medical text dataset into the ERNIE encoder included in the integrated encoder, perform encoding processing on the embedded medical text dataset through multiple Transformer encoders in sequence to obtain a feature vector.

[0168] It should be noted that for the specific implementation manner of step S501, reference can be made correspondingly to steps S401 to S406 in the above method embodiment, which will not be elaborated here.

[0169] S502. Perform knowledge encoding on the feature vector through multiple knowledge encoders in sequence to obtain the target feature vector.

[0170] It should be noted that for the ERNIE encoder, the encoding result of the last layer of the Transformer encoder also needs to be further encoded through multiple layers of knowledge encoders to effectively fuse the text semantic information and the knowledge entity information.

[0171] Optionally, in another embodiment of the present application, a specific implementation manner of step S502 is as Figure 6 shown, including the following steps:

[0172] S601. For each knowledge encoder, calculate the token multi-head self-attention vector by using the first multi-head self-attention mechanism module in the knowledge encoder for the feature vector, and calculate the entity multi-head self-attention vector by using the second multi-head self-attention mechanism module in the knowledge encoder for the entity.

[0173] It should be noted that in the knowledge encoder, the input of each layer of the knowledge encoder is the output of the previous layer of the knowledge encoder, which includes two parts, namely, the token input and the entity input. The token input of the first layer of the knowledge encoder is the output of the last layer of the Transformer encoder. Among them, the entity is pre-recognized by using the ERNIE encoder for the medical text to be classified, encoded by using the TransE model for the recognized entity, and then input into the knowledge encoder to obtain.

[0174] For the i-th layer of the knowledge encoder, its token input and entity input are respectively the token output and entity output of the (i - 1)-th layer of the knowledge encoder, and the formulas are as follows:

[0175]

[0176] Among them, are respectively the token output and entity output of the (i - 1)-th layer of the knowledge encoder.

[0177] It should also be noted that the token input and the entity input need to pass through two different multi-head self-attention mechanism modules to obtain the corresponding token output and entity output, and then align the two.

[0178] Therefore, the calculation formulas of the first multi-head self-attention mechanism module and the second multi-head self-attention mechanism module are as follows:

[0179]

[0180] Then, the token multi-head self-attention vector and the entity multi-head self-attention vector are aligned so that the token output and entity output of this layer of the knowledge encoder are obtained through the heterogeneous information fusion layer.

[0181] S602. Determine whether there is an entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector through the heterogeneous information fusion layer in the knowledge encoder.

[0182] It should be noted that in the embodiment of the present application, when obtaining the target feature vector output by the knowledge encoder through the heterogeneous information fusion layer, there are two ways for the heterogeneous information fusion layer to calculate the target feature vector. One is that the token w j has an aligned entity e kThe calculation method, and the other is the token w j There is no aligned entity e k The calculation method. Therefore, in order for the heterogeneous information fusion layer to accurately calculate the target feature vector, it is necessary to determine whether there is an entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector through the heterogeneous information fusion layer in the knowledge encoder at this time. Therefore, if there is an entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector, it means that the token w j There is an aligned entity e k The calculation method, so step S603 is executed. If there is no entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector, it means that the token w j There is no aligned entity e k The calculation method, so step S604 is executed.

[0183] S603. Use the entity fusion algorithm through the heterogeneous information fusion layer to calculate the token multi-head self-attention vector and the entity vector to obtain the target feature vector.

[0184] Specifically, when there is an entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector, at this time, the heterogeneous information fusion layer can use the entity fusion algorithm to calculate the token multi-head self-attention vector and the entity vector, that is, the calculation formula is:

[0185]

[0186] Among them, h j Is the internal hidden state that integrates tokens and entities, and σ() is the non-linear activation function GELU. In this case, there is a fusion of text token information and knowledge entity information.

[0187] S604. Use the fusion algorithm through the heterogeneous information fusion layer to calculate the token multi-head self-attention vector to obtain the target feature vector.

[0188] Specifically, when there is no entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector, at this time, the heterogeneous information fusion layer can use the fusion algorithm to calculate the token multi-head self-attention vector, that is, the calculation formula is:

[0189]

[0190] Among them, h j Is the internal hidden state of the token, and σ() is the non-linear activation function GELU. In this case, there is no fusion of text token information and knowledge entity information.

[0191] It should be noted that the token output of the last-layer knowledge encoder is the final encoding representation of the medical text by the ERNIE encoder, that is, the target feature vector output by the ERNIE encoder.

[0192] S306. Input the feature vector into the fully connected layer included in the integrated encoder to obtain the first sample classification result, and input the target feature vector into the target fully connected layer included in the integrated encoder to obtain the second sample classification result.

[0193] Specifically, after obtaining the feature vector and the target feature vector, the feature vector output by the BERT encoder and the target feature vector output by the ERNIE encoder need to perform classification learning through two different fully connected layers (linear classifiers) respectively to obtain the final output. The specific formula is as follows:

[0194] output1 = Y1W1 + b1

[0195] output2 = Y2W2 + b2

[0196] Among them, Y1 ∈ R n×d and Y2 ∈ R n×d are the encoding results of the BERT encoder and the ERNIE encoder respectively, W1 ∈ R d×l and b1 ∈ R l , W2 ∈ R d×l and b2 ∈ R l are the weight matrix and bias of the fully connected layers corresponding to the BERT encoder and the ERNIE encoder respectively. Among them, d is the feature dimension and l is the number of categories.

[0197] For output1 ∈ R n×l and output2 ∈ R n×l Take the index corresponding to the maximum element of each row of the matrix to obtain L1 ∈ R n and L2 ∈ R n , which are the first sample classification result and the second sample classification result respectively. The formula is as follows:

[0198] L1 = argmax(output1, axis = 1)

[0199] L2 = argmax(output2, axis = 1)

[0200] S307. Calculate the first loss function between the first sample classification result and the actual classification result corresponding to the medical text dataset, and calculate the second loss function between the second sample classification result and the actual classification result.

[0201] It should be noted that in order to know the difference between the model output and the true result and optimize the performance of the model, the first loss function between the first sample classification result and the actual classification result corresponding to the medical text dataset can be calculated, and the second loss function between the second sample classification result and the actual classification result can be calculated to achieve this.

[0202] S308. Determine whether the first loss function converges and whether the second loss function converges.

[0203] Specifically, in order to know whether the BERT encoder and the ERNIE encoder in the integrated encoder are trained successfully and whether accurate classification results can be obtained, it can be achieved by determining whether the first loss function converges and whether the second loss function converges. Therefore, if the first loss function does not converge and the second loss function does not converge, it means that the BERT encoder and the ERNIE encoder still need to be optimized, so step S309 is executed. Therefore, if the first loss function converges and the second loss function converges, it means that accurate classification results can be obtained through the integrated encoder, so step S310 is executed.

[0204] S309. Use the Adam optimizer to adjust the parameters of the BERT encoder and the parameters of the ERNIE encoder.

[0205] Specifically, when the first loss function does not converge and the second loss function does not converge, it is necessary to use the Adam optimizer to adjust the parameters of the BERT encoder and the parameters of the ERNIE encoder, and return to execute step S305 until the first loss function converges and the second loss function converges.

[0206] S310. Determine that the integrated encoder is the trained integrated encoder.

[0207] It should be noted that the training process of the integrated encoder can be referred to Figure 7 the structural schematic diagram of the medical text classification method based on large language model data augmentation shown.

[0208] S104. Perform a voting process on the first classification result and the second classification result to obtain the final classification result corresponding to the medical text to be classified.

[0209] It should be noted that in order to enable the integrated learning of the BERT encoder and the ERNIE encoder and thus improve the final classification performance, it is also necessary to perform a voting operation on the first classification result and the second classification result, and then obtain the final classification result of the medical text to be classified.

[0210] Specifically, the calculation formula for the voting operation is: L = voting(L1, L2).

[0211] Among them, L is the final classification result of the medical text to be classified, L1 is the first classification result, and L2 is the second classification result.

[0212] A classification method for medical texts provided by this application. By obtaining the medical text to be classified, then performing embedding processing on the medical text to be classified to obtain the embedding of the medical text to be classified, and then inputting the embedding of the medical text to be classified into a pre-trained integrated encoder to obtain the first classification result and the second classification result. Among them, the integrated encoder is composed of a BERT encoder and its corresponding fully connected layer and an ERNIE encoder and its corresponding target fully connected layer. The integrated encoder is pre-trained using an enhanced medical text dataset. The enhanced medical text dataset is pre-obtained by using a large language model to retrieve professional terms from an external knowledge database for the medical text dataset, obtaining the explanations of the professional terms, and writing the explanations of the professional terms into the medical text dataset for data enhancement. Finally, voting processing is performed on the first classification result and the second classification result to obtain the final classification result corresponding to the medical text to be classified. Thus, using multiple encoders to encode the enhanced medical text data and obtaining the corresponding classification results, and voting on all classification results can achieve large-scale and efficient classification of medical texts, and further effectively improve the accuracy of classification.

[0213] Another embodiment of this application provides a classification device for medical texts, as Figure 8 shown, including the following units:

[0214] A text acquisition unit 801, configured to acquire a medical text to be classified.

[0215] An embedding processing unit 802, configured to perform embedding processing on the medical text to be classified to obtain the embedding of the medical text to be classified.

[0216] A text input unit 803, configured to input the embedding of the medical text to be classified into a pre-trained integrated encoder to obtain the first classification result and the second classification result.

[0217] Among them, the integrated encoder is composed of a BERT encoder and its corresponding fully connected layer and an ERNIE encoder and its corresponding target fully connected layer. The integrated encoder is pre-trained using an enhanced medical text dataset. The enhanced medical text dataset is pre-obtained by using a large language model to retrieve professional terms from an external knowledge database for the medical text dataset, obtaining the explanations of the professional terms, and writing the explanations of the professional terms into the medical text dataset for data enhancement.

[0218] A voting unit 804, configured to perform voting processing on the first classification result and the second classification result to obtain the final classification result corresponding to the medical text to be classified.

[0219] It should be noted that the specific working process of the above units in the embodiments of the present application can be correspondingly referred to the steps S101 to S104 in the above method embodiments, and will not be elaborated here.

[0220] Optionally, in a classification device for medical texts provided in another embodiment of the present application, the embedding processing unit 102 includes:

[0221] A token embedding processing unit, configured to perform token embedding processing on the medical text to be classified, and obtain the token embedding of the medical text to be classified.

[0222] A segment embedding processing unit, configured to perform segment embedding processing on the medical text to be classified, and obtain the segment embedding of the medical text to be classified.

[0223] A position embedding processing unit, configured to perform position embedding processing on the medical text to be classified, and obtain the position embedding of the medical text to be classified.

[0224] A summing unit, configured to perform summation calculation on the token embedding, the segment embedding, and the position embedding of the medical text to be classified, and obtain the embedding of the medical text to be classified.

[0225] Optionally, in a classification device for medical texts provided in another embodiment of the present application, it further includes:

[0226] An acquisition unit, configured to acquire a medical text dataset.

[0227] A retrieval unit, configured to use a large language model to retrieve professional terms from an external knowledge database for the medical text dataset, and obtain explanations of multiple professional terms.

[0228] An enhancement processing unit, configured to add the explanation of each professional term to the medical text dataset for data enhancement processing, and obtain an enhanced medical text dataset.

[0229] A processing unit, configured to perform embedding processing on the enhanced medical text dataset, and obtain an embedded medical text dataset.

[0230] A first input unit, configured to input the embedded medical text dataset into the BERT encoder included in the integrated encoder to obtain a feature vector, and input the embedded medical text dataset into the ERNIE encoder included in the integrated encoder to obtain a target feature vector.

[0231] A second input unit, configured to input the feature vector into the fully connected layer included in the integrated encoder to obtain a first sample classification result, and input the target feature vector into the target fully connected layer included in the integrated encoder to obtain a second sample classification result.

[0232] A loss function calculation unit for calculating a first loss function between the first sample classification result and the actual classification result corresponding to the medical text dataset, and calculating a second loss function between the second sample classification result and the actual classification result.

[0233] A judgment unit for judging whether the first loss function converges and whether the second loss function converges.

[0234] An adjustment unit for, if the first loss function does not converge and the second loss function does not converge, using an Adam optimizer to adjust the parameters of the BERT encoder and the parameters of the ERNIE encoder, and returning to execute inputting the embedded medical text dataset into the BERT encoder included in the integrated encoder to obtain a feature vector, and inputting the embedded medical text dataset into the ERNIE encoder included in the integrated encoder to obtain a target feature vector.

[0235] A determination unit for, if the first loss function converges and the second loss function converges, determining that the integrated encoder is a trained integrated encoder.

[0236] Optionally, in a medical text classification device provided in another embodiment of the present application, the first input unit includes:

[0237] A first encoding processing unit for, after inputting the embedded medical text dataset into the BERT encoder included in the integrated encoder, sequentially encoding the embedded medical text dataset through a plurality of Transformer encoders to obtain a feature vector.

[0238] Optionally, in a medical text classification device provided in another embodiment of the present application, the first encoding processing unit includes:

[0239] A first calculation unit for, respectively for each Transformer encoder, calculating the input data through the multi-head self-attention mechanism module in the Transformer encoder to obtain a multi-head self-attention vector.

[0240] A first residual connection unit for performing a residual connection on the multi-head self-attention vector and the input data through the residual connection module in the Transformer encoder to obtain a first connection vector.

[0241] A first normalization calculation unit for performing a normalization calculation on the first connection vector through the first layer normalization module in the Transformer encoder to obtain a normalized vector.

[0242] A second calculation unit for calculating the normalized vector through the feed-forward network in the Transformer encoder to obtain a feed-forward network vector.

[0243] A second residual connection unit, configured to perform a residual connection on the normalized vector and the feed-forward network vector through a residual connection module to obtain a second connection vector.

[0244] A second normalization calculation unit, configured to perform a normalization calculation on the second connection vector through the second layer normalization module in the Transformer encoder to obtain a feature vector.

[0245] Optionally, in a medical text classification device provided in another embodiment of the present application, the first input unit includes:

[0246] A second encoding processing unit, configured to input the embedded medical text dataset into the ERNIE encoder included in the integrated encoder, and then perform encoding processing on the embedded medical text dataset through a plurality of Transformer encoders in sequence to obtain a feature vector.

[0247] A knowledge encoding unit, configured to perform knowledge encoding on the feature vector through a plurality of knowledge encoders in sequence to obtain a target feature vector.

[0248] Optionally, in a medical text classification device provided in another embodiment of the present application, the knowledge encoding unit includes:

[0249] A third calculation unit, configured to respectively for each knowledge encoder, calculate the feature vector through the first multi-head self-attention mechanism module in the knowledge encoder to obtain a token multi-head self-attention vector, and calculate the entity through the second multi-head self-attention mechanism module in the knowledge encoder to obtain an entity multi-head self-attention vector. Wherein, the entity is pre-recognized by the ERNIE encoder for the medical text to be classified, encoded by the TransE model for the recognized entity, and then input into the knowledge encoder.

[0250] An entity judgment unit, configured to judge whether there is an entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector through the heterogeneous information fusion layer in the knowledge encoder.

[0251] A fourth calculation unit, configured to if there is an entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector, calculate the token multi-head self-attention vector and the entity vector through the heterogeneous information fusion layer using an entity fusion algorithm to obtain a target feature vector.

[0252] A fifth calculation unit, configured to if there is no entity vector corresponding to the token multi-head self-attention vector in the entity multi-head self-attention vector, calculate the token multi-head self-attention vector through the heterogeneous information fusion layer using a fusion algorithm to obtain a target feature vector.

[0253] It should be noted that the specific working processes of the respective units provided in the above embodiments of the present application can be correspondingly referred to the corresponding steps in the above method embodiments, and will not be elaborated herein.

[0254] It should also be noted that a classification device for medical texts provided in the embodiments of the present application has the technical effects of any one of the above embodiments, and will not be elaborated herein.

[0255] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0256] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for classifying medical texts, characterized in that: include: Obtain medical text to be classified; Performing embedding processing on the medical text to be classified to obtain embedding of the medical text to be classified; The embedding of the medical text to be classified is input into a pre-trained integrated encoder to obtain a first classification result and a second classification result; wherein the integrated encoder is composed of a BERT encoder and its corresponding fully connected layer and an ERNIE encoder and its corresponding target fully connected layer; the integrated encoder is pre-trained using an enhanced medical text dataset; the enhanced medical text dataset is pre-trained using a large language model to retrieve professional terms from an external knowledge database for a medical text dataset, obtain explanations of professional terms, and write the explanations of professional terms into the medical text dataset for data enhancement; Voting is performed on the first classification result and the second classification result to obtain a final classification result corresponding to the medical text to be classified.

2. The method according to claim 1, characterized in that The embedding process of the medical text to be classified to obtain the embedding of the medical text to be classified includes: Performing word element embedding processing on the medical text to be classified to obtain word element embedding of the medical text to be classified; Performing segment embedding processing on the medical text to be classified to obtain segment embedding of the medical text to be classified; Performing position embedding processing on the medical text to be classified to obtain position embedding of the medical text to be classified; The word element embedding of the medical text to be classified, the segment embedding of the medical text to be classified and the position embedding of the medical text to be classified are summed up to obtain the embedding of the medical text to be classified.

3. The method according to claim 1, characterized in that The training method of the integrated encoder comprises: Get the medical text dataset; Using a large language model to retrieve professional terms from an external knowledge database for the medical text dataset, and obtaining explanations of multiple professional terms; Adding the explanation of each of the professional terms to the medical text dataset for data enhancement processing to obtain an enhanced medical text dataset; Performing embedding processing on the enhanced medical text dataset to obtain an embedded medical text dataset; Inputting the embedded medical text data set into a BERT encoder included in an integrated encoder to obtain a feature vector, and inputting the embedded medical text data set into an ERNIE encoder included in the integrated encoder to obtain a target feature vector; Inputting the feature vector into a fully connected layer included in the integrated encoder to obtain a first sample classification result, and inputting the target feature vector into a target fully connected layer included in the integrated encoder to obtain a second sample classification result; Calculating a first loss function between the first sample classification result and an actual classification result corresponding to the medical text data set, and calculating a second loss function between the second sample classification result and the actual classification result; Determine whether the first loss function converges and whether the second loss function converges; If the first loss function does not converge and the second loss function does not converge, the parameters of the BERT encoder and the parameters of the ERNIE encoder are adjusted using the Adam optimizer, and the embedded medical text data set is input into the BERT encoder included in the integrated encoder to obtain a feature vector, and the embedded medical text data set is input into the ERNIE encoder included in the integrated encoder to obtain a target feature vector; If the first loss function converges and the second loss function converges, it is determined that the integrated encoder is a trained integrated encoder.

4. The method according to claim 3, characterized in that The step of inputting the embedded medical text dataset into the BERT encoder included in the integrated encoder to obtain a feature vector includes: After the embedded medical text dataset is input into the BERT encoder included in the integrated encoder, the embedded medical text dataset is encoded through multiple Transformer encoders in sequence to obtain a feature vector.

5. The method according to claim 4, characterized in that After the embedded medical text dataset is input into the BERT encoder included in the integrated encoder, the embedded medical text dataset is encoded by multiple Transformer encoders in sequence to obtain a feature vector, including: For each of the Transformer encoders, respectively, the input data is calculated by the multi-head self-attention mechanism module in the Transformer encoder to obtain a multi-head self-attention vector; Performing a residual connection on the multi-head self-attention vector and the input data through a residual connection module in the Transformer encoder to obtain a first connection vector; Performing normalization calculation on the first connection vector by a first-layer normalization module in the Transformer encoder to obtain a normalized vector; Calculating the standard vector through a feedforward network in the Transformer encoder to obtain a feedforward network vector; Performing a residual connection on the standard vector and the feedforward network vector through the residual connection module to obtain a second connection vector; The second connection vector is normalized by a second-layer normalization module in the Transformer encoder to obtain a feature vector.

6. The method according to claim 3, characterized in that The step of inputting the embedded medical text dataset into the ERNIE encoder included in the integrated encoder to obtain a target feature vector includes: After the embedded medical text data set is input into the ERNIE encoder included in the integrated encoder, the embedded medical text data set is encoded by multiple Transformer encoders in sequence to obtain a feature vector; The feature vector is knowledge encoded by a plurality of knowledge encoders in turn to obtain a target feature vector.

7. The method according to claim 6, characterized in that The step of sequentially performing knowledge encoding on the feature vector through a plurality of knowledge encoders to obtain a target feature vector includes: For each of the knowledge encoders, the feature vector is calculated by the first multi-head self-attention mechanism module in the knowledge encoder to obtain a word-unit multi-head self-attention vector, and the entity is calculated by the second multi-head self-attention mechanism module in the knowledge encoder to obtain an entity multi-head self-attention vector; wherein, the entity is preliminarily identified by using the ERNIE encoder for the medical text to be classified, and the identified entity is encoded by the TransE model and then input into the knowledge encoder to obtain; Determining, by the heterogeneous information fusion layer in the knowledge encoder, whether there is an entity vector corresponding to the word-unit multi-head self-attention vector in the entity multi-head self-attention vector; If the entity multi-head self-attention vector contains an entity vector corresponding to the word-unit multi-head self-attention vector, the heterogeneous information fusion layer uses an entity fusion algorithm to calculate the word-unit multi-head self-attention vector and the entity vector to obtain a target feature vector; If the entity vector corresponding to the word-unit multi-head self-attention vector does not exist in the entity multi-head self-attention vector, the word-unit multi-head self-attention vector is calculated by the heterogeneous information fusion layer using a fusion algorithm to obtain a target feature vector.

8. A medical text classification device, characterized in that: include: A text acquisition unit, used for acquiring medical text to be classified; An embedding processing unit, used for performing embedding processing on the medical text to be classified to obtain an embedding of the medical text to be classified; A text input unit, used for inputting the embedding of the medical text to be classified into a pre-trained integrated encoder to obtain a first classification result and a second classification result; wherein the integrated encoder is composed of a BERT encoder and its corresponding fully connected layer and an ERNIE encoder and its corresponding target fully connected layer; the integrated encoder is pre-trained using an enhanced medical text dataset; the enhanced medical text dataset is pre-trained using a large language model to retrieve professional terms from an external knowledge database for a medical text dataset, obtain explanations of professional terms, and write the explanations of professional terms into the medical text dataset for data enhancement; The voting unit is used to vote on the first classification result and the second classification result to obtain a final classification result corresponding to the medical text to be classified.

9. The device according to claim 8, characterized in that The embedded processing unit comprises: A word element embedding processing unit, used for performing word element embedding processing on the medical text to be classified to obtain word element embedding of the medical text to be classified; A segment embedding processing unit, used for performing segment embedding processing on the medical text to be classified to obtain segment embedding of the medical text to be classified; A position embedding processing unit, used for performing position embedding processing on the medical text to be classified to obtain position embedding of the medical text to be classified; A summing unit is used to sum the word element embedding of the medical text to be classified, the segment embedding of the medical text to be classified, and the position embedding of the medical text to be classified to obtain the embedding of the medical text to be classified.

10. The device according to claim 8, characterized in that Also includes: An acquisition unit, used for acquiring a medical text dataset; A retrieval unit, used to use a large language model to perform professional term retrieval on the medical text data set from an external knowledge database to obtain explanations of multiple professional terms; An enhancement processing unit, used for adding the explanation of each of the professional terms to the medical text dataset for data enhancement processing to obtain an enhanced medical text dataset; A processing unit, configured to perform embedding processing on the enhanced medical text dataset to obtain an embedded medical text dataset; A first input unit is used to input the embedded medical text data set into a BERT encoder included in the integrated encoder to obtain a feature vector, and input the embedded medical text data set into an ERNIE encoder included in the integrated encoder to obtain a target feature vector; A second input unit, used for inputting the feature vector into a fully connected layer included in the integrated encoder to obtain a first sample classification result, and inputting the target feature vector into a target fully connected layer included in the integrated encoder to obtain a second sample classification result; A loss function calculation unit, used to calculate a first loss function between the first sample classification result and an actual classification result corresponding to the medical text data set, and to calculate a second loss function between the second sample classification result and the actual classification result; A judging unit, configured to judge whether the first loss function converges and whether the second loss function converges; An adjustment unit, configured to adjust the parameters of the BERT encoder and the parameters of the ERNIE encoder using an Adam optimizer if the first loss function has not converged and the second loss function has not converged, and return to execute the steps of inputting the embedded medical text dataset into the BERT encoder included in the integrated encoder to obtain a feature vector, and inputting the embedded medical text dataset into the ERNIE encoder included in the integrated encoder to obtain a target feature vector; A determining unit is used to determine that the integrated encoder is a trained integrated encoder if the first loss function converges and the second loss function converges.