Chinese medical continuous entity recognition method based on continuous learning and knowledge distillation

Through the Chinese medical continuous entity recognition method based on knowledge distillation, the catastrophic forgetting problem of the Chinese medical named entity recognition model in the face of new entity types is solved, efficient entity recognition and continuous learning are achieved, and the adaptability and accuracy of the model are improved.

CN120337925APending Publication Date: 2025-07-18CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510335519.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Chinese medical named entity recognition models are prone to catastrophic forgetting and interference problems when facing new entity types, and it is difficult to continuously learn and adapt to data changes.

Method used

Using the Chinese medical continuous entity recognition method based on knowledge distillation, entity feature representation is generated through the BERT coding layer, combining the bidirectional GRU model to capture context information, introduce relative position coding and specific feedforward network coding entity spans, and use binary cross entropy loss and Bernoulli KL divergence loss to fuse new and old knowledge to ensure that the model does not forget old knowledge in incremental learning.

Benefits of technology

It improves the accuracy and generalization ability of entity recognition, can effectively handle nested or overlapping entities, enhances the cross-level learning and multi-label classification capabilities of the model, adapts to the needs of multi-task learning, reduces interference between new and old knowledge, and improves the continuous learning ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337925A_ABST
    Figure CN120337925A_ABST
Patent Text Reader

Abstract

A Chinese medical continuous entity recognition method based on continuous learning and knowledge distillation comprises the following steps: step 1, inputting a Chinese medical text containing continuous named entities into a CK-CMCNER model, and generating entity feature representation through a BERT coding layer; 2, in the feature extraction layer, capturing forward and backward context information between lexical elements through a bidirectional GRU model, and weighting and combining forward and backward hidden states to obtain more accurate lexical element feature representation; step 3, in a span presentation layer, adopting a feedforward network specific to an entity type to encode the starting and ending positions of each span, enhancing the presentation capability through residual connection, and introducing relative position codes to refine position information between lexical elements, so as to improve the recognition precision of the model on nested or overlapped entities; and 4, in the multi-label loss layer, independently classifying each entity type through a binary cross entropy BCE loss function, and combining a knowledge distillation technology to ensure that new knowledge learning does not cause old knowledge forgetting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, in particular to entity recognition technology, and specifically to a Chinese medical continuous entity recognition method based on continuous learning and knowledge distillation. Background Art

[0002] In many research fields of natural language processing, named entity recognition has always been a core and extremely challenging task. The progress of deep learning technology has promoted the neural network to achieve remarkable achievements in a variety of standard tasks, thus effectively promoting the development of named entity recognition technology.

[0003] In recent years, in the field of named entity recognition, the application of continuous learning has made remarkable progress. The core challenge in this research field is to enable the NER model to continuously learn and adapt to new entity types while maintaining the recognition ability of previously learned entity types. It is worth noting the "Learn and Review" (L&R) method proposed by Xia et al. This method includes a novel two-stage framework aimed at alleviating the inter-type confusion problem in the type-incremental setting of continuous NER. The learning stage of the framework includes knowledge distillation from the teacher model to the student model, with the focus on retaining old knowledge while acquiring new information. De Lange et al. designed a meta-function pre-training algorithm, modeled the PLM as a meta-function, and injected the NER ability in the context into the PLM by comparing the extractor constructed by (instruction, demonstration) and the proxy gold extractor. The SKD-NER model proposed by Chen et al. provides an innovative method for CL-NER. This span-based model incorporates a reinforcement learning strategy to enhance the model's ability to resist catastrophic forgetting. By utilizing knowledge distillation and optimizing the soft label and distillation loss during the learning process, SKD-NER effectively alleviates the forgetting problem, enabling the model to retain the previously learned knowledge while acquiring new knowledge. Zhang et al. proposed a model SpanKL, which uses different losses according to whether different entities are previously learned or currently being learned. This model cleverly balances the trade-off between retaining old entity types and acquiring new entity types, thus more effectively alleviating the catastrophic forgetting problem, especially solving the semantic shift problem of non-entity types in continuous NER.

[0004] There are many polysemous and unannotated entities in medical texts, and many new entities will appear in reality, requiring the model to be able to continuously learn new knowledge. With the arrival of new entities and new tasks, the model will face the problems of "catastrophic forgetting" or interference. For example, Biesialska et al. found that deep learning models suffer from catastrophic forgetting or interference problems when continuously learning a series of tasks. The entity types that have been learned may appear in the training samples of new tasks, but there are no relevant annotations. The non-entities currently being learned may also belong to a certain entity type to be learned in future tasks. These entities will force the model to forget old knowledge to adapt to new conflicting knowledge. The present invention is based on a neural span model, improves and optimizes it for the field of Chinese medical entity recognition, and proposes a Chinese medical continuous entity recognition technology based on continuous learning and knowledge distillation. Summary of the Invention

[0005] The purpose of the present invention is to solve the technical problem of catastrophic forgetting existing in Chinese medical named entity recognition, and to propose a Chinese medical continuous entity recognition method based on knowledge distillation.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0007] A Chinese medical continuous entity recognition method based on knowledge distillation, comprising the following steps:

[0008] Step 1: Input the Chinese medical text containing continuous named entities into the CK-CMCNER model, and generate the feature representation of each entity through the text encoding layer BERT.

[0009] Step 2: In the feature extraction layer, capture the forward and backward context information between tokens through a bidirectional GRU model, and weight and combine the forward and backward hidden states to obtain a more accurate token feature representation.

[0010] Step 3: In the span representation layer, use a feed-forward network specific to the entity type to encode the start and end positions of each span, enhance the representation ability through residual connections, and introduce relative position encoding to refine the position information between tokens, improving the model's recognition accuracy for nested or overlapping entities.

[0011] Step 4: In the multi-label loss layer, independently classify each entity type through the binary cross-entropy BCE loss function, and combine knowledge distillation technology to ensure that the learning of new knowledge will not cause the forgetting of old knowledge, and improve the continuous learning ability of the model.

[0012] In step 1, it specifically includes the following steps:

[0013] Step 1-1: Connect each token with its corresponding representation in the pre-trained model, and use E = [e1, e2,..., e n to represent the embedding vector of the input sentence X, where

[0014] Step 1-2: After being processed by the text encoding layer BERT, the context hidden vectors are obtained. Among them, the representation of each token is:

[0015] E = Embed(X), H = CtxEnc(E)

[0016] where Embed is the embedding layer, CtxEnc is the text encoder, and d e and d h are the embedding dimension and the hidden dimension respectively, and the encoder is shared for all tasks.

[0017] In Step 2, it specifically includes the following steps:

[0018] Step 2-1: For each token x i in the input sequence, the bidirectional GRU model calculates the current hidden state according to the previous hidden state and the current token x i as follows:

[0019] The update of the forward hidden state is:

[0020]

[0021] The update of the backward hidden state is:

[0022]

[0023] Step 2-2: Combine the forward and backward hidden states of each token through a linear transformation to obtain the final feature representation:

[0024]

[0025] where ω f and ω b are the weight matrices of the forward and backward hidden states, and b h is the bias term.

[0026] In Step 3, it specifically includes the following steps:

[0027] Step 3-1: For any token sequence in the given input sentence, define a span, which is composed of a continuous token sequence from the start token h i to the end token h j ;

[0028] Step 3-2: Encode the start and end positions of each span through a feed-forward network specific to the entity type, thereby generating span features for each entity type k. The process of representing the span is as follows:

[0029]

[0030] where i and j represent the start and end of the span respectively, and k represents the k-th entity type;

[0031] Step 3-3: By fusing the hidden states after relative position encoding, the feature representation of the span can be further refined:

[0032]

[0033] where k represents the entity type to be recognized, R i and R j represent the relative position encodings of the start and end tokens; as the number of tasks increases, simply adding more span representation layers (substantially internal feed-forward networks) for new tasks can expand the learning ability of the model.

[0034] In step 4, it specifically includes the following steps:

[0035] Step 4-1: Calculate the prediction results for each span and convert these prediction results into probability form through the sigmoid function;

[0036] Step 4-2: Use the binary cross-entropy BCE loss to compare the difference between the predicted probability and the true label; for the k-th entity type in span s ij the predicted probability is:

[0037]

[0038] Step 4-3: When the model enters the t-th incremental learning step (t > 1), first use the previously trained model M t-1 (as the teacher model) to make a one-time prediction on the entity types learned in the current training set D t up to the previous step in order to generate soft labels for each span, aiming to simulate the teacher model's understanding of the entity types learned in the previous steps; the generated soft labels provide a predicted value of a Bernoulli distribution for each old entity type of each span;

[0039] Step 4-4: Use these soft labels to calculate the Bernoulli KL divergence loss of the current model M t (student model) to evaluate the prediction difference between the student model and the teacher model; the loss calculation formula in this step is:

[0040]

[0041] Step 4-5: Finally, the loss function adopted by the model after multiple training iterations is a weighted combination of binary cross-entropy loss and Bernoulli KL divergence loss:

[0042] L = αL BCE + βL KD

[0043] where α and β are the weight values of the two losses; this ensures that the model can effectively integrate new and old knowledge and maintain accurate recognition of previous entity types.

[0044] Compared with the prior art, the present invention has the following technical effects:

[0045] 1) The method of the present invention has high entity recognition ability. Specifically, by combining BERT and bidirectional GRU, the CK-CMCNER model can deeply capture the context semantics and complex relationships between tokens, thereby improving the accuracy of entity recognition;

[0046] 2) The method of the present invention enhances the cross-level learning ability. Specifically, the introduced relative position encoding and span representation layer not only improve the model's perception ability of entity boundaries, but also can effectively handle nested or overlapping entities, improving the generalization ability of the model;

[0047] 3) The method of the present invention has flexible multi-label classification ability. Specifically, different from the traditional multi-class classification method, CK-CMCNER adopts a multi-label classification strategy and achieves higher flexibility through binary cross-entropy loss, and can adapt to the needs of multi-task learning;

[0048] 4) The method of the present invention combines incremental learning with knowledge fusion. Specifically, through knowledge distillation, CK-CMCNER can effectively integrate new and old knowledge during the incremental learning process, avoid forgetting old knowledge, and at the same time improve the recognition ability of new entity types, and has strong continuous learning ability;

[0049] 5) The method of the present invention has strong generalization ability. Compared with traditional models, this model has achieved good results in both Chinese and English datasets. It is especially optimized for the named entity recognition task in the Chinese medical field and can more accurately recognize complex entity types and context relationships in medical texts;

[0050] 6) The Chinese Medical Continuous Entity Recognition Model based on Knowledge Distillation (CK-CMCNER) proposed by the present invention integrates continuous learning and knowledge distillation on the basis of the neural span model. Through the continuous learning mechanism, the model can effectively adapt to the continuous changes of data, while the knowledge distillation technology is used in the model update process to effectively transfer and retain the previously learned knowledge. The CK-CMCNER model uses the BERT pre-trained model combined with a complex neural network as the text encoder, introduces a bidirectional gated recurrent unit network for feature extraction, and deeply optimizes the input of the model. By integrating continuous learning and knowledge distillation methods, it uses the previously learned (teacher) model to predict distillation labels on the current samples, and then jointly trains the current model (student) through these labels and the current gold labels. Through this operation, it can be backward compatible and more effectively integrate new and old knowledge in continuous learning tasks. At the same time, it is forward compatible through binary classification, modifies the traditional sequence labeling method, and independently models at the span and entity levels, thereby reducing interference with future tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The present invention will be further described below with reference to the drawings and embodiments:

[0052] Figure 1 is the overall flowchart of the present invention;

[0053] Figure 2 is the performance graph of some entities in the ultrasound dataset at each step in the embodiment of the present invention;

[0054] Figure 3 is the performance graph of six entities at each step on the CCKS2020 Chinese medical dataset in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] As Figure 1 shown, a Chinese medical continuous entity recognition method based on knowledge distillation includes the following steps:

[0056] Step 1: Input the Chinese medical text containing continuous named entities into the CK-CMCNER model, and generate the feature representation of each entity through the text encoding layer BERT;

[0057] Step 2: In the feature extraction layer, capture the forward and backward context information between tokens through the bidirectional GRU model, and perform weighted combination of the forward and backward hidden states to obtain a more accurate token feature representation;

[0058] Step 3: In the span representation layer, a feed-forward network specific to the entity type is used to encode the start and end positions of each span, enhance the representation ability through residual connections, and introduce relative position encoding to refine the position information between tokens, improving the model's recognition accuracy for nested or overlapping entities;

[0059] Step 4: In the multi-label loss layer, each entity type is independently classified through the binary cross-entropy (BCE) loss function, and combined with knowledge distillation technology to ensure that the learning of new knowledge does not lead to the forgetting of old knowledge, improving the model's continuous learning ability.

[0060] In Step 1, it specifically includes the following steps:

[0061] Step 1-1: Connect each token with its corresponding representation in the pre-trained model, and use E = [e1, e2,..., e n to represent the embedding vector of the input sentence X, where

[0062] Step 1-2: After being processed by the text encoding layer BERT, the context hidden vectors are obtained where the representation of each token is:

[0063] E = Embed(X), H = CtxEnc(E)

[0064] where Embed is the embedding layer, CtxEnc is the text encoder, d e and d h are the embedding dimension and the hidden dimension respectively, and the encoder is shared for all tasks.

[0065] In Step 2, it specifically includes the following steps:

[0066] Step 2-1: For each token x i in the input sequence, the bidirectional GRU calculates the current hidden state according to the previous hidden state and the current token x i representation:

[0067] The update of the forward hidden state is:

[0068]

[0069] The update of the backward hidden state is:

[0070]

[0071] Step 2-2: Combine the forward and backward hidden states of each token through a linear transformation to obtain the final feature representation:

[0072]

[0073] Among them, ω f and ω b are the weight matrices of the forward and backward hidden states, and b h is the bias term.

[0074] In step 3, it specifically includes the following steps:

[0075] Step 3-1: For any sequence of tokens in the given input sentence, first define a span, which is composed of a continuous sequence of tokens from the starting token h i to the ending token h j inclusive.

[0076] Step 3-2: Encode the starting and ending positions of each span through a feed-forward network specific to the entity type, thereby generating span features for each entity type k. The process of representing a span is described as:

[0077]

[0078] where i and j represent the start and end of the span respectively, and k represents the k-th entity type.

[0079] Step 3-3: By fusing the hidden states encoded with relative position encoding, the feature representation of the span can be further refined:

[0080]

[0081] where k represents the entity type to be recognized, R i and R j represent the relative position encodings of the starting and ending tokens. As the number of tasks increases, simply adding more span representation layers (substantially internal feed-forward networks) for new tasks can expand the learning ability of the model.

[0082] In step 4, it specifically includes the following steps:

[0083] Step 4-1: Calculate the prediction results for each span and convert these prediction results into probability form through the sigmoid function;

[0084] Step 4-2: Use binary cross-entropy (BCE) loss to compare the difference between the predicted probability and the true label (i.e., the gold label). For the k-th entity type in the span s ij , its predicted probability is:

[0085]

[0086] Step 4-3: When the model enters the t-th incremental learning step (t > 1), first use the previously trained model M t-1 (as the teacher model) to make a one-time prediction on the current training set D t for the entity types learned up to the previous step until the previous step, aiming to generate soft labels for each span to simulate the teacher model's perception of the entity types learned in the previous steps. The generated soft labels provide a predicted value of a Bernoulli distribution for each old entity type of each span.

[0087] Step 4-4: Use these soft labels to calculate the Bernoulli KL divergence loss of the current model M t (the student model) to evaluate the prediction difference between the student model and the teacher model. The loss calculation formula in this step is:

[0088]

[0089] Step 4-5: Finally, the loss function adopted by the model after multiple training iterations is a weighted combination of binary cross-entropy loss and Bernoulli KL divergence loss:

[0090] L = αL BCE + βL KD

[0091] where α and β are the weight values of the two losses. This ensures that the model can effectively integrate new and old knowledge and maintain accurate recognition of previous entity types.

[0092] Example:

[0093] The present invention was tested on a breast cancer ultrasound examination report dataset and the CCKS2020 Chinese medical dataset. The breast cancer ultrasound examination report dataset contains 3,100 reports, and the CCKS2020 Chinese medical dataset contains 1,050 detailed annotated medical record reports. The samples were divided into a training set, a validation set, and a test set in a ratio of 8:1:1 for model training, parameter adjustment, and performance evaluation, respectively.

[0094] Table 1 shows the comparison of experimental results on the breast cancer ultrasound examination report dataset. The present invention has the best experimental effect on this dataset, with the precision rate and F1 value reaching the highest level. The recall rate is slightly lower than that of the GraphModel-Dict model, but the precision rate and F1 value are 0.88% and 0.37% higher than the highest values of other models, respectively. Experiments show that the CK-CMCNER model is more excellent in overall performance and the recognition of specific entity types. At the same time, after specific domain adaptation adjustment, the performance of the CK-CMCNER model has been significantly improved, demonstrating its excellent ability in processing complex medical data.

[0095] Table 1 Comparison of Experimental Results of Breast Cancer Ultrasound Examination Report Datasets

[0096]

[0097] To comprehensively evaluate and compare the effectiveness of these NER models in the entity recognition task of breast cancer ultrasound examination reports, this study carefully calculated and analyzed their performance on different entity recognition tasks, as shown in Table 2. Through analysis, it was found that for each entity type, the CK-CMCNER model demonstrated superior or nearly the highest F1 score on the vast majority of entity types. In medical practice, this high-precision entity recognition ability can greatly assist doctors in making diagnoses. Among the six models, the F1 values for the skin entity type were almost all 0. Judging from the statistical information of the provided dataset, the skin entity type had the lowest occurrence frequency in the entire dataset. For entities with relatively low occurrence frequencies, the CK-CMCNER model showed significant performance improvements, which is of great significance for the overall model performance and the accuracy of medical diagnoses.

[0098] Table 2 Performance Evaluation of Entity Types of Each Model in Breast Cancer Ultrasound Examination Report Datasets

[0099]

[0100]

[0101] To verify the applicability of the present invention in common medical fields, the present invention used the CCKS2020 Chinese medical public dataset to evaluate the performance of BERT-BiLSTM-CRF, NFLAT, GraphModel-Dict, SpanKL, and CK-CMCNER. Among them, the CK-CMCNER model had the best performance, with the precision rate, recall rate, and F1 value leading other models, reaching 85.92%, 86.62%, and 86.23% respectively, as shown in Table 3. From the experimental results, it can be seen that the CK-CMCNER model demonstrated strong generalization ability. On different types of medical texts, this model could maintain stable and efficient performance, showing its adaptability to diverse medical data.

[0102] Table 3 Comparison of Experimental Results of CCKS2020 Chinese Medical Datasets

[0103]

[0104] For various entities in the CCKS2020 Chinese medical dataset, the present invention also conducted a detailed computational analysis of the above model to evaluate their performance in different entity recognition tasks, as shown in Table 4. The CK-CMCNER model demonstrated significant superiority in the recognition tasks of various entity types. Although it did not rank first in every entity category, its comprehensive performance had an obvious leading edge compared to other models. Whether for common or relatively complex entity types, the CK-CMCNER model showed high adaptability and accuracy. Through fine model design and optimization, the CK-CMCNER model can better understand and handle the characteristics of professional medical texts.

[0105] Table 4 Performance Evaluation of Entity Types of Each Model in the CCKS2020 Dataset

[0106]

[0107] To deeply analyze the details of the performance improvement of the CK-CMCNER model for different entity types, the present invention conducted a real-time performance analysis. In this analysis, according to a specific learning order, the present invention plotted the F1 value curves for each recognition of entity types on two datasets to visually present these performance changes in the form of charts. The experiments showed that due to the low frequency of occurrence of certain entities in the breast cancer ultrasound examination report dataset. For example, the entity "skin" only appeared 18 times. This low-frequency characteristic led to significant fluctuations in the immediate performance of entity recognition even with a slight change (increase or decrease by one) in the number of correctly recognized entities. To more effectively show the performance differences between the CL method and the non-CL method, the present invention chose to display the real-time performance curves of only some entities in the breast cancer ultrasound examination report dataset, including size, duct changes, echo, margin, and blood flow. The present invention divided the entities in the ultrasound dataset into two groups and plotted their performance charts respectively, as Figure 2 shown.

[0108] On the CCKS2020 Chinese medical dataset, the present invention also conducted the above-mentioned entity performance analysis, as Figure 3 shown.

[0109] Although different entities exhibit different performances due to their inherent difficulties, the advantages of CK-CMCNER can still be seen: 1) The performance curve of CK-CMCNER is stable. During the medical multi-task recognition process, the phenomenon of forgetting is a common challenge, especially when facing diverse entity types. The CK-CMCNER model demonstrates a relatively flat performance curve for most entity types through the effective integration of continuous learning and knowledge distillation techniques. This stability not only means that the model forgets less when learning new tasks but also reflects its reliability and resilience in handling long-term and complex medical data processing tasks. 2) CK-CMCNER is adaptable and stable when the number of recognition tasks increases. Generally, as the number of tasks increases, the accuracy of entity recognition will decrease to a certain extent. For example, in the non-CL case, although the F1 value of laboratory test entities is relatively high at the beginning of the task, its performance significantly decreases as the number of tasks increases, and the real-time performance fluctuations of other entities are also large. After adding CL, its stability has been significantly improved.

Claims

1. A Chinese medical continuous entity recognition method based on continuous learning and knowledge distillation, characterized in that It includes the following steps: Step 1: Input the Chinese medical text containing consecutive named entities into the text encoding layer BERT to generate entity feature representations; Step 2: Capture the forward and backward context information between tokens through a bidirectional GRU model, and combine the forward and backward hidden states to obtain more accurate token feature representations; Step 3: Encode the start and end positions of the span using a feed-forward network specific to the entity type, enhance the representation ability through residual connections, and introduce relative position encoding to improve the recognition accuracy of nested or overlapping entities; Step 4: Independently classify each entity type through the binary cross-entropy BCE loss function, and combine knowledge distillation technology to ensure that learning new knowledge does not cause forgetting of old knowledge.

2. The method according to claim 1, wherein In Step 1, it specifically includes the following steps: Step 1-1: Connect each token with its corresponding representation in the pre-trained model, and use E = [e1, e2,..., e n to represent the embedding vector of the input sentence X, where Step 1-2: After being processed by the text encoding layer BERT, context hidden vectors are obtained. Among them, the representation of each token is: E = Embed(X), H = CtxEnc(E) Among them, Embed is the embedding layer, CtxEnc is the text encoder, d e and d h are the embedding dimension and the hidden dimension respectively, and the encoder is shared for all tasks.

3. The method according to claim 1, wherein In Step 2, it specifically includes the following steps: Step 2-1: For each token x in the input sequence i , the bidirectional GRU calculates the current hidden state based on the previous hidden state and the current token x i as follows: The update of the forward hidden state is: The update of the backward hidden state is: Step 2-2: Combine the forward and backward hidden states of each token through a linear transformation to obtain the final feature representation.

4. The method according to claim 1, wherein In Step 3, it specifically includes the following steps: Step 3-1: For any sequence of tokens in a given input sentence, delimit a span that consists of a consecutive sequence of tokens from a starting token h i to an ending token h j ; Step 3-2: Encode the start and end positions of each span through a feed-forward network specific to the entity type, thereby generating span features for each entity type k; Step 3-3: Further refine the feature representation of the span by fusing the hidden states after relative position encoding.

5. The method according to claim 1, characterized in that, In Step 4, it specifically includes the following steps: Step 4-1: Calculate the prediction results for each span and convert these prediction results into probability form through the sigmoid function; Step 4-2: Use the binary cross-entropy (BCE) loss to compare the difference between the predicted probability and the true label; for the k-th entity type in the span s ij ; Step 4-3: When the model enters the t-th incremental learning step, where t > 1, first use the previously trained model M t-1 , to make a one-time prediction of the entity types learned so far in the current training set D t up to the previous step in order to generate soft labels for each span, so as to simulate the teacher model's perception of the entity types learned in the previous steps; the generated soft labels provide a predicted value of a Bernoulli distribution for each old entity type of each span; Step 4-4: Use these soft labels to calculate the Bernoulli KL divergence loss of the current model M t to evaluate the prediction difference between the student model and the teacher model; Step 4-5: The loss function adopted by the model after multiple training iterations is a weighted combination of the binary cross-entropy loss and the Bernoulli KL divergence loss: L = αL BCE + βL KD where α and β are the weight values of the two losses; ensure that the model can effectively fuse new and old knowledge.

6. The method according to claim 3, characterized in that, In Step 2-2, the final feature representation is obtained: where, ω f and ω b are the weight matrices of the forward and backward hidden states, and b h is the bias term.

7. The method according to claim 4, characterized in that, In Step 3-2, the process of representing the span is described as: where i and j respectively represent the start and end of the span, and k represents the k-th entity type.

8. The method according to claim 4 or 7, characterized in that In Step 3-3, the feature representation of the span can be further refined as: Among them, k represents the entity type to be recognized, and R i and R j represent the relative position encodings of the start and end tokens; as the number of tasks increases, simply add more span representation layers for new tasks.

9. The method according to claim 5, wherein In step 4-2, for the k-th entity type in ij with span s ij the predicted probability is:

10. The method according to claim 5, characterized in that In Step 4-4, the adopted loss calculation formula is: