Entity annotation model training method, device and storage medium based on knowledge transfer
Through a knowledge transfer-based method, the initial entity annotation model is parameter updated and fine-tuned to the target domain, which solves the problem that the model cannot be quickly associated in the new domain, generates a high-precision entity annotation model, improves efficiency and enhances generalization ability.
Patent Information
- Application Number
- CN202510920539.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-04
AI Technical Summary
In entity labeling tasks, when encountering new domains and entity types that are not covered by the source domain, existing models cannot quickly associate them and require a lot of retraining, resulting in low efficiency.
Through a knowledge transfer-based method, a small batch of labeled data is used to update the parameters of the initial model and generate temporary parameters. The generalization ability of the model is enhanced by integrating tasks in multiple fields. The generalization model is fine-tuned using target domain data to generate a high-precision target entity annotation model.
It achieves rapid adaptation and generation of high-precision entity annotation models in new domains, reduces retraining costs, enhances the model's generalization ability to different domains, and captures common knowledge across domains.
Smart Images

Figure CN120430380B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method, device and storage medium for training an entity annotation model based on knowledge transfer. Background Art
[0002] In traditional entity labeling methods, when faced with tasks in new domains, a model architecture based on supervised learning is usually adopted. Using the source domain labeled data, a neural network is used to learn the mapping relationship from input text to entity labels, and a fixed association pattern is established between text features and specific entity types. When encountering a new domain and entity types that are not covered by the source domain, since the model only learned the correspondence rules between the source domain entity types and text features during the training phase, its internal parameters and feature representations are highly focused on the source domain knowledge system, lacking the ability to generalize to new entity types and unable to automatically transfer existing knowledge to new entity types. At this time, in order for the model to be able to recognize new entity types, it is necessary to collect a large amount of similar labeled data containing the new entity types, retrain the model, and adjust the parameters to adapt to the feature distribution of the new entity types. This process requires a lot of time and manpower costs, resulting in low efficiency.
[0003] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide an entity labeling model training method, device and storage medium based on knowledge transfer, aiming to solve the technical problem that when the entity labeling task encounters a new domain and entity types not covered by the source domain appear, the entity labeling model cannot be quickly associated and requires a lot of retraining.
[0005] In order to solve the above problems, the present application provides an entity annotation model training method based on knowledge transfer, which includes:
[0006] The entity annotation model training method based on knowledge transfer includes:
[0007] Based on the annotation data of the entity annotation task, the parameters of the preset initial entity annotation model are updated to obtain temporary parameters;
[0008] According to the total loss value of the temporary parameters of each entity labeling task on the test data, the initial entity labeling model is updated to obtain a generalized entity labeling model;
[0009] The generalized entity annotation model is trained according to the annotation data of the target domain to obtain a target entity annotation model corresponding to the target domain.
[0010] In one embodiment, the step of training the generalized entity annotation model based on the annotation data of the target domain to obtain the target entity annotation model corresponding to the target domain includes:
[0011] Fine-tuning the generalized entity annotation model according to the support set of the target domain to obtain a first entity annotation model;
[0012] Inputting the query set of the target domain into the first entity annotation model to obtain a performance evaluation result of the first entity annotation model;
[0013] When the performance evaluation result meets a preset condition, the first entity annotation model is determined as the target entity annotation model.
[0014] In one embodiment, the step of inputting the query set of the target domain into the first entity annotation model to obtain the performance evaluation result of the first entity annotation model includes:
[0015] Match the unlabeled data of the target domain with the preset prototype features of each domain to determine the similar domain corresponding to the target domain;
[0016] Confirming the entity annotation information corresponding to the similar field as the sample pseudo label of the query set;
[0017] The query set is input into the first entity annotation model to obtain a performance evaluation result of the first entity annotation model.
[0018] In one embodiment, before the step of matching the unlabeled data of the target domain with the preset prototype features of each domain to determine the similar domain corresponding to the target domain, the step further includes:
[0019] Obtain feature representations of labeled data in various fields;
[0020] An average value of the feature representation is obtained to obtain the domain prototype feature corresponding to the domain.
[0021] In one embodiment, the step of updating parameters of a preset initial entity annotation model based on the annotation data of the entity annotation task to obtain temporary parameters includes:
[0022] The entity labeling task is predicted based on the initial entity labeling model to obtain a prediction result.
[0023] A first loss function is determined according to a difference value between the predicted result and the true annotation.
[0024] Parameters are updated based on the first loss function and a gradient descent algorithm to obtain the temporary parameters.
[0025] In one embodiment, the step of updating the initial entity annotation model according to the total loss value of the temporary parameters of each entity annotation task on the test data to obtain the generalized entity annotation model includes:
[0026] Predicting the query set data according to the temporary parameters corresponding to each of the entity labeling tasks, and generating a second loss function according to the prediction results;
[0027] Summing the second loss functions corresponding to the entity labeling tasks to determine the total loss;
[0028] Perform gradient back propagation based on the total loss to obtain the gradient of the total loss with respect to the initial parameters of the initial entity labeling model;
[0029] Update the initial parameters according to the gradient to obtain target parameters;
[0030] The initial entity annotation model is updated based on the target parameters to obtain the generalized entity annotation model.
[0031] In one embodiment, after the step of training the generalized entity annotation model based on the annotation data of the target domain to obtain the target entity annotation model corresponding to the target domain, the method further includes:
[0032] Determine the data type and contextual features of the data to be annotated, and determine the weights of each sub-model in the target entity annotation model, wherein the sub-models include: a rule engine, a statistical model, and a deep learning model;
[0033] A weighted vote is performed based on the entity labeling sub-results of each sub-model and the corresponding weights, and the label with the highest weighted vote is determined as the entity labeling result.
[0034] In one embodiment, the steps of determining the data source type and context features of the data to be annotated and determining the weights of each sub-model in the target entity annotation model include:
[0035] Obtaining a text structuredness score, an entity frequency score, and a context complexity score of the data to be annotated;
[0036] The weight corresponding to each sub-model is determined according to the indicator score corresponding to each sub-model.
[0037] In addition, to achieve the above-mentioned purpose, the present application also proposes a rights issuance device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the entity labeling model training method based on knowledge transfer as described above.
[0038] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the entity labeling model training method based on knowledge transfer as described above are implemented.
[0039] The present application provides a method for training an entity labeling model based on knowledge transfer. When an entity type not covered by the source domain appears in the entity labeling task, a small batch of labeled data is used to update the parameters of the initial model to obtain temporary parameters, so that the model can adapt to the patterns related to the new entity type that may exist in the new domain data. According to the total loss value of the temporary parameters of each entity labeling task on the test data, the initial entity labeling model is updated to obtain a generalized entity labeling model; by integrating multiple domain tasks, the model's generalization ability for various situations in different domains is enhanced, so that it can capture the common knowledge of cross-domain entities. The generalized model is fine-tuned with the target domain labeled data, while retaining the cross-domain commonality, the domain-specific features are strengthened to generate a high-precision target entity labeling model. By gradually guiding the model from contacting new domain data, enhancing generalization ability to accurately adapting to the target domain, the problem that the entity labeling model cannot be quickly associated and requires a lot of retraining when the entity labeling task encounters a new domain and entity types not covered by the source domain appear is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0042] Figure 1 The first flow chart provided for the entity annotation model training method based on knowledge transfer in this application;
[0043] Figure 2 A second flow chart of the entity annotation model training method based on knowledge transfer provided in this application;
[0044] Figure 3 A third flow chart of the entity annotation model training method based on knowledge transfer provided in this application;
[0045] Figure 4This is a structural diagram of the hardware operating environment involved in the entity annotation model training method based on knowledge transfer in an embodiment of the present application.
[0046] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0047] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0048] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0049] To achieve the above-mentioned objectives, the present application proposes a method for training an entity labeling model based on knowledge transfer, in which the parameters of a preset initial entity labeling model are updated based on the labeling data of the entity labeling task to obtain temporary parameters; the initial entity labeling model is updated according to the total loss value of the temporary parameters of each entity labeling task on the test data to obtain a generalized entity labeling model; the generalized entity labeling model is trained according to the labeling data of the target domain to obtain a target entity labeling model corresponding to the target domain.
[0050] In traditional entity labeling methods, when faced with tasks in new domains, a model architecture based on supervised learning is usually adopted. Using the source domain labeled data, a neural network is used to learn the mapping relationship from input text to entity labels, and a fixed association pattern is established between text features and specific entity types. When encountering a new domain and entity types that are not covered by the source domain, since the model only learned the correspondence rules between the source domain entity types and text features during the training phase, its internal parameters and feature representations are highly focused on the source domain knowledge system, lacking the ability to generalize to new entity types and unable to automatically transfer existing knowledge to new entity types. At this time, in order for the model to be able to recognize new entity types, it is necessary to collect a large amount of similar labeled data containing the new entity types, retrain the model, and adjust the parameters to adapt to the feature distribution of the new entity types. This process requires a lot of time and manpower costs, resulting in low efficiency.
[0051] The present application provides a method for training an entity labeling model based on knowledge transfer. When an entity type not covered by the source domain appears in the entity labeling task, a small batch of labeled data is used to update the parameters of the initial model to obtain temporary parameters, so that the model can adapt to the patterns related to the new entity type that may exist in the new domain data. According to the total loss value of the temporary parameters of each entity labeling task on the test data, the initial entity labeling model is updated to obtain a generalized entity labeling model; by integrating multiple domain tasks, the model's generalization ability for various situations in different domains is enhanced, so that it can capture the common knowledge of cross-domain entities. The generalized model is fine-tuned with the target domain labeled data, while retaining the cross-domain commonality, the domain-specific features are strengthened to generate a high-precision target entity labeling model. By gradually guiding the model from contacting new domain data, enhancing generalization ability to accurately adapting to the target domain, the problem that the entity labeling model cannot be quickly associated and requires a lot of retraining when the entity labeling task encounters a new domain and entity types not covered by the source domain appear is solved.
[0052] It should be noted that the execution subject of this embodiment can be a computing service device with network communication and program execution capabilities, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device or device capable of implementing the above functions. The following uses the entity annotation model training device based on knowledge transfer as an example to illustrate this embodiment and the following embodiments.
[0053] Based on this, the embodiment of the present application provides an entity annotation model training method based on knowledge transfer, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the entity annotation model training method based on knowledge transfer in this application.
[0054] In this embodiment, the entity labeling model training method based on knowledge transfer is applied to the entity labeling model training device based on knowledge transfer, and the method includes steps S10 to S30:
[0055] Step S10 : updating the parameters of the preset initial entity annotation model based on the annotation data of the entity annotation task to obtain temporary parameters.
[0056] In this embodiment, training is performed on multiple tasks based on a meta-learning framework, enabling the model to quickly adapt to entity labeling tasks in new domains. The meta-learning framework can be the Model-Agnostic Meta-Learning (MAML) framework. For each task, an inner loop is first performed to quickly adapt the model (i.e., fine-tune parameters) using a small amount of training data (support set). Next, an outer loop is performed to evaluate the performance of the adapted model on multiple tasks and update the model's initial parameters using gradient descent. This allows the model to quickly adapt to new tasks with a small number of gradient steps.
[0057] For each entity labeling task, the labeled data is divided into a training set and a validation set. The training set is used to train the initial model for supervised learning, minimizing the loss function using a backpropagation algorithm (such as the Adam optimizer). After each round of training, the validation set is used to evaluate model performance, and hyperparameters (such as the learning rate and batch size) are adjusted until the preset inner loop termination condition is achieved. The termination condition can be reaching a preset number of single-task iterations, such as completing N gradient updates on a single task; the single-task loss value converges, such as the loss decreases by less than a threshold, or the temporary parameters no longer improve performance on the single-task test set. The temporary parameters corresponding to each task are obtained. Temporary parameters are the parameter state of the model after rapid adjustment for a single task. They can achieve good performance on the query set of that task, but are only applicable to the current task.
[0058] Step S20 , updating the initial entity labeling model according to the total loss value of the temporary parameters of each entity labeling task on the test data to obtain a generalized entity labeling model.
[0059] In this embodiment, the outer loop selects multiple different tasks, repeats the operation of the inner loop, and generates corresponding temporary parameters for each task. For the temporary parameters of each task, the corresponding query set data is used for prediction, and the loss between the prediction result and the true annotation is calculated. The losses of all tasks are aggregated and summed to obtain the total loss. Based on this total loss, the gradient of the total loss to the initial parameters is calculated by the back propagation algorithm, and then the initial parameters are updated. After multiple iterations of the outer loop, the initial parameters are gradually optimized to parameters with good generalization ability. The termination condition of the outer loop can be that the preset meta-training rounds are reached, the total loss values of all tasks converge, or the rapid adaptability of the initial parameters to the new task, such as the accuracy of this learning reaches the expected threshold.
[0060] Through a multi-task learning mechanism, the model learns shared features across different tasks (such as entity boundary recognition and type generalization). The resulting generalized entity annotation model possesses cross-task and cross-domain entity recognition capabilities, capturing common features of entities across different domains and reducing repetitive training costs. By repeatedly adapting the inner loop to the task and optimizing the initial parameters in the outer loop, the model learns how to quickly adjust parameters using small amounts of data and how to make the initial parameters more generalizable across tasks, ultimately achieving the meta-learning goal of rapidly learning new tasks.
[0061] Step S30 : training the generalized entity annotation model according to the annotation data of the target domain to obtain a target entity annotation model corresponding to the target domain.
[0062] In this example, we collect annotated data from the target domain, such as an annotated corpus of electronic medical record entities in the medical field. The data volume must cover all entity types in the target domain. Using the generalized model as initialization parameters, we perform domain-adaptive training using the annotated data from the target domain. We evaluate the model using an in-domain validation set, and optimize parameters until convergence. By training on domain-specific data, the model becomes more robust to in-domain textual noise (such as unstructured text in medical records).
[0063] In this embodiment, when an entity type not covered by the source domain appears in the entity labeling task, a small batch of labeled data is used to update the parameters of the initial model to obtain temporary parameters, so that the model can adapt to the patterns related to the new entity type that may exist in the new domain data. According to the total loss value of the temporary parameters of each entity labeling task on the test data, the initial entity labeling model is updated to obtain a generalized entity labeling model; by integrating multiple domain tasks, the model's generalization ability for various situations in different domains is enhanced, so that it can capture the common knowledge of cross-domain entities. The generalized model is fine-tuned with the target domain labeled data, while retaining the cross-domain commonality, the domain-specific features are strengthened to generate a high-precision target entity labeling model. By gradually guiding the model from contacting new domain data, enhancing generalization ability to accurately adapting to the target domain, the problem that the entity labeling model cannot be quickly associated and requires a lot of retraining when the entity labeling task encounters a new domain and entity types not covered by the source domain appear is solved.
[0064] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 2 , step S30 may include steps S31 to S33:
[0065] Step S31 : fine-tuning the preset initial entity annotation model according to the support set of the target domain to obtain a first entity annotation model.
[0066] It should be noted that the support set refers to the data used to provide task-specific training data in the adaptation phase (inner loop) in meta-learning, covering the core entities, sentence structures, and terminology of the target domain.
[0067] In this embodiment, the initial entity annotation model is based on the Model-Agnostic Meta-Learning (MAML) framework, and performs inner and outer loops on multiple historical tasks (domains). The obtained entity annotation model captures common cross-domain knowledge, such as entity boundary features and semantic patterns. Only a small amount of target domain data needs to be fine-tuned to quickly adapt to the entity annotation model of the new domain.
[0068] In this embodiment, the support set of the target domain can be a small number of samples obtained by crawling public data (such as social media text, public image library) through legal channels, or by calling data service APIs, and manually annotated. Before model training, hyperparameter configuration is performed, such as setting the learning rate to 0.001 and the number of fine-tuning steps to 3. The sample data in the support set are input into the initial entity annotation model in batches according to the preset quantity. Each sample data includes a text sequence and a corresponding entity annotation. The sample data is input into the model, and a word vector representation is generated through the word embedding layer. The context representation is obtained through the feature extraction layer. The entity label of each token is predicted through the classification layer, and the loss value between the predicted label and the true label is calculated. The loss value reflects the degree of inadaptability of the model to the target domain data. Secondly, the gradient of the loss to the trainable parameters is calculated through backpropagation. The optimizer updates the parameters according to the gradient and adjusts the parameters of the word embedding layer and the domain-related feature extraction layer. If the loss does not decrease significantly after 2-3 consecutive steps, fine-tuning is stopped to prevent overfitting. The model parameters are fine-tuned to reduce the support set loss, focusing on optimizing feature representations relevant to the target domain while retaining the general knowledge in the initial parameters (such as entity boundary detection). Only incremental adjustments are made to the target domain characteristics. The forward propagation and loss calculation, as well as the gradient calculation and parameter update steps described above, are repeated for multiple rounds on the support set. After reaching the preset number of iterations, the first entity annotation model is obtained. The model parameters of the fine-tuned first entity annotation model retain the cross-domain general knowledge in the initial parameters while incorporating knowledge specific to the target domain.
[0069] In an optional implementation, a regularization method can be used to introduce additional penalty terms into the model loss function to constrain the complexity of the model parameters, thereby avoiding overfitting of the model due to the small amount of support set data. The loss function after regularization is introduced as L'=L=λ*R(θ), where λ is the regularization coefficient used to control the penalty strength; θ is the regularization term used to measure the parameter complexity. Optionally, the regularization method can be L2 regularization, and the regularization term R(θ)=∑ w∈θ w 2When computing the loss, the squared values of each parameter w are summed, multiplied by the regularization factor λ, and added to the total loss. During backpropagation, the gradient includes a penalty term for w. Alternatively, the regularization method can be weight decay, which directly decays the weights when the optimizer updates the parameters, without explicitly modifying the loss function.
[0070] Step S32: inputting the query set of the target domain into the first entity annotation model to obtain a performance evaluation result of the first entity annotation model.
[0071] It should be noted that the query set is labeled data that is screened from the target domain and does not participate in model training. It is used to evaluate model performance.
[0072] In this example, the model is set to inference mode, and gradient calculation is disabled. The query set data is divided into batches according to a preset number of samples. Each batch of sample data contains a text sequence and its corresponding ground-truth annotations. Each batch of text data is fed into the fine-tuned model. The model passes through the word embedding layer, feature extraction layer (such as LSTM, Transformer), and classification layer, outputting predicted entity labels for each token. The predicted results for each batch are stored in a one-to-one correspondence with the ground-truth annotations. The model makes predictions for the query set data, generating preliminary entity annotation results that reflect the model's performance on unseen data.
[0073] Specifically, the text data in the query set is split into token sequences based on whitespace, punctuation, or specific tokenizers (such as the BERT tokenizer). [CLS] and [SEP] tags are added to the beginning and end of the token sequences to convert them into an input format acceptable to the model. All samples are divided into batches of a fixed size (e.g., 32 samples per batch). If the batch size is less than one, zero padding is performed to bring the length to a uniform value. The actual length of each sample is recorded, and each token is mapped to an index in the vocabulary. The word embedding layer converts each token index into a corresponding word vector. By mapping the token into a semantic vector, it incorporates a domain-specific representation. The feature extraction layer captures sequential dependencies in the text, extracts features relevant for entity recognition, and converts the word vector sequence into a contextual feature matrix. The classification layer maps the contextual feature matrix to the label space through a fully connected layer. A softmax function is applied to the linear output to convert the scores into probability distributions. The features are then mapped to the entity label space, and the predicted category for each token is output. After mapping the predicted label index back to a string label, the prediction results are stored in a one-to-one correspondence with the ground-truth annotation.
[0074] Step S33: When the performance evaluation result meets a preset condition, the first entity annotation model is determined as the target entity annotation model corresponding to the target domain.
[0075] In this embodiment, the preset condition can be precision, recall or F1 value. When the preset condition is precision, the statistical prediction is the proportion of entities that are truly entities in the positive samples (identified entities). For example, the model predicts 100 entities, 80 of which are consistent with the true annotations, then the precision = 80 / 100 = 0.8. The recall rate counts the proportion of real entities that are correctly identified by the model. Assuming that there are 120 entities in the true annotations and the model correctly identifies 80, then the recall rate = 80 / 12≈0.67. The F1 value is the harmonic mean of the precision and recall rates, and the calculation formula is F1=2*(precision*recall) / (precision+recall), which is used to balance the two indicators to avoid a single indicator being too high and the other indicator being too low. The entity recognition ability of the model on the target domain query set is quantified through precision, recall and F1 value, and the model performance is comprehensively evaluated. The precision rate reflects the accuracy of the model's prediction, the recall rate reflects the model's ability to capture real entities, and the F1 value gives a comprehensive score to help determine whether the model has over-prediction or under-reporting problems.
[0076] Compare the calculated precision, recall, and F1 value against pre-set business performance thresholds. If performance meets the requirements (e.g., F1 value ≥ 0.85), the model is considered to meet requirements and can proceed to deployment. If precision is low, there may be misidentifications. The model's accuracy in determining entity boundaries or categories needs to be optimized, such as by adjusting classification layer parameters or increasing regularization strength. If recall is low, there may be missed identifications. Increase the support set data size or add special samples, or adjust the number of fine-tuning steps to better learn the characteristics of the target domain. Based on the analysis results, re-evaluate the quality and size of the support set or adjust the fine-tuning strategy (e.g., increasing the number of fine-tuning steps or optimizing regularization parameters). Then, return to the fine-tuning step and retrain the model, then re-verify until performance meets the requirements. Determine the model's usability based on objective metrics to ensure that the deployed model meets business requirements. If performance does not meet the requirements, continuously optimize the model through targeted adjustments to improve its generalization and practicality in the target domain.
[0077] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 3 Step S32 may further include steps A10 to A30:
[0078] Step A10: feature matching is performed on the unlabeled data of the target domain with the preset prototype features of each domain to determine a similar domain corresponding to the target domain.
[0079] In this embodiment, for the query set sample x in the target domain j , get its feature representation fφ(xj ) with all class prototypes c k The Euclidean distance between them is calculated, and the category with the smallest Euclidean distance or the smallest cosine distance is selected as the similar domain. The support set samples of the target domain containing a small amount of manually annotated data are input into the feature extractor fφ to obtain the query set samples x j When the query set samples are image data, models such as ResNet and MobileNet can be used to extract image features; when the query set samples are text data, models such as BERT and RoBERTa can be used to extract text features.
[0080] In a feasible implementation, supervised learning or meta-learning (such as prototype network) is used to optimize the parameter φ of the feature extractor so that the features of similar samples are tightly clustered in the embedding space and heterogeneous samples are fully separated, thereby improving the classification accuracy of the model in small sample scenarios. Samples from different fields are classified as categories, so that the feature extractor learns domain-discriminative features. For example, K fields are randomly selected from M fields each time, and each field D m There is a labeled sample set S m {(x i ,y i )}. Take N samples from each field to form a field classification task, train the feature extractor to classify the samples in these K fields, and optimize the cross entropy loss:
[0081] ;
[0082] Among them, p(y m |fφ(x i ))is the eigenvector fφ(x i ) belongs to field D m probability.
[0083] Step A20: Confirm the entity annotation information corresponding to the similar field as the sample pseudo label of the query set.
[0084] Step A30: input the query set into the first entity annotation model to obtain a performance evaluation result of the first entity annotation model.
[0085] In this embodiment, if the entity labels corresponding to each domain are clearly defined when the prototype of the preset domain is constructed, such as: domain A is the animal domain, and the entity labels are all animal categories, then the pseudo labels can be directly assigned to the entity label set of the similar domain. For example, if the similar domain is the animal domain, the pseudo label of the target domain may be animal-cat (assuming that the prototype of the domain is constructed by samples such as cats and dogs). If the domain prototype is composed of the feature mean of multiple category samples in the domain, the pseudo label can be generated based on the label distribution probability of each category in the domain. For example, the prototype of domain A is obtained by averaging the features of 50% cat samples and 50% dog samples. When the target domain matches domain A, its pseudo label is assigned to cats with a probability of 50% and to dogs with a probability of 50%.
[0086] Optionally, a similarity threshold is set. If the feature similarity between the target field and the similar field is less than the preset similarity threshold, the label of the similar field is not used as a pseudo label of the target field to avoid mismatching.
[0087] In this embodiment, pseudo labels can be used to generate pseudo annotations for the prediction results of the query set using the model, thereby expanding the training data and avoiding the problem that the samples of the query set are usually unlabeled or insufficiently labeled.
[0088] In a feasible implementation manner, before step A10, steps A40 to A50 are further included:
[0089] Step A40: Obtain feature representations of the labeled data in various fields.
[0090] Step A50: Obtain an average value of the feature representation to obtain the domain prototype feature corresponding to the domain.
[0091] In this implementation, for each domain, a feature extractor generates sample features based on the labeled data for that domain, and then statistically aggregates the features to obtain a prototype vector. For image data, a CNN model can be selected as a feature extractor to output the input image as a feature vector; for text data, a Ransformer encoder, LSTM, etc. can be used to output the input text sequence as a semantic feature vector. The feature vectors of each domain are aggregated to obtain the corresponding domain prototype features.
[0092] Optionally, the aggregation method can be to take the mean of the feature vectors of all samples in the domain; it can also be to take the median of the feature vectors of all samples in the domain; or it can be weighted according to the importance of the samples and perform a weighted average calculation. The resulting domain prototype feature can be regarded as the feature center of the domain, representing the commonality of samples in the domain.
[0093] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above introduction and will not be repeated hereafter. On this basis, step S10 may include steps S11 to S13:
[0094] Step S11: predicting the entity labeling task based on the initial entity labeling model to obtain a prediction result.
[0095] Step S12: determining a first loss function according to the difference between the predicted result and the true annotation.
[0096] Step S13: Update parameters based on the first loss function and the gradient descent algorithm to obtain the temporary parameters.
[0097] In this example, each entity labeling task corresponds to a specific domain, such as healthcare, finance, or law. The labeled data for this entity labeling task can come from a text corpus in that domain. A domain task is sampled from the task distribution, and its labeled data is divided into a support set and a query set. The support set consists of a small amount of annotated text, each containing an entity type. The query set consists of unannotated or unverified text, which is used to evaluate the performance of the adapted model.
[0098] Input the sample data of the support set into the initial entity annotation model one by one or in batches, and the support set S={(x1, y1), (x2, y2), ..., (x n ,y n )}, 0 where x i is the input data (such as text, image), y i is the true label (category index in classification tasks and numerical value in regression tasks). i Input the initial entity annotation model and output the logits vector [z1, z2, ..., z K ], where K is the number of categories. For each sample’s logits vector z i, Compute the class probability distribution:
[0099] ;
[0100] Among them, (k=1, 2, ..., K), p i =[p1(x i ), p2(x i ),…,p K (x i )] is the predicted probability vector, 0≤p K ≤1 and∑p K = 1. Set the true label y i Convert to one-hot vector .
[0101] Calculate the loss value of a single sample. The loss function can be cross entropy loss or focal loss. For sample xi, its cross entropy loss is:
[0102] ;
[0103] If the model is not aware of the true category y i The predicted probability p yi (x i ) is higher, the loss L i The smaller it is, the loss is 0 when P=1.
[0104] Average the losses over all n examples in the support set:
[0105] ;
[0106] Among them, the loss function L 总 is a function of the initial parameter θ, whose gradient Used for parameter update. The gradient formula of cross entropy loss is:
[0107] .
[0108] After obtaining the loss function, from the loss L 总 Start backpropagation, calculate the gradient layer by layer, and update the initial parameters and temporary parameters according to the calculated gradient: θ′=θ-α*∇θLtotal(θ).
[0109] In this embodiment, the inner loop processes only one specific task at a time. For example, if the outer loop contains tasks A, B, and C, the inner loop might process only task A in one round, task B in the next round, and so on. The inner loop generates temporary parameters through gradient calculation on a single task, simulating the model's rapid adaptation to new tasks. Temporary parameters are the model's rapidly adjusted parameter state for that single task, resulting in better performance on the query set for that task, but are only suitable for the current task. For example, temporary parameters are only suitable for the current domain (such as healthcare) and do not affect parameters in other domains (such as finance). This isolation ensures that the model can quickly switch between different domains without requiring global retraining.
[0110] Based on the first embodiment of the present application, in the fourth embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to above and will not be described in detail. On this basis, step S20 may include steps S21 to S25:
[0111] Step S21 , predicting the query set data according to the temporary parameters corresponding to each of the entity labeling tasks, and generating a second loss function according to the prediction results.
[0112] Step S22: summing the second loss functions corresponding to the entity labeling tasks to determine the total loss.
[0113] In this example, for each task, the corresponding temporary parameters obtained from the inner loop are used to predict the query set data, and the loss between the prediction result and the true annotation is calculated. The loss values of all tasks are summed to obtain the total loss of meta-learning.
[0114] Step S23: performing gradient back propagation based on the total loss to obtain the gradient of the total loss with respect to the initial parameters of the initial entity labeling model.
[0115] Step S24: Update the initial parameters according to the gradient to obtain target parameters.
[0116] Step S25 : updating the initial entity annotation model based on the target parameters to obtain the generalized entity annotation model.
[0117] In this embodiment, starting from the total loss, the gradient of the initial parameter θ is calculated by the chain rule. , the gradient reflects the impact of the initial parameters on the comprehensive performance of all tasks. Using the meta-learning rate β, the initial parameters θ are updated based on the gradient:
[0118] .
[0119] Move the initial parameters in a direction that improves the inner loop performance for all tasks. For example, if θ performs well in the inner loop on tasks 1-50 (low loss) but poorly on tasks 51-100 (high loss), the gradient will guide the update of θ towards improving the performance of tasks 51-100 while minimizing the performance of tasks 1-50.
[0120] Repeat the above steps of loss calculation, backpropagation, and parameter update for multiple rounds of outer loop iterations. After multiple rounds of iterations, the initial parameters are gradually optimized into a meta-knowledge carrier, which is no longer targeted at a specific type of task, but instead contains a general strategy for how to quickly learn various tasks. When the total loss L no longer decreases significantly, or the model's ability to quickly adapt to new tasks (such as few-shot accuracy) reaches a preset threshold, the outer loop terminates and the final generalization parameters are obtained. Test data usually contains different types of data samples, including various situations that may arise in new domains. By calculating the total loss value of the temporary parameters on the test data, the performance of the model in different data scenarios can be evaluated. The total loss value reflects the gap between the model's prediction results and the true label. Based on this loss value, the first entity labeling model is updated, guiding the model to adjust parameters in the direction of better performance on a wider range of data.
[0121] In this embodiment, through the loop of multitask simulation, loss aggregation, and parameter tuning, the initial parameters are enabled to have the ability of rapid cross-task learning. For example: after being optimized by the outer loop, when facing the entity annotation task in a new field, only a small amount of support set data of this task needs to be used to execute the inner loop, and it can converge quickly, while traditional models may require a large amount of data for retraining.
[0122] Based on the first embodiment of the present application, in the fifth embodiment of the present application, the same or similar content as that in the above-mentioned embodiment 1 can be referred to the above introduction and will not be repeated hereinafter. On this basis, after step S30, steps S40 to S50 may further be included:
[0123] Step S40, determining the data type and context features of the data to be annotated, and determining the weights of each sub-model in the target entity annotation model, where the sub-models include: a rule engine, a statistical model, and a deep learning model.
[0124] In this embodiment, a rule engine, a statistical model, and a deep learning model are integrated in the target entity annotation model to form a heterogeneous annotation system, and the decision weights of the rule engine, the statistical model, and the deep learning model are dynamically adjusted according to the data type and context features of the data to be annotated.
[0125] The rule engine performs entity annotation based on predefined rules. By writing structured rules to match strings that conform to specific patterns in the text, entity recognition is achieved. It can be based on dictionary matching. Pre-organize an entity dictionary (such as a person name library, a place name library), and directly identify the entities in the text through string matching. For example: if the dictionary contains Place A and Place B, when these words appear in the text, they are labeled as place names. It can be based on regular expression rules to define rules based on syntactic patterns to capture entities with fixed formats. For example: the ID card number rule: \d{17}[\d|X], and it is labeled as a certificate number after matching. It can also be labeled based on syntactic rules, and rules are defined based on词性标注 (such as nouns, proper nouns) and syntactic structures (such as subject-verb-object). For example: the pattern of place name + university is labeled as an institution name.
[0126] The statistical model can be a Hidden Markov Model (HMM), a Conditional Random Field (CRF), etc. The statistical model regards entity annotation as a sequence annotation problem, learns the probability distribution of words in the text sequence through the statistical model, and predicts the entity label of each word. For example, B-PER and I-PER represent the start and continuation of a person name. When the statistical model performs entity annotation, it first performs word segmentation and词性标注 to generate labeled training data, such as: [Xiaoming] / B-PER [is] / O [student] / O. Secondly, it constructs context features for each word, such as a word sequence with a window size of 3, and determines the label transition probability and feature weights through statistical learning to calculate the probabilities of each label for the new text, and generates an optimal annotation sequence.
[0127] Deep learning models automatically learn semantic representations of text using neural networks, eliminating the need for manual feature design and enabling end-to-end entity labeling. These models are suitable for processing complex contexts and massive amounts of data. These models can be BiLSTM-CRF or Transformer-based models. Language models (such as BERT) are pre-trained using large-scale unlabeled data to obtain semantic representations. The model is then fine-tuned using annotated corpus to learn the mapping between entity labels and semantic representations. Finally, when new text is input, the model directly outputs the location and type of the entity.
[0128] In a feasible implementation, step S40 may include steps S41 and S42:
[0129] Step S41: Obtain the text structure degree score, entity frequency score, and context complexity score of the data to be annotated.
[0130] In this embodiment, the data type of the data to be labeled may include the degree of text structure and entity frequency distribution, among which the degree of structure is used to measure whether the data to be labeled contains structured fields, whether it is a tabular text, etc. For texts with a high degree of structure, the rule engine is more effective; the entity frequency represents the frequency of occurrence of the token in the training set. High-frequency entities are suitable for statistical models, and low-frequency or new words are suitable for deep models; context complexity is used to analyze the length of syntactic dependencies, nesting levels, and degree of ambiguity. Deep learning models are more suitable for complex contexts.
[0131] To score the degree of text structure, we traverse the text and count the number of matching features. We can detect structural features in the text through rule matching and calculate the proportion of matching features to the total number of features. Structural features can include delimiters, key-value pairs, list markers, and table borders. We divide the number of matching structural features by the total number of features to calculate the proportion of the total number of matching features. The total number of features can be defined as the size of a preset feature set (with a fixed denominator) or dynamically count the total number of features that may appear in the text.
[0132] For entity frequency scores, we calculate a standardized score (Z-score) based on the frequency of the entity in the training data. We traverse the training data, count the number of occurrences of each entity, and normalize it to a frequency. We then calculate the mean μ and standard deviation σ of all entity frequencies. The entity frequency score Z = (f - μ) / σ, where f is the entity frequency.
[0133] The contextual complexity score can be calculated by taking a weighted average of indicators such as the lexical difficulty index, average sentence length, and syntactic complexity index to obtain the contextual complexity score. Using a predefined vocabulary level table (such as the CEFR level or Lexile level), each word in the text is mapped to a corresponding difficulty level. The number of words at different difficulty levels is counted, and the proportion of high-difficulty words to the total number of words is calculated to obtain the lexical difficulty index. The number of words in all sentences in the text is counted and averaged to obtain the average sentence length. A syntactic analysis tool (such as a dependency parser) is used to analyze the syntactic structure of each sentence, counting the number of complex syntactic structures (such as nested clauses and parallel structures) and calculating the proportion of complex syntactic structures to the total number of sentences to obtain the syntactic complexity index.
[0134] Step S42: determining the weight corresponding to each sub-model according to the indicator score corresponding to each sub-model.
[0135] In this embodiment, the rules for weight allocation are defined in advance based on the impact of each score on the applicability of each model. The rule engine weight is positively correlated with the degree of architecture. The higher the degree of architecture, the higher the weight of the rule engine. The rule engine weight can be a linear or nonlinear function, for example: W1 = α * text architecture score, where α is a scaling factor. The statistical model weight is positively correlated with entity frequency: the higher the entity frequency, the higher the weight of the statistical model. The statistical model weight can be a linear or nonlinear function, W2 = β * entity frequency score, where β is a scaling factor. The deep learning model weight is positively correlated with context complexity. The more complex the context, the higher the weight of the deep learning model. At the same time, the deep learning model weight is negatively correlated with entity frequency. The lower the entity frequency, the higher the weight of the deep learning model. The weight of the deep learning model W3 = γ * context complexity score + δ * (1 - entity frequency score), where γ and δ are scaling factors. The obtained weights are normalized to ensure that the sum of the weights is 1.
[0136] Alternatively, a neural network may be used to automatically learn a weight assignment function, or similar texts may be grouped based on clustering and a fixed weight may be assigned to each group.
[0137] Step S50 , performing weighted voting based on the entity labeling sub-results of each sub-model and the corresponding weights, and determining the label with the highest weighted vote as the entity labeling result.
[0138] In this embodiment, the final label is determined through weighted voting. For each candidate label, the voting results of each model are tallied and multiplied by the weight. The label with the highest weighted votes is selected as the final result. The annotation results are verified by comparing them with manual annotations or knowledge graphs, and the error rate of each model under different characteristics is calculated. For scenarios with high error rates, the corresponding model weight is automatically reduced. For example, if the rule engine's error rate in unstructured text exceeds a preset threshold, the upper limit of its weight is lowered.
[0139] Optionally, a gating mechanism can be used to select a rule engine, statistical model, or deep learning model for entity annotation, resetting the weights of other models to zero. When the characteristics of the data to be annotated are clear, such as pure table text, using the rule engine directly can avoid interference from other models.
[0140] In this implementation, the dominant model is dynamically switched based on text characteristics, and weight distribution enables complementary model strengths. This allows the rule engine to handle structured parts, the statistical model to cover high-frequency scenarios, and deep learning to solve complex problems, ultimately improving overall annotation accuracy. This approach, rather than relying on the limitations of a single model, enhances entity annotation accuracy.
[0141] The present application provides an entity labeling model training device based on knowledge transfer, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the entity labeling model training method based on knowledge transfer in the above-mentioned embodiment one.
[0142] Reference below Figure 4 , which shows a schematic diagram of the structure of an entity annotation model training device based on knowledge transfer suitable for implementing the embodiments of the present application. The entity annotation model training device based on knowledge transfer in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptops, personal digital assistants (PDAs), tablet computers (portable Android devices), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The entity annotation model training device based on knowledge transfer shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0143] like Figure 4As shown, the entity labeling model training device based on knowledge transfer may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1003 to the random access memory (RAM) 1004. The random access memory 1004 also stores various programs and data required for the operation of the entity labeling model training device based on knowledge transfer. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the knowledge transfer-based entity annotation model training device to communicate wirelessly or wired with other devices to exchange data. Although the figure shows a knowledge transfer-based entity annotation model training device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems can be implemented or have instead.
[0144] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0145] The entity annotation model training device based on knowledge transfer provided by the present application adopts the entity annotation model training method based on knowledge transfer in the above embodiment, which can solve the technical problem that when the entity annotation task encounters a new domain and the source domain does not cover the entity type, the entity annotation model cannot be quickly associated and requires a lot of retraining. Compared with the existing technology, the beneficial effects of the entity annotation model training device based on knowledge transfer provided by the present application are the same as the beneficial effects of the entity annotation model training method based on knowledge transfer provided by the above embodiment, and the other technical features of the entity annotation model training device based on knowledge transfer are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0146] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0147] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0148] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, and the computer-readable program instructions are used to execute the entity labeling model training method based on knowledge transfer in the above-mentioned embodiment.
[0149] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.
[0150] The computer-readable storage medium may be included in the entity labeling model training device based on knowledge transfer, or may exist independently without being assembled into the entity labeling model training device based on knowledge transfer. The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the entity labeling model training device based on knowledge transfer, the entity labeling model training device based on knowledge transfer: updates the parameters of the preset initial entity labeling model based on the labeling data of the entity labeling task to obtain temporary parameters; updates the initial entity labeling model based on the total loss value of the temporary parameters of each entity labeling task on the test data to obtain a generalized entity labeling model; and trains the generalized entity labeling model based on the labeling data of the target domain to obtain a target entity labeling model corresponding to the target domain.
[0151] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0152] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0153] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0154] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned entity annotation model training method based on knowledge transfer. This computer-readable storage medium can solve the technical problem that when the entity annotation task encounters a new domain and entity types not covered by the source domain appear, the entity annotation model cannot be quickly associated and requires a large amount of retraining. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the entity annotation model training method based on knowledge transfer provided in the above-mentioned embodiment, and will not be repeated here.
[0155] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A method for training an entity annotation model based on knowledge transfer, characterized in that: The entity annotation model is applied to the text field, and the entity annotation model training method based on knowledge transfer includes: Based on the annotation data of the entity annotation task, the parameters of the preset initial entity annotation model are updated to obtain temporary parameters; According to the total loss value of the temporary parameters of each entity labeling task on the test data, the initial entity labeling model is updated to obtain a generalized entity labeling model; Training the generalized entity annotation model based on the annotation data of the target domain to obtain a target entity annotation model corresponding to the target domain; Determine the data type and contextual features of the data to be annotated, and determine the weights of each sub-model in the target entity annotation model, wherein the sub-models include: a rule engine, a statistical model, and a deep learning model; Perform weighted voting based on the entity labeling sub-results of each sub-model and the corresponding weights, and determine the label with the highest weighted votes as the entity labeling result; The steps of determining the data source type and contextual features of the data to be annotated, and determining the weights of each sub-model in the target entity annotation model include: obtaining the text structured degree score, entity frequency score, and context complexity score of the data to be annotated; determining the weight corresponding to each sub-model based on the indicator score corresponding to each sub-model, wherein the rule engine weight is positively correlated with the text structured degree score, and the higher the text structured degree score, the higher the weight of the rule engine; the statistical model weight is positively correlated with the entity frequency score, and the higher the entity frequency score, the higher the weight of the statistical model; the deep learning model weight is positively correlated with the context complexity score, and the higher the context complexity score, the higher the weight of the deep learning model; the deep learning model weight is negatively correlated with the entity frequency score, and the lower the entity frequency score, the higher the weight of the deep learning model.
2. The entity annotation model training method based on knowledge transfer according to claim 1, characterized in that: The step of training the generalized entity annotation model according to the annotation data of the target domain to obtain the target entity annotation model corresponding to the target domain includes: Fine-tuning the generalized entity annotation model according to the support set of the target domain to obtain a first entity annotation model; Inputting the query set of the target domain into the first entity annotation model to obtain a performance evaluation result of the first entity annotation model; When the performance evaluation result meets a preset condition, the first entity annotation model is determined as the target entity annotation model.
3. The entity annotation model training method based on knowledge transfer according to claim 2 is characterized in that: The step of inputting the query set of the target domain into the first entity annotation model to obtain the performance evaluation result of the first entity annotation model includes: Match the unlabeled data of the target domain with the preset prototype features of each domain to determine the similar domain corresponding to the target domain; Confirming the entity annotation information corresponding to the similar field as the sample pseudo label of the query set; The query set is input into the first entity annotation model to obtain a performance evaluation result of the first entity annotation model.
4. The entity annotation model training method based on knowledge transfer according to claim 3 is characterized in that: Before the step of matching the unlabeled data of the target domain with the preset prototype features of each domain to determine the similar domain corresponding to the target domain, the method further includes: Obtain feature representations of labeled data in various fields; An average value of the feature representation is obtained to obtain the domain prototype feature corresponding to the domain.
5. The entity annotation model training method based on knowledge transfer according to claim 1, characterized in that: The step of updating parameters of a preset initial entity annotation model based on the annotation data of the entity annotation task to obtain temporary parameters includes: Predicting the entity labeling task based on the initial entity labeling model to obtain a prediction result; Determine a first loss function based on the difference between the predicted result and the true annotation; Parameters are updated based on the first loss function and a gradient descent algorithm to obtain the temporary parameters.
6. The entity annotation model training method based on knowledge transfer according to claim 1, characterized in that: The step of updating the initial entity annotation model according to the total loss value of the temporary parameters of each entity annotation task on the test data to obtain a generalized entity annotation model includes: Predicting the query set data according to the temporary parameters corresponding to each of the entity labeling tasks, and generating a second loss function according to the prediction results; Summing the second loss functions corresponding to the entity labeling tasks to determine the total loss; Perform gradient back propagation based on the total loss to obtain the gradient of the total loss with respect to the initial parameters of the initial entity labeling model; Update the initial parameters according to the gradient to obtain target parameters; The initial entity annotation model is updated based on the target parameters to obtain the generalized entity annotation model.
7. A knowledge transfer-based entity annotation model training device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the entity labeling model training method based on knowledge transfer as described in any one of claims 1 to 6.
8. A storage medium, characterized in that: The storage medium is a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the entity labeling model training method based on knowledge transfer as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Model training method and device, computer equipment and computer readable storage medium
CN115457572A
Model training method and device suitable for large language model, equipment and medium
CN116976424A