Chinese Ancient Book Entity Extraction Method Based on Deep Active Learning Strategy
By adopting deep active learning strategies in the physical extraction of ancient Chinese books, combining deep neural network models and active learning query strategies, the problems of low identification accuracy and insufficient corpus in the existing technology are solved, and more efficient entity extraction and learning efficiency are achieved.
Patent Information
- Application Number
- CN202510247071.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The prior art has problems in the extraction of Chinese ancient books with low recognition accuracy and relying on a large amount of artificial field knowledge in the extraction of Chinese ancient books, and the deep learning-based models have performance bottlenecks due to insufficient corpus.
Using a method based on deep active learning strategy, we use original ancient text resources, perform entity annotation, and build a deep neural network model. We also use active learning query strategies, including query strategies that integrate uncertainty and minimum boundary, to select data pools and update the annotation data, and gradually improve model performance.
It effectively improves the recognition accuracy and learning efficiency of the Chinese ancient book entity extraction model, alleviating the problems of insufficient labeling data and high expert labeling costs.
Smart Images

Figure CN119761372B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and particularly to a method for extracting Chinese ancient book entities based on a deep active learning strategy. Background Art
[0002] Entity extraction technology has important research value and application prospects in the fields of natural language processing and artificial intelligence. With the development of machine learning, especially deep learning technology, entity extraction technology based on machine learning methods has been widely used in various basic disciplines and practical applications, such as the construction of various humanities and social science databases, database projects, and digital humanities research. Striving to accelerate the digital construction of Chinese ancient book resources has important dual values in academic innovation and social development.
[0003] As one of the key methods for the construction of ancient book digital resources, the entity extraction technology for Chinese ancient book texts is an effective method for automatically extracting entities with clear entity objects in Chinese ancient book rare editions, where entity types include personal names, place names, organization names, and other definable entity types such as official positions, book names, etc. However, such technical solutions face huge challenges in research and implementation, which mainly include two aspects of limitations: (1) Most current ancient book entity extraction methods generally adopt traditional rule matching methods or machine recognition models based on feature templates (such as the conditional random field CRF model), etc. These methods have low entity recognition accuracy and rely on a large amount of prior artificial domain knowledge, so the model is not easy to be extended to other related historical fields; (2) Although the ancient book entity extraction model based on machine learning methods (especially deep learning) is efficient, it requires a large number of standard training samples. Ancient books belong to a typical low-resource field, and the cost of annotating ancient Chinese (annotation difficulty and time cost) is significantly higher than that of modern Chinese corpus annotation, and the professional level of different annotators will directly affect the quality of training data. Therefore, directly using deep learning to build an entity extraction model may have a performance bottleneck due to "insufficient corpus".
[0004] However, in the above methods, some models are currently only used for optimization in the decision boundary sampling of traditional machine classification models (such as the conditional random field CRF model, support vector machine SVM model) (such as the method based on committee multi-model voting); the method complexity of some models is too high - for example, the Fisher information matrix needs to calculate the covariance matrix containing the squares of the gradients of all feature parameters; at the same time, the existing models all ignore how to balance the feature representativeness differences in different sample hypothesis spaces and the redundancy problem of pooling sampling in the active sampling method. Therefore, the above model methods all have certain performance bottlenecks and cannot be well applied to the task of extracting Chinese ancient book entities, especially being more scientifically and effectively integrated into the network design of deep learning models. Summary of the Invention
[0005] (1) Technical problems to be solved
[0006] In view of the deficiencies of the prior art, the present invention provides a method for extracting Chinese ancient book entities based on a deep active learning strategy, which solves the technical problem that existing model methods are not well applicable to the task of extracting Chinese ancient book entities.
[0007] (2) Technical solutions
[0008] To achieve the above object, the present invention is realized through the following technical solutions:
[0009] A method for extracting Chinese ancient book entities based on a deep active learning strategy, comprising:
[0010] Obtain the original ancient book text resources, select a part of the text for entity annotation to obtain an annotated data set, and use the other unannotated text as an unannotated data set; wherein the number of unlabeled text is greater than the labeled text;
[0011] Configure the hyperparameters for model operation and calculation, and based on the initial annotated data set, train, predict and evaluate the performance of a pre-constructed entity extraction network model based on a deep neural network to obtain an initial evaluation score;
[0012] If the current evaluation score exceeds the passing threshold or the number of iterations reaches the maximum number of iterations, output the final entity extraction network model; otherwise, repeatedly execute the following active learning query strategy, including:
[0013] Perform pool sampling on the current unlabeled data set, and use the current entity extraction network model to predict the corresponding composite score by adopting a query strategy that combines uncertainty and minimum boundary;
[0014] Based on the composite score, adopt a query strategy with multi-space feature representativeness to select a batch of unlabeled text with information value from the current unlabeled data set and use it as a query sample pool;
[0015] Accept the manual annotation of the expert for the query sample pool to obtain the annotated query sample pool;
[0016] Add the annotated query sample pool to the current labeled data set and remove it from the current unlabeled data set;
[0017] Train, predict and evaluate the performance of the current entity extraction network model with the current labeled data set to obtain an updated model and evaluation score for the next iteration process.
[0018] Preferably, the entity extraction network model adopts a deep neural sequence annotation network with a dual encoder-decoder, including a main learning channel for entity sequence annotation and a class learning channel for sentence representation;
[0019] Configure the hyperparameters for model operation and calculation, and based on the initial annotation dataset, train, predict, and evaluate the performance of the pre-constructed entity extraction network model based on a deep neural network to obtain an initial evaluation score, including:
[0020] For the main learning channel:
[0021] Adopt a global encoder based on pre-trained learning to perform sentence embedding encoding on the input tokenized text to obtain a global vector representation;
[0022] Adopt a local encoder based on character-level embedding to perform character-level context-related word embedding on the input tokenized text and perform smoothing optimization to obtain a local vector representation;
[0023] Concatenate the global vector representation and the local vector representation and perform multi-layer encoding mapping to obtain the final semantic representation;
[0024] For the class learning channel:
[0025] Introduce and randomly assign a clustering latent variable to enhance the cohesion of sentence encoding features;
[0026] Define a main channel loss function based on the final semantic representation, and define a class channel loss function based on the clustering latent variable and the final semantic representation; based on the main channel loss function and the class channel loss function, construct a hybrid loss function for the model;
[0027] After configuring the hyperparameters for model operation and calculation, use the mini-batch gradient descent method to solve the hybrid loss function to train, predict, and evaluate the performance of the model, and obtain an initial entity extraction network model, an initial clustering latent variable, and an initial evaluation score.
[0028] Preferably, assume that the input text is x , the annotation result y , where x Adopt an embedded representation based on Chinese characters and satisfy , y Adopt a one-hot representation of 0-1, the sentence completion length is l , the representation dimension is d ;
[0029] Adopt a global encoder based on the BERT model to perform sentence embedding encoding and output the global vector representation at the first [CLS] position G x ;
[0030] Use a local encoder based on the GloVe model for character-level context-related word embedding, and perform smoothing optimization through the weighted sentence vector representation model WR algorithm to obtain the local vector representation L x , expressed as:
[0031]
[0032] Among them, represents the sentence representation based on the GloVe model; is E x projected onto the first principal vector of;
[0033] represents each Chinese character x s probability of the Unigram of the unit model;
[0034] α represents a fixed scalar, used to smooth the frequency size of a single Chinese character and α (0, 1), so that under a smaller weight is biased towards Chinese characters with a larger word frequency;
[0035] The final semantic representation obtained by splicing and performing multi-layer coding mapping is as follows:
[0036]
[0037] Among them, h represents the stacking of multi-layer encoders, and concat represents the splicing operation.
[0038] Preferably, the hybrid loss function is expressed as:
[0039]
[0040] Among them, represents the hybrid loss function;
[0041] Main channel loss function is measured using the relative path score function of the CRF model; log is the logarithmic function; represents the total predicted score of all paths of the CRF model; represents the true path score;
[0042] Category channel loss function uses the mean squared error function to describe the clustering latent variable c iThe difference from the final semantic representation x E ;
[0043] γ is a weighted scalar and γ (0, 1).
[0044] Preferably, the current unlabeled dataset is subjected to pool sampling, and using the current entity extraction network model, a query strategy that combines uncertainty and minimum boundary is adopted to predict the corresponding composite score, including:
[0045] Assume an unlabeled text randomly selected from the current unlabeled dataset X U and used as the sample to be queried , then the composite score X score is expressed as follows:
[0046]
[0047] Wherein, and respectively represent the predicted probability scores of the most likely prediction result and the second most likely prediction result ;
[0048] θ represents a tuning hyperparameter and θ (0, +∞).
[0049] Preferably, based on the composite score, a query strategy with multi - spatial feature representation is adopted to select a batch of unlabeled texts with information value from the current unlabeled dataset and used as the query sample pool, including:
[0050] The maximum marginal relevance (MMR) method is used to calculate the importance degree of combinations of multiple unlabeled texts in the current unlabeled dataset, as shown in the following formula:
[0051]
[0052] Wherein, represents the importance degree score of combinations of multiple unlabeled texts including X U in the current unlabeled dataset ;
[0053] max represents the maximization function;
[0054] x uIndicates the current unlabeled dataset X U Any non- sample in;
[0055] c i Indicates any clustering latent variable; Indicates the current clustering latent variable, with the superscript t Indicating the current iteration number;
[0056] Indicates the sample to be queried The correlation with the unlabeled data space; Indicates the unlabeled dataset X U The number of texts; sim Indicates the correlation, calculated using the cosine similarity function;
[0057] Indicates the sample to be queried The correlation with the labeled training data space; c L Indicates the number of clustering latent variables;
[0058] Indicates a hyperparameter and (0, 1);
[0059] For X U All samples to be queried in Calculate the importance score, and obtain the top k samples with the highest scores as the query sample pool , and its formula is as follows:
[0060]
[0061] Wherein, argmax Indicates the variable value that maximizes the objective function.
[0062] A Chinese ancient book entity extraction system based on a deep active learning strategy, comprising:
[0063] A data acquisition module, used to acquire the original ancient book text resources, select a part of the texts for entity annotation, obtain the annotated dataset, and use the other unannotated texts as the unlabeled dataset; wherein the number of unlabeled texts is greater than the number of labeled texts;
[0064] A model initialization module, configured to configure hyperparameters for model operation and calculation, and based on an initial labeled dataset, train, predict, and evaluate the performance of a pre-constructed entity extraction network model based on a deep neural network to obtain an initial evaluation score;
[0065] An active learning query module, configured to output the final entity extraction network model if the current evaluation score exceeds the passing threshold or the number of iterations reaches the maximum number of iterations, otherwise repeatedly execute the following active learning query strategy, including:
[0066] A score prediction unit, configured to perform pool sampling on the current unlabeled dataset, and use the current entity extraction network model to predict the corresponding composite score by adopting a query strategy that combines uncertainty and minimum boundary;
[0067] A text selection unit, configured to select a batch of unlabeled texts with information value from the current unlabeled dataset as a query sample pool based on the composite score by adopting a query strategy with multi-space feature representation;
[0068] A labeling acceptance unit, configured to accept the manual labeling of the expert for the query sample pool to obtain the labeled query sample pool;
[0069] A data update unit, configured to add the labeled query sample pool to the current labeled dataset and remove it from the current unlabeled dataset;
[0070] A model update unit, configured to train, predict, and evaluate the performance of the current entity extraction network model with the current labeled dataset to obtain an updated model and an evaluation score for the next iteration process.
[0071] A storage medium stores a computer program for Chinese ancient book entity extraction based on a deep active learning strategy, wherein the computer program enables a computer to execute the Chinese ancient book entity extraction method as described above.
[0072] An electronic device includes:
[0073] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include those for executing the Chinese ancient book entity extraction method as described above.
[0074] (III) Beneficial effects
[0075] The present invention provides a Chinese ancient book entity extraction method based on a deep active learning strategy. Compared with the prior art, it has the following beneficial effects:
[0076] In the present invention, an active learning strategy is integrated with an entity extraction model based on a deep neural network, which can not only give full play to the model fitting ability of the deep network learner, but also make full use of the improved active learning strategy to further alleviate the problems of insufficient labeled data and high expert annotation costs. At the same time, a query strategy that combines uncertainty, minimum boundary, and multi-space feature representation is used to select the data pool, effectively improving the learning efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0078] Figure 1 It is a block diagram of a Chinese ancient book entity extraction method based on a deep active learning strategy provided by an embodiment of the present invention;
[0079] Figure 2 It is a schematic diagram of the steps of a Chinese ancient book entity extraction based on a deep active learning strategy provided by an embodiment of the present invention;
[0080] Figure 3 It is a schematic structural diagram of an entity extraction network model based on a deep neural network provided by an embodiment of the present invention;
[0081] Figure 4 It is a flowchart of a Chinese ancient book entity extraction method based on a deep active learning strategy provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0082] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0083] By providing a Chinese ancient book entity extraction method based on a deep active learning strategy in the embodiments of the present application, the technical problem that the existing model methods cannot be well applied to the Chinese ancient book entity extraction task is solved, and the recognition accuracy of the ancient book entity extraction model and the efficiency of the related annotation process are further improved.
[0084] The overall idea of the technical solutions in the embodiments of the present application to solve the above technical problems is as follows:
[0085] The applicant recognizes that: Active learning methods are a type of strategy method that uses limited artificial expert knowledge to label a certain number of sample candidates with key feature information from a large number of unlabeled data samples, and then continuously imports and optimizes the model training process. On the one hand, this method can improve the prediction accuracy of related machine learning tasks and can effectively reduce the cost of expert annotation as much as possible, and is particularly suitable for various sequence annotation tasks in natural language processing (such as entity extraction tasks). Currently, the active learning field mainly adopts the query sampling mode based on the pool. Specifically, the query methods in entity extraction tasks mainly include four aspects: (1) Sampling based on uncertainty estimation, such as Culotta's minimum confidence model, Scheffer's minimum marginal difference model, and Kim's maximum sequence entropy, etc.; (2) Sample extraction based on committee query, such as Dan's voting entropy model based on Hidden Markov Model (HMM) and McCallum's Kullback-Leibler divergence consistency detection model, etc.; (3) Estimation sampling method based on gradient variation, such as Lin's representative distribution measure model based on information density and Burr's model parameter uncertainty estimation model based on Fisher Information, etc.
[0086] Accordingly, a Chinese ancient book entity extraction method based on a deep active learning strategy is proposed, which can be summarized as:
[0087] First, obtain unlabeled Chinese ancient book text resources and establish a certain number of Chinese ancient book annotation training data sets, where the number of unlabeled samples is much larger than the number of labeled samples;
[0088] Second, determine the supervised learning method solution, construct and train a deep network entity extraction model, and obtain the initial model and some hyperparameters;
[0089] Third, in the active learning stage, configure the hyperparameters for model operation and calculation, perform pool-based query sampling on the unlabeled data resources, then estimate the corresponding prediction score results according to the above model, and determine a set of unlabeled samples with information value according to the above data results and query strategy function;
[0090] Then, perform expert annotation on the above sample set, and add the labeled sample set back to the annotation data set;
[0091] Finally, repeatedly iterate the above operations and stop the calculation process of the entire deep active learning model until a specific termination condition is met.
[0092] To better understand the above technical solution, the above technical solution will be described in detail below in combination with the specification drawings and specific implementation manners.
[0093] Example 1:
[0094] As Figure 1 shown, the embodiment of the present invention provides a method for extracting Chinese ancient book entities based on a deep active learning strategy, including:
[0095] S1. Obtain the original ancient book text resources, select a part of the text for entity annotation to obtain an annotated data set, and use the other unannotated text as an unannotated data set; where the number of unlabeled text is greater than the labeled text;
[0096] S2. Configure the hyperparameters for model operation and calculation, and based on the initial annotated data set, train, predict and evaluate the performance of a pre-constructed entity extraction network model based on a deep neural network to obtain an initial evaluation score;
[0097] S3. If the current evaluation score exceeds the passing threshold or the number of iterations reaches the maximum number of iterations, output the final entity extraction network model; otherwise, repeatedly execute the following active learning query strategy, including:
[0098] S31. Perform pool sampling on the current unlabeled data set, and use the current entity extraction network model to predict the corresponding composite score by adopting a query strategy that combines uncertainty and minimum boundary;
[0099] S32. Based on the composite score, adopt a query strategy with multi-space feature representation to select a batch of unlabeled text with information value from the current unlabeled data set as the query sample pool;
[0100] S33. Accept the manual annotation of the expert on the query sample pool to obtain the annotated query sample pool;
[0101] S34. Add the annotated query sample pool to the current labeled data set and remove it from the current unlabeled data set;
[0102] S35. Use the current labeled data set to train, predict and evaluate the performance of the current entity extraction network model to obtain an updated model and evaluation score for the next iteration process.
[0103] In the embodiment of the present invention, the active learning strategy is integrated with the entity extraction model based on a deep neural network, which can not only give full play to the model fitting ability of the deep network learner, but also make full use of the improved active learning strategy to further alleviate the problems of insufficient labeled data and high expert annotation cost. At the same time, a query strategy that combines uncertainty, minimum boundary and multi-space feature representation is adopted for data pool selection, effectively improving the learning efficiency.
[0104] Furthermore, for the entity extraction network model, the embodiment of the present invention designs a dual-encoder structure, which integrates a global encoder based on pre-trained learning and a local encoder based on character-level embedding to obtain a fused semantic representation of the sentence, so as to comprehensively capture more comprehensive context semantic knowledge of the sentence, thereby improving the performance of the entity extraction model based on the deep neural network.
[0105] As Figure 2 shown, Figure 2 a schematic diagram of the steps of Chinese ancient book entity extraction based on a deep active learning strategy is disclosed.
[0106] Next, each step of the above solution will be described in detail in combination with Figure 2 :
[0107] In step S1, the original ancient book text resources are obtained, a part of the text is selected for entity annotation to obtain an annotated data set, and the other unannotated text is used as an unannotated data set; where the number of unlabeled texts is greater than the number of labeled texts.
[0108] In this step, a certain number of text sets are randomly selected from the original ancient book text resources for entity annotation and used as the initial annotated data set X L , and the other initial unannotated data set is denoted as X U .
[0109] It should be noted that, preferably, the number of texts in the unannotated data set here is much larger than the number of texts in the annotated data set .
[0110] In step S2, the hyperparameters for model operation and calculation are configured, and based on the initial annotated data set, the pre-constructed entity extraction network model based on the deep neural network is trained, predicted, and performance evaluated to obtain an initial evaluation score.
[0111] As Figure 3 shown, the entity extraction network model provided by the embodiment of the present invention adopts a deep neural sequence annotation network with a dual encoder-decoder, including a main learning channel for entity sequence annotation and a class learning channel for sentence representation.
[0112] Correspondingly, in this step, the hyperparameters for model operation and calculation are configured, and based on the initial annotated data set, the pre-constructed entity extraction network model based on the deep neural network is trained, predicted, and performance evaluated to obtain an initial evaluation score, including:
[0113] First, it is explained that the hyperparameters for configuring model operation and calculation may include the annotated data set X L, the unlabeled dataset is denoted as X U and the number of training categories C L , and can also include active learning hyperparameters reaching thresholds S , maximum number of iterations T , the number of query samples ranked at the top of the sample query sample pool k , adjust hyperparameters θ , hyperparameters , weighted scalar γ wait.
[0114] As mentioned above, in the stage of building a supervised deep learning model, we first construct a "dual encoder-decoder" structure based on the general end-to-end sequence learning network topology. This structure includes two information channels, namely a main learning channel for entity sequence labeling and a class learning channel for sentence representation; among them:
[0115] 1. For the main learning channel:
[0116] Assume the input text is x , marking results y ,in x Adopting embedded representation based on Chinese characters and satisfying , y Using 0-1 unique hot representation, the sentence completion length is l , the representation dimension is d .
[0117] On the one hand, a global encoder based on pre-training learning is used to perform sentence embedding encoding on the input labeled text to obtain a global vector representation.
[0118] For example, a global encoder based on the BERT (Bidirectional Encoder Representation from Transformers) model is used to perform sentence embedding encoding, and the global vector representation of the first "[CLS]" position is output. G x .
[0119] On the other hand, a local encoder based on character-level embedding is used to embed character-level context-related words into the input labeled text, and smooth optimization is performed to obtain a local vector representation.
[0120] Exemplarily, a local encoder based on the GloVe (Global Vectors for Word Representation) model is adopted for character-level context-related word embedding, and the WR algorithm of the weighted sentence vector representation model is used for smoothing optimization. Compared with traditional word vector algorithms such as the Skip-thought model and the GloVe model, a large amount of sentence semantic information will be lost when representing sentences. The WR optimization model takes the word frequency distribution as the weighting factor, calculates the weighted mean of the character vector representations in all sentences, and removes irrelevant semantics through the SVD method to obtain the core semantic representation of the sentence. Considering that GloVe+WR has more stable and excellent performance in related sentence representation experiments. Assume that the sentence representation based on the GloVe model is E x ( ) E x projected onto the first principal vector of , which can be calculated by the singular value decomposition algorithm. Then the local vector representation based on GloVe+WR L x can be expressed as follows:
[0121]
[0122] where, represents the sentence representation based on the GloVe model, that is, using the statistical language model NGRAM (N = 1) for the sequence x representation;
[0123] represents the probability of the Unigram of the unit model for each Chinese character x s ;
[0124] α represents a fixed scalar used to smooth the frequency of a single Chinese character and α (0, 1), so that tends to the Chinese character with a larger word frequency under a smaller weight.
[0125] Immediately afterwards, the global vector representation and the local vector representation are concatenated and multi-layer coding mapping is performed to obtain the final semantic representation, as shown below:
[0126]
[0127] where, h represents the stacking of multi-layer encoders, and concat represents the concatenation operation.
[0128] It is understandable that for each layer of encoding in the above multi-layer encoder, any one (one or more layers) of the neural encoding networks for sequence data (such as the common perceptron network MLP, recurrent neural network RNN, convolutional neural network CNN, attention network Attention, and their combined network forms) can be selected for combined encoding.
[0129] (2) For the class learning channel:
[0130] Introduce and randomly assign a clustering latent variable to enhance the cohesion of sentence encoding features.
[0131] Therefore, based on the final semantic representation, define the main channel loss function, and based on the clustering latent variable and the final semantic representation, define the class channel loss function; based on the main channel loss function and the class channel loss function, construct the hybrid loss function of the model, which is defined as:
[0132]
[0133] Among them, represents the hybrid loss function;
[0134] Main channel loss function Use the relative path score function of the CRF model for measurement; log is the logarithmic function; represents the total predicted score of all paths of the CRF model; represents the true path score;
[0135] Class channel loss function Use the mean squared error function to describe the difference between the clustering latent variable c i and the final semantic representation x E ;
[0136] γ is a weighted scalar and γ (0, 1).
[0137] Finally, after configuring the hyperparameters for model operation and calculation, use the mini-batch gradient descent method to solve the hybrid loss function to train, predict, and evaluate the performance of the model, and obtain the initial entity extraction network model M t Initial clustering latent variable And the initial evaluation score φ t (in the initial case t = 0, φ tIt can be a model evaluation metric for any entity extraction task, such as recall, precision, or F F1 value).
[0138] In step S3, if the current evaluation score exceeds the passing threshold or the number of iterations reaches the maximum number of iterations, the final entity extraction network model is output; otherwise, the active learning query strategy is repeatedly executed.
[0139] Such as Figure 4 shown, Figure 4 A flowchart of a Chinese ancient book entity extraction method based on a deep active learning strategy is disclosed, which fully covers the processing process in this step, specifically including:
[0140] S31. Perform pool sampling on the current unlabeled dataset, and use the current entity extraction network model to predict the corresponding composite score by adopting a query strategy that combines uncertainty and least margin.
[0141] In the active learning query strategy stage, first based on the model M t Predict the unlabeled dataset X U Labels ( X U The number of samples of ) to obtain the decoded score X score .
[0142] Specifically, integrate two common query strategies, namely uncertainty and least margin. The uncertainty used in the embodiments of the present invention measures the output score of the decoding layer. The smaller the score value, the more uncertain the prediction of the model for the unlabeled text, so it is more valuable to be annotated; while the least margin measures the difference in the prediction score distribution of the label sequence. The smaller this value, the more difficult it is to distinguish the category of the unlabeled text, and thus it should be more worthy of expert annotation.
[0143] Assume that an unlabeled text randomly selected from the current unlabeled dataset X U is used as the sample to be queried , then the composite score X score is expressed as follows:
[0144]
[0145] Wherein, and respectively represent the predicted probability scores of the most likely prediction result and the predicted probability scores of the second most likely prediction result ;
[0146] θ denote the tuning hyperparameters and θ (0, +∞).
[0147] It should be noted that since the minimum boundary is generally much smaller than the index of the uncertainty itself and its importance is relatively weaker than the latter, therefore, the exponential decay adopted in the embodiments of the present invention can well represent the above requirements, and by adjusting the above tuning hyperparameters θ to flexibly adapt to the task of ancient book entity extraction in specific business scenarios.
[0148] S32. Based on the composite score, adopt a query strategy with multi-space feature representativeness to select a batch of unlabeled texts with information value from the current unlabeled dataset and use them as a query sample pool.
[0149] In the past, the research on active learning methods sorted and screened each individual sample to be queried one by one, which often brought a large amount of redundant sample information and greatly reduced the efficiency of active learning. The key reason is that since the overall space is unknown, the above sampling function often selects data points on the edge of the decision facet, ignoring the representativeness of the global hypothesis space. Some methods use diversification metrics to solve the redundancy problem without considering the representativeness in the training space.
[0150] In contrast, the embodiments of the present invention use the Maximal Marginal Relevance (MMR) method common in the field of information retrieval to calculate the importance of multiple combinations of unlabeled texts in the current unlabeled dataset, as shown in the following formula:
[0151]
[0152] where represents the importance score of multiple combinations of unlabeled texts included in the current unlabeled dataset X U included in ;
[0153] max represents the maximization function;
[0154] x u represents any sample in the current unlabeled dataset X U not ;
[0155] c i represents any clustering latent variable; Denote the current clustering latent variable, the superscript t denotes the current iteration number;
[0156] denotes the sample to be queried and the correlation with the unlabeled data space; denotes the unlabeled data set X U the number of texts; sim denotes the correlation, calculated using the cosine similarity function;
[0157] denotes the sample to be queried and the correlation with the labeled training data space; c L denotes the number of clustering latent variables;
[0158] denotes a hyperparameter and ϵ ∈ (0, 1);
[0159] For X U all the samples to be queried in calculate the importance scores, and obtain the top k samples with the highest scores as the query sample pool , and its formula is as follows:
[0160]
[0161] where, argmax denotes the variable value that maximizes the objective function.
[0162] S33. Accept the manual annotation of the expert for the query sample pool, and obtain the annotated query sample pool.
[0163] Utilize the annotation expert to manually annotate the above query sample pool to generate the annotated query sample pool .
[0164] S34. Add the annotated query sample pool to the current labeled data set and remove it from the current unlabeled data set.
[0165] Add the generated annotated query sample pool to the current labeled data set X L and at the same time X U remove this data set from
[0166] S35. Train, predict, and evaluate the performance of the current entity extraction network model using the current labeled dataset to obtain an updated model and an evaluation score for use in the next iteration process.
[0167] Use the updated labeled dataset, i.e., the current labeled dataset X L to train, predict, and evaluate the performance of the current entity extraction network model M t and thereby obtain an updated model for predicting the composite score in the next iteration M t+1 and an evaluation score for determining the termination condition in the next iteration φ t+1 .
[0168] The evaluation score of the iteration termination condition model in the embodiments of the present invention exceeds the passing threshold S and S ∈(0, 1), or the number of iterations exceeds the maximum number of iterations T ( T is a positive integer). It is not difficult to understand that the former represents the minimum requirement for the model performance to meet the standard, and the latter represents the cost of the maximum label, and the values of the above hyperparameters can be confirmed by specific business scenarios.
[0169] In particular, those skilled in the art should be aware that the provided entity extraction method can be applied to the field of ancient Chinese entity information extraction and can also be extended to related natural language processing fields.
[0170] Embodiment 2:
[0171] The embodiments of the present invention provide a Chinese ancient book entity extraction system based on a deep active learning strategy, including:
[0172] A data acquisition module for acquiring the original ancient book text resources, selecting a part of the text for entity annotation to obtain an annotated dataset, and using the other unannotated text as an unannotated dataset; wherein the number of unlabeled text is greater than the number of labeled text;
[0173] A model initialization module for configuring the hyperparameters for model operation and calculation, and based on the initial annotated dataset, training, predicting, and evaluating the performance of a pre-constructed entity extraction network model based on a deep neural network to obtain an initial evaluation score;
[0174] An active learning query module for outputting the final entity extraction network model if the current evaluation score exceeds the passing threshold or the number of iterations reaches the maximum number of iterations, otherwise repeatedly executing the following active learning query strategy, including:
[0175] A score prediction unit, which is used to perform pool sampling on the current unlabeled dataset, and utilize the current entity extraction network model to adopt a query strategy that combines uncertainty and minimum boundary to predict the corresponding composite score;
[0176] A text selection unit, which is used to select a batch of unlabeled texts with information value from the current unlabeled dataset based on the composite score and adopt a query strategy with multi-spatial feature representation as the query sample pool;
[0177] An annotation acceptance unit, which is used to accept the manual annotation of the expert on the query sample pool to obtain the annotated query sample pool;
[0178] A data update unit, which is used to add the annotated query sample pool to the current labeled dataset and remove it from the current unlabeled dataset;
[0179] A model update unit, which is used to train, predict and evaluate the performance of the current entity extraction network model with the current labeled dataset to obtain the updated model and evaluation score for the next iteration process.
[0180] Embodiment 3:
[0181] The embodiment of the present invention provides a storage medium, which stores a computer program for Chinese ancient book entity extraction based on a deep active learning strategy, wherein the computer program enables a computer to execute the Chinese ancient book entity extraction method as described in Embodiment 1.
[0182] Embodiment 4:
[0183] The embodiment of the present invention provides an electronic device, including:
[0184] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include those for executing the Chinese ancient book entity extraction method as described in Embodiment 1.
[0185] It can be understood that the Chinese ancient book entity extraction system, storage medium and electronic device provided by the embodiment of the present invention correspond to the Chinese ancient book entity extraction method provided by the embodiment of the present invention. The explanations, examples, beneficial effects, etc. of the relevant content can refer to the corresponding parts in the Chinese ancient book entity extraction method, and will not be elaborated here.
[0186] To sum up, compared with the prior art, the following beneficial effects are achieved:
[0187] 1. In the embodiments of the present invention, an active learning strategy is integrated with an entity extraction model based on a deep neural network, which can not only give full play to the model fitting ability of the deep network learner, but also make full use of the improved active learning strategy to further alleviate the problems of insufficient labeled data and high expert annotation costs.
[0188] 2. In the embodiments of the present invention, a dual-encoder structure is designed to integrate a global encoder based on pre-trained learning and a local encoder based on character-level embedding to obtain a fused semantic representation of the sentence, so as to comprehensively capture more comprehensive context semantic knowledge of the sentence, thereby improving the performance of the entity extraction model based on the deep neural network.
[0189] 3. In the supervised learning process, the embodiments of the present invention innovatively introduce a clustering latent variable that does not occupy the running space to reconstruct a hybrid loss function, thereby obtaining the encoding information of the training sample hypothesis space, and adopting a certain query strategy to capture the representational differences from the labeled space and the unlabeled space, and then more high-quality unlabeled query samples can be obtained, thereby improving the efficiency of pool-based active selection.
[0190] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including an..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0191] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A Chinese ancient book entity extraction method based on deep active learning strategy, characterized in that: include: Obtain the original ancient book text resources, select a part of the text for entity annotation, obtain the annotated data set, and use other unannotated texts as unannotated data sets; the number of unannotated texts is greater than the number of labeled texts; Configure the hyperparameters for model operation and calculation, and train, predict, and evaluate the performance of the pre-built deep neural network-based entity extraction network model based on the initial annotated data set to obtain the initial evaluation score; If the current evaluation score exceeds the threshold or the number of iterations reaches the maximum number of iterations, the final entity extraction network model is output. Otherwise, the following active learning query strategies are repeatedly executed, including: Pool sampling is performed on the current unlabeled dataset, and the corresponding composite score is predicted using the current entity extraction network model and a query strategy that integrates uncertainty and minimum boundary. Based on the composite score, a query strategy based on multi-space feature representation is adopted to select a batch of unlabeled texts with information value from the current unlabeled data set and use them as a query sample pool; Accepting manual annotation of the query sample pool by experts to obtain the annotated query sample pool; Adding the labeled query sample pool to the current labeled data set and removing it from the current unlabeled data set; Use the current labeled data set to train, predict, and evaluate the performance of the current entity extraction network model, and obtain an updated model and evaluation score for the next iteration process; The entity extraction network model adopts a dual encoder-decoder deep neural sequence labeling network, including a main learning channel for entity sequence labeling and a class learning channel for sentence representation; The configuration model runs and calculates hyperparameters, and based on the initial annotated data set, trains, predicts, and evaluates the performance of the pre-built deep neural network-based entity extraction network model to obtain an initial evaluation score, including: For the main learning channel: A global encoder based on pre-training learning is used to perform sentence embedding encoding on the input labeled text to obtain a global vector representation; A local encoder based on character-level embedding is used to embed character-level context-related words into the input labeled text, and then smooth optimization is performed to obtain a local vector representation. splicing the global vector representation and the local vector representation and performing multi-layer encoding mapping to obtain a final semantic representation; For class study channels: Introduce and randomly assign a clustering latent variable to enhance the cohesion of sentence encoding features; Based on the final semantic representation, a main channel loss function is defined, and based on the clustering latent variables and the final semantic representation, a category channel loss function is defined; based on the main channel loss function and the category channel loss function, a mixed loss function of the model is constructed; After configuring the hyperparameters for model operation and calculation, the mini-batch gradient descent method is used to solve the hybrid loss function to train, predict and evaluate the performance of the model, and obtain the initial entity extraction network model, initial clustering latent variables and initial evaluation scores.
2. The Chinese ancient book entity extraction method as claimed in claim 1, characterized in that: Assume the input text is x , marking results y ,in x Adopting embedded representation based on Chinese characters and satisfying , y Using 0-1 unique hot representation, the sentence completion length is l , the representation dimension is d ; A global encoder based on the BERT model is used to perform sentence embedding encoding and output the global vector representation of the first [CLS] position G x ; The local encoder based on the GloVe model is used to embed character-level context-related words, and the weighted sentence vector representation model WR algorithm is used for smooth optimization to obtain the local vector representation. L x , expressed as: in, Represents sentence representation based on the GloVe model; for E x Projection onto the first principal vector of ; Represents each Chinese character x s The probability of the unit model Unigram; α Represents a fixed scalar used to smooth the frequency of a single Chinese character. α (0,1) so that the smaller the weight, the Prefer Chinese characters with higher word frequency; The final semantic representation obtained by concatenating and performing multi-layer encoding mapping is as follows: in, h Represents the stacking of multiple layers of encoders, and concat represents the concatenation operation.
3. The Chinese ancient book entity extraction method as claimed in claim 2, characterized in that: The hybrid loss function is expressed as: in, represents the mixed loss function; Main channel loss function The relative path score function of the CRF model is used for measurement; log is a logarithmic function; Represents the total predicted score of all paths of the CRF model; represents the true path score; Category channel loss function The square difference function is used to describe the clustering latent variables. c i With the final semantic representation x E The difference between γ is a weighted scalar and γ (0,1).
4. The Chinese ancient book entity extraction method as claimed in claim 1, characterized in that: The method of performing pool sampling on the current unlabeled data set, using the current entity extraction network model, and adopting a query strategy that integrates uncertainty and minimum boundary to predict the corresponding composite score includes: Assume that the current dataset is unlabeled X U An unlabeled text is selected from the query sample , then the composite fraction X score It is expressed as follows: in, and Represent the most likely prediction results The predicted probability score and the second most likely predicted outcome The predicted probability score of θ represents the hyperparameters to be adjusted and θ (0,+∞).
5. The Chinese ancient book entity extraction method as claimed in claim 4, characterized in that: The method of using a query strategy based on the composite score and a multi-space feature representation to select a batch of unlabeled texts with information value from the current unlabeled data set as a query sample pool includes: The maximum marginal relevance (MMR) method is used to calculate the importance of multiple unlabeled text combinations in the current unlabeled dataset, as shown in the following formula: in, Indicates the current unlabeled dataset X U Included The importance score of multiple unlabeled text combinations; max represents the maximization function; x u Indicates the current unlabeled dataset X U Any non Samples of c i represents any clustering latent variable; represents the current clustering latent variable, the superscript t Indicates the current iteration number; Indicates the sample to be queried Relevance to the unlabeled data space; Represents an unlabeled dataset X U The number of texts; sim Represents correlation, calculated using the cosine similarity function; Indicates the sample to be queried Correlation with the labeled training data space; c L represents the number of clustering latent variables; represents a hyperparameter and (0,1); right X U All the samples to be queried in Calculate the importance score and get the top ranking with the largest score k Samples as query sample pool , the formula is as follows: in, argmax Represents the variable value that maximizes the objective function.
6. A Chinese ancient book entity extraction system based on deep active learning strategy, characterized in that: The method for extracting entities from ancient Chinese books as claimed in claim 1 comprises: The data acquisition module is used to obtain the original ancient book text resources, select a part of the text for entity annotation, obtain the annotated data set, and use other unannotated texts as unannotated data sets; the number of unannotated texts is greater than the number of marked texts; The model initialization module is used to configure the hyperparameters of model operation and calculation, and to train, predict and evaluate the performance of the pre-built entity extraction network model based on the deep neural network based on the initial annotated data set to obtain the initial evaluation score; The active learning query module is used to output the final entity extraction network model if the current evaluation score exceeds the threshold or the number of iterations reaches the maximum number of iterations. Otherwise, the following active learning query strategies are repeatedly executed, including: The score prediction unit is used to perform pool sampling on the current unlabeled data set, and predict the corresponding composite score by using the query strategy of integrating uncertainty and minimum boundary using the current entity extraction network model; A text selection unit, configured to select a batch of unlabeled texts with information value from the current unlabeled data set based on the composite score and adopt a query strategy of multi-space feature representation, and use the unlabeled texts as a query sample pool; A labeling receiving unit, used for receiving manual labeling of the query sample pool by experts, and obtaining the labeled query sample pool; A data updating unit, used to add the labeled query sample pool to the current labeled data set and remove it from the current unlabeled data set; The model updating unit is used to train, predict and evaluate the performance of the current entity extraction network model using the current labeled data set, and obtain the updated model and evaluation score for the next iteration process.
7. A storage medium, characterized in that: It stores a computer program for extracting entities from ancient Chinese books based on a deep active learning strategy, wherein the computer program enables a computer to execute the method for extracting entities from ancient Chinese books as described in any one of claims 1 to 5.
8. An electronic device, characterized in that: include: one or more processors; Memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include programs for executing the Chinese ancient book entity extraction method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Biomedical relationship extraction method and system fusing iterative active learning
CN116070700A