Tang poetry named entity interactive visual labeling method and device and storage medium
By using a two-way long and short-term memory network model embedded in the self-attention mechanism in the named entity recognition task, combining intelligent recommendation and interactive visual interface, the problem of low labeling efficiency and quality in traditional active learning methods is solved, and more efficient and high-quality named entity recognition is achieved.
Patent Information
- Application Number
- CN202510175686.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional active learning methods have problems with low labeling efficiency and quality in naming entity recognition tasks, especially in the absence of visual labeling tools.
The bidirectional long and short-term memory network model (AtnBiLSTM) embedded with the self-attention mechanism is combined with the BERT dictionary, and samples are recommended through the minimum confidence algorithm and the maximum normalized logarithmic probability algorithm, and the labeling process is optimized through an interactive visual interface.
It improves the accuracy of naming entity recognition, reduces the workload of manual labeling, improves the quality and coverage of labeling data, and enhances the efficiency and transparency of model training and data set production.
Smart Images

Figure CN120106067A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine learning, and in particular relates to a method and device for interactively visually annotating named entities in Tang poetry, and a storage medium. Background Art
[0002] Active learning is a machine learning method that uses intelligent algorithms to guide the model to select the most valuable samples for learning in order to improve model performance. Compared with the traditional passive learning model, active learning can achieve higher accuracy and efficiency by selectively annotating and learning the samples that contribute the most to the model when the amount of data is limited. Currently, active learning has been widely used in text classification, image classification, medical diagnosis and other fields, especially when the cost of data acquisition is high or the annotation resources are limited.
[0003] The active learning process is an iterative cycle, including training models, generating prediction probabilities, recommending sample sets, and labeling samples. Traditional active learning methods mainly include two modules: one is the machine learning model used to predict sample labels. However, traditional machine learning methods such as support vector machines and decision trees usually require manual feature design when dealing with named entity recognition (NER) tasks. Manual feature design is not only time-consuming and labor-intensive, but also difficult to cover all possible features. The second is the query algorithm that prioritizes the labeled data. The current mainstream query methods include uncertainty-based queries, committee queries, and information density queries. Through the collaboration of these two modules, the active learning method can select the sample set that is most worth labeling in each round. The labeler forms a new labeled sample set by labeling the sample set generated in each round. In this process, the labeler needs to label every sample in the sample set, even if these samples have been completely predicted correctly by the model.
[0004] In summary, traditional active learning methods have the following defects: on the one hand, although traditional machine learning methods are still used in some small-scale or simple tasks, they generally perform worse than neural network models in NER tasks; on the other hand, in the absence of visual annotation tools, the effectiveness and efficiency of active learning may be significantly limited. Annotators face a higher cognitive burden and find it difficult to understand the model selection and data characteristics, resulting in reduced annotation quality, reduced efficiency, and difficulty in discovering potential problems. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a method and device for interactive visual annotation of named entities in Tang poetry, and a storage medium, which can ensure the quality of annotated samples while ensuring the annotation speed.
[0006] To achieve the above object, the present invention adopts the following technical solution:
[0007] An interactive visual annotation method for named entities in Tang poetry includes:
[0008] Step 1: According to the Tang poetry JSON dataset, obtain Tang poetry samples that need to be recommended to the annotation interface;
[0009] Step 2: Present the Tang poetry samples that need to be recommended to the annotation interface on the interface.
[0010] Preferably, step 1 comprises:
[0011] Parse the Tang poetry JSON dataset and store it in a formatted format;
[0012] Use the BERT dictionary to get the word ID matrix for each text and feed it into the Embedding layer to get the corresponding word vector;
[0013] Construct a bidirectional long short-term memory network model AtnBiLSTM embedded with self-attention mechanism and use softmax for classification;
[0014] The minimum confidence algorithm and the maximum normalized logarithmic probability algorithm are used to calculate the score of each sample on the corresponding algorithm, and then sort them accordingly to determine the Tang poetry samples that need to be recommended to the annotation interface.
[0015] Preferably, in step 2, a Tang poetry data statistics view, a recommended sample view and an annotation view are presented on the interface.
[0016] The present invention also provides an interactive visual annotation device for named entities in Tang poetry, comprising:
[0017] The first processing module is used to obtain Tang poetry samples that need to be recommended to the annotation interface according to the Tang poetry JSON data set;
[0018] The second processing module is used to present the Tang poetry samples that need to be recommended to the annotation interface on the interface.
[0019] Preferably, the first processing module comprises:
[0020] The first processing unit is used to parse the Tang poetry JSON data set and store it in a formatted manner;
[0021] The second processing unit is used to obtain the word ID matrix of each text using the BERT dictionary and feed it into the Embedding layer to obtain the corresponding word vector;
[0022] The third processing unit is used to construct a bidirectional long short-term memory network model AtnBiLSTM embedded with a self-attention mechanism and use softmax for classification;
[0023] The fourth processing unit is used to calculate the score of each sample on the corresponding algorithm using the minimum confidence algorithm and the maximum normalized logarithmic probability algorithm, and sort them accordingly to determine the Tang poetry samples that need to be recommended to the annotation interface.
[0024] Preferably, the second processing module presents a Tang poetry data statistics view, a recommended sample view and an annotation view on the interface.
[0025] An embodiment of the present invention further provides a storage medium, on which a computer program is stored, and the computer program executes the Tang poetry entity interactive annotation method when running.
[0026] The present invention can ensure the quality of annotation while ensuring the speed of named entity annotation of Tang poetry. Specifically, the bidirectional long short-term memory network model (AtnBiLSTM) embedded with a self-attention mechanism integrates the representative features of Tang poetry words and their contexts. By improving the BiLSTM model, the model's recognition ability of Tang poetry text features is enhanced, thereby improving the accuracy of named entity recognition. In addition, an intelligent algorithm is used to select the most representative and informative samples from a large-scale unlabeled data set for annotation, effectively reducing the workload of manual annotation and improving the quality and coverage of labeled data. This method enables the model to learn from limited labeled data more efficiently, thereby improving overall recognition performance. In addition, the present invention intuitively presents the results and processes of the intelligent model through visualization technology, and incorporates people into the intelligent recognition closed loop through interactive design, optimizing the entire entity annotation and model training process, thereby enhancing the efficiency and transparency of model training and data set production. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0028] Figure 1 is a flow chart of the Tang poetry entity interactive annotation method of the present invention;
[0029] Figure 2 It is a comparative schematic diagram of the interactive annotation method of Tang poetry entities based on active learning of the present invention in three selection strategies and two annotation methods;
[0030] Figure 3 Compare the performance of the dataset annotated using the method of the present invention and the randomly annotated dataset on the same model. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0032] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0033] Embodiment 1:
[0034] The embodiment of the present invention provides an interactive annotation method for Tang poetry entities. According to the characteristics of Tang poetry texts, a set of annotation specifications for guiding entity annotation is formulated; the specifications define entity types and entity boundaries; wherein entity annotation types are divided into three entity types: person names, place names, and time, and the entity types are clarified; for person name entities, named references, partial noun references, names of groups of people, and phrases of "place + direction" are annotated, and noun references borrowed to refer to occupations, official positions or titles used alone, and names of characters in mythology are not annotated; for place entities, named references and noun references such as place names, country names, mountain names, and water names are annotated, and noun references without clear reference meanings and place names in mythology are not annotated; for time entities, years, seasons, months, reign titles, etc. are annotated. Phrases describing images are not annotated.
[0035] like Figure 1 As shown, the Tang poetry entity interactive annotation method of the embodiment of the present invention includes:
[0036] Step 1: According to the Tang poetry JSON dataset, obtain Tang poetry samples that need to be recommended to the annotation interface, including the following steps:
[0037] Step 1.1, upload the annotated text set: the text set to be annotated is organized into a JSON file, and each piece of data has a fixed format (attribute). The formatted JSON file is sent to the server, which parses the JSON file and stores the data one by one into the database.
[0038] Step 1.2, text word vector generation: Use the vocabulary table established by the "Bidirectional Encoder Representations from Transformers" (BERT) model to replace each word in the Tang poetry text with its corresponding ID in the vocabulary table. Input the generated ID sequence into the Embedding layer, and finally obtain the context-sensitive word vector of each token.
[0039] Step 1.3, model training: The word vector matrix is input into the Bidirectional Long Short-Term Memory with Attention Mechanism (AtnBiLSTM) model for training.
[0040] In this step, the model training includes the following steps:
[0041] Step 1.3.1: Convert each user's labeled sample and initial training set into a corresponding word ID matrix through BERT.
[0042] Step 1.3.2: Establish a bidirectional long short-term memory network model AtnBiLSTM embedded with self-attention mechanism
[0043] Step 1.3.3: The word ID matrix is calculated through the Embedding layer to obtain the word vector matrix L. The word vector matrix is first calculated through the forward LSTM layer of AtnBiLSTM to obtain the L1 matrix. Then L1 is sent to the self-attention layer for calculation to obtain La1. L1 and La1 are concatenated to obtain H1. Flip L and send it to the reverse LSTM to calculate L2. L2 is also calculated through self-attention to obtain La2. L2 and La2 are concatenated to obtain H2. The final output of the AtnBiLSTM layer is [H1, H2].
[0044] Step 1.3.4: Since the number of labels is 8, the output of AtnBiLSTM needs to pass through a fully connected layer with an output dimension of 8 and softmax as the activation function. Softmax can calculate the probability value of each label.
[0045] Step 1.3.5: Use the argmax function to calculate the position of the maximum probability value in the label matrix corresponding to each word, thereby calculating the label corresponding to each word.
[0046] Step 1.4, text score calculation: Calculate the minimum confidence score and the maximum normalized logarithmic probability score of each text. When recommending samples, sort the corresponding scores according to the score type and sample number selected by the user, and select the corresponding number of samples to return to the client.
[0047] In this step, the text score calculation includes the following steps:
[0048] Step 1.4.1, calculate the predicted probability matrix of the sample through the reduce_max function, and calculate the maximum probability value corresponding to each word; through such calculation, the probability matrix of each sample is obtained.
[0049] Step 1.4.2, then calculate the minimum confidence score of the sample, where the minimum probability of a sample is used as the minimum confidence score of the sample. Specifically, the probability matrix of the sample is reduced by min to calculate the minimum probability value.
[0050] Step 1.4.3, calculate the maximum normalized logarithmic probability score of the sample again. Specifically, find the logarithmic value of each position in the probability matrix of the sample, and then calculate the sum of the logarithmic values of each sample. Finally, divide the value of each sample by the total number of samples.
[0051] Step 2: Presenting Tang poetry samples that need to be recommended to the annotation interface on the interface, including the following steps:
[0052] Step 2.1, Tang poetry data statistics: After the data uploaded by the user is processed by the server, the basic information of the Tang poetry data will be returned, including the data volume, the labeled volume, the unlabeled volume, the number of each labeled entity, etc.
[0053] Step 2.2: Present recommended samples: Each sample or entity is presented in a square or circle on the interface, and different colors or border styles are used to distinguish different entity types or annotation states.
[0054] Step 2.3, presenting annotated samples: The server sends the samples to be annotated to the client in the form of text and labels. The client presents the text on the interface and highlights the corresponding entity content in the text according to the label of the corresponding text. A variety of interactive methods are designed to enable users to operate the text and annotate the entities in the text.
[0055] In the embodiment of the present invention, first, a Tang poetry server is built to support uploading of a Tang poetry JSON data set file. The server parses the file and formats it and stores it in a database. Secondly, a word vector of the Tang poetry text is generated, and the word ID matrix of each text is obtained using the BERT dictionary, and it is sent to the Embedding layer to obtain the corresponding word vector. Next, a bidirectional long short-term memory network model AtnBiLSTM embedded with a self-attention mechanism is established, and softmax is used for classification to complete model training. Then, the score of each sample is calculated using the minimum confidence algorithm and the maximum normalized logarithmic probability algorithm, and sorted accordingly to determine the Tang poetry samples that need to be recommended to the annotation interface. Finally, an interactive annotation visualization interface is built, which mainly includes a Tang poetry data statistical view, a recommended sample view, and an annotation view. This method can intelligently recommend the text annotation samples that require the most manual intervention, intuitively visualize the Tang poetry text, model, and annotation process, thereby effectively improving the efficiency and transparency of Tang poetry named entity annotation.
[0056] Below through Figure 2 and Figure 3 Further understand the advantages of the method of the present invention compared with other traditional methods:
[0057] Figure 2 The figure is a comparative diagram of the three selection strategies and two annotation methods of the interactive annotation method of Tang poetry entities based on active learning described in the present invention. The three selection strategies are random selection, selection method based on minimum confidence and selection method based on maximum normalized logarithmic probability. The two annotation methods are interactive system assisted annotation and plain text annotation. Figure 2 It can be seen that the minimum confidence method based on the interactive system has the shortest annotation time and achieves the best effect.
[0058] Figure 3 The performance of the dataset annotated using the method of the present invention and the randomly annotated dataset on the same model was compared. Figure 3 The results show that the dataset annotated by the method of the present invention has a higher recognition effect on the model than the other three randomly annotated datasets, thereby proving that the quality of the dataset annotated by the method of the present invention is higher than that of the randomly selected dataset.
[0059] Embodiment 2:
[0060] The embodiment of the present invention further provides a Tang poetry entity interactive annotation device, comprising:
[0061] The first processing module is used to obtain Tang poetry samples that need to be recommended to the annotation interface according to the Tang poetry JSON data set;
[0062] The second processing module is used to present the Tang poetry samples that need to be recommended to the annotation interface on the interface.
[0063] As an implementation of the embodiment of the present invention, the first processing module includes:
[0064] The first processing unit is used to parse and format the Tang poetry JSON data set;
[0065] The second processing unit is used to obtain the word ID matrix of each text using the BERT dictionary and feed it into the Embedding layer to obtain the corresponding word vector;
[0066] The third processing unit is used to construct a bidirectional long short-term memory network model AtnBiLSTM embedded with a self-attention mechanism and use softmax for classification;
[0067] The fourth processing unit is used to calculate the score of each sample on the corresponding algorithm using the minimum confidence algorithm and the maximum normalized logarithmic probability algorithm, and sort them accordingly to determine the Tang poetry samples that need to be recommended to the annotation interface.
[0068] As an implementation method of the embodiment of the present invention, the second processing module presents a Tang poetry data statistics view, a recommended sample view and an annotation view on the interface.
[0069] Embodiment 3:
[0070] An embodiment of the present invention further provides a storage medium, on which a computer program is stored, and the computer program executes the Tang poetry entity interactive annotation method when running.
[0071] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.
Claims
1. An interactive visual annotation method for named entities in Tang poetry, characterized in that: include: Step 1: According to the Tang poetry JSON dataset, obtain Tang poetry samples that need to be recommended to the annotation interface; Step 2: Present the Tang poetry samples that need to be recommended to the annotation interface on the interface.
2. The method for interactive visual annotation of named entities in Tang poetry as claimed in claim 1, characterized in that: Step 1 includes: Parse the Tang poetry JSON dataset and store it in a formatted format; Use the BERT dictionary to get the word ID matrix for each text and feed it into the Embedding layer to get the corresponding word vector; Construct a bidirectional long short-term memory network model AtnBiLSTM embedded with self-attention mechanism and use softmax for classification; The minimum confidence algorithm and the maximum normalized logarithmic probability algorithm are used to calculate the score of each sample on the corresponding algorithm, and then sort them accordingly to determine the Tang poetry samples that need to be recommended to the annotation interface.
3. The method for interactive visual annotation of named entities in Tang poetry as claimed in claim 2, characterized in that: In step 2, the Tang poetry data statistics view, recommended sample view and annotation view are presented on the interface.
4. An interactive visual annotation device for named entities in Tang poetry, characterized in that: include: The first processing module is used to obtain Tang poetry samples that need to be recommended to the annotation interface according to the Tang poetry JSON data set; The second processing module is used to present the Tang poetry samples that need to be recommended to the annotation interface on the interface.
5. The device for interactive visual annotation of named entities in Tang poetry as claimed in claim 4, characterized in that: The first processing module includes: The first processing unit is used to parse the Tang poetry JSON data set and store it in a formatted manner; The second processing unit is used to obtain the word ID matrix of each text using the BERT dictionary and feed it into the Embedding layer to obtain the corresponding word vector; The third processing unit is used to construct a bidirectional long short-term memory network model AtnBiLSTM embedded with a self-attention mechanism and use softmax for classification; The fourth processing unit is used to calculate the score of each sample on the corresponding algorithm using the minimum confidence algorithm and the maximum normalized logarithmic probability algorithm, and sort them accordingly to determine the Tang poetry samples that need to be recommended to the annotation interface.
6. The device for interactive visual annotation of named entities in Tang poetry as claimed in claim 5, characterized in that: The second processing module presents the Tang poetry data statistics view, recommended sample view and annotation view on the interface.
7. A storage medium, characterized in that: The storage medium stores a computer program, which, when running, executes the method for interactive visual annotation of named entities in Tang poetry as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Data labeling method, device and equipment and computer readable storage medium
CN111859854A
Method and device for identifying named entities in text and storage medium
CN117875328A
Large model data intelligent labeling method and system
CN119378564A