Method for labeling and verification of textual data
The method addresses the limitations of existing data labeling by using a deep learning model with interactive labeling and expert verification, ensuring accurate and time-efficient data labeling through iterative uncertainty-based selection.
Patent Information
- Application Number
- PCT/RU2024/050285
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-11-11
- Publication Date
- 2025-07-03
AI Technical Summary
Existing data labeling methods lack statistical language modeling tools for handling text context, leading to incomplete entity extraction under semantic ambiguity, and lack expert verification mechanisms for quality assurance.
A method involving pre-training a deep learning language model, interactive labeling of textual fragments, preprocessing, and using a classifier model to predict categories with uncertainty metrics, allowing expert verification at each iteration until consensus is reached.
Facilitates efficient and accurate data labeling by automatically selecting data for labeling while ensuring expert verification, reducing time and improving labeling accuracy.
Abstract
Description
[0001] Method for labeling and verification of textual data
[0002] The present invention relates to the scope of interactive labeling of textual data.
[0003] The invention seeks to provide the method of fast automatic selection of candidates-data examples for textual data labeling while labeling by a user (while maintaining an acceptable level of accuracy, compared to a user).
[0004] To assess the novelty of the claimed solution, let us consider a number of known technical means of similar purpose. Interactive machine learning system for automated annotation of information in text is known [US 2005027664 Al, published 03.02.2005], designed for labeling textual data interactively. The system starts with partially annotated training data or alternatively unannotated training data and a set of examples of what is to be learned. Through iterative interactive (user-participated) training sessions, the system trains annotators, and these are in turn used to discover more annotations. Once all of the text data or a sufficient amount of the text data is annotated, at a user's discretion, the system learns a final module (annotator or annotators) which can then be exported and used to label new textual data. As the iterative training process occurs, a user is selectively presented for review and appropriate action, allowing a user to check the labeled data and (if necessary) make corrections.
[0005] A disadvantage of the present invention is the lack of statistical language modeling tools, represented, inter alia, by deep learning language models. The lack of language modeling tools affects the ability to effectively handle the context of text fragments and leads to entity extraction based on incomplete information under conditions of semantic ambiguity.
[0006] A service for data labeling based on closed-loop active learning is known [US11048979 Bl, published 29.06.2023], which describes methods for data labeling based on active learning. Data labeling service based on active learning allows a user to create and manage large datasets for use in various machine learning systems. Machine learning-based schemes can be used to automate the labeling and management of datasets, improving labeling efficiency and reducing the associated timecost. Implementations use active learning techniques to reduce the amount of a dataset that requires manual labeling. As subsets of the dataset are labeled, this label data is used to train a model which can then identify additional objects in the dataset without manual intervention. The process may continue iteratively until the model converges; this scheme ensures a dataset to be labeled without the need for manual labeling for each object.
[0007] A disadvantage of the present invention is the determination of the quality score of the labeled data according to a predetermined threshold, without the possibility of verification by an expert, as well as putting the results of AUTO LABELING into a set of labeled data (LABEL STORE) in the absence of a reliable mechanism for verification.
[0008] The closest analogue is a Deep learning engine and methods for content and context aware data classification [US20200279105 Al, publ. 09.032020], which describes a deep learning framework and techniques for categorizing data into business categories and privacy levels, taking content and context into account. The deep learning engine includes a feature extraction module and a classification and labeling module; the feature extraction module extracts both context features and document features from documents and the classification and labeling module is configured for content and context aware data classification of the documents by business category and confidentiality level using neural networks.
[0009] The disadvantages include the following:
[0010] 1. Limited number of levels (tags) related to the level of data privacy;
[0011] 2. Lack of a policy of role allocation and labeling verification by an expert. The objective of the claimed invention is to provide an interactive data labeling method that saves time on data labeling, while preserving the possibility of verification. The technical result of the invention is to provide time-saving data labeling by automatically selecting data for labeling based on uncertainty metrics, while preserving the possibility to verify the results of labeling by monitoring the data at each iteration.
[0012] The technical result is achieved as follows.
[0013] The method for labeling and verification of textual data is implemented in several consecutive stages:
[0014] 1. Pre-training the deep learning language model on a trained data corpus that includes collections of texts with a broad thematic coverage;
[0015] 2. Labeling of task-relevant textual data using the software interface by selecting textual fragments of arbitrary length, assigning various user- defined categories to textual fragments, which are used as an additional training sample for the language model;
[0016] 3. Pre-processing of labeled textual data;
[0017] 4. Training of the language model on the newly labeled (tagged) data and vectorization of the tagged data;
[0018] 5. Category prediction on a set of unlabelled textual fragments using a classifier model paired with a language model; including calculation of metrics reflecting the degree of model uncertainty for each category, and sampling strategy: each object is assigned a degree of informativeness based on metrics, after which the most informative objects are selected for expert judgment, with maximum entropy and minimum confidence used as uncertainty metrics, ‘duplicates’ metric reflects the degree of uncertainty in assigning data belonging to one category to another category; the metric is formed by calculating the average confidence of the model for ‘category#l’, according to the labeled data for category ‘category#2’.
[0019] The sequence of steps 2-5 is repeated until consensus is reached between an expert's judgment and the uncertainty metrics for all objects provided for evaluation and their predicted categories, with the choice of the moment of consensus being determined by an expert.
[0020] The claimed technical solution is characterized by a number of additional optional features:
[0021] 1. Arbitrary definition of categories, for example, ‘legal entity’ («iop. JIHIIO»), ‘details’ («peKBH3HTBi»), ‘condition’ («ycjiOBne»), ‘measure of responsibility’ («Mepa OTBeTCTBeHHOCTn»), etc. by a user. The set of the defined categories is formed depending on the stylistic affiliation of a text and the specifics of the business task;
[0022] 2. Labeling of textual fragments with arbitrary depth of nesting (selection of sub-entities within entities); within this hierarchy, word combinations within the boundaries of the sentence-entity and words within the boundaries of the word combination-entity are selected as sub-entities.
[0023] The invention is explained in detail by the diagram.
[0024] The claimed method is implemented as follows.
[0025] • Pre-training a deep learning language model based on the Transformer architecture (RoBERTa implementation) (1);
[0026] • Data labeling by a user (2); during the labeling, a user selects arbitrary textual fragments with different syntactic organization (words, phrases, sentences, paragraphs, etc.) that are of interest from the point of view of the business task;
[0027] • Adding labeled data to the training sample (3); there is a possibility of expert evaluation of labeled data. Based on the evaluation results, labeled data can be sent back for relabeling;
[0028] • Pre-processing of labeled data (tokenization, stop-word removal, lemmatization, normalization to lower case, preserving the formatting of the original textual fragment) (4);
[0029] • Vectorization of tagged textual fragments (textual fragments tagged with a category) using a language model; with additional training, the tagged data and the contextual environment of the tagged fragments allow the language model to generate the necessary feature description to solve the category search and prediction task (5);
[0030] • Prediction of categories on a set of unlabelled data using a classifier model based on gradient boosting (6). For each category, metrics reflecting the degree of model uncertainty for each category are calculated (maximum entropy and minimum confidence are used as uncertainty metrics); based on the metrics, the degree of informativeness is determined, after which the most informative objects are selected for expert evaluation. In addition, the metric ‘category overlap’ is calculated, reflecting the degree of uncertainty in attributing data belonging to one category to another category; the calculation of this metric is formed by calculating the average confidence of the model for ‘ category#! ’, according to the labeled data for category ‘category#2’.
[0031] The above steps 2-5 are repeated until a consensus is reached between an expert's assessment and uncertainty metrics for all objects and their categories provided for assessment; the choice of the moment of consensus is determined by an expert.
[0032] The Industrial application of the claimed technical solution is possible using known technical and technological means.
[0033] As an example, the results at each step of the method implementation are given:
[0034] (1): the deep learning language model has been pre-trained;
[0035] (2): in the case of the task of extracting liability clauses in contractual documentation, the following text fragments and the rules of their joint arrangement in the text of the document can be used as labeled textual data (at a user's discretion):
[0036] - “... bear full responsibility ”;
[0037] - “... total amount of penalties ... should not exceed ...”;
[0038] - “... the only measure of responsibility”;
[0039] - “... for each violation”; The presence of the phrase ‘a fine / penalty in the amount of’ and the absence of the phrase ‘for each’ (limited liability marker). Marking of the corresponding textual fragment by a user may mean the definition of the category ‘Limited Liability’ or any other category (at the discretion of a user) for this fragment.
[0040] (3): labeled data has been added to the training sample;
[0041] (4): the labeled data has been preprocessed;
[0042] (5): vector representation (array of numerical vectors) of the labeled textual fragments has been obtained;
[0043] (6): category prediction on the unlabeled data set has been obtained.
Claims
Claims1. Method for labeling and verification of textual data is implemented in several consecutive steps: a) Pre-training of a language model of deep learning on the prepared data corpus including collections of texts of wide thematic orientation; b) Labeling of textual data relevant to the task by means of selection of textual fragments of arbitrary length, assigning of various arbitrarily defined user-defined data to the labeled data (used as an additional training sample for the language model); c) Preprocessing of the labeled data; d) Training of the language model including the newly labeled data and vectorization of the labeled data; e) Prediction of categories on a set of unlabelled data using a classifier model, paired with the language model; for each category, metrics reflecting the degree of model uncertainty for each category are calculated. Based on the metrics, the degree of informativeness is determined, after which the most informative objects are selected for expert evaluation. In addition, the metric ‘category overlap’ is calculated, reflecting the degree of uncertainty in attributing data belonging to one category to another category; the calculation of this metric is formed by calculating the average confidence of the model for ‘category#! ’, according to the labeled data for category ‘category#2’; the above steps are repeated until a consensus is reached between an expert's assessment and uncertainty metrics for all objects and their categories provided for assessment; the choice of the moment of consensus is determined by an expert.
2. The method according to claim 1, characterized as follows: arbitrary definition of categories, for example, ‘legal entity’, ‘details’, ‘condition’, ‘measure of responsibility’; the set of the defined categories is formeddepending on the stylistic affiliation of a text and the specifics of the business task.
3. The method according to claim 1, characterized as follows: labeling of textual fragments with arbitrary depth of nesting; within this hierarchy, word combinations within the boundaries of the sentence-entity and words within the boundaries of the word combination-entity are selected as sub-entities.
Citation Information
Patent Citations
Training classifiers used to extract information from natural language texts
RU2691855C1
Automatic determination of set of categories for document classification
RU2701995C2
Deep learning engine and methods for content and context aware data classification
US20200279105A1
Analyzing documents using machine learning
US20210065041A1