Automated Text Labeling for NLP Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Processing textual descriptions of complex technical systems, such as utility and industrial systems, is challenging due to their length, complexity, and lack of standardization, requiring significant effort and time to create labeled training sets for natural language processing techniques.
Innovation Solution
A method that automatically labels textual descriptions using clustering based on semantic distances, generating an annotated training batch for training an AI model to process textual system descriptions, which can then be used for system configuration and commissioning tasks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If natural language processing techniques are used to process textual descriptions of technical systems, then automated processing capability is improved, but the requirement for extensive labeled training sets increases processing complexity and time consumption
Solution Approach 1:
The system enables self-service by allowing the processing logic generation system to automatically generate and refine processing logic for textual descriptions without requiring extensive manual labeling of training sets. The system serves itself by iteratively improving its own processing capabilities through automated feedback mechanisms.
Solution Approach 2:
The patent applies preliminary action by pre-defining a controlled vocabulary and structured data models before processing textual descriptions. This preliminary structuring of the processing logic framework enables more efficient automated processing without requiring extensive labeled training data, as the system already has a prepared structure to work within.
2Measurement precision
If traditional labeled training sets are used for NLP model training, then processing accuracy can be improved, but the time and effort required to create labeled training sets increases
Solution Approach 1:
The system applies partial action by focusing labeling efforts only on critical portions of textual descriptions that require high accuracy, rather than labeling entire documents. The controlled vocabulary approach allows the system to achieve sufficient processing accuracy for specific tasks without the need for comprehensive labeling of all text, thereby reducing time consumption while maintaining adequate precision.
Solution Approach 2:
The patent implements local quality by applying different levels of processing rigor to different parts of textual descriptions. Critical sections requiring high accuracy receive more focused processing attention and labeling, while less critical sections use automated processing with acceptable accuracy levels. This selective approach maintains overall processing accuracy while significantly reducing the time required for training set creation.
3Productivity
If standardized processing approaches are used for textual descriptions, then processing efficiency is improved, but the ability to handle non-standardized and varied textual formats decreases
Solution Approach 1:
The system achieves universality through its controlled vocabulary and structured data models that can handle multiple textual formats and domains. The processing logic generation system is designed to work with various types of technical descriptions (system specifications, operational procedures, maintenance manuals) using a unified approach, enabling it to process diverse text formats efficiently while maintaining adaptability through the flexible controlled vocabulary structure.
Data Source
Figure 1
Figure 2~3
Figure 4~5
AI summary
A processing logic generation system is operative to generate a processing logic that can automatically process a textual system description related to a system. The processing logic generation system may perform a method that comprises performing an automatic labeling of textual descriptions included in a training batch of textual descriptions to generate an annotated training batch, wherein performing the automatic labeling may comprise performing a clustering technique to annotate a text item included in the training batch with a label from a set of labels based on semantic distances of the text item from text items previously assigned to labels of the set of labels.