Industry Text Expansion Using Distant Supervision for Low-Resource Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing information extraction technologies struggle to effectively train models for low-resource industry fields due to insufficient amounts of training data, leading to inadequate performance in tasks such as entity recognition and relationship extraction.
Innovation Solution
The method employs a distant supervision approach to incrementally increase the volume of industry texts by leveraging associations between nouns in original texts and additional data sources, using techniques like synonym substitution and back translation to generate new samples, followed by error removal and model training to enhance the precision of subject-predicate-object triple extraction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional machine learning models are used for information extraction, then manually defined advanced features are required, but this increases the complexity and time consumption of feature engineering
Solution Approach 1:
The patent replaces manual feature engineering (mechanical process) with deep learning models that automatically learn features from data. The neural network structures automatically extract relevant features without human intervention, substituting the manual mechanical feature definition process with an automated learning-based system.
Solution Approach 2:
The deep learning model performs self-service by automatically learning and extracting features from the input data without requiring external manual feature engineering. The model autonomously identifies patterns and relationships in the data, eliminating the need for human experts to manually define features.
2Measurement precision
If deep learning models are trained on large-scale labeled data, then high accuracies and recall rates can be achieved, but this increases data collection and labeling costs
Solution Approach 1:
The patent applies preliminary action by using distant supervision to pre-label data before formal training. Knowledge graphs and entity relationships are constructed in advance to provide preliminary labels for training data, reducing the need for extensive manual labeling while still enabling effective deep learning model training.
Solution Approach 2:
The patent introduces knowledge graphs as an intermediary between raw text data and training labels. The knowledge graph serves as a mediator that automatically provides structural information and relationships, enabling the system to generate training labels without direct human annotation while maintaining high quality.
3Manufacturing precision
If manual feature engineering is performed to achieve high extraction accuracy, then the process becomes complex and time-consuming, but reducing feature complexity decreases extraction performance
Solution Approach 1:
The patent replaces manual feature engineering complexity with automated deep learning feature extraction. The neural network automatically learns optimal features from data without human intervention, eliminating the complexity of manual feature design while maintaining or improving extraction precision through data-driven feature learning.
Data Source
AI summary
A method for an industry text increment, as well as an electronic device and a computer readable storage medium for the same are provided. The method may include: acquiring an original industry text in a target industry field, an order of magnitude of a number of the original industry text being smaller than a preset first order of magnitude; and performing a sample incremental processing on the original industry text by using a distant supervision method, to obtain increased industry texts, an order of magnitude of a number of the increased industry texts is greater than a preset second order of magnitude, wherein the preset second order of magnitude is not smaller than the preset first order of magnitude.


