Parser Training Using Pre-existing Structural Descriptions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current natural language processing parsers require extensive and costly human-annotated corpora for training, making them expensive to develop and maintain.
Innovation Solution
A computer-implemented method that uses the output from a pre-existing parser to automatically generate training data, eliminating the need for human input, by transforming structural descriptions into training data for a new parser.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If hand-annotated corpora are used to train parsers, then parser accuracy is improved, but development cost and time increase significantly
Solution Approach 1:
The patent applies preliminary action by using a pre-existing parser to automatically generate training data before training the new parser. The pre-existing parser processes the corpus and generates structural descriptions that serve as training data, eliminating the need for time-consuming manual annotation while maintaining training quality.
Solution Approach 2:
The patent uses copying by creating training data that replicates the quality and structure of hand-annotated corpora through automatic generation. The system copies the functional outcome of human annotation (structural descriptions) without requiring the actual human labor, thereby reducing time and cost while preserving accuracy.
2Measurement precision
If hand-annotated corpora are used to train parsers, then parser accuracy is improved, but development cost increases significantly
Solution Approach 1:
The patent applies copying by generating training data automatically through a pre-existing parser rather than through expensive manual annotation processes. This creates a cost-effective approach that replicates the quality of hand-annotated data without the associated high costs of human labor.
Solution Approach 2:
The patent replaces the mechanical system of manual human annotation with an automated computational system. The pre-existing parser mechanically processes text and generates structural descriptions, substituting human cognitive labor with algorithmic processing, thereby significantly reducing development costs.
3Ease of manufacture
If automatic generation of training data is used, then development cost is reduced, but training data quality may deteriorate
Solution Approach 1:
The patent applies feedback by using the output of a pre-existing parser to generate training data for a new parser. The pre-existing parser's structural descriptions serve as feedback that guides the training process, ensuring that the generated data maintains high quality and accuracy while reducing development costs.
Solution Approach 2:
The patent uses copying to create training data that replicates the quality characteristics of hand-annotated corpora. By copying the structural description format and quality standards from the pre-existing parser's output, the system maintains data quality while achieving cost reduction through automation.
Data Source
AI summary
A computer-implemented method for developing a parser is provided. The method includes accessing a corpus of sentences and parsing the sentences to generate a structural description of each sentence. The parser is trained based on the structural description of each sentence.


