Automatic Template Generation for Information Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current information extraction from natural language texts is complex and typically requires manual creation of extraction rules by specialized ontology experts, limiting accessibility and efficiency.
Innovation Solution
A method and system for automatically generating templates for information extraction rules, allowing regular users to extract information objects by selecting sample tokens with desired semantic structures, without assistance from ontology engineers, using a template-generating graphical user interface and semantico-syntactic analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual creation of extraction rules by specialized ontology experts is used, then extraction precision is improved, but device complexity and ease of operation deteriorate
Solution Approach 1:
The system enables users to create extraction rules themselves by selecting sample tokens and defining templates through a graphical interface, eliminating the need for specialized ontology experts. The automated template generation system performs semantico-syntactic analysis and generates production rules based on user-provided samples, making the rule creation process self-service oriented.
Solution Approach 2:
The patent introduces an automated template generation system as an intermediary between users and the information extraction process. This intermediary handles the complex tasks of semantico-syntactic analysis, template construction, and production rule generation, translating simple user inputs into sophisticated extraction rules without requiring users to understand the underlying complexity.
2Measurement precision
If manual creation of extraction rules by specialized ontology experts is used, then extraction precision is improved, but ease of operation deteriorates
Solution Approach 1:
The system enables users to create extraction rules themselves by selecting sample tokens and defining templates through a graphical interface, eliminating the need for specialized ontology experts. The automated template generation system performs semantico-syntactic analysis and generates production rules based on user-provided samples, making the rule creation process self-service oriented.
Solution Approach 2:
The system allows users to create extraction rules by providing sample tokens that represent desired extraction patterns. The automated system analyzes these samples and generates templates that can be copied and applied to extract similar information objects from different texts, making rule creation intuitive and accessible to regular users.
3Productivity
If automated template generation is used, then ease of operation and productivity are improved, but manufacturing precision may deteriorate
Solution Approach 1:
The system incorporates a feedback mechanism where users can review generated templates and production rules, provide corrections, and refine the extraction patterns. The automated template generation system uses this feedback to improve subsequent template creations, balancing automation efficiency with extraction precision through iterative refinement.
Solution Approach 2:
The template generation system is designed to be dynamic and adaptive, adjusting its analysis depth and rule generation strategies based on the complexity of input samples and user preferences. The system can switch between fully automated generation and user-guided refinement modes, optimizing the balance between productivity and precision for different use cases.
Data Source
AI summary
Systems and methods for automatic generation of templates for information extraction rules to extract information objects from natural language texts. An example method of automatic generation of templates for information extraction comprises receiving, by a computer device, a first text fragment, comprising a first identifier of a first text token, wherein the first token comprises one or more words in a natural language and wherein the first token references a first information object from a first category of information objects; displaying, using a template-generating graphical user interface, a plurality of linguistic characteristics of the first token; receiving, via the template-generating graphical user interface, a first input identifying template attributes from the plurality of the linguistic characteristics of the first information object; generating a first template based at least in part on the identified template attributes; creating a first production rule for the first template; applying the first production rule to portions of a first natural language text matching the first template; displaying, using the template-generating graphical user interface, a second information object identified in the first natural language text using the first production rule.


