Unsupervised Event Extraction From Unstructured Text
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing amount of unstructured data poses a challenge for machine-interpretable information extraction, as most data is initially unstructured and described in natural language, limiting its machine-interpretable nature and hindering automation in decision-making processes.
Innovation Solution
The system employs unsupervised machine learning techniques, specifically using abstract meaning representation (AMR) parsing and graph embeddings, to identify candidate event components, cluster them into event types, and label argument roles, generating structured event schema from unstructured text without requiring prior knowledge or annotated datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If supervised machine learning methods are used for event extraction, then extraction accuracy can be improved, but the requirement for annotated datasets and prior knowledge increases system complexity and reduces ease of deployment
Solution Approach 1:
The system performs self-service by automatically learning event extraction patterns from unstructured text without requiring external annotated datasets or prior knowledge. The unsupervised machine learning model autonomously identifies event triggers, arguments, and relationships, eliminating the need for manual annotation and reducing system complexity while maintaining extraction capability
Solution Approach 2:
The patent introduces an intermediary unsupervised machine learning model that bridges the gap between unstructured text and structured event extraction. This intermediary layer processes raw text and automatically generates event schemas without direct supervision, reducing the complexity associated with supervised methods while preserving extraction accuracy
2Measurement precision
If manual annotation of datasets is performed to train event extraction models, then extraction quality improves, but time consumption and productivity decrease
Solution Approach 1:
The system eliminates the need for manual annotation by employing unsupervised machine learning that automatically learns from unstructured text. This self-service approach removes the time-consuming annotation process entirely while maintaining extraction quality through autonomous pattern recognition and event schema generation
Solution Approach 2:
The unsupervised model performs preliminary action by pre-processing unstructured text and automatically identifying event patterns before any extraction task. This preliminary unsupervised learning enables the system to quickly process new text without requiring time-consuming manual annotation or retraining, thereby improving processing speed and productivity
3Ease of operation
If unstructured text is processed without structured event schema, then ease of operation is maintained, but machine-interpretable information extraction capability is reduced
Solution Approach 1:
The patent introduces an intermediary structured event schema layer that translates unstructured text into machine-interpretable format. This intermediary structure preserves operational simplicity by automatically generating schemas from unstructured input while enabling effective machine interpretation and information extraction through organized event representations
Solution Approach 2:
The system applies parameter changes by transforming unstructured text parameters into structured event schema parameters automatically. This transformation changes the state of the data from unstructured to structured format, enabling machine-interpretable information extraction while maintaining ease of operation through automated parameter transformation without manual intervention
Data Source
AI summary
Computer-implemented techniques for unsupervised event extraction are provided. In one instance, a computer implemented method can include parsing, by a system operatively coupled to a processor, unstructured text comprising event information to identify candidate event components. The computer implemented method can further include employing, by the system, one or more unsupervised machine learning techniques to generate structured event information defining events represented in the unstructured text based on the candidate event components.


