Semantic Annotation System for Unstructured Data Knowledge Graphs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The growth of unstructured data poses challenges in deriving actionable information due to its rapid expansion, human-generated nature, and imprecision, leading to difficulties in timely action in industries like finance and intelligence, where large volumes of data hinder meaningful connections and timely responses.

Innovation Solution

A computer-implemented method and system that uses a trained statistical language model to create semantic annotations for text data, aggregates messages, and performs global analytics functions, including error identification and correction, to improve language model training and generate annotated messages for better data understanding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If unstructured data is collected and stored to capture comprehensive information, then data completeness is improved, but data processing complexity and difficulty in deriving actionable information increase

Engineering Contradiction:
Improvedata completenessVSAvoiddata processing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts key information from unstructured data by training language models to predict and identify important entities, events, and relationships. The system selectively extracts only the meaningful patterns and semantic annotations from the vast amount of unstructured data, filtering out noise and redundant information to reduce processing complexity while maintaining data completeness.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary layer of semantic annotations and knowledge representations between the raw unstructured data and the final actionable insights. This intermediary knowledge graph structure organizes and pre-processes the data, making it easier to query and analyze without directly processing the entire unstructured data volume, thus reducing processing complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If more data is collected to improve understanding, then knowledge accuracy is improved, but time required for analysis and response increases

Engineering Contradiction:
Improveknowledge accuracyVSAvoidanalysis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-training language models on large datasets and pre-building knowledge graphs with semantic annotations. This preliminary processing enables the system to quickly query and analyze new data without reprocessing the entire dataset, thus maintaining high knowledge accuracy while reducing analysis time for specific queries.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the data processing task by dividing the knowledge representation into manageable components (entities, events, relationships) and organizing them in a structured knowledge graph. This segmentation allows the system to retrieve and analyze only the relevant segments needed for specific questions, reducing overall analysis time while maintaining accuracy through comprehensive pre-processing.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If statistical language models are trained on diverse data to improve annotation accuracy, then semantic understanding is improved, but training time and computational resources increase

Engineering Contradiction:
Improveannotation accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSDuration of action of moving object

Solution Approach 1:

The patent applies partial action by training language models on strategically selected representative samples rather than requiring processing of all available data. The system identifies and trains on the most informative and diverse data portions needed to achieve sufficient annotation accuracy, avoiding the excessive computational resources and time that would be required to process every data point exhaustively.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12026455B1Systems and methods for construction, maintenance, and improvement of knowledge representations
Publication Date: 2024.07.02 DIGITAL REASONING SYSTEMS INC
  • US12026455B1 patent drawing
  • US12026455B1 patent drawing
  • US12026455B1 patent drawing

AI summary

In one aspect, the present disclosure relates to a method which, in one example embodiment, can include reading text data corresponding to messages and creating semantic annotations to the text data to generate annotated messages. Creating the semantic annotations can include generating, at least in part by at least one trained statistical language model, predictive labels as annotations corresponding to language patterns associated with the text data. The method further includes aggregating the annotated messages and storing information associated with the aggregated annotated messages in a message store, and performing, based on information from the message store and associated with the messages, global analytics functions. The global analytics functions can include identifying an annotation error in the created semantic annotations, updating the respective semantic annotation to correct the annotation error, to form an updated semantic annotation, and back-propagating the updated semantic annotation into training data for further language model training.