NLP Unstructured Data Segmentation and Repository Linking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Natural language processing of unstructured data poses challenges in understanding human language, particularly in identifying logical segments and linking relevant data from repositories, due to irregularities and ambiguities in text-heavy, unorganized data.

Innovation Solution

The method employs natural language processing (NLP) to analyze unstructured data, partition it into logical segments based on predefined criteria, and link relevant data from a repository, using techniques such as machine learning and data structures like pointers and linked lists to enhance data organization and compliance with jurisdiction-specific obligations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If natural language processing is used to analyze unstructured data, then data processing efficiency is improved, but the complexity of the system increases

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides the unstructured data into logical segments based on predefined criteria (such as jurisdiction, topic, or content type). This segmentation allows the NLP system to process data in manageable portions, improving efficiency while reducing the complexity of handling the entire data corpus at once.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary layer between the raw unstructured data and the final processed output. This intermediary involves using NLP techniques as a mediator to transform and interpret the data, enabling efficient processing while managing system complexity through structured transformation layers.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Stability of the object's composition

If unstructured data is partitioned into logical segments, then data organization is improved, but the time required for processing increases

Engineering Contradiction:
Improvedata organizationVSAvoidprocessing time
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-defining the logical segment criteria and data structures before processing begins. This allows the system to quickly partition data according to predetermined rules rather than determining organization structure during processing, improving data organization while minimizing processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual or mechanical data organization methods with automated NLP-based segmentation. This substitution enables the system to automatically partition data into logical segments based on content analysis, improving organization efficiency while reducing the time required compared to manual processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of information

If data from repository is linked to unstructured data, then information completeness is improved, but the complexity of data linking increases

Engineering Contradiction:
Improveinformation completenessVSAvoiddata linking complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent uses logical segments as an intermediary structure that bridges the repository data and unstructured data. This intermediary layer simplifies the linking process by providing a structured framework for associating related information, improving information completeness while reducing the complexity of direct data linking.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a universal linking mechanism that can associate multiple types of data (policy documents, guidance files, and unstructured data) through common logical segment identifiers. This multi-functional approach improves information completeness while reducing linking complexity by using a unified association method.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11914597B2Natural language processing of unstructured data
Publication Date: 2024.02.27 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11914597B2 patent drawing
  • US11914597B2 patent drawing
  • US11914597B2 patent drawing

AI summary

A computer system for processing unstructured data, the computing system comprising a computer processor, a computer memory operatively coupled to the computer processor and the computer memory having disposed within it computer program instructions that, when executed by the processor, cause the computing system to carry out the steps of receiving unstructured data input from a client device, analyzing the unstructured data for features that satisfy logical segment criteria by using natural language processing (NLP), partitioning the unstructured data into logical segments based on satisfaction of the logical segment criteria, and linking data from a repository to the unstructured data based on the logical segments.