Unstructured Data Risk Analysis via Text Mining

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Analyzing risk using unstructured data is challenging due to its irregular structure and ambiguity, making it difficult for computers to understand and summarize large volumes of text-based information effectively.

Innovation Solution

A system and method that employs text mining techniques to deconstruct unstructured data into individual terms, convert them into a structured form, categorize, and quantify them, facilitating risk analysis and identification across an organization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If unstructured data is used for risk analysis, then the analysis can capture diverse and nuanced information from multiple sources, but the data becomes difficult to quantify and automatically analyze

Engineering Contradiction:
Improveability to capture diverse informationVSAvoiddifficulty to quantify and analyze
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments unstructured text data into individual terms through tokenization, then converts each term into a structured numerical representation (vector) that can be systematically analyzed. This segmentation transforms the unquantifiable text into quantifiable data points while preserving the diverse information content.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary transformation layer that converts unstructured text into structured numerical vectors through techniques like term frequency-counting and term-document matrix construction. This intermediary representation bridges the gap between diverse unstructured information and quantifiable analytical data.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If readers manually read and summarize many bodies of text, then they can understand the intended message, but it becomes impracticable to process large volumes of text in reasonable time

Engineering Contradiction:
Improveunderstanding of intended messageVSAvoidtime to process text
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces the mechanical human reading and summarizing process with an automated computational system. The system uses algorithms to automatically extract, convert, and analyze text data, eliminating the time-consuming manual process while maintaining accurate understanding through structured data representation and analysis.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service analysis where the text data automatically transforms into structured formats that can be directly analyzed without human intervention. The automated pipeline processes large volumes of text independently, providing timely results without requiring manual summarization efforts.

Inventive Principle:
Principle #25Self-service

3Difficulty of detecting and measuring

If unstructured data is transformed into structured form, then the data becomes easier to quantify and analyze, but the transformation process adds complexity

Engineering Contradiction:
Improveease to quantify and analyzeVSAvoidcomplexity of transformation process
Core Design Contradiction:
Difficulty of detecting and measuringVSDevice complexity

Solution Approach 1:

The patent changes the parameters of the data representation by transforming text from unstructured form to structured numerical vectors. Through parameter transformation (converting words to numbers, creating term-document matrices), the data becomes quantifiable and analyzable while the transformation process, though complex, follows systematic algorithms.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9141686B2Risk analysis using unstructured data
Publication Date: 2015.09.22 BANK OF AMERICA CORP
  • US9141686B2 patent drawing
  • US9141686B2 patent drawing
  • US9141686B2 patent drawing

AI summary

Unstructured data is received from a plurality of sources to facilitate risk analysis. The unstructured data comprises a plurality of bodies of text. Each body of text from the unstructured data is deconstructed into individual terms. The individual terms from each body of text are converted into a structured form. The individual terms in the structured form are categorized according to a comparison of the structured form to another structured form. The individual terms in the structured form are quantified according to at least the categorization of the individual terms.